Method and apparatus for signaling subpicture partitioning information

The method optimizes video encoding by efficiently signaling sub-picture segmentation information, reducing redundancy and improving resource utilization through clear rules for sub-picture identifier signaling in video encoding standards.

JP2025107257AActive Publication Date: 2025-07-17ALIBABA GROUP HOLDING LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025074342
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-12-27
Filing Date
2025-04-28
Publication Date
2025-07-17
Estimated Expiration
2040-12-18

AI Technical Summary

Technical Problem

Existing video encoding standards face inefficiencies in signaling sub-picture segmentation information, leading to redundant and unnecessary signaling of sub-picture identifiers, which can result in increased bitstream complexity and resource utilization.

Method used

A method and apparatus for signaling sub-picture segmentation information by determining the presence of sub-picture information in a bitstream and providing clear rules for signaling sub-picture identifiers within sequence parameter sets, picture parameter sets, or picture headers, ensuring efficient and unambiguous identification of sub-pictures.

Benefits of technology

This approach reduces bitstream redundancy and signaling overhead, enhancing encoding efficiency and resource utilization by minimizing unnecessary sub-picture identifier signaling, thus optimizing video processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025107257000001_ABST
    Figure 2025107257000001_ABST
Patent Text Reader

Abstract

To provide a method and apparatus for signaling subpicture partitioning information.SOLUTION: An exemplary method includes: determining, according to a subpicture information present flag signaled in a bitstream, whether the bitstream comprises subpicture information; and, in response to the bitstream comprising the subpicture information, signaling in the bitstream at least one of a number of subpictures in a picture, a width, a height, a position and an identifier (ID) mapping of a target subpicture, a subpic_treated_as_pic_flag, and a loop_filter_across_subpic_enabled_flag.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications

[0001] This disclosure claims priority to U.S. Provisional Patent Application No. 62 / 954,014, filed on December 27, 2019, which is hereby incorporated by reference in its entirety.

[0002] Technical Field

[0002] This disclosure generally relates to video processing, and more particularly, to methods and apparatus for signaling sub - picture segmentation information.

Background Art

[0003] Background

[0003] Video is a set of static pictures (or "frames") that capture visual information. To reduce memory storage and transmission bandwidth, video can be compressed before storage or transmission and restored before display. The compression process is usually referred to as encoding, and the restoration process is usually referred to as decoding. Most commonly, there are various video encoding formats that use standardized video encoding techniques based on prediction, transformation, quantization, entropy encoding, and in - loop filtering. Video encoding standards such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, which specify a particular video encoding format, have been developed by standardization organizations. As evolving video encoding technologies are successively adopted by video standards, the encoding efficiency of new video encoding standards becomes even higher.

Summary of the Invention

Means for Solving the Problems

[0004]

[0004] In some embodiments, an exemplary method for signaling sub-picture segmentation information includes determining whether a bitstream includes sub-picture information according to a sub-picture information presence flag signaled in the bitstream, and in response to the bitstream including sub-picture information, signaling in the bitstream at least one of the number of sub-pictures in a picture, the width, height, position, and identifier (ID) mapping of a target sub-picture, subpic_treated_as_pic_flag, and loop_filter_across_subpic_enabled_flag.

[0005]

[0005] In some embodiments, an exemplary video processing device includes at least one memory for storing instructions and at least one processor. The at least one processor is configured to execute instructions to cause the device to determine whether a bitstream includes sub-picture information according to a sub-picture information presence flag signaled in the bitstream, and in response to the bitstream including sub-picture information, signal in the bitstream at least one of the number of sub-pictures in a picture, the width, height, position, and identifier (ID) mapping of a target sub-picture, subpic_treated_as_pic_flag, and loop_filter_across_subpic_enabled_flag.

[0006]

[0006] In some embodiments, an exemplary non - transitory computer - readable storage medium stores a set of instructions. The set of instructions is executable by one or more processing devices to cause a video processing device to determine whether a bitstream includes sub - picture information according to a sub - picture information presence flag signaled within the bitstream, and in response to the bitstream including sub - picture information, signal within the bitstream at least one of the number of sub - pictures within a picture, the width, height, position, and identifier (ID) mapping of a target sub - picture, subpic_treated_as_pic_flag, and loop_filter_across_subpic_enabled_flag.

[0007] Brief Description of the Drawings

[0007] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings. The various features shown in the drawings are not drawn to scale.

Brief Description of the Drawings

[0008]

Figure 1

[0008] It is a schematic diagram showing the structure of an exemplary video sequence according to some embodiments of the present disclosure.

Figure 2A

[0009] A schematic diagram showing an exemplary encoding process of a hybrid video encoding system according to an embodiment of the present disclosure is shown.

Figure 2B

[0010] A schematic diagram showing another exemplary encoding process of a hybrid video encoding system according to an embodiment of the present disclosure is shown.

Figure 3A

[0011] A schematic diagram showing an exemplary decoding process of a hybrid video encoding system according to an embodiment of the present disclosure is shown.

Figure 3B

[0012] A schematic diagram showing another exemplary decoding process of a hybrid video encoding system according to an embodiment of the present disclosure is shown.

Figure 4

[0013] A block diagram of an exemplary device for encoding or decoding video according to some embodiments of the present disclosure is shown.

Figure 5

[0014] A schematic diagram showing an example of a picture divided into coding tree units (CTUs) according to some embodiments of the present disclosure.

Figure 6

[0015] A schematic diagram showing an example of a picture divided into tiles and raster scan lines according to some embodiments of the present disclosure.

Figure 7

[0016] A schematic diagram showing an example of a picture divided into tiles and rectangular slices according to some embodiments of the present disclosure.

Figure 8

[0017] A schematic diagram showing another example of a picture divided into tiles and rectangular slices according to some embodiments of the present disclosure.

Figure 9

[0018] A schematic diagram showing an example of a picture divided into sub-pictures according to some embodiments of the present disclosure.

Figure 10

[0019] Exemplary Table 1 showing an exemplary sequence parameter set (SPS) syntax for sub-picture division according to some embodiments of the present disclosure.

Figure 11

[0020] Exemplary Table 2 showing an exemplary SPS syntax for sub-picture identifiers according to some embodiments of the present disclosure.

Figure 12

[0021] Exemplary Table 3 showing an exemplary picture parameter set (PPS) syntax for sub-picture identifiers according to some embodiments of the present disclosure.

Figure 13

[0022] Exemplary Table 4 showing an exemplary picture header (PH) syntax for sub-picture identifiers according to some embodiments of the present disclosure.

Figure 14

[0023] A schematic diagram showing exemplary bitstream compliance constraints according to some embodiments of the present disclosure.

Figure 15

[0024] Exemplary Table 5 showing another exemplary PH syntax of sub-picture identifiers according to some embodiments of the present disclosure is shown.

Figure 16

[0025] Exemplary Table 6 showing another exemplary PH syntax of sub-picture identifiers according to some embodiments of the present disclosure is shown.

Figure 17A

[0026] Exemplary Table 7A showing an exemplary SPS syntax according to some embodiments of the present disclosure is shown.

Figure 17B

[0027] Exemplary Table 7B showing another exemplary SPS syntax according to some embodiments of the present disclosure is shown.

Figure 18

[0028] Exemplary Table 8 showing another exemplary SPS syntax according to some embodiments of the present disclosure is shown.

Figure 19

[0029] Exemplary Table 9 showing another exemplary SPS syntax according to some embodiments of the present disclosure is shown.

Figure 20

[0030] Exemplary Table 10 showing another exemplary SPS syntax according to some embodiments of the present disclosure is shown.

Figure 21

[0031] A flowchart of an exemplary video processing method according to some embodiments of the present disclosure is shown.

Figure 22

[0032] A flowchart of another exemplary video processing method according to some embodiments of the present disclosure is shown.

Figure 23

[0033] A flowchart of another exemplary video processing method according to some embodiments of the present disclosure is shown.

Figure 24

[0034] A flowchart of another exemplary video processing method according to some embodiments of the present disclosure is shown.

Figure 25

[0035] A flowchart of another exemplary video processing method according to some embodiments of the present disclosure is shown.

DETAILED DESCRIPTION OF THE INVENTION

[0009] Detailed Description

[0036] Here, reference is made in detail to exemplary embodiments illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which like numerals in different drawings represent the same or similar elements unless otherwise indicated. The implementations shown in the following description of the exemplary embodiments do not represent all implementations in accordance with the present invention. Rather, they are merely examples of apparatuses and methods in accordance with aspects related to the present invention as set forth in the appended claims. Specific aspects of the present disclosure are described in more detail below. Where terms and / or definitions incorporated by reference conflict, the terms and definitions provided herein shall prevail.

[0010]

[0037] The Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.

[0011]

[0038] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET has been developing technologies beyond HEVC using the Joint Exploration Model (JEM) reference software. Since the coding technology was incorporated into JEM, JEM has achieved substantially higher coding performance than HEVC.

[0012]

[0039] The VVC standard has been recently developed and continues to incorporate more coding technologies that bring better compression performance. VVC is based on the same hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.

[0013]

[0040] A video is a set of static pictures (or "frames") arranged in time series for storing visual information. A video capture device (e.g., a camera) can be used to capture and store those pictures in time series, and a video playback device (e.g., a TV, computer, smartphone, tablet computer, video player, or any end-user terminal having a display function) can be used to display such pictures in time series. Also, depending on the application, the video capture device can transmit the captured video in real time to a video playback device (e.g., a computer having a monitor) for supervision, holding a meeting, or live broadcast.

[0014]

[0041] In order to reduce the memory space and transmission bandwidth required by such applications, the video can be compressed before being stored and transmitted, and restored before being displayed. Compression and restoration can be implemented by software executed by a processor (for example, the processor of a general-purpose computer) or by special hardware. The module for compression is generally referred to as an "encoder", and the module for restoration is generally referred to as a "decoder". The encoder and decoder can be collectively referred to as a "codec". The encoder and decoder can be implemented as any of various suitable hardware, software, or combinations thereof. For example, the hardware implementation of the encoder and decoder can include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. The software implementation of the encoder and decoder can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and restoration can be implemented by various algorithms or standards such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, or the like. Depending on the application, the codec can restore the video from a first encoding standard and re-compress the restored video using a second encoding standard. In this case, the codec can be referred to as a "transcoder".

[0015]

[0042] The video encoding process can identify and retain the useful information that can be used to reconstruct the picture and ignore the information that is not important for reconstruction. If the ignored, unimportant information cannot be completely reconstructed, such an encoding process can be referred to as "irreversible". Otherwise, it can be referred to as "reversible". Most encoding processes are irreversible, which is a trade-off for reducing the required memory space and transmission bandwidth.

[0016]

[0043] Useful information of a picture being encoded (referred to as the "current picture") includes changes with respect to a reference picture (e.g., a previously encoded and reconstructed picture). Such changes can include changes in pixel position, brightness, or color, among which the change in position is the most important. The change in the position of a group of pixels representing an object can reflect the movement of the object between the reference picture and the current picture.

[0017]

[0044] A picture encoded without referring to another picture (i.e., it is its own reference picture) is referred to as an "I picture". A picture encoded using a previous picture as a reference picture is referred to as a "P picture". A picture encoded using both a previous picture and a future picture as reference pictures (i.e., the reference is "bidirectional") is referred to as a "B picture".

[0018]

[0045] FIG. 1 shows the structure of an exemplary video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 can be a live video or a captured and archived video. The video 100 can be a real-world video, a computer-generated video (e.g., a computer game video), or a combination thereof (e.g., a real-world video with an augmented reality effect). The video sequence 100 can be input from a video capture device (e.g., a camera), a video archive including previously captured videos (e.g., video files stored in a storage device), or a video supply interface (e.g., a video broadcast transceiver) for receiving videos from a video content provider.

[0019]

[0046] As shown in FIG. 1, video sequence 100 can include a series of pictures temporally arranged along a timeline, including pictures 102, 104, 106, and 108. Pictures 102 to 106 are consecutive, and there are additional pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I picture, and its reference picture is picture 102 itself. Picture 104 is a P picture, and its reference picture is picture 102, as indicated by the arrow. Picture 106 is a B picture, and its reference pictures are pictures 104 and 108, as indicated by the arrows. In some embodiments, the reference picture of a picture (e.g., picture 104) may not be immediately before or after that picture. For example, the reference picture of picture 104 can be a picture before picture 102. Note that the reference pictures of pictures 102 to 106 are merely examples, and the present disclosure does not limit the embodiments of the reference pictures to the examples shown in FIG. 1.

[0020]

[0047] Typically, due to the computational complexity of such tasks, video codecs do not encode or decode an entire picture at once. Instead, they can divide the picture into basic segments and encode or decode the picture segment by segment. Such a basic segment is referred to as a basic processing unit ("BPU") in the present disclosure. For example, the structure 110 in FIG. 1 shows an exemplary structure of a picture (e.g., any of pictures 102-108) of the video sequence 100. In structure 110, the picture is divided into 4×4 basic processing units, and their boundaries are shown as dashed lines. In some embodiments, the basic processing unit may be referred to as a "macroblock" in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as a "coding tree unit" ("CTU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit can have a variable size in the picture, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, etc., or any shape and size of pixels. The size and shape of the basic processing unit can be selected based on a balance between coding efficiency and the level of detail to be maintained in the basic processing unit for the picture.

[0021]

[0048] The basic processing unit can be a logical unit that can include groups of different types of video data stored in a computer memory (e.g., within a video frame buffer). For example, the basic processing unit of a color picture can include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luma and chroma components can have the same size as the basic processing unit. The luma and chroma components may be referred to as "coding tree blocks" ("CTB") in some video coding standards (e.g., H.265 / HEVC or H.266 / VVC). Any operation performed on the basic processing unit can be repeatedly performed on each of its luma and chroma components.

[0022]

[0049] Video encoding has multiple operation stages, and examples thereof are shown in FIGS. 2A-2B and FIGS. 3A-3B. At each stage, the size of the basic processing unit can still be too large for processing, and thus can be further divided into segments referred to as "basic processing subunits" in the present disclosure. In some embodiments, the basic processing subunits can be referred to as "blocks" in some video encoding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as "coding units" ("CUs") in some other video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunits can have a size that is the same as or smaller than the basic processing unit. Similar to the basic processing unit, the basic processing subunits are also logical units that can include groups of different types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in a computer memory (e.g., within a video frame buffer). Any operation performed on a basic processing subunit can be repeatedly performed for each of its luma and chroma components. Note that such division can be carried out to further levels as required by the processing. Also note that different stages can divide the basic processing unit in different ways.

[0023]

[0050] For example, in the mode decision stage (an example of which is shown in FIG. 2B), the encoder can determine which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for the basic processing unit, but the basic processing unit can be too large to make such a decision. The encoder can divide the basic processing unit into a plurality of basic processing subunits (e.g., CUs as in the case of H.265 / HEVC or H.266 / VVC) and determine the type of prediction for each individual basic processing subunit.

[0024]

[0051] As another example, in the prediction stage (an example thereof is shown in FIGS. 2A to 2B), the coder can perform a prediction operation at the level of a basic processing subunit (e.g., a CU). However, in some cases, the basic processing subunit may still be too large to process. The coder can further divide the basic processing subunit into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), and the prediction operation can be performed at that level.

[0025]

[0052] As another example, in the transform stage (an example thereof is shown in FIGS. 2A to 2B), the coder can perform a transform operation for a residual basic processing subunit (e.g., a CU). However, in some cases, the basic processing subunit may still be too large to process. The coder can further divide the basic processing subunit into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), and the transform operation can be performed at that level. It should be noted that the division method of the same basic processing subunit may be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.

[0026]

[0053] In the structure 110 of FIG. 1, the basic processing unit 112 is further divided into 3×3 basic processing subunits, and their boundaries are shown as dotted lines. Different basic processing units of the same picture can be divided into basic processing subunits in different ways.

[0027]

[0054] Depending on the implementation form, in order to bring parallel processing and error tolerance capabilities to video encoding and decoding, a picture can be divided into areas for processing. As a result, the encoding or decoding process does not have to depend on information from any other area of the picture with respect to the area of the picture. In other words, each area of the picture can be processed independently. By doing so, the codec can process different areas of the picture in parallel, thus increasing the encoding efficiency. Also, when the data of an area is damaged during processing or lost during network transmission, the codec can correctly encode or decode other areas of the same picture without relying on the damaged or lost data, thus bringing an error tolerance capability. In some video encoding standards, a picture can be divided into different types of areas. For example, H.265 / HEVC and H.266 / VVC provide two types of areas: "slices" and "tiles". It should also be noted that different pictures of video sequence 100 can have different partitioning methods for dividing the picture into areas.

[0028]

[0055] For example, in FIG. 1, structure 110 is divided into three areas 114, 116, and 118, and their boundaries are shown as solid lines inside structure 110. Area 114 contains four basic processing units. Each of areas 116 and 118 contains six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and areas of structure 110 in FIG. 1 are merely examples, and the present disclosure does not limit its embodiments.

[0029]

[0056] FIG. 2A shows a schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure. For example, the encoding process 200A can be performed by an encoder. As shown in FIG. 2A, the encoder can encode a video sequence 202 into a video bitstream 228 according to process 200A. Similar to the video sequence 100 in FIG. 1, the video sequence 202 can include a set of pictures (referred to as “original pictures”) arranged in chronological order. Similar to the structure 110 in FIG. 1, each original picture of the video sequence 202 can be divided by the encoder into basic processing units, basic processing subunits, or regions for processing. In some embodiments, the encoder can perform process 200A at the level of basic processing units for each original picture of the video sequence 202. For example, the encoder can perform process 200A in an iterative manner, in which case the encoder can encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder can perform process 200A in parallel for regions (e.g., regions 114-118) of each original picture of the video sequence 202.

[0030]

[0057] In FIG. 2A, the coder can supply the basic processing unit of the original picture of video sequence 202 (referred to as "original BPU") to prediction stage 204 and generate prediction data 206 and prediction BPU 208. The coder can subtract prediction BPU 208 from the original BPU to generate residual BPU 210. The coder can supply residual BPU 210 to transform stage 212 and quantization stage 214 and generate quantized transform coefficients 216. The coder can supply prediction data 206 and quantized transform coefficients 216 to binary encoding stage 226 and generate video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 can be referred to as the "forward path". During process 200A, after quantization stage 214, the coder can supply quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 and generate reconstructed residual BPU 222. The coder can add reconstructed residual BPU 222 to prediction BPU 208 and generate prediction reference 224 for use in prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A can be referred to as the "reconstruction path". The reconstruction path can be used to ensure that both the coder and the decoder use the same reference data for prediction.

[0031]

[0058] The coder can iteratively perform process 200A to encode each original BPU of the original picture (within the forward path) and generate prediction reference 224 for encoding the next original BPU of the original picture (within the reconstruction path). After encoding all the original BPUs of the original picture, the coder can proceed to encode the next picture within video sequence 202.

[0032]

[0059] Referring to process 200A, the coder can receive video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" can refer to receiving, inputting, acquiring, obtaining, getting, reading, accessing, or any act by any means for inputting data.

[0033]

[0060] In prediction stage 204, in the current iteration, the coder can receive the original BPU and prediction criterion 224, perform a prediction operation, and generate prediction data 206 and prediction BPU 208. Prediction criterion 224 can be generated from the reconstruction path of a previous iteration of process 200A. The purpose of prediction stage 204 is to reduce information redundancy by extracting prediction data 206, and prediction data 206 can be used to reconstruct the original BPU as prediction BPU 208 from prediction data 206 and prediction criterion 224.

[0034]

[0061] Ideally, prediction BPU 208 can be the same as the original BPU. However, due to non-ideal prediction and reconstruction operations, prediction BPU 208 generally differs slightly from the original BPU. To record such a difference, after generating prediction BPU 208, the coder can subtract it from the original BPU to generate residual BPU 210. For example, the coder can subtract the pixel values (e.g., grayscale values or RGB values) of prediction BPU 208 from the corresponding pixel values of the original BPU. Each pixel of residual BPU 210 can have a residual value as a result of such subtraction between the corresponding pixels of the original BPU and prediction BPU 208. Compared with the original BPU, prediction data 206 and residual BPU 210 can have fewer bits, but they can be used to reconstruct the original BPU without significant quality degradation. Therefore, the original BPU is compressed.

[0035]

[0062] To further compress the residual BPU 210, in the transformation stage 212, the coder can reduce the spatial redundancy of the residual BPU 210 by decomposing it into a set of two-dimensional “basis patterns”, with each basis pattern associated with a “transformation coefficient”. The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a frequency component of the change in the residual BPU 210 (e.g., the frequency of luminance change). None of the basis patterns can be reproduced from any combination (e.g., linear combination) of any other basis patterns. In other words, the decomposition can decompose the change in the residual BPU 210 into the frequency domain. Such a decomposition is similar to the discrete Fourier transform of a function, where in this case the basis patterns are similar to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the transformation coefficients are similar to the coefficients associated with the basis functions.

[0036]

[0063] Different transformation algorithms can use different basis patterns. For example, various transformation algorithms such as the discrete cosine transform, the discrete sine transform, or the like can be used in the transformation stage 212. The transformation in the transformation stage 212 is invertible. That is, the coder can recover the residual BPU 210 by means of the inverse operation of the transformation (referred to as “inverse transformation”). For example, to recover the pixels of the residual BPU 210, the inverse transformation can multiply the corresponding pixel values of the basis patterns by their respective associated coefficients and add up the products to generate a weighted sum. For a video coding standard, both the coder and the decoder can use the same transformation algorithm (and thus the same basis patterns). Therefore, the coder can record only the transformation coefficients, and the decoder can reconstruct the residual BPU 210 from the transformation coefficients without receiving the basis patterns from the coder. Compared with the residual BPU 210, the transformation coefficients can have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Therefore, the residual BPU 210 is further compressed.

[0037]

[0064] The coder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns can represent different change frequencies (e.g., luminance change frequencies). Since the human eye is generally more adept at recognizing low-frequency changes, the coder can ignore the information of high-frequency changes without causing significant quality degradation in decoding. For example, in the quantization stage 214, the coder can divide each transform coefficient by an integer value (referred to as the "quantization parameter") and round the quotient to the nearest integer to generate the quantized transform coefficient 216. After such an operation, some of the transform coefficients of the high-frequency basis pattern can be converted to 0, and the transform coefficients of the low-frequency basis pattern can be converted to smaller integers. The coder can ignore the quantized transform coefficients 216 with a value of 0, thereby further compressing the transform coefficients. The quantization process is also invertible, and in this case, the quantized transform coefficient 216 can be reconstructed into the transform coefficient in the inverse operation of quantization (referred to as "inverse quantization").

[0038]

[0065] Since the coder ignores the remainder of such division in the rounding operation, the quantization stage 214 can be non-invertible. Typically, the quantization stage 214 can contribute to the largest information loss in the process 200A. The greater the information loss, the fewer bits required for the quantized transform coefficient 216. To obtain different information loss levels, the coder can use different values of the quantization parameter or any other parameter of the quantization process.

[0039]

[0066] In the binary encoding stage 226, the coder can encode the predicted data 206 and the quantized transform coefficients 216 using binary encoding techniques such as, for example, entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context adaptive binary arithmetic encoding, or any other reversible or irreversible compression algorithm. In some embodiments, in addition to the predicted data 206 and the quantized transform coefficients 216, the coder can encode other information in the binary encoding stage 226, such as, for example, the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the type of transformation in the transformation stage 212, the parameters of the quantization process (e.g., quantization parameters), the coder control parameters (e.g., bitrate control parameters), or the like. The coder can generate a video bitstream 228 using the output data of the binary encoding stage 226. In some embodiments, the video bitstream 228 can be further packetized for network transmission.

[0040]

[0067] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the coder can perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In the inverse transformation stage 220, the coder can generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The coder can add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224 that will be used in the next iteration of process 200A.

[0041]

[0068] Note that other variations of process 200A can also be used to encode video sequence 202. In some embodiments, the steps of process 200A can be performed by an encoder in a different order. In some embodiments, one or more steps of process 200A can be combined into a single step. In some embodiments, a single step of process 200A can be divided into multiple steps. For example, the transform step 212 and the quantization step 214 can be combined into a single step. In some embodiments, process 200A can include additional steps. In some embodiments, process 200A can omit one or more steps in FIG. 2A.

[0042]

[0069] FIG. 2B shows a schematic diagram of another exemplary encoding process 200B according to an embodiment of the present disclosure. Process 200B can be changed from process 200A. For example, process 200B can be used by an encoder compliant with a hybrid video encoding standard (e.g., H.26x series). Compared with process 200A, the forward path of process 200B additionally includes a mode decision step 230, and the prediction step 204 is divided into a spatial prediction step 2042 and a temporal prediction step 2044. The reconstruction path of process 200B additionally includes a loop filter step 232 and a buffer 234.

[0043]

[0070] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra prediction") can use pixels from one or more already-encoded adjacent BPUs within the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include adjacent BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter prediction") can use regions from one or more already-encoded pictures to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include encoded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.

[0044]

[0071] Referring to process 200B, within the forward path, the coder performs prediction operations at spatial prediction stage 2042 and temporal prediction stage 2044. For example, at spatial prediction stage 2042, the coder can perform intra prediction. For the original BPU of the picture being encoded, the prediction reference 224 can include one or more adjacent BPUs that are encoded (within the forward path) and reconstructed (within the reconstruction path) in the same picture. The coder can generate the predicted BPU 208 by extrapolating the adjacent BPUs. The extrapolation technique can include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, or the like. In some embodiments, the coder can perform the extrapolation at the pixel level, such as by extrapolating the value of the corresponding pixel for each pixel of the predicted BPU 208. The adjacent BPUs used for extrapolation can be located with respect to the original BPU in various directions, such as a vertical direction (e.g., above the original BPU), a horizontal direction (e.g., to the left of the original BPU), a diagonal direction (e.g., bottom left, bottom right, top left, or top right of the original BPU), or any direction defined in the video coding standard being used. For intra prediction, the prediction data 206 can include, for example, the location (e.g., coordinates) of the adjacent BPUs used, the size of the adjacent BPUs used, the parameters of the extrapolation, the direction of the adjacent BPUs used with respect to the original BPU, or the like.

[0045]

[0072] As another example, in the temporal prediction stage 2044, the coder can perform inter prediction. For the original BPU of the current picture, the prediction reference 224 can include one or more pictures (referred to as "reference pictures") that are encoded (within the forward path) and reconstructed (within the reconstruction path). In some embodiments, the reference pictures can be encoded and reconstructed for each BPU. For example, the coder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all the reconstructed BPUs of the same picture have been generated, the coder can generate the reconstructed picture as a reference picture. The coder can perform an operation of "motion estimation" to search for a matching region within the range of the reference pictures (referred to as the "search window"). The location of the search window within the reference picture can be determined based on the location of the original BPU of the current picture. For example, the search window can be centered at a location within the reference picture that has the same coordinates as the original BPU within the current picture and can be extended outward over a predetermined distance. When the coder identifies a region similar to the original BPU within the search window (e.g., by using a pixel recursion algorithm, a block matching algorithm, or the like), the coder can determine such a region as the matching region. The matching region can have dimensions that are different from those of the original BPU (e.g., smaller than, equal to, larger than, or of a different shape than the original BPU). Since the reference picture and the current picture are temporally separated within the timeline (e.g., as shown in FIG. 1), the matching region can be considered to "move" towards the location of the original BPU as time passes. The coder can record such a direction and distance of motion as a "motion vector". When multiple reference pictures are used (e.g., as picture 106 in FIG. 1), the coder can search for a matching region for each reference picture and determine its associated motion vector. In some embodiments, the coder can assign weights to the pixel values of the matching regions of the respective matching reference pictures.

[0046]

[0073] Motion estimation can be used to identify various types of motion, such as translation, rotation, zooming, or the like. For inter prediction, the prediction data 206 can include, for example, the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference pictures, the weight associated with the reference pictures, or the like.

[0047]

[0074] To generate the prediction BPU 208, the coder can perform an operation of "motion compensation". Motion compensation can be used to reconstruct the prediction BPU 208 based on the prediction data 206 (e.g., motion vector) and the prediction reference 224. For example, the coder can move the matching region of the reference picture according to the motion vector, in which case the coder can predict the original BPU of the current picture. When multiple reference pictures are used (e.g., picture 106 in FIG. 1), the coder can move the matching regions of the reference pictures according to their respective motion vectors and average the pixel values of the matching regions. In some embodiments, when the coder assigns weights to the pixel values of the matching regions of each matching reference picture, the coder can add the weighted sum of the pixel values to the moved matching region.

[0048]

[0075] In some embodiments, inter prediction can be unidirectional or bidirectional. Unidirectional inter prediction can use one or more reference pictures in the same temporal direction with respect to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter prediction picture where the reference picture (i.e., picture 102) precedes picture 104. Bidirectional inter prediction can use one or more reference pictures in both temporal directions with respect to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter prediction picture where the reference pictures (i.e., pictures 104 and 108) are in both temporal directions with respect to picture 104.

[0049]

[0076] Referring still to the forward path of process 200B, after the spatial prediction 2042 and the temporal prediction stage 2044, at the mode decision stage 230, the coder can select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of process 200B. For example, the coder can perform a rate-distortion optimization technique. In this technique, the coder can select a prediction mode to minimize the value of a cost function that depends on the bitrate of the candidate prediction modes and the distortion of the reconstructed reference pictures under such candidate prediction modes. Depending on the selected prediction mode, the coder can generate the corresponding prediction BPU 208 and prediction data 206.

[0050]

[0077] In the reconstruction path of process 200B, when the intra prediction mode is selected in the forward path, after generating a prediction reference 224 (e.g., the currently encoded and reconstructed current BPU in the current picture), the coder can directly supply the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). When the inter prediction mode is selected in the forward path, after generating a prediction reference 224 (e.g., the current picture in which all BPUs are encoded and reconstructed), the coder can supply the prediction reference 224 to the loop filter stage 232, where the coder can apply a loop filter to the prediction reference 224 to reduce or eliminate the distortion (e.g., blocking artifacts) introduced by inter prediction. The coder can apply various loop filter techniques, such as deblocking, sample adaptive offset, adaptive loop filter, or the like, in the loop filter stage 232. The loop-filtered reference picture can be stored in buffer 234 (or "decoded picture buffer") for later use (e.g., for use as an inter prediction reference picture for future pictures of video sequence 202). The coder can store one or more reference pictures in buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the coder can encode loop filter parameters (e.g., loop filter strength) together with quantization transform coefficients 216, prediction data 206, and other information in the binary encoding stage 226.

[0051]

[0078] FIG. 3A shows a schematic diagram of an exemplary decoding process 300A according to an embodiment of the present disclosure. Process 300A may be a restoration process corresponding to the compression process 200A in FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder can decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 may be very similar to the video sequence 202. However, due to information loss in the compression and restoration processes (e.g., the quantization stage 214 in FIGS. 2A-2B), generally, the video stream 304 is not identical to the video sequence 202. Similar to processes 200A and 200B in FIGS. 2A-2B, the decoder can perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in the video bitstream 228. For example, the decoder can perform process 300A in an iterative manner, in which case the decoder can decode the basic processing unit in one iteration of process 300A. In some embodiments, the decoder can perform process 300A in parallel for each region (e.g., regions 114-118) of each picture encoded in the video bitstream 228.

[0052]

[0079] In FIG. 3A, the decoder can supply a portion of the video bitstream 228 associated with the basic processing unit of the encoded picture (referred to as the "encoded BPU") to the binary decoding stage 302. In the binary decoding stage 302, the decoder can decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder supplies the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220, and can generate a reconstructed residual BPU 222. The decoder supplies the prediction data 206 to the prediction stage 204 and can generate a prediction BPU 208. The decoder adds the reconstructed residual BPU 222 to the prediction BPU 208 and can generate a prediction reference 224. In some embodiments, the prediction reference 224 can be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder can supply the prediction reference 224 to the prediction stage 204 to perform prediction operations in the next iteration of process 300A.

[0053]

[0080] The decoder can repeatedly perform process 300A to decode each encoded BPU of the encoded picture and generate a prediction reference 224 for encoding the next encoded BPU of the encoded picture. After decoding all the encoded BPUs of the encoded picture, the decoder can output the picture to the video stream 304 for display and proceed to decode the next encoded picture in the video bitstream 228.

[0054]

[0081] In the binary decoding stage 302, the decoder can perform the inverse operation of the binary encoding technique (e.g., entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context adaptive binary arithmetic encoding, or any other reversible compression algorithm) used by the encoder. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder can also decode other information in the binary decoding stage 302, such as, for example, the prediction mode, the parameters of the prediction operation, the type of transformation, the parameters of the quantization process (e.g., quantization parameters), the encoder control parameters (e.g., bitrate control parameters), or the like. In some embodiments, when the video bitstream 228 is transmitted in the form of packets through a network, the decoder can depacketize the video bitstream 228 before supplying it to the binary decoding stage 302.

[0055]

[0082] FIG. 3B shows a schematic diagram of another exemplary decoding process 300B according to an embodiment of the present disclosure. The process 300B can be different from the process 300A. For example, the process 300B can be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with the process 300A, the process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.

[0056]

[0083] In process 300B, for an encoding basic processing unit (referred to as the "current BPU") of an encoded picture being decoded (referred to as the "current picture"), prediction data 206 decoded by a decoder from a binary decoding stage 302 can include various types of data depending on which prediction mode was used by an encoder to encode the current BPU. For example, if intra prediction was used by the encoder to encode the current BPU, the prediction data 206 can include an intra prediction, parameters of an intra prediction operation, or a prediction mode indicator (e.g., a flag value) indicating the like. The parameters of the intra prediction operation can include, for example, the location (e.g., coordinates) of one or more adjacent BPUs used as references, the size of the adjacent BPUs, extrapolation parameters, the direction of the adjacent BPUs with respect to the original BPU, or the like. As another example, if inter prediction was used by the encoder to encode the current BPU, the prediction data 206 can include an inter prediction, parameters of an inter prediction operation, or a prediction mode indicator (e.g., a flag value) indicating the like. The parameters of the inter prediction operation can include, for example, the number of reference pictures associated with the current BPU, the weights respectively associated with the reference pictures, the location (e.g., coordinates) of one or more matching regions in each reference picture, one or more motion vectors respectively associated with the matching regions, or the like.

[0057]

[0084] Based on the prediction mode indicator, the decoder can determine whether to perform spatial prediction (e.g., intra prediction) in a spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in a temporal prediction stage 2044. Details of performing such spatial or temporal prediction are described in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder can generate a predicted BPU 208. The decoder can add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as described in FIG. 3A.

[0058]

[0085] In process 300B, the decoder can supply the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing the prediction operation in the next iteration of process 300B. For example, when the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can directly supply the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). When the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture with all BPUs decoded), the coder can supply the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can apply the loop filter to the prediction reference 224 in the manner described in FIG. 2B. The loop-filtered reference picture can be stored in buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., for use as an inter prediction reference picture for future encoded pictures of video bitstream 228). The decoder can store one or more reference pictures in buffer 234 for use in the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter prediction was used to encode the current BPU, the prediction data can further include loop filter parameters (e.g., loop filter strength).

[0059]

[0086] FIG. 4 is a block diagram of an exemplary device 400 for encoding or decoding video according to an embodiment of the present disclosure. As shown in FIG. 4, device 400 can include a processor 402. When processor 402 executes the instructions described herein, device 400 can become a special machine for video encoding or decoding. Processor 402 can be any kind of circuitry having the ability to manipulate or process information. For example, processor 402 can include any number of any combination of a central processing unit (or “CPU”), a graphics processing unit (or “GPU”), a neural processing unit (“NPU”), a microcontroller unit (“MCU”), an optical processor, a programmable logic controller, a microcontroller, a microprocessor, a digital signal processor, an intellectual property (IP) core, a programmable logic array (PLA), a programmable array logic (PAL), a generic array logic (GAL), a complex programmable logic device (CPLD), a field programmable gate array (FPGA), a system on chip (SoC), an application specific integrated circuit (ASIC), or the like. In some embodiments, processor 402 can also be a set of processors grouped as a single logical component. For example, as shown in FIG. 4, processor 402 can include a plurality of processors, including processor 402a, processor 402b, and processor 402n.

[0060]

[0087] The machine 400 can also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, or the like). For example, as shown in FIG. 4, the stored data can include program instructions (e.g., program instructions for implementing the steps in processes 200A, 200B, 300A, or 300B), as well as data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 can access the program instructions and data for processing (e.g., via bus 410), execute the program instructions, and perform operations or manipulations on the data for processing. The memory 404 can include a high-speed random access memory device or a non-volatile memory device. In some embodiments, the memory 404 can include any number of any combination of random access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (registered trademark) (CF) cards, or the like. The memory 404 can also be a group of memories grouped as a single logical component (not shown in FIG. 4).

[0061]

[0088] The bus 410 can be a communication device that transfers data between components inside the machine 400, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a universal serial bus port, a peripheral component interconnect express port), or the like.

[0062]

[0089] To facilitate explanation without creating ambiguity, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits" in this disclosure. The data processing circuits can be implemented entirely as hardware, or as a combination of software, hardware, or firmware. Additionally, the data processing circuits can be a single stand-alone module, or can be fully or partially integrated with any other component of the device 400.

[0063]

[0090] The device 400 can further include a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, or the like). In some embodiments, the network interface 406 can include any number of any combination of a network interface controller (NIC), a radio frequency (RF) module, a transponder, a transceiver, a modem, a router, a gateway, a wired network adapter, a wireless network adapter, a Bluetooth® adapter, an infrared adapter, a near field communication ("NFC") adapter, a cellular network chip, or the like.

[0064]

[0091] In some embodiments, optionally, the device 400 can further include a peripheral interface 408 for providing connection to one or more peripheral devices. As shown in FIG. 4, the peripheral devices can include, but are not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or an input interface coupled to a video archive), or the like.

[0065]

[0092] Note that the video codec (e.g., the codec that performs processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules within device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instructions that can be loaded into memory 404. As another example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as special data processing circuits (e.g., FPGA, ASIC, NPU, or the like).

[0066]

[0093] In the quantization and inverse quantization function blocks (e.g., quantization 214 and inverse quantization 218 in FIGS. 2A or 2B, inverse quantization 218 in FIGS. 3A or 3B), quantization parameters (QPs) are used to determine the amount of quantization (and inverse quantization) applied to the prediction residuals. The initial QP value used for encoding a picture or slice can be signaled at a high level, for example, using the init_qp_minus26 syntax element within the picture parameter set (PPS) and the slice_qp_delta syntax element within the slice header. Further, the QP value can be adapted at a local level for each CU using the delta QP values sent at the granularity of the quantization group.

[0067]

[0094] In the disclosed embodiments, to encode a frame, a picture is divided into a sequence of coding tree units (CTUs). A plurality of CTUs can form a tile, slice, or subpicture. A picture is divided into a sequence of CTUs. In a picture having three sample arrays, a CTU is composed of an N×N block of luma samples along with corresponding two blocks of chroma samples. FIG. 5 shows an example of a picture divided into a plurality of CTUs according to some embodiments of the present disclosure.

[0068]

[0095] According to some embodiments, the maximum allowable size of a luma block within a CTU is specified to be 128×128 (where the maximum size of a luma transform block can be 64×64), and the minimum allowable size of a luma block within a CTU is specified to be 32×32.

[0069]

[0096] A picture is divided into one or more tile rows and one or more tile columns. A tile is a sequence of CTUs that covers a rectangular region of the picture. A slice contains an integer number of complete tiles, or an integer number of consecutive complete CTU rows within a tile of the picture. Two modes of slices, namely the raster scan slice mode and the rectangular slice mode, can be supported. In the raster scan slice mode, a slice contains a sequence of complete tiles within the tile raster scan of the picture. In the rectangular slice mode, a slice contains some complete tiles that collectively form a rectangular region of the picture, or some consecutive complete CTU rows of one tile that collectively form a rectangular region of the picture. The tiles within a rectangular slice are scanned in the tile raster scan order within the rectangular region corresponding to that slice.

[0070]

[0097] A subpicture contains one or more slices that collectively cover a rectangular region of the picture.

[0071]

[0098] FIG. 6 shows an example of a picture divided into tiles and raster scan slices according to some embodiments of the present disclosure. As shown in FIG. 6, the picture is divided into 12 tiles (4 tile rows and 3 tile columns) and 3 raster scan slices.

[0072]

[0099] FIG. 7 shows an example of a picture divided into tiles and rectangular slices according to some embodiments of the present disclosure. As shown in FIG. 7, the picture is divided into 20 tiles (5 tile rows and 4 tile columns) and 9 rectangular slices.

[0073]

[0100] Figure 8 shows another example of a picture that is divided into tiles and rectangular slices according to some embodiments of the present disclosure. As shown in Figure 8, the picture is divided into 4 tiles (2 tile rows and 2 tile columns) and 4 rectangular slices.

[0074]

[0101] Figure 9 shows an example of a picture that is divided into sub-pictures according to some embodiments of the present disclosure. As shown in Figure 9, the picture is divided into 20 tiles (5 tile columns and 4 tile rows), with 12 covering one slice of a 4×4 CTU on the left, and 8 tiles covering two vertically stacked slices of a 2×2 CTU on the right, resulting in a total of 28 slices and 28 sub-pictures of various sizes (each slice is a sub-picture).

[0075]

[0102] According to some embodiments of the disclosure, sub-picture segmentation information is signaled within a sequence parameter set (SPS). Figure 10 shows an exemplary Table 1 showing an exemplary SPS syntax for sub-picture segmentation according to some embodiments of the present disclosure.

[0076]

[0103] In Table 1, the syntax element sps_num_subpics_minus1 plus 1 specifies the number of sub-pictures within one picture, the syntax elements subpic_ctu_top_left_x[i] and subpic_ctu_top_left_y[i] specify the position of the top-left CTU of the i-th sub-picture in units of CtbSizeY, and the syntax elements subpic_width_minus1[i] plus 1 and subpic_height_minus1[i] plus 1 specify the width and height of the i-th sub-picture in units of CtbSizeY, respectively. The semantics of these syntax elements are as follows.

[0077]

[0104] A subpics_present_flag equal to 1 indicates that the subpicture parameters are within the SPS RBSP syntax, and a subpics_present_flag equal to 0 indicates that the subpicture parameters are not within the SPS RBSP syntax.

[0078]

[0105] sps_num_subpics_minus1 plus 1 specifies the number of subpictures. The syntax element sps_num_subpics_minus1 is in the range of 0 to 254. If not present, the value of the syntax element sps_num_subpics_minus1 is inferred to be equal to 0.

[0079]

[0106] subpic_ctu_top_left_x[i] specifies the horizontal position of the top-left CTU of the i-th subpicture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_ctu_top_left_x[i] is inferred to be equal to 0.

[0080]

[0107] subpic_ctu_top_left_y[i] specifies the vertical position of the top-left CTU of the i-th subpicture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_ctu_top_left_y[i] is inferred to be equal to 0.

[0081]

[0108] subpic_width_minus1[i] plus 1 specifies the width of the i-th subpicture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_width_minus1[i] is inferred to be equal to Ceil(pic_width_max_in_luma_samples÷CtbSizeY)-1.

[0082]

[0109] subpic_height_minus1[i] plus 1 specifies the height of the i-th subpicture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_height_minus1[i] is inferred to be equal to Ceil(pic_height_max_in_luma_samples÷CtbSizeY)-1.

[0083]

[0110] subpic_treated_as_pic_flag[i] equal to 1 specifies that the i-th subpicture of each coded picture within the coded layer video sequence (CLVS) is to be treated as a picture within the decoding process that excludes in-loop filtering operations. The syntax element subpic_treated_as_pic_flag[i] equal to 0 specifies that the i-th subpicture of each coded picture within the CLVS is not to be treated as a picture within the decoding process that excludes in-loop filtering operations. If not present, the value of the syntax element subpic_treated_as_pic_flag[i] is inferred to be equal to 0.

[0084]

[0111] A loop_filter_across_subpic_enabled_flag[i] equal to 1 specifies that in-loop filtering operations can be performed across the boundaries of the i-th sub-picture within each coded picture in the CLVS. A syntax element loop_filter_across_subpic_enabled_flag[i] equal to 0 specifies that in-loop filtering operations are not performed across the boundaries of the i-th sub-picture within each coded picture in the CLVS. If not present, the value of the syntax element loop_filter_across_subpic_enabled_pic_flag[i] is inferred to be equal to 1.

[0085]

[0112] According to the disclosed embodiments of the present disclosure, an identifier can be assigned to each sub-picture. The sub-picture identifier information can be signaled within a sequence parameter set (SPS), a picture parameter set (PPS), or a picture header (PH). FIG. 11 shows an exemplary Table 2 showing an exemplary SPS syntax of sub-picture identifiers according to some embodiments of the present disclosure. FIG. 12 shows an exemplary Table 3 showing an exemplary PPS syntax of sub-picture identifiers according to some embodiments of the present disclosure. FIG. 13 shows an exemplary Table 4 showing an exemplary PH syntax of sub-picture identifiers according to some embodiments of the present disclosure.

[0086]

[0113] As shown in Tables 2 to 4, the syntax element sps_subpic_id_present_flag indicates whether the sub-picture ID mapping is within the SPS. The syntax elements sps_subpic_id_signaling_present_flag, pps_subpic_id_signaling_present_flag, and ph_subpic_id_signaling_present_flag indicate whether the sub-picture ID mapping is signaled within the SPS, PPS, or PH respectively. The syntax element sps_subpic_id_len_minus1 plus 1, the syntax element pps_subpic_id_len_minus1 plus 1, and the syntax element ph_subpic_id_len_minus1 plus 1 specify the number of bits used to present the syntax elements sps_subpic_id[i], pps_subpic_id[i], and ph_subpic_id[i] respectively, and they are the sub-picture IDs signaled within the SPS, PPS, and PH respectively.

[0087]

[0114] The semantics of the above syntax elements and the related bitstream compliance requirements are described as follows.

[0088]

[0115] A sps_subpic_id_present_flag equal to 1 specifies that the sub-picture ID mapping is within the SPS, and a syntax element sps_subpic_id_present_flag equal to 0 specifies that the sub-picture ID mapping is not within the SPS.

[0089]

[0116] A sps_subpic_id_signaling_present_flag equal to 1 specifies that the sub-picture ID mapping is signaled within the SPS, and a syntax element sps_subpic_id_signaling_present_flag equal to 0 specifies that the sub-picture ID mapping is not signaled within the SPS. If not, the value of the syntax element sps_subpic_id_signaling_present_flag is inferred to be equal to 0.

[0090]

[0117] sps_subpic_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax element sps_subpic_id[i]. The value of the syntax element sps_subpic_id_len_minus1 can be in the range of 0 to 15.

[0091]

[0118] sps_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the syntax element sps_subpic_id[i] is sps_subpic_id_len_minus1 + 1 bits. If not present and if the syntax element sps_subpic_id_present_flag is equal to 0, the value of the syntax element sps_subpic_id[i] is inferred to be equal to i for each i in the range of 0 to sps_num_subpics_minus1.

[0092]

[0119] A pps_subpic_id_signaling_present_flag equal to 1 specifies that the sub-picture ID mapping is signaled within the PPS. A syntax element pps_subpic_id_signaling_present_flag equal to 0 specifies that the sub-picture ID mapping is not signaled within the PPS. The syntax element pps_subpic_id_signaling_present_flag can be equal to 0 if the syntax element sps_subpic_id_present_flag is 0 or if the syntax element sps_subpic_id_signaling_present_flag is equal to 1.

[0093]

[0120] pps_num_subpics_minus1 plus 1 specifies the number of sub-pictures in the coded picture that refers to the PPS. It can be a bitstream conformity requirement that the value of the syntax element pps_num_subpic_minus1 is equal to the syntax element sps_num_subpics_minus1.

[0094]

[0121] pps_subpic_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax element pps_subpic_id[i]. The value of the syntax element pps_subpic_id_len_minus1 is in the range of 0 to 15. It may be a requirement for bitstream compliance that the value of the syntax element pps_subpic_id_len_minus1 be the same for all PPSs referenced by the coded pictures within the CLVS.

[0095]

[0122] pps_subpic_id[i] specifies the subpicture ID of the i-th subpicture. The length of the syntax element pps_subpic_id[i] is pps_subpic_id_len_minus1 + 1 bits.

[0096]

[0123] ph_subpic_id_signaling_present_flag equal to 1 specifies that the subpicture ID mapping is signaled within the PH, and the syntax element ph_subpic_id_signaling_present_flag equal to 0 specifies that the subpicture ID mapping is not signaled within the PH.

[0097]

[0124] ph_subpic_id_len_minus1 plus 1 specifies the number of bits used to represent the syntax element ph_subpic_id[i]. The value of the syntax element pic_subpic_id_len_minus1 can be in the range of 0 to 15. It may be a requirement for bitstream compliance that the value of the syntax element ph_subpic_id_len_minus1 be the same for all PHs referenced by the coded pictures within the CLVS.

[0098]

[0125] ph_subpic_id[i] specifies the subpicture ID of the i-th subpicture. The length of the syntax element ph_subpic_id[i] is ph_subpic_id_len_minus1 + 1 bits.

[0099]

[0126] After parsing these syntax elements related to the sub-picture ID, the sub-picture ID list SubpicIdList is derived using the following syntax (1): for(i = 0; i <= sps_num_subpics_minus1; i++) SubpicIdList[i]=sps_subpic_id_present_flag? syntax (1) (sps_subpic_id_signaling_present_flag?sps_subpic_id[i]: (ph_subpic_id_signaling_present_flag?ph_subpic_id[i]:pps_subpic_id[i])):i

[0100]

[0127] However, there are some problems with the above signaling of sub-picture partitioning. First, semantically, both syntax elements sps_subpic_id_present_flag and sps_subpic_id_signaling_present_flag specify whether the sub-picture ID is within the SPS. According to Table 2, the sub-picture ID information is signaled within the SPS only when both of these two syntax elements are true. Therefore, there is redundancy in the signaling. Second, even if sps_subpics_id_present_flag is true, the sub-picture ID can still be signaled within the PPS or PH if pps_subpic_id_signalling_present_flag or ph_subpic_id_signalling_present_flag is true. Third, according to syntax (1), if sps_subpic_id_present_flag is false, a default ID equal to the sub-picture index is specified for each sub-picture. If sps_subpic_id_present_flag is true, SubpicIdList[i] is derived as sps_subpic_id[i], ph_subpic_id[i], or pps_subpic_id[i]. However, if sps_subpic_id_present_flag is true, the sub-picture ID may not be within the SPS, PPS, or PH, and in that case, the undefined value of the syntax element pps_subpic_id is specified in SubpicIdList. In syntax (1), when sps_subpic_id_present_flag is true, both syntax elements sps_subpic_id_signaling_present_flag and ph_subpic_id_signaling_present_flag are false, and syntax element pps_subpic_id[i] is specified in SubpicIdList[i] regardless of the value of syntax element pps_subpic_id_signaling_present_flag. When syntax element ps_subpic_id_signaling_present_flag is false, syntax element pps_subpic_id[i] is undefined.

[0101]

[0128] Further, as shown in Table 1 of FIG. 10, when the syntax element subpics_present_flag is true, the number of sub-pictures is signaled first, and then the upper-left position, width, and height of each sub-picture, as well as the two control flags subpic_treated_as_pic_flag and loop_filter_across_subpic_enabled_flag, are signaled. Even when there is only one sub-picture (when the syntax element sps_num_subpics_minus1 is equal to 0), the upper-left position, width, and height, as well as these two control flags, are signaled. However, when there is only one sub-picture within a picture, the sub-picture is equal to the picture, so the information to be signaled can be derived from the picture itself, and it is not necessary to indicate these matters.

[0102]

[0129] Further, a sub-picture is obtained by dividing a picture, and a picture is formed by merging all sub-pictures. The position and size of the last sub-picture can be derived from the size of the entire picture and the positions and sizes of all previous sub-pictures. Therefore, it is not necessary to signal the position, width, and height information of the last sub-picture.

[0103]

[0130] Further, as shown in Table 2 of FIG. 11, the syntax element sps_subpic_id_present_flag is always signaled regardless of the value of the syntax element subpics_present_flag. Therefore, in the above signaling method, even if there is no sub-picture, the sub-picture identifier may still be signaled, which is meaningless.

[0104]

[0131] The present disclosure provides a signaling method for solving the above problems. Some exemplary embodiments will be described in detail below.

[0105]

[0132] In some embodiments of the present disclosure, when the syntax element sps_subpic_id_present_flag is true but both syntax elements sps_subpic_id_signaling_present_flag and pps_subpic_id_signaling_present_flag are false, signaling of the sub-picture ID in the picture header can be enforced. This can avoid cases where the sub-picture ID is not defined.

[0106]

[0133] For example, bitstream compliance constraints can be imposed in the following two ways. In the first way, the semantics for bitstream compliance constraints are as follows (emphasized in italics): ph_subpic_id_signaling_present_flag equal to 1 specifies that the sub-picture ID mapping is signaled within the PH. The syntax element ph_subpic_id_signaling_present_flag equal to 0 specifies that the sub-picture ID mapping is not signaled within the PH. When the syntax element sps_subpic_id_present_flag is equal to 1, the syntax element sps_subpic_id_signaling_present_flag is equal to 0, and the syntax element pps_subpic_id_signaling_present_flag is equal to 0, the value of the syntax element ph_subpic_id_signaling_present_flag being 1 can be a bitstream compliance requirement.

[0107]

[0134] In the second method, the semantics for bitstream compliance constraints are as follows (emphasized in italics): A ph_subpic_id_signaling_present_flag equal to 1 specifies that the subpicture ID mapping is signaled within the PH. A syntax element ph_subpic_id_signaling_present_flag equal to 0 specifies that the subpicture ID mapping is not signaled within the PH. When the syntax element sps_subpic_id_present_flag is equal to 1, the syntax element sps_subpic_id_signaling_present_flag is equal to 0, and the syntax element pps_subpic_id_signaling_present_flag in all of the PPSs referenced by the coded pictures within the CLVS is equal to 0, it may be a bitstream compliance requirement that there is at least one PH among all of the PHs referenced by the coded pictures within the CLVS reference where the value of the syntax element ph_subpic_id_signaling_present_flag is equal to 1.

[0108]

[0135] FIG. 14 is a schematic diagram showing exemplary bitstream compliance constraints of this second method according to some embodiments of the present disclosure.

[0109]

[0136] The semantics of the syntax element sps_subpic_id_present_flag are not clearly defined in the current VVC draft and can be changed as follows (emphasized in italics).

[0110]

[0137] A sps_subpic_id_present_flag equal to 1 specifies that the subpicture ID mapping is within the SPS, PPS, or PH. A syntax element sps_subpic_id_present_flag equal to 0 specifies that the subpicture ID mapping is not within the SPS, PPS, and PH.

[0111]

[0138] This can ensure that the syntax element ph_subpic_id is signaled when the syntax element sps_subpic_id_present_flag is true but both syntax elements sps_subpic_id_signaling_present_flag and pps_subpic_id_signaling_present_flag are false. This syntax is shown in Tables 2 to 4 of FIGS. 11 to 13.

[0112]

[0139] In the above embodiment, when the sub-picture ID presence flag is true (syntax element sps_subpic_id_present_flag = 1), the sub-picture ID used is signaled within the bitstream (within SPS, PPS, or PH), and no inference rules are required. When the sub-picture ID presence flag is true (syntax element sps_subpic_id_present_flag = 1), since the syntax element sps_subpic_id_present_flag indicates the presence of the sub-picture ID, forcing the signaling of the sub-picture ID in one of SPS, PPS, or PH may be better than deriving the sub-picture ID using inference rules without signaling the sub-picture ID within the bitstream.

[0113]

[0140] As another example, FIG. 15 shows an exemplary Table 5 showing another exemplary PH syntax of a sub-picture identifier according to some embodiments of the present disclosure. Table 5 shows the modification of the PH syntax shown in Table 4 (shown in box 1501 and emphasized in italics). Referring to Table 5, when the syntax element sps_subpic_id_present_flag is true but both syntax elements sps_subpic_id_signaling_present_flag and pps_subpic_id_signaling_present_flag are false, the signaling of the syntax element ph_subpic_id is forced by inferring that the syntax element ph_sub_pic_id_signaling_present_flag is true.

[0114]

[0141] The syntax element ph_subpic_id_signaling_present_flag can have the following two alternative semantics (emphasized in italics).

[0115]

[0142] The first semantics includes the following (emphasized in italics): A ph_subpic_id_signaling_present_flag equal to 1 specifies that the subpicture ID mapping is signaled within the PH. A syntax element ph_subpic_id_signaling_present_flag equal to 0 specifies that the subpicture ID mapping is not signaled within the PH. If not present, the value of the syntax element ph_subpic_id_signaling_present_flag is inferred to be 1.

[0116]

[0143] The second semantics includes the following (emphasized in italics): A ph_subpic_id_signaling_present_flag equal to 1 specifies that the subpicture ID mapping is signaled within the PH, and a syntax element ph_subpic_id_signaling_present_flag equal to 0 specifies that the subpicture ID mapping is not signaled within the PH. When not present, if the syntax element sps_subpic_id_present_flag is equal to 1 and the syntax element sps_subpic_id_signaling_present_flag is equal to 0, the value of the syntax element ph_subpic_id_signaling_present_flag is inferred to be 1.

[0117]

[0144] The semantics of the syntax element sps_subpic_id_present_flag can be changed as follows (emphasized in italics).

[0118]

[0145] The sps_subpic_id_present_flag equal to 1 specifies that the subpicture ID mapping is within the SPS, PPS, or PH. The syntax element sps_subpic_id_present_flag equal to 0 specifies that the subpicture ID mapping is not within the SPS, PPS, and PH.

[0119]

[0146] When the syntax element sps_subpic_id_present_flag is true, the inference rule for the subpicture ID is not necessary in some embodiments. Therefore, when the subpicture ID presence flag is true (syntax element sps_subpic_id_present_flag = 1) but the subpicture ID is not signaled within the SPS or PPS (syntax elements sps_subpic_id_signaling_present_flag = 0 and pps_subpic_id_signaling_present_flag = 0), the signaling of the syntax element ph_subpic_id_signaling_present_flag is skipped. Then 1 bit can be saved.

[0120]

[0147] As another example, when the syntax element sps_subpic_id_present_flag is true but both of the syntax elements sps_subpic_id_signaling_present_flag and pps_subpic_id_signaling_present_flag are false (i.e., the subpicture ID is not signaled within the SPS or PPS), the signaling of the syntax element ph_subpic_id is forced, which means that in this case the subpicture ID is signaled within the PH. When the syntax element pps_subpic_id is signaled (the syntax element sps_subpic_id_present_flag is true, the syntax element sps_subpic_id_signaling_present_flag is false, and the syntax element pps_subpic_id_signaling_present_flag is true), the syntax element ph_subpic_id cannot be signaled. FIG. 16 shows an exemplary Table 6 according to some embodiments of the present disclosure, which shows another exemplary PH syntax of the subpicture identifier (emphasis is shown within box 1601 and highlighted in italics).

[0121]

[0148] The list SubpicIdList[i] is derived according to syntax (2) as follows. for(i=0;i<=sps_num_subpics_minus1;i++) SubpicIdList[i]=sps_subpic_id_present_flag? syntax (2) (sps_subpic_id_signaling_present_flag?sps_subpic_id[i]: (pps_subpic_id_signaling_present_flag?pps_subpic_id[i]:ph_subpic_id[i])):i

[0122]

[0149] When the syntax element sps_subpic_id_present_flag is true, the inference rule for the subpicture ID is not necessary in some embodiments. The syntax element ph_subpic_id_signaling_present_flag can be removed. Therefore, when the syntax element sps_subpic_id_present_flag is equal to 1, the syntax element sps_subpic_id_signaling_present_flag is equal to 0, and the syntax element pps_subpic_id_signaling_present_flag is equal to 1, 1 bit is saved.

[0123]

[0150] When the subpicture ID has already been signaled within the PPS, some embodiments may give the coder the option to override the subpicture ID within the PPS by re-signaling the subpicture ID within the PH. This is a more flexible way for the coder.

[0124]

[0151] In some embodiments of the present disclosure, inference rules are provided to ensure that a subpicture ID list SubpicIdList can be derived.

[0125]

[0152] As an example, when the subpicture ID is not signaled within the SPS, PPS, or PH, the inference rule is provided to the syntax element pps_subpic_id to derive the subpicture ID list SubpicIdList using the default value of the syntax element pps_subpic_id inferred by the inference rule.

[0126]

[0153] The semantics of the syntax element pps_subpic_id are as follows (emphasized in italics): pps_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the syntax element pps_subpic_id[i] is pps_subpic_id_len_minus1 + 1 bits. If not, the value of the syntax element pps_subpic_id[i] is inferred to be i for each i in the range of 0 to pps_num_subpics_minus1.

[0127]

[0154] As another example, the inference rule is given within the derivation process of SubpicIdList. When the syntax element sps_subpic_id_present_flag is true and the syntax elements sps_subpic_id_signaling_present_flag, pps_subpic_id_signaling_present_flag, and ph_subpic_id_signaling_present_flag are all false, the default value is specified for SubpicIdList[i].

[0128]

[0155] The derivation of SubpicIdList follows the syntax (3) as follows (emphasized in italics): for(i = 0; i <= sps_num_subpics_minus1; i++) SubpicIdList[i] = sps_subpic_id_present_flag? syntax (3) (sps_subpic_id_signaling_present_flag? sps_subpic_id[i] : (ph_subpic_id_signaling_present_flag? ph_subpic_id[i] : (pps_subpic_id_signaling_present_flag? pps_subpic_id[i] : i))): i

[0129]

[0156] In some embodiments, by imposing inference rules on any pps_subpic_id, it can be guaranteed that the SubpicIdList can be derived even if the subpicture ID is not signaled at all in the bitstream. Therefore, the bits allocated for signaling the subpicture ID can be saved.

[0130]

[0157] In some embodiments of the present disclosure, in order to give a higher priority to the subpicture ID signaled within PPS than PH, the derivation rule of the subpicture ID list SubpicIdList can be changed. Therefore, the syntax element pps_subpic_id_signaling_present_flag is checked before the syntax element ph_subpic_id_signaling_present_flag.

[0131]

[0158] As an example, the derivation rule of the SubpicIdList follows the syntax (4) shown below (emphasized in italics), and an inference rule is given for the syntax element ph_subpic_id. for(i=0;i<=sps_num_subpics_minus1;i++) SubpicIdList[i]=sps_subpic_id_present_flag? Syntax (4) (sps_subpic_id_signaling_present_flag?sps_subpic_id[i]: (pps_subpic_id_signaling_present_flag?pps_subpic_id[i]:ph_subpic_id[i])):i

[0132]

[0159] The semantics following the syntax (4) (emphasized in italics) are as follows: ph_subpic_id[i] specifies the sub-picture ID of the i-th sub-picture. The length of the syntax element ph_subpic_id[i] is ph_subpic_id_len_minus1 + 1 bits. If not, the value of the syntax element ph_subpic_id[i] is inferred to be i for each i in the range of 0 to ph_num_subpics_minus1.

[0133]

[0160] As another example, the derivation rule SubpicIDList follows the syntax (5) shown below (emphasized in italics). In this example, there is no additional inference rule for the syntax element ph_subpic_id. for(i = 0; i <= sps_num_subpics_minus1; i++) SubpicIdList[i] = sps_subpic_id_present_flag? Syntax (5) (sps_subpic_id_signaling_present_flag? sps_subpic_id[i] : (pps_subpic_id_signaling_present_flag? pps_subpic_id[i] : (ph_subpic_id_signaling_present_flag? ph_subpic_id[i] : i))): i

[0134]

[0161] To ensure that SubpicIdList can be correctly derived even if no sub-picture ID is signaled in the bitstream, inference rules are given for the derivation process of the syntax element ph_subpic_id or SubpicIdList. Therefore, when the default sub-picture ID inferred by the inference rule is used, the bits allocated for sub-picture signaling can be saved.

[0135]

[0162] In some embodiments of the present disclosure, redundant information signaled for the sub-picture can be removed when the number of sub-pictures is equal to 1.

[0136]

[0163] As an example, the SPS syntax is shown in Table 7A of FIG. 17A (emphasis is shown within boxes 1701 to 1702 and is emphasized in italics) or Table 7B of FIG. 17B (emphasis is shown within boxes 1711 to 1712 and is emphasized in italics). It will be understood that Table 7A and Table 7B are equivalent. The semantics (emphasized in italics) according to the syntax in Table 7A and Table 7B are shown below.

[0137]

[0164] subpic_ctu_top_left_x[i] specifies the horizontal position of the top-left CTU of the i-th subpicture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_ctu_top_left_x[i] is inferred to be equal to 0.

[0138]

[0165] subpic_ctu_top_left_y[i] specifies the vertical position of the top-left CTU of the i-th subpicture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_ctu_top_left_y[i] is inferred to be equal to 0.

[0139]

[0166] subpic_width_minus1[i] plus 1 specifies the width of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_width_minus1[i] is inferred to be equal to Ceil(pic_width_max_in_luma_samples÷CtbSizeY)-1. Here, "Ceil()" is a function for rounding up to the nearest integer. Therefore, Ceil(pic_width_max_in_luma_samples÷CtbSizeY)-1 is equal to (pic_width_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1, where " / " is integer division.

[0140]

[0167] subpic_height_minus1[i] plus 1 specifies the height of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_height_minus1[i] is inferred to be equal to Ceil(pic_height_max_in_luma_samples÷CtbSizeY)-1. Here, "Ceil()" is a function for rounding up to the nearest integer. Therefore, Ceil(pic_height_max_in_luma_samples÷CtbSizeY)-1 is equal to (pic_height_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1, where " / " is integer division.

[0141]

[0168] A subpic_treated_as_pic_flag[i] equal to 1 specifies that the i-th subpicture of each coded picture in the CLVS is to be treated as a picture in the decoding process that excludes in-loop filtering operations. A syntax element subpic_treated_as_pic_flag[i] equal to 0 specifies that the i-th subpicture of each coded picture in the CLVS is not to be treated as a picture in the decoding process that excludes in-loop filtering operations. When absent, if the syntax element subpics_present_flag is equal to 1 and the syntax element sps_num_subpics_minus1 is equal to 0, the value of the syntax element subpic_treated_as_pic_flag[i] is inferred to be equal to 1; otherwise, the value of the syntax element subpic_treated_as_pic_flag[i] is inferred to be equal to 0.

[0142]

[0169] A loop_filter_across_subpic_enabled_flag[i] equal to 1 specifies that in-loop filtering operations can be performed across the boundaries of the i-th subpicture within each coded picture in the CLVS. A syntax element loop_filter_across_subpic_enabled_flag[i] equal to 0 specifies that in-loop filtering operations are not to be performed across the boundaries of the i-th subpicture within each coded picture in the CLVS. When absent, if the syntax element subpics_present_flag is equal to 1 and the syntax element sps_num_subpics_minus1 is equal to 0, the value of the syntax element loop_filter_across_subpic_enabled_flag[i] is inferred to be equal to 0; otherwise, the value of the syntax element loop_filter_across_subpic_enabled_pic_flag[i] is inferred to be equal to 1.

[0143]

[0170] As another example, Table 8 in FIG. 18 (with emphasis shown within boxes 1801 to 1802 and emphasized in italics) shows the SPS syntax. The semantics (emphasized in italics) according to the syntax in Table 8 are shown below.

[0144]

[0171] subpic_ctu_top_left_x[i] specifies the horizontal position of the top - left CTU of the i - th sub - picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_ctu_top_left_x[i] is inferred to be equal to 0.

[0145]

[0172] subpic_ctu_top_left_y[i] specifies the vertical position of the top - left CTU of the i - th sub - picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_ctu_top_left_y[i] is inferred to be equal to 0.

[0146]

[0173] subpic_width_minus1[i] plus 1 specifies the width of the i - th sub - picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_width_minus1[i] is inferred to be equal to Ceil((pic_width_max_in_luma_samples÷CtbSizeY)-1). Here, "Ceil()" is a function for rounding up to the nearest integer. Thus, Ceil(pic_width_max_in_luma_samples÷CtbSizeY)-1 is equal to (pic_width_max_in_luma_samples + CtbSizeY - 1) / CtbSizeY - 1, where " / " is integer division.

[0147]

[0174] subpic_height_minus1[i] plus 1 specifies the height of the i-th subpicture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_height_minus1[i] is inferred to be equal to Ceil(pic_height_max_in_luma_samples÷CtbSizeY)-1. Here, "Ceil()" is a function for rounding up to the nearest integer. Therefore, Ceil(pic_height_max_in_luma_samples÷CtbSizeY)-1 is equal to (pic_height_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1, where " / " is integer division.

[0148]

[0175] In some embodiments of the present disclosure, the position and / or size information of the last subpicture can be skipped and derived from the size of the entire picture and the sizes and positions of all previous subpictures. FIG. 19 shows an exemplary Table 9 showing another exemplary SPS syntax according to some embodiments of the present disclosure. In Table 9 (highlighted in boxes 1901 to 1902 and emphasized in italics), the width and height of the last subpicture, which is the subpicture having an index equal to the syntax element sps_num_subpics_minus1, are skipped.

[0149]

[0176] The width and height of the last subpicture are derived from the width and height of the entire picture and the upper left position of the last subpicture.

[0150]

[0177] The following are the semantics (emphasized in italics) according to the semantics of Table 9.

[0151]

[0178] subpic_ctu_top_left_x[i] specifies the horizontal position of the top-left CTU of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_ctu_top_left_x[i] is inferred to be equal to 0.

[0152]

[0179] subpic_ctu_top_left_y[i] specifies the vertical position of the top-left CTU of the i-th sub-picture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_ctu_top_left_y[i] is inferred to be equal to 0.

[0153]

[0180] subpic_width_minus1[i] plus 1 specifies the width of the i-th subpicture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_width_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_width_minus1[i] is inferred to be equal to Ceil((pic_width_max_in_luma_samples)÷CtbSizeY)-1-(i==sps_num_subpics_minus1?subpic_ctu_top_left_x[sps_num_subpics_minus1]:0). Here, "Ceil()" is a function for rounding up to the nearest integer. That is, Ceil(pic_width_max_in_luma_samples÷CtbSizeY) is equal to (pic_width_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1, where " / " is integer division. "sps_num_subpics_minus1" is the number of subpictures in the picture. For the last subpicture in the picture, i is equal to sps_num_subpics_minus1, and in this case, subpic_width_minus1[i] is inferred to be equal to (pic_width_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1-subpic_ctu_top_left_x[sps_num_subpics_minus1]. If there is only one subpicture in the picture, i can only be 0, and subpic_ctu_top_left_x[0] is 0. Therefore, subpic_width_minus1[i] is inferred to be equal to (pic_width_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1-subpic_ctu_top_left_x[0] or (pic_width_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1.

[0154]

[0181] subpic_height_minus1[i] plus 1 specifies the height of the i-th subpicture in units of CtbSizeY. The length of this syntax element is Ceil(Log2(pic_height_max_in_luma_samples÷CtbSizeY)) bits. If not present, the value of the syntax element subpic_height_minus1[sps_num_subpics_minus1] is inferred to be equal to (Ceil(pic_height_max_in_luma_samples)÷CtbSizeY)-1-(i==sps_num_subpics_minus1?subpic_ctu_top_left_y[i]:0). Here, "Ceil()" is a function for rounding up to the nearest integer. That is, Ceil(pic_height_max_in_luma_samples÷CtbSizeY) is equal to (pic_height_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1, where " / " is integer division. "sps_num_subpics_minus1" is the number of subpictures in the picture. For the last subpicture in the picture, i is equal to sps_num_subpics_minus1, and in this case, subpic_height_minus1[i] is inferred to be equal to (pic_height_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1-subpic_ctu_top_left_y[sps_num_subpics_minus1]. If there is only one subpicture in the picture, i can only be 0, and subpic_ctu_top_left_y[0] is 0. Therefore, subpic_width_minus1[i] is inferred to be equal to (pic_height_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1-subpic_ctu_top_left_y[0] or (pic_height_max_in_luma_samples+CtbSizeY-1) / CtbSizeY-1.

[0155]

[0182] In some embodiments of the present disclosure, the sub-picture ID is signaled only if there is a sub-picture. FIG. 20 shows an exemplary Table 10 showing another exemplary SPS syntax (highlighted and italicized within boxes 2001-2002) according to some embodiments of the present disclosure.

[0156]

[0183] FIG. 21 shows a flowchart of an exemplary video processing method 2100 according to some embodiments of the present disclosure. Method 2100 can be executed by an encoder (e.g., by process 200A of FIG. 2A or process 200B of FIG. 2B), by a decoder (e.g., by process 300A of FIG. 3A or process 300B of FIG. 3B), or by one or more software or hardware components of a device (e.g., device 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) can execute method 2100. In some embodiments, method 2100 can be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions such as program code executed by a computer (e.g., device 400 of FIG. 4).

[0157]

[0184] In step 2101, it can be determined whether a sub-picture ID mapping is in the bitstream. In some embodiments, method 2100 can include signaling a flag indicating whether a sub-picture ID mapping is in the bitstream. For example, this flag can be the sps_subpic_id_present_flag shown in Table 2 of FIG. 11, Table 4 of FIG. 13, Table 5 of FIG. 15, or Table 6 of FIG. 16.

[0158]

[0185] In step 2103, it can be determined whether one or more sub-picture IDs are signaled within the first syntax or the second syntax. In step 2105, in response to determining that there is a sub-picture ID mapping and that one or more sub-picture IDs are not signaled within the first syntax and the second syntax, the one or more sub-picture IDs are signaled within the third syntax. The first syntax, the second syntax, or the third syntax is one of SPS, PPS, and PH. For example, the first syntax, the second syntax, and the third syntax are respectively SPS, PPS, and PH. Thus, if there is a sub-picture ID mapping (e.g., sps_subpic_id_present_flag = 1), the sub-picture ID may be forced to be signaled within SPS, PPS, or PH.

[0159]

[0186] In some embodiments, method 2100 may include signaling a first flag (e.g., sps_subpic_id_signaling_present_flag shown in Table 2 of FIG. 11, Table 5 of FIG. 15, or Table 6 of FIG. 16) indicating that one or more sub-picture IDs are signaled within the first syntax (e.g., SPS shown in Table 2 of FIG. 11). In some embodiments, method 2100 may include signaling a second flag (e.g., pps_subpic_id_signaling_present_flag shown in Table 3 of FIG. 12, Table 5 of FIG. 15, or Table 6 of FIG. 16) indicating that one or more sub-picture IDs are signaled within the second syntax (e.g., PPS shown in Table 3 of FIG. 12). In some embodiments, method 2100 may include signaling a third flag (e.g., ph_subpic_id_signaling_present_flag shown in Table 5 of FIG. 15) indicating that one or more sub-picture IDs are signaled within the third syntax (e.g., PH shown in Table 5 of FIG. 15).

[0160]

[0187] In some embodiments, method 2100 may include determining whether the bitstream includes a third flag indicating that one or more sub-picture IDs are signaled within a third syntax, and signaling one or more sub-picture IDs within the third syntax in response to the bitstream not including the third flag. For example, the third syntax may be PH, and the third flag may be ph_subpic_id_signaling_present_flag. If ph_subpic_id_signaling_present_flag is not signaled within PH, it can be inferred that ph_subpic_id_signaling_present_flag is 1, and one or more sub-picture IDs are signaled within PH.

[0161]

[0188] In some embodiments, method 2100 may include signaling one or more sub-picture IDs within a second syntax and a third syntax (e.g., Table 5 of FIG. 15) in response to determining that the one or more sub-picture IDs are not signaled within a first syntax.

[0162]

[0189] FIG. 22 shows a flowchart of an exemplary video processing method 2200 according to some embodiments of the present disclosure. Method 2200 may be executed by an encoder (e.g., by process 200A of FIG. 2A or process 200B of FIG. 2B), by a decoder (e.g., by process 300A of FIG. 3A or process 300B of FIG. 3B), or by one or more software or hardware components of a device (e.g., device 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) may execute method 2200. In some embodiments, method 2200 may be implemented by a computer program product embodied in a computer-readable medium including computer-executable instructions such as program code executed by a computer (e.g., device 400 of FIG. 4).

[0163]

[0190] In step 2201, it can be determined whether one or more sub-picture IDs are signaled in at least one of SPS, PH, or PPS. In some embodiments, method 2200 may include determining whether one or more sub-picture IDs are signaled in PH before determining whether one or more sub-picture IDs are signaled in PPS. In some embodiments, method 2200 may include determining whether one or more sub-picture IDs are signaled in PPS before determining whether one or more sub-picture IDs are signaled in PH.

[0164]

[0191] In step 2203, in response to determining that one or more sub-picture IDs are not signaled in SPS, PH, and PPS, it can be determined that one or more sub-picture IDs have default values.

[0165]

[0192] FIG. 23 shows a flowchart of an exemplary video processing method 2300 according to some embodiments of the present disclosure. Method 2300 can be executed by an encoder (e.g., by process 200A of FIG. 2A or process 200B of FIG. 2B), by a decoder (e.g., by process 300A of FIG. 3A or process 300B of FIG. 3B), or by one or more software or hardware components of a device (e.g., device 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) can execute method 2300. In some embodiments, method 2300 can be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions such as program code executed by a computer (e.g., device 400 of FIG. 4).

[0166]

[0193] In step 2301, it can be determined whether the number of sub-pictures of an encoded picture is equal to 1. For example, as shown in Table 8 of FIG. 18, it can be determined whether sps_num_subpic_minus1 is greater than 0.

[0167]

[0194] In step 2303, in response to determining that the number of sub-pictures is equal to 1, the sub-pictures of the coded picture can be treated as pictures within the decoding process. For example, in response to determining that the number of sub-pictures is equal to 1, the flag subpic_treated_as_pic_flag[i] can be inferred to be equal to 1. In some embodiments, in response to determining that the number of sub-pictures is equal to 1, in-loop filtering operations can be excluded.

[0168]

[0195] FIG. 24 shows a flowchart of an exemplary video processing method 2400 according to some embodiments of the present disclosure. Method 2400 can be executed by an encoder (e.g., by process 200A of FIG. 2A or process 200B of FIG. 2B), by a decoder (e.g., by process 300A of FIG. 3A or process 300B of FIG. 3B), or by one or more software or hardware components of a device (e.g., device 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) can execute method 2400. In some embodiments, method 2400 can be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions such as program code executed by a computer (e.g., device 400 of FIG. 4).

[0169]

[0196] In step 2401, it can be determined whether the sub-picture is the last sub-picture of the picture. In step 2403, in response to determining that the sub-picture is the last sub-picture, the information on the position or size of the sub-picture can be derived from the size of the picture and the sizes and positions of the previous sub-pictures of the picture.

[0170]

[0197] FIG. 25 shows a flowchart of an exemplary video processing method 2500 according to some embodiments of the present disclosure. The method 2500 can be executed by an encoder (e.g., by process 200A of FIG. 2A or process 200B of FIG. 2B), by a decoder (e.g., by process 300A of FIG. 3A or process 300B of FIG. 3B), or by one or more software or hardware components of a device (e.g., device 400 of FIG. 4). For example, a processor (e.g., processor 402 of FIG. 4) can execute the method 2500. In some embodiments, the method 2500 can be implemented by a computer program product embodied in a computer-readable medium that includes computer-executable instructions such as program code executed by a computer (e.g., device 400 of FIG. 4).

[0171]

[0198] In step 2501, it can be determined whether a subpicture is within a picture. For example, this determination can be made based on a flag (e.g., subpics_present_flag shown in Table 10 of FIG. 20).

[0172]

[0199] In step 2503, in response to determining that one or more subpictures are within a picture, a first flag indicating whether a subpicture ID mapping is within the SPS can be signaled. For example, the first flag can be sps_subpic_id_present_flag shown in Table 10 of FIG. 20.

[0173]

[0200] In some embodiments, method 2500 may include signaling a second flag indicating whether a subpicture ID mapping is signaled within the SPS in response to a first flag indicating that the subpicture ID mapping is within the SPS. In response to the second flag indicating that the subpicture ID mapping is signaled within the SPS, the subpicture IDs of one or more subpictures can be signaled within the SPS. For example, the second flag may be the sps_subpic_id_signaling_present_flag shown in Table 10 of FIG. 20. When sps_subpic_id_signaling_present_flag is true, sps_subpic_id[i] can be signaled.

[0174]

[0201] The embodiments can be further described using the following clauses: 1. A video processing method, comprising: determining whether a subpicture ID mapping is in a bitstream; determining whether one or more subpicture IDs are signaled within a first syntax or a second syntax; and signaling one or more subpicture IDs within a third syntax in response to determining that there is a subpicture ID mapping and that one or more subpicture IDs are not signaled within the first syntax and the second syntax. The video processing method as described above. 2. The method according to clause 1, wherein the first syntax, the second syntax, or the third syntax is one of a sequence parameter set (SPS), a picture parameter set (PPS), and a picture header (PH). 3. The method according to clauses 1 and 2, further comprising signaling a first flag indicating that one or more subpicture IDs are signaled within the first syntax, or signaling a second flag indicating that one or more subpicture IDs are signaled within the second syntax. The method according to clauses 1 and 2, further comprising the above. Signaling a third flag indicating that one or more sub-picture IDs are signaled within a third syntax The method according to any one of clauses 1 to 3, further comprising. 5. Determining whether the bitstream contains a third flag indicating that one or more sub-picture IDs are signaled within a third syntax, and Signaling one or more sub-picture IDs within the third syntax in response to the bitstream not containing the third flag The method according to any one of clauses 1 to 4, further comprising. 6. Signaling one or more sub-picture IDs within the second syntax and the third syntax in response to determining that one or more sub-picture IDs are not signaled within the first syntax The method according to any one of clauses 1 to 5, further comprising. 7. Signaling a fourth flag indicating whether a sub-picture ID mapping is within the bitstream The method according to any one of clauses 1 to 6, further comprising. 8. A video processing device, comprising At least one memory for storing instructions, and At least one processor, the at least one processor being configured to Determine whether a sub-picture ID mapping is within the bitstream, Determine whether one or more sub-picture IDs are signaled within the first syntax or the second syntax, and In response to determining that there is a sub-picture ID mapping and that one or more sub-picture IDs are not signaled within the first syntax and the second syntax, signal one or more sub-picture IDs within the third syntax A video processing device configured to execute instructions to cause the device to perform. 9. The first syntax, the second syntax, or the third syntax is the device according to clause 8, which is one of a sequence parameter set (SPS), a picture parameter set (PPS), and a picture header (PH). 10. At least one processor signals a first flag indicating that one or more sub-picture IDs are signaled within the first syntax, or signals a second flag indicating that one or more sub-picture IDs are signaled within the second syntax The device according to clauses 8 and 9, which is configured to execute instructions to cause the device to do so. 11. At least one processor signals a third flag indicating that one or more sub-picture IDs are signaled within the third syntax The device according to any one of clauses 8 to 10, which is configured to execute instructions to cause the device to do so. 12. At least one processor determines whether the bitstream includes a third flag indicating that one or more sub-picture IDs are signaled within the third syntax, and signals one or more sub-picture IDs within the third syntax in response to the bitstream not including the third flag The device according to any one of clauses 8 to 11, which is configured to execute instructions to cause the device to do so. 13. At least one processor signals one or more sub-picture IDs within the second syntax and the third syntax in response to a determination that one or more sub-picture IDs are not signaled within the first syntax The device according to any one of clauses 8 to 12, which is configured to execute instructions to cause the device to do so. 14. At least one processor signals a fourth flag indicating whether a sub-picture ID mapping is within the bitstream A device according to any one of clauses 8 to 13, configured to execute instructions to cause the device to perform an operation. 15. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions comprising: Determining whether a sub-picture ID mapping is present in a bitstream; Determining whether one or more sub-picture IDs are signaled within a first syntax or a second syntax; and In response to determining that a sub-picture ID mapping is present and that one or more sub-picture IDs are not signaled within the first syntax and the second syntax, signaling one or more sub-picture IDs within a third syntax The non-transitory computer-readable storage medium is executable by one or more processing devices to cause a video processing device to perform a method including the above steps. 16. The non-transitory computer-readable storage medium according to clause 15, wherein the first syntax, the second syntax, or the third syntax is one of a sequence parameter set (SPS), a picture parameter set (PPS), and a picture header (PH). 17. The set of instructions further comprises: Signaling a first flag indicating that one or more sub-picture IDs are signaled within the first syntax; or Signaling a second flag indicating that one or more sub-picture IDs are signaled within the second syntax The non-transitory computer-readable storage medium according to clauses 15 and 16 is executable by one or more processing devices to cause a video processing device to perform the above operations. 18. The set of instructions further comprises: Signaling a third flag indicating that one or more sub-picture IDs are signaled within the third syntax The non-transitory computer-readable storage medium according to any one of clauses 15 to 17 is executable by one or more processing devices to cause a video processing device to perform the above operation. 19. The set of instructions further comprises: Determining whether the bitstream includes a third flag indicating that one or more sub-picture IDs are signaled within a third syntax, and in response to the bitstream not including the third flag, signaling one or more sub-picture IDs within the third syntax A non-transitory computer-readable storage medium according to any one of clauses 15 to 18, which is executable by one or more processing devices to cause the video processing device to perform the above. 20. The set of instructions in response to determining that one or more sub-picture IDs are not signaled within the first syntax, signaling one or more sub-picture IDs within the second and third syntaxes A non-transitory computer-readable storage medium according to any one of clauses 15 to 19, which is executable by one or more processing devices to cause the video processing device to perform the above. 21. The set of instructions signaling a fourth flag indicating whether a sub-picture ID mapping is within the bitstream A non-transitory computer-readable storage medium according to any one of clauses 15 to 20, which is executable by one or more processing devices to cause the video processing device to perform the above. 22. A video processing method, comprising: determining whether one or more sub-picture IDs are signaled in at least one of a sequence parameter set (SPS), a picture header (PH), or a picture parameter set (PPS); and in response to determining that one or more sub-picture IDs are not signaled in the SPS, PH, and PPS, determining that one or more sub-picture IDs have default values A video processing method including the above. 23. The method according to clause 22, wherein before determining whether one or more sub-picture IDs are signaled in the PPS, it is determined whether one or more sub-picture IDs are signaled in the PH. 24. The method according to clause 22, wherein before determining whether one or more sub-picture IDs are signaled within the PH, it is determined whether one or more sub-picture IDs are signaled within the PPS. 25. A video processing method, determining whether the number of sub-pictures of an encoded picture is equal to 1, and in response to determining that the number of sub-pictures is equal to 1, treating the sub-picture of the encoded picture as a picture within a decoding process A video processing method including the above. 26. Further including excluding in-loop filtering operations in response to determining that the number of sub-pictures is equal to 1 The method according to clause 25. 27. A video processing method, determining whether a sub-picture is the last sub-picture of a picture, and in response to determining that the sub-picture is the last sub-picture, deriving information on the position or size of the sub-picture from the size of the picture and the sizes and positions of the previous sub-pictures of the picture A video processing method including the above. 28. A video processing method, determining whether a sub-picture is within a picture, and in response to determining that one or more sub-pictures are within the picture, signaling a first flag indicating whether a sub-picture ID mapping is within a sequence parameter set (SPS) A video processing method including the above. 29. In response to the first flag indicating that the sub-picture ID mapping is within the SPS, signaling a second flag indicating whether the sub-picture ID mapping is signaled within the SPS The method according to clause 28, further including the above. In response to a second flag indicating that sub-picture ID mapping is signaled within the SPS, signaling the sub-picture IDs of one or more sub-pictures within the SPS The method according to clause 29, further comprising. 31. A video processing device, At least one memory for storing instructions, and At least one processor, the at least one processor Determining whether one or more sub-picture IDs are signaled in at least one of a sequence parameter set (SPS), a picture header (PH), or a picture parameter set (PPS), and Determining that one or more sub-picture IDs have default values in response to determining that the one or more sub-picture IDs are not signaled in the SPS, PH, and PPS A video processing device configured to execute instructions to cause the device to perform. 32. The at least one processor Determining whether one or more sub-picture IDs are signaled in the PH before determining whether one or more sub-picture IDs are signaled in the PPS The device according to clause 31, configured to execute instructions to cause the device to perform. 33. The at least one processor Determining whether one or more sub-picture IDs are signaled in the PPS before determining whether one or more sub-picture IDs are signaled in the PH The device according to clause 31, configured to execute instructions to cause the device to perform. 34. A video processing device, At least one memory for storing instructions, and At least one processor, the at least one processor Determining whether the number of sub-pictures of an encoded picture is equal to 1, and In response to determining that the number of sub - pictures is equal to 1, treating the sub - pictures of the coded picture as pictures within the decoding process A video processing device configured to execute instructions to cause the device to do so. 35. At least one processor is configured to In response to determining that the number of sub - pictures is equal to 1, exclude in - loop filtering operations The device according to clause 34, configured to execute instructions to cause the device to do so. 36. A video processing device, comprising At least one memory for storing instructions, and At least one processor, the at least one processor being configured to Determine whether a sub - picture is the last sub - picture of a picture, and In response to determining that the sub - picture is the last sub - picture, derive information on the position or size of the sub - picture from the size of the picture and the sizes and positions of the previous sub - pictures of the picture The video processing device is configured to execute instructions to cause the device to do so. 37. A video processing device, comprising At least one memory for storing instructions, and At least one processor, the at least one processor being configured to Determine whether a sub - picture is within a picture, and In response to determining that one or more sub - pictures are within the picture, signal a first flag indicating whether a sub - picture ID mapping is within a sequence parameter set (SPS) The video processing device is configured to execute instructions to cause the device to do so. 38. At least one processor is configured to In response to a first flag indicating that a sub-picture ID mapping is within the SPS, signaling a second flag indicating whether the sub-picture ID mapping is signaled within the SPS The apparatus according to clause 37, configured to execute an instruction to cause the apparatus to perform 39. At least one processor In response to a second flag indicating that a sub-picture ID mapping is signaled within the SPS, signaling the sub-picture ID of one or more sub-pictures within the SPS The apparatus according to clause 38, configured to execute an instruction to cause the apparatus to perform 40. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions Determining whether one or more sub-picture IDs are signaled in at least one of a sequence parameter set (SPS), a picture header (PH), or a picture parameter set (PPS), and In response to determining that one or more sub-picture IDs are not signaled within the SPS, PH, and PPS, determining that the one or more sub-picture IDs have default values A non-transitory computer-readable storage medium executable by one or more processing devices to cause a video processing device to perform a method including 41. The set of instructions Determining whether one or more sub-picture IDs are signaled within the PH before determining whether the one or more sub-picture IDs are signaled within the PPS The non-transitory computer-readable storage medium according to clause 40, executable by one or more processing devices to cause a video processing device to perform 42. The set of instructions Determining whether one or more sub-picture IDs are signaled within the PPS before determining whether the one or more sub-picture IDs are signaled within the PH A non - transitory computer - readable storage medium according to clause 40, which is executable by one or more processing devices to cause a video processing device to perform. 43. A non - transitory computer - readable storage medium storing a set of instructions, the set of instructions determining whether the number of sub - pictures of an encoded picture is equal to 1, and in response to determining that the number of sub - pictures is equal to 1, treating the sub - picture of the encoded picture as a picture within a decoding process A non - transitory computer - readable storage medium that is executable by one or more processing devices to cause a video processing device to perform a method including the above. 44. The set of instructions excluding in - loop filtering operations in response to determining that the number of sub - pictures is equal to 1 A non - transitory computer - readable storage medium according to clause 43, which is executable by one or more processing devices to cause a video processing device to perform. 45. A non - transitory computer - readable storage medium storing a set of instructions, the set of instructions determining whether a sub - picture is the last sub - picture of a picture, and in response to determining that the sub - picture is the last sub - picture, deriving information on the position or size of the sub - picture from the size of the picture and the size and position of the previous sub - picture of the picture A non - transitory computer - readable storage medium that is executable by one or more processing devices to cause a video processing device to perform a method including the above. 46. A non - transitory computer - readable storage medium storing a set of instructions, the set of instructions determining whether a sub - picture is within a picture, and in response to determining that one or more sub - pictures are within the picture, signaling a first flag indicating whether a sub - picture ID mapping is within a sequence parameter set (SPS) A non - transitory computer - readable storage medium executable by one or more processing devices to cause a video processing device to perform a method including 47. The set of instructions in response to a first flag indicating that a sub - picture ID mapping is within the SPS, signaling a second flag indicating whether the sub - picture ID mapping is signaled within the SPS A non - transitory computer - readable storage medium according to clause 46, executable by one or more processing devices to cause a video processing device to perform 48. The set of instructions in response to a second flag indicating that a sub - picture ID mapping is signaled within the SPS, signaling the sub - picture IDs of one or more sub - pictures within the SPS A non - transitory computer - readable storage medium according to clause 47, executable by one or more processing devices to cause a video processing device to perform 49. A video processing method comprising determining whether a bitstream includes sub - picture information according to a sub - picture information presence flag signaled within the bitstream, and in response to the bitstream including sub - picture information, signaling within the bitstream at least one of the number of sub - pictures within a picture, the width, height, position, and identifier (ID) mapping of a target sub - picture, subpic_treated_as_pic_flag, and loop_filter_across_subpic_enabled_flag A video processing method including 50. The method according to clause 49, wherein signaling of at least one of the width, height, and position of a target sub - picture is based on the number of sub - pictures within a picture. 51. If there are at least two sub-pictures in a picture, signaling at least one of subpic_treated_as_pic_flag, loop_filter_across_subpic_enabled_flag, and the width, height, and position of the target sub-picture further comprising If there is only one sub-picture in a picture, signaling of at least one of the width, height, and position of the target sub-picture is skipped, subpic_treated_as_pic_flag indicates whether the sub-picture of each coded picture in the coded layer video sequence (CLVS) is treated as a picture in the decoding process that excludes in-loop filtering operations, and loop_filter_across_subpic_enabled_flag is the method described in clause 50 that indicates whether in-loop filtering operations across the boundaries of the sub-pictures of each coded picture in the CLVS are enabled. 52. If the width of the target sub-picture is not signaled, determining the value of the width of the target sub-picture as the width of the picture, and If the height of the target sub-picture is not signaled, determining the value of the height of the target sub-picture as the height of the picture The method according to clause 51, further comprising. 53. If the width of the target sub-picture is not signaled, determining the value of the width of the target sub-picture in units of the coding tree block (CTB) size as the width of the picture in units of the CTB size, and If the height of the target sub-picture is not signaled, determining the value of the height of the target sub-picture in units of the CTB size as the height of the picture in units of the CTB size The method according to clause 52, further comprising. When 54.subpic_treated_as_pic_flag is not signaled in the bitstream, determining that subpic_treated_as_pic_flag has a value of 1, and when loop_filter_across_subpic_enabled_flag is not signaled in the bitstream, determining that loop_filter_across_subpic_enabled_flag has a value of 0 The method according to clause 51, further comprising. 55. When the target subpicture is the last subpicture in the picture, skipping the signaling of at least one of the width and height of the target subpicture The method according to clause 49, further comprising. 56. When the width of the target subpicture is not signaled, determining the value of the width of the target subpicture in coding tree block (CTB) size units as the width of the picture in CTB size units minus the horizontal position of the top-left coding tree unit (CTU) of the target subpicture in CTB size units, or determining the value of the width of the target subpicture as the width of the picture minus the horizontal position of the top-left CTU of the target subpicture, and when the height of the target subpicture is not signaled, determining the value of the height of the target subpicture in CTB size units as the height of the picture in CTB size units minus the vertical position of the top-left CTU of the target subpicture in CTB size units, or determining the value of the height of the target subpicture as the height of the picture minus the vertical position of the top-left CTU of the target subpicture The method according to clause 55, further comprising. 57. The signaling of the ID mapping of the target subpicture is signaling a first flag in the bitstream, and In response to the first flag being equal to 1, signaling the ID mapping of the target sub-picture within the first data unit or the second data unit, further comprising The method according to clause 49, wherein the first flag equal to 0 indicates that the ID mapping of the target sub-picture is not signaled within the bitstream. 58. In response to the first flag being equal to 1 and the ID mapping of the target sub-picture not being signaled within the first data unit, signaling the ID mapping of the target sub-picture within the second data unit, or skipping signaling the ID mapping of the target sub-picture within the second data unit in response to the first flag being equal to 0 or the ID mapping of the target sub-picture being signaled within the first data unit The method according to clause 57, further comprising 59. The method according to clause 58, wherein each of the first data unit and the second data unit is one of a sequence parameter set (SPS), a picture parameter set (PPS), or a picture header (PH). 60. A video processing device, at least one memory for storing instructions, and at least one processor, the at least one processor determining whether the bitstream contains sub-picture information according to a sub-picture information presence flag signaled within the bitstream, and in response to the bitstream containing sub-picture information, the number of sub-pictures within the picture, the width, height, position and identifier (ID) mapping of the target sub-picture, subpic_treated_as_pic_flag, and loop_filter_across_subpic_enabled_flag signaling at least one of them within a bitstream A video processing device configured to execute instructions to cause a device to do so. 61. The device according to clause 60, wherein signaling of at least one of the width, height, and position of a target subpicture is based on the number of subpictures in a picture. 62. At least one processor when there are at least two subpictures in a picture, signaling at least one of subpic_treated_as_pic_flag, loop_filter_across_subpic_enabled_flag, and the width, height, and position of a target subpicture configured to execute instructions to cause a device to do so, when there is only one subpicture in a picture, signaling of at least one of the width, height, and position of the target subpicture is skipped, subpic_treated_as_pic_flag indicates whether the subpictures of each coded picture in a coded layer video sequence (CLVS) are treated as pictures in a decoding process excluding in-loop filtering operations, and loop_filter_across_subpic_enabled_flag indicates whether in-loop filtering operations across subpicture boundaries are enabled for the subpictures of each coded picture in a CLVS, the device according to clause 61. 63. At least one processor when the width of the target subpicture is not signaled, determining the value of the width of the target subpicture as the width of the picture, and when the height of the target subpicture is not signaled, determining the value of the height of the target subpicture as the height of the picture configured to execute instructions to cause a device to do so, the device according to clause 62. 64. At least one processor When the width of the target sub-picture is not signaled, determining the value of the width of the target sub-picture in units of coding tree block (CTB) size as the width of the picture in units of CTB size, and When the height of the target sub-picture is not signaled, determining the value of the height of the target sub-picture in units of CTB size as the height of the picture in units of CTB size The apparatus according to clause 63, configured to execute an instruction to cause the apparatus to do so. 65. At least one processor When subpic_treated_as_pic_flag is not signaled in the bitstream, determining that subpic_treated_as_pic_flag has a value of 1, and When loop_filter_across_subpic_enabled_flag is not signaled in the bitstream, determining that loop_filter_across_subpic_enabled_flag has a value of 0 The apparatus according to clause 62, configured to execute an instruction to cause the apparatus to do so. 66. At least one processor When the target sub-picture is the last sub-picture in the picture, skipping at least one signaling of the width and height of the target sub-picture The apparatus according to clause 60, configured to execute an instruction to cause the apparatus to do so. 67. At least one processor When the width of the target sub-picture is not signaled, determining the value of the width of the target sub-picture in units of CTB size as the value obtained by subtracting the horizontal position of the top-left coding tree unit (CTU) of the target sub-picture in units of CTB size from the width of the picture in units of CTB size, or determining the value of the width of the target sub-picture as the value obtained by subtracting the horizontal position of the top-left coding tree unit (CTU) of the target sub-picture from the width of the picture, and When the height of the target sub-picture is not signaled, determine the value of the height of the target sub-picture in CTB size units as the height of the picture in CTB size units minus the vertical position of the top-left CTU of the target sub-picture in CTB size units, or determine the value of the height of the target sub-picture as the height of the picture minus the vertical position of the top-left CTU of the target sub-picture. The apparatus according to clause 66, configured to execute instructions to cause the apparatus to perform. 68. A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions determine whether the bitstream contains sub-picture information according to a sub-picture information presence flag signaled in the bitstream, and in response to the bitstream containing sub-picture information, the number of sub-pictures in the picture, the width, height, position, and identifier (ID) mapping of the target sub-picture, subpic_treated_as_pic_flag, and loop_filter_across_subpic_enabled_flag signal at least one of in the bitstream. A non-transitory computer-readable storage medium executable by one or more processing devices to cause a video processing device to perform a method including.

[0175]

[0202] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by a device (such as the encoder and decoder of the present disclosure) to perform the above-described method. Examples of common forms of non-transitory media include, for example, floppy (registered trademark) disks, flexible disks, hard disks, solid state drives, magnetic tapes, or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media having a pattern of holes, RAM, PROM, and EPROM, FLASH (registered trademark)-EPROM, or any other flash memory, NVRAM, caches, registers, any other memory chips, or cartridges, as well as networked versions of these. The device can include one or more processors (CPUs), an input / output interface, a network interface, and / or a memory.

[0176]

[0203] It should be noted that relative terms in this specification such as "first" and "second" are merely used to distinguish one entity or operation from another entity or operation, and do not require, imply, or suggest any actual relationship or order between these entities or operations. Further, the words "comprising", "having", "containing", "including", and other similar forms are of equal meaning, and the elements or groups of elements following any one of these words are not meant to be a limiting enumeration of such elements or groups of elements, or to be limited to only the enumerated elements or groups of elements, and are intended to be open-ended.

[0177]

[0204] As used herein, unless otherwise specifically stated, the term "or" includes all possible combinations, except when the combination is infeasible. For example, if it is stated that a database can include A or B, then, unless otherwise specifically stated or infeasible, the database can include A or B or both A and B. As a second example, if it is stated that a database can include A, B, or C, then, unless otherwise specifically stated or infeasible, the database can include A, B, or C, or A and B, A and C, or B and C, or A, B, and C.

[0178]

[0205] It is understood that the above-described embodiments can be implemented by hardware or software (program code) or a combination of hardware and software. When implemented by software, it can be stored in the above-described computer-readable medium. The software can perform the method of the present disclosure when executed by a processor. The computing unit and other functional units described in the present disclosure can be implemented by hardware or software or a combination of hardware and software. Those skilled in the art will understand that a plurality of the above-described modules / units can be combined into one module / unit, and each of the above-described modules / units can be further divided into a plurality of sub-modules / sub-units.

[0179]

[0206] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Specific adaptations and modifications of the foregoing embodiments may be made. From the consideration of this specification and the practice of the invention disclosed herein, other embodiments may become apparent to those skilled in the art. This specification and examples are intended to be considered only as examples, and the true scope and spirit of the invention are indicated by the appended claims. Also, the arrangement of steps shown in the figures is for illustrative purposes only and is not intended to be limited to any particular arrangement of steps. Therefore, those skilled in the art can understand that these steps can be performed in a different order while implementing the same method.

[0180]

[0207] In the drawings and in this specification, exemplary embodiments have been disclosed. However, many variations and modifications can be made to these embodiments. Therefore, even if specific terms are employed, they are used only in a general descriptive sense and not for purposes of limitation.

Claims

1. A video processing method, comprising: determining whether the bitstream includes sub-picture information according to a sub-picture information presence flag signaled within the bitstream; and in response to the bitstream including the sub-picture information, signaling at least one of the number of sub-pictures within the picture, the width, height, position, and identifier (ID) mapping of a target sub-picture, subpic_treated_as_pic_flag, and loop_filter_across_subpic_enabled_flag within the bitstream.

2. The method according to claim 1, wherein the signaling of at least one of the width, height, and position of the target sub-picture is based on the number of sub-pictures within the picture.

3. further comprising, when there are at least two sub-pictures within the picture, signaling at least one of the subpic_treated_as_pic_flag, the loop_filter_across_subpic_enabled_flag, and the width, height, and position of the target sub-picture; wherein, when there is only one sub-picture within the picture, the signaling of at least one of the width, height, and position of the target sub-picture is skipped; wherein the subpic_treated_as_pic_flag indicates whether sub-pictures of each coded picture within a coded layer video sequence (CLVS) are treated as pictures within a decoding process that excludes in-loop filtering operations; and wherein the loop_filter_across_subpic_enabled_flag indicates whether in-loop filtering operations across sub-picture boundaries are possible across sub-picture boundaries of each coded picture within the CLVS, the method according to claim 2.

4. determining the value of the width of the target sub-picture as the width of the picture when the width of the target sub-picture is not signaled; and determining the value of the height of the target sub-picture as the height of the picture when the height of the target sub-picture is not signaled ​ The method according to claim 3, further comprising

5. When the width of the target sub-picture is not signaled, determining the value of the width of the target sub-picture in units of coding tree blocks (CTBs) as the width of the picture in units of CTB sizes, and When the height of the target sub-picture is not signaled, determining the value of the height of the target sub-picture in units of CTB sizes as the height of the picture in units of CTB sizes The method according to claim 4, further comprising

6. When the subpic_treated_as_pic_flag is not signaled in the bitstream, determining that the subpic_treated_as_pic_flag has a value of 1, and When the loop_filter_across_subpic_enabled_flag is not signaled in the bitstream, determining that the loop_filter_across_subpic_enabled_flag has a value of 0 The method according to claim 3, further comprising

7. When the target sub-picture is the last sub-picture in the picture, skipping the signaling of at least one of the width and the height of the target sub-picture The method according to claim 1, further comprising

8. When the width of the target sub-picture is not signaled, determining the value of the width of the target sub-picture in units of coding tree blocks (CTBs) as the value obtained by subtracting the horizontal position of the top-left coding tree unit (CTU) of the target sub-picture in units of CTB sizes from the width of the picture in units of CTB sizes, or determining the value of the width of the target sub-picture as the value obtained by subtracting the horizontal position of the top-left coding tree unit (CTU) of the target sub-picture from the width of the picture, and When the height of the target sub-picture is not signaled, determine the value of the height of the target sub-picture in units of CTB size as the height of the picture in units of CTB size minus the vertical position of the top-left CTU of the target sub-picture in units of CTB size, or determine the value of the height of the target sub-picture as the height of the picture minus the vertical position of the top-left CTU of the target sub-picture. The method according to claim 7, further comprising.

9. Signaling the ID mapping of the target sub-picture includes signaling a first flag in the bitstream, and in response to the first flag being equal to 1, signaling the ID mapping of the target sub-picture within a first data unit or a second data unit. further comprising The method according to claim 1, wherein the first flag equal to 0 indicates that the ID mapping of the target sub-picture is not signaled in the bitstream.

10. In response to the first flag being equal to 1 and the ID mapping of the target sub-picture not being signaled within the first data unit, signaling the ID mapping of the target sub-picture within the second data unit, or skipping signaling the ID mapping of the target sub-picture within the second data unit in response to the first flag being equal to 0 or the ID mapping of the target sub-picture being signaled within the first data unit. The method according to claim 9, further comprising.

11. Each of the first data unit and the second data unit is one of a sequence parameter set (SPS), a picture parameter set (PPS), or a picture header (PH). The method according to claim 10.

12. A video processing device, comprising at least one memory for storing instructions, and at least one processor, wherein the at least one processor determines whether the bitstream contains sub-picture information according to a sub-picture information presence flag signaled in the bitstream, and In response to the bitstream including the subpicture information, the number of subpictures within a picture, the width, height, position, and identifier (ID) mapping of a target subpicture, subpic_treated_as_pic_flag, and loop_filter_across_subpic_enabled_flag at least one of which is signaled within the bitstream configured to execute the instruction to cause the device to perform such, a video processing device. **Claim 13** The device according to claim 12, wherein signaling at least one of the width, height, and position of the target subpicture is based on the number of subpictures within the picture. **Claim 14** The at least one processor is configured to execute the instruction to cause the device to signal at least one of the subpic_treated_as_pic_flag, the loop_filter_across_subpic_enabled_flag, and the width, height, and position of the target subpicture when there are at least two subpictures within the picture, wherein signaling of at least one of the width, height, and position of the target subpicture is skipped when there is only one subpicture within the picture, wherein the subpic_treated_as_pic_flag indicates whether subpictures of each coded picture within a coded layer video sequence (CLVS) are treated as pictures within a decoding process excluding in-loop filtering operations, and wherein the loop_filter_across_subpic_enabled_flag indicates whether in-loop filtering operations across subpicture boundaries are possible across subpicture boundaries of subpictures of each coded picture within the CLVS, the device according to claim 13. **Claim 15** The at least one processor is configured to determine the value of the width of the target subpicture as the width of the picture when the width of the target subpicture is not signaled, and ​ When the height of the target sub-picture is not signaled, determining the value of the height of the target sub-picture as the height of the picture The apparatus according to claim 14, configured to execute the instruction to cause the apparatus to perform the above

16. The at least one processor When the width of the target sub-picture is not signaled, determining the value of the width of the target sub-picture in units of Coding Tree Block (CTB) size as the width of the picture in units of CTB size, and When the height of the target sub-picture is not signaled, determining the value of the height of the target sub-picture in units of CTB size as the height of the picture in units of CTB size The apparatus according to claim 15, configured to execute the instruction to cause the apparatus to perform the above

17. The at least one processor When the subpic_treated_as_pic_flag is not signaled in the bitstream, determining that the subpic_treated_as_pic_flag has a value of 1, and When the loop_filter_across_subpic_enabled_flag is not signaled in the bitstream, determining that the loop_filter_across_subpic_enabled_flag has a value of 0 The apparatus according to claim 14, configured to execute the instruction to cause the apparatus to perform the above

18. The at least one processor When the target sub-picture is the last sub-picture in the picture, skipping at least one signaling of the width and the height of the target sub-picture The apparatus according to claim 12, configured to execute the instruction to cause the apparatus to perform the above

19. The at least one processor When the width of the target sub-picture is not signaled, determine the value of the width of the target sub-picture in units of coding tree blocks (CTBs) as the value obtained by subtracting the horizontal position of the top-left coding tree unit (CTU) of the target sub-picture in units of CTB size from the width of the picture in units of CTB size, or determine the value of the width of the target sub-picture as the value obtained by subtracting the horizontal position of the top-left CTU of the target sub-picture from the width of the picture, and when the height of the target sub-picture is not signaled, determine the value of the height of the target sub-picture in units of CTB size as the value obtained by subtracting the vertical position of the top-left CTU of the target sub-picture in units of CTB size from the height of the picture in units of CTB size, or determine the value of the height of the target sub-picture as the value obtained by subtracting the vertical position of the top-left CTU of the target sub-picture from the height of the picture The apparatus according to claim 18, configured to execute the instructions to cause the apparatus to perform the above. **Claim 20** A non-transitory computer-readable storage medium storing a set of instructions, the set of instructions comprising determining whether the bitstream includes sub-picture information according to a sub-picture information presence flag signaled in the bitstream, and in response to the bitstream including the sub-picture information, the number of sub-pictures in a picture, the width, height, position, and identifier (ID) mapping of a target sub-picture, subpic_treated_as_pic_flag, and loop_filter_across_subpic_enabled_flag signaling at least one of the above in the bitstream A non-transitory computer-readable storage medium executable by one or more processing devices to cause a video processing device to perform a method including the above.

Citation Information

Patent Citations

  • Tile partitions including subtiles in video coding

    JP2021528003A

  • Intra-prediction-based image encoding / decoding method and apparatus

    JP2022515992A

  • Video coding with subpicture, slice, and tile support

    JP2022553599A

  • Tile partitions with sub-tiles in video coding

    WO2019243539A1