Methods and systems for performing gradual decoding refresh processing on pictures

Gradual decoding refresh processing enhances video coding efficiency and error resilience by enabling flexible encoding and decoding strategies, addressing the challenges of advanced video standards like VVC/H.266.

JP2025123466APending Publication Date: 2025-08-22ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2025104495
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2019-12-27
Filing Date
2025-06-20
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

Existing video coding standards face challenges in achieving high compression efficiency while maintaining subjective quality, particularly with the development of advanced standards like VVC/H.266, which require efficient methods for encoding and decoding video sequences.

Method used

The implementation of gradual decoding refresh (GDR) processing, including encoding and decoding flag data to enable or disable GDR for video sequences, and applying loop filters based on virtual boundaries to manage pixel regions, enhances the encoding and decoding processes.

Benefits of technology

This approach improves coding efficiency by allowing parallel processing and error resilience, reducing computational complexity, and maintaining video quality during decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025123466000001_ABST
    Figure 2025123466000001_ABST
Patent Text Reader

Abstract

To provide methods and systems for performing gradual decoding refresh (GDR) processing on pictures.SOLUTION: Methods and apparatuses video processing include: in response to receiving a video sequence, encoding first flag data in a parameter set associated with the video sequence, where the first flag data represents whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence; when the first flag data represents that the GDR is disabled for the video sequence, encoding a picture header associated with a picture in the video sequence to indicate that the picture is a non-GDR picture; and encoding the non-GDR picture.SELECTED DRAWING: Figure 5
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This disclosure claims priority to U.S. Provisional Patent Application No. 62 / 954,011, filed December 27, 2019, the entire contents of which are incorporated herein by reference.

[0002] Technical Field FIELD OF THE DISCLOSURE

[0002] This disclosure relates generally to video processing, and more particularly to methods and systems for performing gradual decoding refresh (GDR) processing on pictures. [Background technology]

[0003] background

[0003] A video is a set of static pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, a video can be compressed before storage or transmission and decompressed before display. The compression process is commonly referred to as encoding, and the decompression process is commonly referred to as decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transform, quantization, entropy coding, and in-loop filtering. Standardization organizations have developed video coding standards that specify specific video coding formats, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard. As more and more advanced video coding techniques are adopted into video standards, the coding efficiency of new video coding standards becomes higher and higher. Summary of the Invention [Means for solving the problem]

[0004] Disclosure Overview

[0004] Embodiments of the present disclosure provide a method and apparatus for video processing. In one aspect, a non-transitory computer-readable medium is provided. The non-transitory computer-readable medium stores a set of instructions, the set of instructions being executable by at least one processor of the apparatus to cause the apparatus to perform a method. The method includes, in response to receiving a video sequence, encoding first flag data in a parameter set associated with the video sequence, the first flag data indicating whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence; if the first flag data indicates that GDR is disabled for the video sequence, encoding a picture header associated with a picture in the video sequence to indicate that the picture is a non-GDR picture; and encoding the non-GDR picture.

[0005] In another aspect, a non-transitory computer-readable medium is provided. The non-transitory computer-readable medium stores a set of instructions, the set of instructions being executable by at least one processor of the device to cause the device to perform a method including: in response to receiving a video bitstream, decoding first flag data in a parameter set associated with a sequence of the video bitstream, the first flag data indicating whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence; if the first flag data indicates that GDR is disabled for the sequence, decoding a picture header associated with a picture in the sequence, the picture header indicating that the picture is a non-GDR picture; and decoding the non-GDR picture.

[0006] In yet another aspect, a non-transitory computer-readable medium is provided. The non-transitory computer-readable medium stores a set of instructions executable by at least one processor of the device to cause the device to perform a method, including: in response to receiving a picture of a video, determining whether the picture is a gradual decoding refresh (GDR) picture based on flag data associated with the picture; based on the determination that the picture is a GDR picture, determining a first region and a second region of the picture using a virtual boundary; disabling a loop filter for a first pixel in the first region if filtering of the first pixel uses information of a second pixel in the second region or applying a loop filter for the first pixel using only information of the pixel in the first region; and applying a loop filter for a pixel in the second region using pixels in at least one of the first region or the second region.

[0007] In yet another aspect, provided is an apparatus including: a memory configured to store a set of instructions; and one or more processors communicatively coupled to the memory, the one or more processors configured to execute the set of instructions to cause the apparatus to: in response to receiving a video sequence, encode first flag data in a parameter set associated with the video sequence, the first flag data indicating whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence, if the first flag data indicates that GDR is disabled for the video sequence, encode a picture header associated with a picture in the video sequence to indicate that the picture is a non-GDR picture, and encode the non-GDR picture.

[0008] In yet another aspect, provided is an apparatus including: a memory configured to store a set of instructions; and one or more processors communicatively coupled to the memory, the one or more processors configured to execute the set of instructions to cause the apparatus, in response to receiving a video bitstream, to: decode first flag data in a parameter set associated with a sequence of the video bitstream, the first flag data indicating whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence, if the first flag data indicates that GDR is disabled for the sequence, decode picture headers associated with pictures in the sequence, the picture headers indicating that the pictures are non-GDR pictures, and decode the non-GDR pictures.

[0009] In yet another aspect, provided is an apparatus including: a memory configured to store a set of instructions; and one or more processors communicatively coupled to the memory, wherein the one or more processors are configured to execute the set of instructions to cause the apparatus to: in response to receiving a picture of a video, determine whether the picture is a gradual decoding refresh (GDR) picture based on flag data associated with the picture, determine a first region and a second region of the picture using a virtual boundary based on a determination that the picture is a GDR picture, disable a loop filter for a first pixel in the first region if filtering of the first pixel uses information of a second pixel in the second region or apply a loop filter for the first pixel using only information of the pixel in the first region, and apply a loop filter for the pixel in the second region using pixels in at least one of the first region or the second region.

[0010] In yet another aspect, a method is provided that includes, in response to receiving a video sequence, encoding first flag data in a parameter set associated with the video sequence, the first flag data indicating whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence, encoding a picture header associated with a picture in the video sequence to indicate that the picture is a non-GDR picture if the first flag data indicates that GDR is disabled for the video sequence, and encoding the non-GDR picture.

[0011] In yet another aspect, a method is provided that includes, in response to receiving a video bitstream, decoding first flag data in a parameter set associated with a sequence of the video bitstream, the first flag data indicating whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence, decoding picture headers associated with pictures in the sequence if the first flag data indicates that GDR is disabled for the sequence, the picture headers indicating that the pictures are non-GDR pictures, and decoding the non-GDR pictures.

[0012] In yet another aspect, provided is a method that includes, in response to receiving a picture of a video, determining whether the picture is a gradual decoding refresh (GDR) picture based on flag data associated with the picture, determining a first region and a second region of the picture using a virtual boundary based on the determination that the picture is a GDR picture, disabling a loop filter for a first pixel in the first region if filtering of the first pixel uses information of a second pixel in the second region or applying a loop filter for the first pixel using only information of the pixel in the first region, and applying a loop filter for the pixel in the second region using pixels in at least one of the first region or the second region.

[0013] BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings, in which various features are not drawn to scale. [Brief explanation of the drawings]

[0014] [Figure 1]

[0014] FIG. 1 is a schematic diagram illustrating the structure of an exemplary video sequence, in accordance with some embodiments of the present disclosure. [Figure 2A]

[0015] 1 shows a schematic diagram of an exemplary encoding process of a hybrid video encoding system, according to an embodiment of the present disclosure. [Figure 2B]

[0016] 3 shows a schematic diagram of another exemplary encoding process of a hybrid video encoding system, according to an embodiment of the present disclosure. [Figure 3A]

[0017] 1 shows a schematic diagram of an exemplary decoding process of a hybrid video coding system, according to an embodiment of the present disclosure. [Figure 3B]

[0018] 1 shows a schematic diagram of another exemplary decoding process of a hybrid video coding system, according to an embodiment of the present disclosure. [Figure 4]

[0019] 1 shows a block diagram of an exemplary device for encoding or decoding video in accordance with some embodiments of the present disclosure. [Figure 5]

[0020] 1 is a schematic diagram illustrating an example operation of gradual decoding refresh (GDR) according to some embodiments of the present disclosure. [Figure 6]

[0021] Table 1 illustrates an example syntax structure of a sequence parameter set (SPS) that enables GDR, according to some embodiments of the present disclosure. [Figure 7]

[0022] Table 2 illustrates an example syntax structure of a picture header that enables GDR, according to some embodiments of the present disclosure. [Figure 8]

[0023] Table 3 illustrates an example syntax structure of an SPS that enables virtual boundaries, according to some embodiments of the present disclosure. [Figure 9]

[0024] Table 4 illustrates an example syntax structure of a picture header that enables virtual boundaries, according to some embodiments of the present disclosure. [Figure 10]

[0025] Table 5 shows an example syntax structure of a modified picture header according to some embodiments of the present disclosure. [Figure 11]

[0026] Table 6 illustrates an example syntax structure of a modified SPS that enables virtual boundaries, according to some embodiments of the present disclosure. [Figure 12]

[0027] Table 7 shows an example syntax structure of a modified picture header that enables virtual boundaries, according to some embodiments of the present disclosure. [Figure 13]

[0028] 1 illustrates a flowchart of an exemplary process for video processing according to some embodiments of the present disclosure. [Figure 14]

[0029] 10 shows a flowchart of another exemplary process for video processing according to some embodiments of the present disclosure. [Figure 15]

[0030] 10 shows a flowchart of yet another exemplary process for video processing according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0015] Detailed Description

[0031] Reference may now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings in which like reference numerals in different drawings represent the same or similar elements unless otherwise indicated. The implementations set forth in the following description of exemplary embodiments do not represent all implementations in accordance with the present invention. Rather, they are merely examples of apparatus and methods in accordance with aspects related to the present invention as recited in the appended claims. Certain aspects of the present disclosure are described in more detail below. In the event of a conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall control.

[0016]

[0032] The ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) Joint Video Experts Team (JVET) are currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 while using half the bandwidth.

[0017]

[0033] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET is developing technology beyond HEVC using the Joint Search Model (JEM) reference software. Because the coding technology has been incorporated into JEM, JEM has achieved substantially higher coding performance than HEVC.

[0018]

[0034] The VVC standard is a recent development and continues to incorporate more coding techniques that result in better compression performance. VVC is based on the same hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.

[0019]

[0035] Video is a set of static pictures (or "frames") arranged in time sequence to store visual information. A video capture device (e.g., a camera) can be used to capture and store these pictures in time sequence, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with display capabilities) can be used to display such pictures in time sequence. In some applications, a video capture device can also transmit the captured video in real time to a video playback device (e.g., a computer with a monitor) for purposes such as supervision, conferencing, or live broadcasting.

[0020]

[0036] To reduce the storage space and transmission bandwidth required by such applications, video can be compressed before storage and transmission and decompressed before display. Compression and decompression can be performed by software executed by a processor (e.g., a processor in a general-purpose computer) or by specialized hardware. A module for compression is commonly referred to as an “encoder,” and a module for decompression is commonly referred to as a “decoder.” Collectively, the encoder and decoder can be referred to as a “codec.” The encoder and decoder can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of the encoder and decoder can include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of the encoder and decoder can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression may be performed by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, or the like. In some applications, a codec may decompress video from a first encoding standard and recompress the decompressed video using a second encoding standard. In this case, the codec may be referred to as a "transcoder."

[0021]

[0037] A video coding process can identify and retain useful information that can be used to reconstruct a picture and ignore information that is not important for reconstruction. If the ignored, unimportant information cannot be perfectly reconstructed, such a coding process may be called "lossy." Otherwise, it may be called "lossless." Most coding processes are lossy; this is a tradeoff to reduce the required storage space and transmission bandwidth.

[0022]

[0038] Useful information about the picture being coded (called the "current picture" or "target picture") includes changes relative to a reference picture (e.g., a previously coded and reconstructed picture). Such changes may include changes in pixel position, brightness, or color, with position changes being the most important. Changes in the position of a group of pixels representing an object may reflect the movement of the object between the reference picture and the target picture.

[0023]

[0039] A picture that is coded without reference to another picture (i.e., it is its own reference picture) is called an "I-picture." A picture that is coded using a previous picture as a reference picture is called a "P-picture." A picture that is coded using both a previous picture and a future picture as a reference picture (i.e., the references are "bidirectional") is called a "B-picture."

[0024]

[0040] 1 illustrates the structure of an exemplary video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 may be live video or captured and archived video. The video 100 may be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video supply interface (e.g., a video broadcast transceiver) for receiving video from a video content provider.

[0025]

[0041] As shown in FIG. 1, video sequence 100 may include a series of pictures arranged temporally along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with additional pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I-picture, and its reference picture is picture 102 itself. Picture 104 is a P-picture, and its reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture, and its reference pictures are pictures 104 and 108, as indicated by the arrows. In some embodiments, the reference picture of a picture (e.g., picture 104) need not immediately precede or follow that picture. For example, the reference picture of picture 104 may be the picture before picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and this disclosure does not limit the reference picture embodiments to the examples shown in FIG.

[0026]

[0042] Typically, video codecs do not encode or decode an entire picture at once due to the computational complexity of the task. Rather, video codecs may divide a picture into elementary segments and encode or decode the picture segment by segment. Such elementary segments are referred to as basic processing units ("BPUs") in this disclosure. For example, structure 110 in FIG. 1 illustrates an example structure for a picture (e.g., any of pictures 102-108) of video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, the boundaries of which are shown as dashed lines. In some embodiments, basic processing units may be referred to as "macroblocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) or as "coding tree units" ("CTUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing units can have any arbitrary shape and size of variable size or pixels in a picture, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, etc. The size and shape of the basic processing unit can be selected based on a balance between coding efficiency and the level of detail to be maintained in the basic processing unit for the picture.

[0027]

[0043] A basic processing unit may be a logical unit that can include groups of different types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color picture may include a luma component (Y) that represents colorless luminance information, one or more chroma components (e.g., Cb and Cr) that represent color information, and related syntax elements, where the luma and chroma components may have the same size of the basic processing unit. The luma and chroma components may be referred to as "coding tree blocks" ("CTBs") in some video coding standards (e.g., H.265 / HEVC or H.266 / VVC). Any operation performed on a basic processing unit may be performed repeatedly on each of its luma and chroma components.

[0028]

[0044] Video coding has multiple computational stages, examples of which are shown in Figures 2A-2B and 3A-3B. At each stage, the size of the basic processing unit may still become too large for processing and therefore may be further divided into segments referred to as "basic processing subunits" in this disclosure. In some embodiments, the basic processing subunits may be referred to as "blocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) or as "coding units" ("CUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunits may have the same or smaller size than the basic processing units. Similar to the basic processing units, the basic processing subunits are also logical units that may contain groups of different types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing sub-unit may be repeatedly performed on each of its luma and chroma components. Note that such division may be performed to further levels as needed for processing. Also, note that different stages may use different schemes to divide the basic processing units.

[0029]

[0045] For example, in a mode decision stage (an example of which is shown in FIG. 2B ), an encoder can decide what prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a basic processing unit, but the basic processing unit may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing sub-units (e.g., CUs, as in the case of H.265 / HEVC or H.266 / VVC) and decide the type of prediction for each individual basic processing sub-unit.

[0030]

[0046] As another example, in the prediction stage (an example of which is shown in FIGS. 2A-2B), the encoder may perform prediction operations at the level of basic processing sub-units (e.g., CUs). However, in some cases, the basic processing sub-units may still be too large to process. The encoder may further divide the basic processing sub-units into smaller segments (e.g., referred to as "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), at which level the prediction operations may be performed.

[0031]

[0047] As another example, in the transform stage (an example of which is shown in FIGS. 2A-2B), the encoder may perform transform operations for residual basic processing sub-units (e.g., CUs). However, in some cases, the basic processing sub-units may still be too large to process. The encoder may further divide the basic processing sub-units into smaller segments (e.g., referred to as "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), at which levels the transform operations may be performed. Note that the division scheme of the same basic processing sub-unit may be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.

[0032]

[0048] 1, the basic processing units 112 are further divided into 3x3 basic processing sub-units, the boundaries of which are shown as dotted lines. Different basic processing units of the same picture may be divided into basic processing sub-units in different ways.

[0033]

[0049] In some implementations, to provide parallel processing and error resilience capabilities to video encoding and decoding, a picture can be divided into regions for processing, so that the encoding or decoding process does not rely on information about a picture region from any other region of the picture. In other words, each region of a picture can be processed independently. Doing so allows a codec to process different regions of a picture in parallel, thereby increasing coding efficiency. Also, when data for a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thereby providing error resilience. Some video coding standards allow a picture to be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles." It should also be noted that different pictures in video sequence 100 can have different partitioning schemes for dividing the picture into regions.

[0034]

[0050] For example, in Figure 1, structure 110 is divided into three regions 114, 116, and 118, the boundaries of which are shown as solid lines within structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that the basic processing units, basic processing subunits, and regions of structure 110 in Figure 1 are merely examples, and the present disclosure does not limit the embodiments thereof.

[0035]

[0051] FIG. 2A shows a schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure. For example, encoding process 200A may be performed by an encoder. As shown in FIG. 2A, the encoder may encode a video sequence 202 into a video bitstream 228 according to process 200A. Similar to video sequence 100 in FIG. 1, video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in a temporal order. Similar to structure 110 in FIG. 1, each original picture in video sequence 202 may be divided into basic processing units, basic processing sub-units, or regions for processing by the encoder. In some embodiments, the encoder may perform process 200A at the level of basic processing units for each original picture in video sequence 202. For example, the encoder may perform process 200A in an iterative manner, in which case the encoder may encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for regions of each original picture of video sequence 202 (eg, regions 114-118).

[0036]

[0052] 2A , an encoder may provide a fundamental processing unit (referred to as an “original BPU”) of an original picture of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a prediction BPU 208. The encoder may subtract the prediction BPU 208 from the original BPU to generate a residual BPU 210. The encoder may provide the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may provide the prediction data 206 and the quantized transform coefficients 216 to a binary coding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a “forward path.” During process 200A, after quantization stage 214, the encoder may provide quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate reconstructed residual BPU 222. The encoder may add reconstructed residual BPU 222 to prediction BPU 208 to generate prediction reference 224, which is used in prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.

[0037]

[0053] The encoder may perform process 200A iteratively to encode each original BPU of the original picture (in the forward path) and generate (in the reconstruction path) a prediction reference 224 for encoding the next original BPU of the original picture. After encoding all original BPUs of the original picture, the encoder may proceed to encode the next picture in video sequence 202.

[0038]

[0054] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to receiving, inputting, acquiring, obtaining, getting, reading, accessing, or any act in any manner to input data.

[0039]

[0055] In the prediction step 204, in the current iteration, the encoder may receive the original BPU and a prediction reference 224, perform a prediction operation, and generate predicted data 206 and a predicted BPU 208. The prediction reference 224 may be generated from a reconstruction path of a previous iteration of the process 200A. The purpose of the prediction step 204 is to reduce information redundancy by extracting predicted data 206, which can be used to reconstruct the original BPU from the prediction data 206 and the prediction reference 224 as a predicted BPU 208.

[0040]

[0056] Ideally, predicted BPU 208 would be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, predicted BPU 208 generally differs slightly from the original BPU. To record such differences, after generating predicted BPU 208, the encoder can subtract it from the original BPU to generate residual BPU 210. For example, the encoder can subtract pixel values ​​(e.g., grayscale or RGB values) of predicted BPU 208 from corresponding pixel values ​​of the original BPU. Each pixel of residual BPU 210 can have a residual value that is the result of such a subtraction between the corresponding pixel of the original BPU and predicted BPU 208. Compared to the original BPU, predicted data 206 and residual BPU 210 can have fewer bits, which can be used to reconstruct the original BPU without significant quality degradation. Therefore, the original BPU is compressed.

[0041]

[0057] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy in the residual BPU 210 by decomposing it into a set of two-dimensional "basis patterns," each associated with a "transform coefficient." The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a change frequency (e.g., a frequency of luminance change) component of the residual BPU 210. No basis pattern can be reconstructed from any combination (e.g., a linear combination) of any other basis patterns. In other words, such a decomposition can decompose the changes in the residual BPU 210 into the frequency domain. Such a decomposition is similar to a discrete Fourier transform of a function, where the basis patterns are similar to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are similar to the coefficients associated with the basis functions.

[0042]

[0058] Different transform algorithms can use different basis patterns. For example, various transform algorithms can be used in transform stage 212, such as a discrete cosine transform, a discrete sine transform, or the like. The transform in transform stage 212 is invertible. That is, the encoder can recover residual BPU 210 by inverting the transform (referred to as an "inverse transform"). For example, to recover pixels of residual BPU 210, the inverse transform can multiply the values ​​of corresponding pixels in the basis pattern by their associated coefficients and add the products to generate a weighted sum. For video coding standards, both the encoder and decoder can use the same transform algorithm (and therefore the same basis pattern). Therefore, the encoder can record only the transform coefficients, and the decoder can reconstruct residual BPU 210 from the transform coefficients without receiving the basis pattern from the encoder. Compared to residual BPU 210, the transform coefficients can have fewer bits, which can be used to reconstruct residual BPU 210 without significant quality degradation. Therefore, the residual BPU 210 is further compressed.

[0043]

[0059] The encoder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns can represent different change frequencies (e.g., luminance change frequencies). Because the human eye is generally better at perceiving low-frequency changes, the encoder can ignore high-frequency change information without significant quality degradation during decoding. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (referred to as a "quantization parameter") and rounding the quotient to its nearest integer. After such an operation, some transform coefficients of high-frequency basis patterns can be converted to zero, and transform coefficients of low-frequency basis patterns can be converted to smaller integers. The encoder can ignore zero-valued quantized transform coefficients 216, thereby further compressing the transform coefficients. The quantization process can also be inverted, in which case the quantized transform coefficients 216 can be reconstructed into transform coefficients in the inverse operation of quantization (referred to as "dequantization").

[0044]

[0060] Because the encoder ignores the remainder of such a division in a rounding operation, quantization stage 214 may be lossy. Typically, quantization stage 214 may contribute the greatest information loss in process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To achieve different levels of information loss, the encoder may use different values ​​of the quantization parameter or any other parameter of the quantization process.

[0045]

[0061] In binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in binary encoding stage 226, such as, for example, a prediction mode used in prediction stage 204, parameters of the prediction operation, the type of transform in transform stage 212, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), or the like. The encoder may generate a video bitstream 228 using the output data of binary encoding stage 226. In some embodiments, the video bitstream 228 may be further packetized for network transmission.

[0046]

[0062] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to a prediction BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.

[0047]

[0063] It should be noted that other variations of process 200A may be used to encode video sequence 202. In some embodiments, the stages of process 200A may be performed in a different order by the encoder. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be split into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit one or more stages in FIG. 2A.

[0048]

[0064] 2B shows a schematic diagram of another exemplary encoding process 200B according to an embodiment of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B additionally includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.

[0049]

[0065] Generally, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can use pixels from one or more already-encoded neighboring BPUs within the same picture to predict a target BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can use regions from one or more already-encoded pictures to predict a target BPU. That is, the prediction reference 224 in temporal prediction can include an encoded picture. Temporal prediction can reduce the inherent temporal redundancy of a picture.

[0050]

[0066] Referring to process 200B, within the forward path, the encoder performs prediction operations in a spatial prediction step 2042 and a temporal prediction step 2044. For example, in the spatial prediction step 2042, the encoder may perform intra prediction. For an original BPU of a picture being encoded, the prediction reference 224 may include one or more neighboring BPUs within the same picture that are coded (in the forward path) and reconstructed (in the reconstruction path). The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, or the like. In some embodiments, the encoder may perform extrapolation at the pixel level, such as by extrapolating, for each pixel of the predicted BPU 208, the value of the corresponding pixel. The neighboring BPUs used for extrapolation can be located relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., below-left, below-right, above-left, or above-right of the original BPU), or any direction defined in the video coding standard used. For intra prediction, the prediction data 206 may include, for example, the locations (e.g., coordinates) of the neighboring BPUs used, the sizes of the neighboring BPUs used, parameters of the extrapolation, the orientations of the neighboring BPUs used relative to the original BPU, or the like.

[0051]

[0067] As another example, in the temporal prediction stage 2044, the encoder may perform inter-prediction. For an original BPU of a target picture, the prediction reference 224 may include one or more pictures (referred to as "reference pictures") that have been coded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be coded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs of the same picture have been generated, the encoder may generate the reconstructed picture as a reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within a range (referred to as a "search window") of the reference picture. The location of the search window in the reference picture may be determined based on the location of the original BPU of the target picture. For example, the search window may be centered in the reference picture at a location having the same coordinates as the original BPU in the target picture and may extend outward over a predetermined distance. When the encoder identifies a region similar to the original BPU within the search window (e.g., by using a pixel-recursive algorithm, a block-matching algorithm, or the like), the encoder can determine such a region as a matching region. The matching region can have dimensions different from those of the original BPU (e.g., smaller than, equal to, larger than, or a different shape than the original BPU). Because the reference picture and the target picture are temporally separated in a timeline (e.g., as shown in FIG. 1), the matching region can be considered to "move" toward the location of the original BPU over time. The encoder can record the direction and distance of such movement as a "motion vector." When multiple reference pictures are used (e.g., as picture 106 in FIG. 1), the encoder can search for the matching region for each reference picture and determine its associated motion vector. In some embodiments, the encoder can assign weights to the pixel values ​​of the matching region in each matching reference picture.

[0052]

[0068] Motion estimation can be used to identify various types of motion, such as, for example, translation, rotation, zooming, or the like. For inter prediction, prediction data 206 can include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, the number of reference pictures, weights associated with the reference pictures, or the like.

[0053]

[0069] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., a motion vector) and the prediction reference 224. For example, the encoder may shift a matching region of a reference picture according to a motion vector, from which the encoder can predict the original BPU of the target picture. When multiple reference pictures are used (e.g., as picture 106 in FIG. 1), the encoder may shift the matching region of the reference picture according to each motion vector and average the pixel values ​​of the matching region. In some embodiments, if the encoder weights the pixel values ​​of the matching region of each matching reference picture, the encoder may add a weighted sum of the pixel values ​​to the shifted matching region.

[0054]

[0070] In some embodiments, inter-prediction can be unidirectional or bidirectional. Unidirectional inter-prediction can use one or more reference pictures in the same temporal direction relative to the target picture. For example, picture 104 in FIG. 1 is a unidirectional inter-predicted picture in which a reference picture (i.e., picture 102) precedes picture 104. Bidirectional inter-prediction can use one or more reference pictures in both temporal directions relative to the target picture. For example, picture 106 in FIG. 1 is a bidirectional inter-predicted picture in which reference pictures (i.e., pictures 104 and 108) are in both temporal directions relative to picture 104.

[0055]

[0071] Still referring to the forward path of process 200B, after spatial prediction step 2042 and temporal prediction step 2044, in mode decision step 230, the encoder may select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of process 200B. For example, the encoder may perform a rate-distortion optimization technique. In this technique, the encoder may select a prediction mode to minimize the value of a cost function that depends on the bitrate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode. Depending on the selected prediction mode, the encoder may generate a corresponding predicted BPU 208 and predicted data 206.

[0056]

[0072] Within the reconstruction path of process 200B, if an intra-prediction mode is selected within the forward path, after generating the prediction reference 224 (e.g., the target BPU coded and reconstructed in the target picture), the encoder can directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the target picture). If an inter-prediction mode is selected within the forward path, after generating the prediction reference 224 (e.g., the target picture coded and reconstructed for all BPUs), the encoder can provide the prediction reference 224 to the loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortions (e.g., blocking artifacts) introduced by the inter prediction. The encoder can apply various loop filter techniques within the loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, or the like. The loop-filtered reference picture may be stored in a buffer 234 (or "decoded picture buffer") for later use (e.g., to be used as an inter-prediction reference picture for a future picture in the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) along with the quantized transform coefficients 216, the prediction data 206, and other information in the binary encoding stage 226.

[0057]

[0073] FIG. 3A shows a schematic diagram of an exemplary decoding process 300A according to an embodiment of the present disclosure. Process 300A may be a decompression process corresponding to compression process 200A in FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder may follow process 300A to decode video bitstream 228 into video stream 304. Video stream 304 may be similar to video sequence 202. However, due to information loss in the compression and decompression processes (e.g., quantization stage 214 in FIGS. 2A-2B), video stream 304 is generally not identical to video sequence 202. Similar to processes 200A and 200B in FIGS. 2A-2B, a decoder may perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, in which case the decoder may decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for regions (e.g., regions 114-118) of each picture encoded in video bitstream 228.

[0058]

[0074] In FIG. 3A , a decoder may provide a portion of a video bitstream 228 associated with a basic processing unit (referred to as a “coding BPU”) of a coded picture to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may provide the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may provide the prediction data 206 to a prediction stage 204 to generate a prediction BPU 208. The decoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder can provide the prediction reference 224 to the prediction stage 204 to perform the prediction operation in the next iteration of the process 300A.

[0059]

[0075] The decoder may perform process 300A iteratively to decode each coded BPU of a coded picture and generate a prediction reference 224 for encoding the next coded BPU of the coded picture. After decoding all coded BPUs of a coded picture, the decoder may output the picture to video stream 304 for display and proceed to decode the next coded picture in video bitstream 228.

[0060]

[0076] In binary decoding step 302, the decoder may perform the inverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to prediction data 206 and quantized transform coefficients 216, the decoder may decode other information in binary decoding step 302, such as, for example, a prediction mode, parameters of the prediction operation, type of transform, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), or the like. In some embodiments, if video bitstream 228 is transmitted in packets over a network, the decoder may depacketize video bitstream 228 before providing it to binary decoding step 302.

[0061]

[0077] 3B shows a schematic diagram of another exemplary decoding process 300B according to an embodiment of the present disclosure. Process 300B may be modified from process 300A. For example, process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B additionally divides prediction stage 204 into spatial prediction stage 2042 and temporal prediction stage 2044, and additionally includes loop filter stage 232 and buffer 234.

[0062]

[0078] In process 300B, prediction data 206 decoded by the decoder from binary decoding stage 302 for a coding basic processing unit (referred to as a "current BPU" or a "target BPU") of a coding picture being decoded (referred to as a "current picture" or a "target picture") may include various types of data, depending on what prediction mode was used by the encoder to encode the target BPU. For example, if intra prediction was used by the encoder to encode the target BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, or the like. Parameters of the intra prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as references, the size of the neighboring BPUs, parameters of extrapolation, the orientation of the neighboring BPUs relative to the original BPU, or the like. As another example, if inter-prediction is used by the encoder to encode the target BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) that indicates inter-prediction, parameters of the inter-prediction operation, or the like. The parameters of the inter-prediction operation may include, for example, the number of reference pictures associated with the target BPU, weights respectively associated with the reference pictures, locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors respectively associated with the matching regions, or the like.

[0063]

[0079] Based on the prediction mode indicator, the decoder may determine whether to perform spatial prediction (e.g., intra prediction) in spatial prediction step 2042 or temporal prediction (e.g., inter prediction) in temporal prediction step 2044. Details of performing such spatial or temporal prediction are described in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate a predicted BPU 208. The decoder may add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as described in FIG. 3A.

[0064]

[0080] In process 300B, the decoder may provide the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 to perform the prediction operation in the next iteration of process 300B. For example, if the target BPU is decoded using intra prediction in spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded target BPU), the decoder may provide the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the target picture). If the target BPU is decoded using inter prediction in temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture from which all BPUs are decoded), the encoder may provide the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B . The loop-filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., to be used as an inter-prediction reference picture for a future coded picture of the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-prediction was used to encode the target BPU, the prediction data may further include parameters of a loop filter (e.g., loop filter strength).

[0065]

[0081] FIG. 4 is a block diagram of an exemplary device 400 for encoding or decoding video in accordance with an embodiment of the present disclosure. As shown in FIG. 4, device 400 may include a processor 402. When processor 402 executes instructions described herein, device 400 can become a specialized machine for video encoding or decoding. Processor 402 can be any type of circuitry capable of manipulating or processing information. For example, processor 402 can include any number and combination of a central processing unit (or "CPU"), a graphics processing unit (or "GPU"), a neural processing unit ("NPU"), a microcontroller unit ("MCU"), an optical processor, a programmable logic controller, a microcontroller, a microprocessor, a digital signal processor, an intellectual property (IP) core, a programmable logic array (PLA), a programmable array logic (PAL), a generic array logic (GAL), a complex programmable logic device (CPLD), a field programmable gate array (FPGA), a system-on-chip (SoC), an application-specific integrated circuit (ASIC), or the like. In some embodiments, processor 402 may be a set of processors grouped as a single logical entity. For example, as shown in Figure 4, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0066]

[0082] Device 400 may also include memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, or the like). For example, as shown in FIG. 4, the stored data may include program instructions (e.g., program instructions for performing steps in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 may access (e.g., via bus 410) the program instructions and data for processing, execute the program instructions, and perform operations or manipulations on the data for processing. Memory 404 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, memory 404 may include any number or combination of random access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, or the like. Memory 404 may also be a group of memories (not shown in FIG. 4) grouped as a single logical entity.

[0067]

[0083] Bus 410 may be a communication device that transfers data between components internal to device 400, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Express port), or the like.

[0068]

[0084] For ease of explanation and without ambiguity, the processor 402 and other data processing circuitry will be collectively referred to in this disclosure as "data processing circuitry." The data processing circuitry may be implemented entirely in hardware or as a combination of software, hardware, or firmware. In addition, the data processing circuitry may be a single, independent module or may be fully or partially combined with any other component of the device 400.

[0069]

[0085] Device 400 may further include a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, or the like). In some embodiments, network interface 406 may include any number and combination of a network interface controller (NIC), a radio frequency (RF) module, a transponder, a transceiver, a modem, a router, a gateway, a wired network adapter, a wireless network adapter, a Bluetooth® adapter, an infrared adapter, a near field communication ("NFC") adapter, a cellular network chip, or the like.

[0070]

[0086] In some embodiments, apparatus 400 may optionally further include a peripheral interface 408 for providing connection to one or more peripheral devices. As shown in Figure 4, the peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touchscreen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., an input interface communicatively coupled to a camera or a video archive), or the like.

[0071]

[0087] It should be noted that a video codec (e.g., a codec performing process 200A, 200B, 300A, or 300B) may be implemented as any combination of software or hardware modules within device 400. For example, some or all of the stages of process 200A, 200B, 300A, or 300B may be implemented as one or more software modules of device 400, such as program instructions that may be loaded into memory 404. As another example, some or all of the stages of process 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of device 400, such as specialized data processing circuitry (e.g., FPGA, ASIC, NPU, or the like).

[0072]

[0088] In the quantization and inverse quantization functional blocks (e.g., quantization 214 and inverse quantization 218 in FIG. 2A or 2B, inverse quantization 218 in FIG. 3A or 3B), a quantization parameter (QP) is used to determine the amount of quantization (and inverse quantization) applied to the prediction residual. The initial QP value used for coding a picture or slice can be signaled at a high level, for example, using the init_qp_minus26 syntax element in the picture parameter set (PPS) and the slice_qp_delta syntax element in the slice header. Furthermore, the QP value can be adapted at a local level per CU using delta QP values ​​sent at the granularity of the quantization group.

[0073]

[0089] In some real-time applications (e.g., videoconferencing or remote control systems), system latency can be a significant issue that significantly affects the user experience and reliability of the system. For example, ITU-T G.114 specifies an acceptable latency limit of 150 milliseconds for two-way audio-video communication. In another example, virtual reality applications typically require ultra-low latency of less than 20 milliseconds to prevent motion sickness caused by timing mismatches between head movements and the visual effects produced by those movements.

[0074]

[0090] In a real-time video application system, total latency includes the time from when a frame is captured to when it is displayed. That is, total latency is the sum of the encoding time at the encoder, the transmission time in the transmission channel, the decoding time at the decoder, and the output delay at the decoder. Generally, transmission time contributes most to total latency. The transmission time of a coded picture is typically equal to the capacity of the coded picture buffer (CPB) divided by the bit rate of the video sequence.

[0075]

[0091] In this disclosure, "random access" refers to the ability to start the decoding process at any random access point in a video sequence or stream and recover the correct decoded picture in the content. To support random access and prevent error propagation, intra-coded random access point (IRAP) pictures can be inserted periodically in a video sequence. However, for high coding efficiency, the size of a coded I-picture (e.g., an IRAP picture) is typically larger than the size of a P-picture or a B-picture. The larger size of an IRAP picture may result in a higher than average transmission delay. Therefore, periodically inserting IRAP pictures may not meet the requirements of low-delay video applications.

[0076]

[0092] According to an embodiment of the present disclosure, for low-latency encoding, a gradual decoding refresh (GDR) technique, also referred to as a gradual intra refresh (PIR) technique, can be used to reduce the latency caused by the insertion of IRAP pictures while enabling random access within a video sequence. GDR can gradually refresh pictures by distributing intra-coded regions within non-intra-coded regions (e.g., B pictures or P pictures). By doing so, the sizes of hybrid-coded pictures can be similar to each other, thereby reducing or minimizing the size of the CPB (e.g., to a value equal to the bit rate of the video sequence divided by the picture rate), and shortening the encoding and decoding times within the total delay.

[0077]

[0093] By way of example, Figure 5 is a schematic diagram illustrating an exemplary operation of gradual decoding refresh (GDR) in accordance with some embodiments of this disclosure. Figure 5 illustrates a GDR period 502 that includes multiple pictures (e.g., pictures 504, 506, 508, 510, and 512) in a video sequence (e.g., video sequence 100 of Figure 1). The first picture in the GDR period 502 is referred to as a GDR picture 504, which may be a random access picture, and the last picture in the GDR period 502 is referred to as a recovery point picture 512. Each picture in the GDR period 502 includes an intra-coded region (represented by a vertical box labeled "INTRA" in each picture in Figure 5). Each intra-coded region may cover various portions of a complete picture. As shown in Figure 5, the intra-coded region may progressively span the entire picture within the GDR period 502. It should be noted that although Figure 5 shows the intra-coded regions as rectangular slices, these regions can be implemented as various shapes and sizes and are not limited by the examples provided in this disclosure.

[0078]

[0094] A picture other than the GDR picture 504 and the recovery point picture 512 within the GDR period 502 (e.g., any of the pictures 506-510), separated by an intra-coded region, can include two regions: a "clean region" containing pixels that have already been refreshed, and a "dirty region" containing pixels that were possibly corrupted by transmission errors in a previous picture and have not yet been refreshed (e.g., can be refreshed in a subsequent picture). The clean region of the current picture (e.g., picture 510) can include pixels that are reconstructed using as reference at least one of the clean regions or intra-coded regions of a previous picture (e.g., pictures 508, 506, and the GDR picture 504). The clean region of the current picture (e.g., picture 510) can include pixels that are reconstructed using as reference at least one of the dirty regions, clean regions, or intra-coded regions of a previous picture (e.g., pictures 508, 506, and the GDR picture 504).

[0079]

[0095] The principle of GDR techniques is to ensure that pixels in clean regions are reconstructed without using any information from any dirty regions (e.g., dirty regions of the current picture or any previous picture). As an example, in Figure 5, GDR picture 504 includes dirty region 514. Picture 506 includes clean region 516, which may be reconstructed using intra-coded regions of GDR picture 504 as reference, and dirty region 518, which may be reconstructed using any portion of GDR picture 504 (e.g., at least one of the intra-coded regions or dirty region 514) as reference. Picture 508 includes a clean region 520 that can be reconstructed using at least one of the intra-coded regions of pictures 504-506 (e.g., the intra-coded regions of pictures 504 or 506) or clean regions (e.g., clean region 516) as reference, and a dirty region 522 that can be reconstructed using at least one of a portion of GDR picture 504 (e.g., the intra-coded region or dirty region 514) or a portion of picture 506 (e.g., the clean region 516, the intra-coded region or dirty region 518) as reference. Picture 510 includes a clean region 524 that can be reconstructed using as reference at least one of the intra-coded regions of pictures 504-508 (e.g., the intra-coded regions of pictures 504, 506, or 508) or clean regions (e.g., clean region 516 or 520), and a dirty region 526 that can be reconstructed using as reference at least one of a portion of GDR picture 504 (e.g., the intra-coded region or dirty region 514), a portion of picture 506 (e.g., the clean region 516, the intra-coded region, or the dirty region 518), or a portion of picture 508 (e.g., the clean region 520, the intra-coded region, or the dirty region 522).The recovery point picture 512 includes a clean region 528 that can be reconstructed using as reference at least one of a portion of the GDR picture 504 (e.g., an intra-coded or dirty region 514), a portion of the picture 506 (e.g., a clean region 516, an intra-coded or dirty region 518), a portion of the picture 508 (e.g., a clean region 520, an intra-coded or dirty region 522), or a portion of the picture 510 (e.g., a clean region 524, an intra-coded or dirty region 526).

[0080]

[0096] As shown, all pixels of recovery point picture 512 have been refreshed. Decoding pictures in output order using GDR techniques after recovery point picture 512 may be equivalent to decoding the pictures using (as if there was an IRAP picture) before GDR picture 504, where the IRAP picture spans all intra-coded regions of pictures 504-512.

[0081]

[0097] As an example, Figure 6 shows Table 1 illustrating an example syntax structure of a sequence parameter set (SPS) that enables GDR, according to some embodiments of the present disclosure. As shown in Table 1, a sequence-level enablement flag for GDR, "gdr_enabled_flag," may be signaled within an SPS of a video sequence to indicate whether any GDR-capable pictures are present within the video sequence (e.g., any pictures within GDR period 502 of Figure 5). In some embodiments, in the VVC / H.266 standard, a "gdr_enabled_flag" that is true (e.g., equal to "1") may specify that a GDR-capable picture is present within a coding layer video sequence (CLVS) that references the SPS, and a "gdr_enabled_flag" that is false (e.g., equal to "0") may specify that no GDR-capable pictures are present within the CLVS.

[0082]

[0098] As an example, FIG. 7 shows Table 2 illustrating an example syntax structure of a picture header that enables GDR, according to some embodiments of the present disclosure. As shown in Table 2, for a picture of a video sequence, a GDR picture-level enablement flag "gdr_pic_flag" may be signaled in the picture header of the picture to indicate whether the picture is GDR-capable (e.g., any picture within GDR period 502 of FIG. 5). If the picture is GDR-capable, a parameter "recovert_poc_cnt" may be signaled to specify a recovery point picture (e.g., recovery point picture 512 of FIG. 5) in output order. In some embodiments, in the VVC / H.266 standard, a "gdr_pic_flag" that is true (e.g., equal to "1") may specify that the picture associated with the picture header is a GDR-capable picture, and a "gdr_pic_flag" that is false (e.g., equal to "0") may specify that the picture associated with the picture header is not a GDR-capable picture.

[0083]

[0099] As an example, in the VVC / H.266 standard, if the "gdr_enabled_flag" shown in Table 1 is true and the parameter "PicOrderCntVal" of the current picture (not shown in Figure 7) is greater than or equal to the sum of "PicOrderCntVal" and "recovery_poc_cnt" of the GDR-compatible picture (or multiple GDR-compatible pictures) related to the current picture, the current picture and subsequent pictures in output order can be decoded as if they were decoded by starting the decoding process from the IRAP picture preceding the GDR-compatible picture (or multiple GDR-compatible pictures).

[0084]

[0100] According to an embodiment of the present disclosure, a virtual boundary technique can be used to implement GDR (e.g., in the VVC / H.266 standard). In some applications (e.g., 360-degree video), the layout of a particular projection format may typically have multiple planes. When these projection formats include multiple planes, discontinuities may occur between two or more adjacent planes in a frame-packed picture, regardless of what kind of compact frame packing configuration is used. If an in-loop filtering operation is performed across this discontinuity, surface seam artifacts may become visible in the rendered reconstructed image.

[0085]

[0101] To mitigate surface seam artifacts, in-loop filtering operations (e.g., deblocking filtering, sample adaptive offset filtering, or adaptive loop filtering) can be disabled across discontinuities in frame-packed pictures, which can be referred to as a virtual boundary technique (e.g., a concept adopted by VVC Draft 7). For example, the encoder can set the discontinuity boundary as a virtual boundary and disable any loop filtering operations across the virtual boundary. By doing so, loop filtering across the discontinuity can be disabled.

[0086]

[0102] For GDR, loop filtering operations should not be applied across the boundary between a clean region (e.g., clean region 520 of picture 508 in FIG. 5) and a dirty region (e.g., dirty region 522 of picture 508). The encoder can set the boundary between the clean region and the dirty region as a virtual boundary and disable loop filtering operations across the virtual boundary. In this way, the virtual boundary can be used as a way to implement GDR.

[0087]

[0103] In some embodiments, the VVC / H.266 standard (e.g., in VVC Draft 7) may signal virtual boundaries in an SPS or picture header. By way of example, Figure 8 shows Table 3 illustrating an example syntax structure of an SPS that enables virtual boundaries, according to some embodiments of the present disclosure. Figure 9 shows Table 4 illustrating an example syntax structure of a picture header that enables virtual boundaries, according to some embodiments of the present disclosure.

[0088]

[0104] As shown in Table 3, a sequence-level virtual boundary presence flag "sps_virtual_boundaries_present_flag" may be signaled in the SPS. For example, "sps_virtual_boundaries_present_flag" that is true (e.g., equal to "1") may specify that virtual boundary information is signaled in the SPS, and "sps_virtual_boundaries_present_flag" that is false (e.g., equal to "0") may specify that virtual boundary information is not signaled in the SPS. If one or more virtual boundaries are signaled in the SPS, in-loop filtering operations across the virtual boundaries in pictures that reference the SPS may be disabled.

[0089]

[0105] As shown in Table 3, if the flag "sps_virtual_boundaries_present_flag" is true, the number of virtual boundaries (represented by the parameters "sps_num_ver_virtual_boundaries" and "sps_num_hor_virtual_boundaries" in Table 3) and their positions (represented by the arrays "sps_virtual_boundaries_pos_x" and "sps_virtual_boundaries_pos_y" in Table 3) may be signaled in the SPS. The parameters "sps_num_ver_virtual_boundaries" and "sps_num_hor_virtual_boundaries" may specify the lengths of the arrays "sps_virtual_boundaries_pos_x" and "sps_virtual_boundaries_pos_y", respectively, in the SPS. In some embodiments, if "sps_num_ver_virtual_boundaries" (or "sps_num_hor_virtual_boundaries") is not present in the SPS, its value may be inferred to be 0. The arrays "sps_virtual_boundaries_pos_x" and "sps_virtual_boundaries_pos_y" may specify the position of the ith vertical or horizontal virtual boundary, respectively, in units of luma samples divided by 8. For example, the value of "sps_virtual_boundaries_pos_x[i]" may be in the closed interval from 1 to Ceil(pic_width_in_luma_samples÷8)-1, where "Ceil" represents the ceiling function and "pic_width_in_luma_samples" is a parameter representing the width of the picture in units of luma samples. The value of "sps_virtual_boundaries_pos_y[i]" may be in the closed interval from 1 to Ceil(pic_height_in_luma_samples÷8)-1, where "pic_height_in_luma_samples" is a parameter representing the height of the picture in units of luma samples.

[0090]

[0106] In some embodiments, if the flag "sps_virtual_boundaries_present_flag" is false (e.g., equal to "0"), a picture-level virtual boundary presence flag "ph_virtual_boundaries_present_flag" may be signaled in the picture header, as shown in Table 4. For example, a true (e.g., equal to "1") "ph_virtual_boundaries_present_flag" may specify that virtual boundary information is signaled in the picture header, and a false (e.g., equal to "0") "ph_virtual_boundaries_present_flag" may specify that virtual boundary information is not signaled in the picture header. If one or more virtual boundaries are signaled in a picture header, in-loop filtering operations across the virtual boundaries in the picture that includes the picture header may be disabled. In some embodiments, if "ph_virtual_boundaries_present_flag" is not present in the picture header, its value may be inferred to represent "false."

[0091]

[0107] As shown in Table 4, if the flag "ph_virtual_boundaries_present_flag" is true (e.g., equal to "1"), the number of virtual boundaries (represented by the parameters "ph_num_ver_virtual_boundaries" and "ph_num_hor_virtual_boundaries" in Table 4) and their positions (represented by the arrays "ph_virtual_boundaries_pos_x" and "ph_virtual_boundaries_pos_y" in Table 4) may be signaled in the picture header. The parameters "ph_num_ver_virtual_boundaries" and "ph_num_hor_virtual_boundaries" may specify the lengths of the arrays "ph_virtual_boundaries_pos_x" and "ph_virtual_boundaries_pos_y", respectively, in the picture header. In some embodiments, if 'ph_virtual_boundaries_pos_x' (or 'ph_virtual_boundaries_pos_y') is not present in the picture header, its value can be inferred to be 0. The arrays 'ph_virtual_boundaries_pos_x' and 'ph_virtual_boundaries_pos_y' may specify the position of the i-th vertical or horizontal virtual boundary, respectively, in units of luma samples divided by 8. For example, the value of 'ph_virtual_boundaries_pos_x[i]' can be in the closed interval from 1 to Ceil(pic_width_in_luma_samples÷8)-1, the value of 'ph_virtual_boundaries_pos_y[i]' can be in the closed interval from 1 to Ceil(pic_height_in_luma_samples÷8)-1, and 'pic_height_in_luma_samples' is a parameter representing the height of the picture in units of luma samples.

[0092]

[0108] In some embodiments, in the VVC / H.266 standard (eg, in VVC Draft 7), a variable "VirtualBoundariesDisabledFlag" may be defined as equation (1): VirtualBoundariesDisabledFlag=sps_virtual_boundaries_present_flag || ph_virtual_boundaries_present_flag expression (1)

[0093]

[0109] However, the implementation of GDR by using virtual boundaries may cause two problems in existing technical solutions. For example, as described above, in existing technical solutions, the picture-level flag "gdr_pic_flag" is always signaled in the picture header regardless of the value of the sequence-level flag "gdr_enabled_flag." That is, even if GDR is disabled for a sequence, the picture header of each picture in the sequence may still indicate whether the picture is GDR-enabled. Therefore, a conflict may occur between the SPS level and the picture level. For example, a conflict may occur when "gdr_enabled_flag" is false and "gdr_pic_flag" is true.

[0094]

[0110] As another example, when a virtual boundary is used as the boundary between a clean region and a dirty region to implement GDR, existing technical solutions do not apply loop filtering operations across the virtual boundary. However, as a requirement of GDR, decoding a pixel in the clean region cannot refer to a pixel in the dirty region, but decoding a pixel in the dirty region can refer to a pixel in the clean region. In this case, completely disabling loop filtering across the virtual boundary may impose an overly strict restriction, and such a restriction may degrade encoding or decoding performance.

[0095]

[0111] To solve the above problems, the present disclosure provides a method, an apparatus, and a system for processing pictures. According to some embodiments of the present disclosure, to eliminate the potential conflict between the GDR indication flag at the SPS level and the picture level, the syntax structure of the picture header can be modified so that the picture-level GDR indication flag can be signaled only when GDR is enabled at the sequence level.

[0096]

[0112] As an example, Figure 10 shows Table 5, which illustrates an example syntax structure of a modified picture header according to some embodiments of the present disclosure. As shown in Table 5, element 1002 (enclosed in a solid box) indicates a syntax modification compared to Table 2 of Figure 7. For example, a "gdr_pic_flag" that is true (e.g., equal to "1") may specify that the picture associated with the picture header is a GDR-compliant picture, and a "gdr_pic_flag" that is false (e.g., equal to "0") may specify that the picture associated with the picture header is not a GDR-compliant picture. In some embodiments, if "gdr_pic_flag" is not present in the picture header, its value can be inferred to represent "false."

[0097]

[0113] According to some embodiments of the present disclosure, to eliminate potential inconsistencies between the GDR indication flags at the SPS level and the picture level, the syntactic structure of the picture header can be kept unchanged (e.g., as shown in Table 2 of FIG. 7 ), and a bitstream conformance requirement (e.g., bitstream conformance defined in the VVC / H.266 standard) can be implemented so that the picture-level GDR indication flag is not true (e.g., invalid or false) if the sequence-level GDR indication flag is not true (e.g., invalid or false). As used herein, a bitstream conformance requirement can refer to an operation that can ensure that a bitstream subset associated with an operation point complies with a video coding standard (e.g., the VVC / H.266 standard). An "operation point" can refer to a first bitstream created from a second bitstream by a sub-bitstream extraction process, in which, for a network abstraction layer (NAL) unit of the second bitstream, if it does not belong to a target set determined by a list of target temporal identifiers and target layer identifiers, it can be removed. For example, the bitstream conformance requirement may be implemented such that if "gdr_enabled_flag" is false, then "gdr_pic_flag" is also set to false.

[0098]

[0114] According to some embodiments of the present disclosure, to increase flexibility in disabling loop filtering operations across a virtual boundary, the syntax structure of the SPS and picture header may be modified to allow partial disabling of loop filtering operations across the virtual boundary. By doing so, pixels on one side of the virtual boundary may not be filtered, but pixels on the other side of the virtual boundary may be filtered. For example, if a virtual boundary vertically divides a picture into a left side and a right side, the encoder or decoder may partially disable the loop filter on the right side so that pixels there are not filtered (e.g., information from pixels on the left side is not used for loop filtering of pixels on the right side), and enable the loop filter on the left side so that pixels there are filtered (e.g., information from pixels on at least one of the left side and the right side may be used for loop filtering).

[0099]

[0115] As an example, FIG. 11 shows Table 6 illustrating an example syntax structure of a modified SPS that enables virtual boundaries, according to some embodiments of the present disclosure. FIG. 12 shows Table 7 illustrating an example syntax structure of a modified picture header that enables virtual boundaries, according to some embodiments of the present disclosure. As shown in the accompanying drawings of the present disclosure, a dashed box indicates that the enclosed content or element is deleted or removed (shown as a strikethrough). As shown in FIGS. 11-12, the sequence-level GDR indication flag "sps_virtual_boundaries_present_flag" and the picture-level GDR indication flag "ph_virtual_boundaries_present_flag" are replaced by GDR control parameters "sps_virtual_boundaries_loopfilter_disable" and "ph_virtual_boundaries_loopfilter_disable," respectively, which are extended to support partially disabled loop filtering operations at the sequence level and picture level, respectively.

[0100]

[0116] If the GDR direction (e.g., left-to-right, right-to-left, top-to-bottom, bottom-to-top, or any combination thereof) is fixed across the entire sequence, the GDR control parameter (e.g., "sps_virtual_boundaries_loopfilter_disable") can be set in the SPS, which can save bits. If the GDR direction needs to be changed within a sequence, the GDR control parameter (e.g., "ph_virtual_boundaries_loopfilter_disable") can be set in the picture header, which can provide more flexibility in low-level control.

[0101]

[0117] According to some embodiments of the present disclosure, the GDR control parameters "sps_virtual_boundaries_loopfilter_disable" and "ph_virtual_boundaries_loopfilter_disable" can be configured to be multiple values ​​(e.g., beyond a "true" or "false" representation) to represent various implementation schemes.

[0102]

[0118] For example, "sps_virtual_boundaries_loopfilter_disable" equal to "0" may specify that virtual boundary information is not signaled in the SPS. "sps_virtual_boundaries_loopfilter_disable" equal to "1" may specify that virtual boundary information is signaled in the SPS and that in-loop filtering operations are disabled across the virtual boundaries. "sps_virtual_boundaries_loopfilter_disable" equal to "2" may specify that virtual boundary information is signaled in the SPS and one of the following: (1) in-loop filtering operations to the left of the virtual boundaries are disabled, (2) the in-loop filtering operations on the left do not use information for any pixels to the right of the virtual boundaries, (3) in-loop filtering operations above the virtual boundaries are disabled, or (4) the in-loop filtering operations above do not use information for any pixels below the virtual boundaries. "sps_virtual_boundaries_loopfilter_disable" being "3" may specify that virtual boundary information is signaled within the SPS and that (1) in-loop filtering operations to the right of the virtual boundary are disabled, (2) in-loop filtering operations on the right do not use information from any pixels to the left of the virtual boundary, (3) in-loop filtering operations below the virtual boundary are disabled, or (4) in-loop filtering operations below do not use information from any pixels above.

[0103]

[0119] Similarly, in another example, "ph_virtual_boundaries_loopfilter_disable" equal to "0" may specify that virtual boundary information is not signaled in the picture header. "ph_virtual_boundaries_loopfilter_disable" equal to "1" may specify that virtual boundary information is signaled in the picture header and that in-loop filtering operations are disabled across the virtual boundary. "ph_virtual_boundaries_loopfilter_disable" equal to "2" may specify that virtual boundary information is signaled in the picture header and one of: (1) in-loop filtering operations on the left side of the virtual boundary are disabled; (2) the in-loop filtering operations on the left side do not use information of any pixels on the right side of the virtual boundary; (3) in-loop filtering operations on the above side of the virtual boundary are disabled; or (4) the in-loop filtering operations on the above side do not use information of any pixels below the virtual boundary. "ph_virtual_boundaries_loopfilter_disable" being "3" may specify that virtual boundary information is signaled in the picture header and one of the following: (1) in-loop filtering operations to the right of the virtual boundary are disabled, (2) the in-loop filtering operations on the right do not use information of any pixels to the left of the virtual boundary, (3) the in-loop filtering operations below the virtual boundary are disabled, or (4) the in-loop filtering operations below do not use information of any pixels above. In some embodiments, if "ph_virtual_boundaries_loopfilter_disable" is not present in the picture header, its value can be inferred to be 0.

[0104]

[0120] In some embodiments, the variable "VirtualBoundariesLoopfilterDisabled" may be defined as equation (2). VirtualBoundariesLoopfilterDisabled=sps_virtual_boundaries_loopfilter_disable ?sps_virtual_boundaries_loopfilter_disable: ph_virtual_boundaries_loopfilter_disable expression (2)

[0105]

[0121] According to some embodiments of the present disclosure, the loop filter may be an adaptive loop filter (ALF). When the ALF is partially disabled on a first side (e.g., left side, right side, upper side, or lower side), pixels on the first side may be padded in filtering, and pixels on the second side (e.g., right side, left side, lower side, or upper side) are not used in filtering.

[0106]

[0122] In some embodiments, the ALF boundary positions may be derived as described below: In the ALF boundary position derivation process, the variables "clipLeftPos", "clipRightPos", "clipTopPos" and "clipBottomPos" may be set as "-128".

[0107]

[0123] Compared to the VVC / H.266 standard (e.g., in VVC Draft 7), the variable "clipTopPos" can be determined as follows: If (y-(CtbSizeY-4)) is greater than or equal to 0, then the variable "clipTopPos" can be set as (yCtb+CtbSizeY-4). If (y-(CtbSizeY-4)) is negative, "VirtualBoundariesLoopfilterDisabled" is equal to 1, and (yCtb+y-VirtualBoundariesPosY[n]) is in the half-open interval [1,3) for any n=0,1,...,(VirtualBoundariesNumHor-1), then "clipTopPos" can be set as "VirtualBoundariesPosY[n]" (i.e., clipTopPos=VirtualBoundariesPosY[n]). If (y-(CtbSizeY-4)) is negative, "VirtualBoundariesLoopfilterDisabled" is equal to 3, and (yCtb+y-VirtualBoundariesPosY[n]) is in the half-open interval [1,3) for any n=0,1,...,(VirtualBoundariesNumHor-1), then "clipTopPos" can be set as "VirtualBoundariesPosY[n]" (i.e., clipTopPos=VirtualBoundariesPosY[n]).

[0108]

[0124] 'clipTopPos' can be set as 'yCtb' if (y-(CtbSizeY-4)) is negative, y is less than 3, and one or more of the following conditions are true: (1) the top boundary of the current coding tree block is the top boundary of a tile and 'loop_filter_across_tiles_enabled_flag' is equal to 0, (2) the top boundary of the current coding tree block is the top boundary of a slice and 'loop_filter_across_slices_enabled_flag' is equal to 0, or (3) the top boundary of the current coding tree block is the top boundary of a subpicture and 'loop_filter_across_subpic_enabled_flag[SubPicIdx]' is equal to 0.

[0109]

[0125] Compared to the VVC / H.266 standard (e.g., in VVC Draft 7), the variable "clipBottomPos" can be determined as follows: if "VirtualBoundariesLoopfilterDisabled" is equal to 1, and "VirtualBoundariesPosY[n]" is not equal to (pic_height_in_luma_samples-1) or 0, and (VirtualBoundariesPosY[n]-yCtb-y) is in the open interval (0,5) for any n=0,...,(VirtualBoundariesNumHor-1), then "clipBottomPos" can be set as "VirtualBoundariesPosY[n]" (i.e., clipBottomPos=VirtualBoundariesPosY[n]).

[0110]

[0126] If "VirtualBoundariesLoopfilterDisabled" is equal to 2, "VirtualBoundariesPosY[n]" is not equal to (pic_height_in_luma_samples-1) or 0, and (VirtualBoundariesPosY[n]-yCtb-y) is in the open interval (0,5) for any n=0,...,(VirtualBoundariesNumHor-1), then "clipBottomPos" can be set as "VirtualBoundariesPosY[n]" (i.e., clipBottomPos=VirtualBoundariesPosY[n]).

[0111]

[0127] Otherwise, if (CtbSizeY-4-y) is in the open interval (0,5), then "clipBottomPos" may be set as "yCtb+CtbSizeY-4". Otherwise, if (CtbSizeY-y) is less than 5 and one or more of the following conditions are true, then "clipBottomPos" may be set as "(yCtb+CtbSizeY)", namely: (1) the bottom boundary of the current coding tree block is the bottom boundary of a tile and "loop_filter_across_tiles_enabled_flag" is equal to 0, (2) the bottom boundary of the current coding tree block is the bottom boundary of a slice and "loop_filter_across_slices_enabled_flag" is equal to 0, or (3) the bottom boundary of the current coding tree block is the bottom boundary of a subpicture and "loop_filter_across_subpic_enabled_flag[SubPicIdx]" is equal to 0.

[0112]

[0128] Compared to the VVC / H.266 standard (e.g., in VVC Draft 7), the variable "clipLeftPos" can be determined as follows: if "VirtualBoundariesLoopfilterDisabled" is equal to 1 and (xCtb+x-VirtualBoundariesPosX[n]) is in the half-open interval [1,3) for any n=0,...,(VirtualBoundariesNumVer-1), then "clipLeftPos" can be set as "VirtualBoundariesPosX[n]" (i.e., clipLeftPos=VirtualBoundariesPosX[n]). If "VirtualBoundariesLoopfilterDisabled" is equal to 3 and "xCtb+x-VirtualBoundariesPosX[n]" is in the half-open interval [1,3) for any n=0,...,(VirtualBoundariesNumVer-1), then "clipLeftPos" can be set as "VirtualBoundariesPosX[n]" (i.e., clipLeftPos=VirtualBoundariesPosX[n]).

[0113]

[0129] Otherwise, if x is less than 3 and one or more of the following conditions are true, then 'clipLeftPos' can be set as 'xCtb', namely: (1) the left boundary of the current coding tree block is the left boundary of a tile and 'loop_filter_across_tiles_enabled_flag' is equal to 0; (2) the left boundary of the current coding tree block is the left boundary of a slice and 'loop_filter_across_slices_enabled_flag' is equal to 0; (3) the left boundary of the current coding tree block is the left boundary of a subpicture and 'loop_filter_across_subpic_enabled_flag[SubPicIdx]' is equal to 0.

[0114]

[0130] Compared to the VVC / H.266 standard (e.g., in VVC Draft 7), the variable "clipRightPos" can be determined as follows: if "VirtualBoundariesLoopfilterDisabled" is equal to 1 and "(VirtualBoundariesPosX[n]-xCtb-x)" is in the open interval (0,5) for any n=0,...,(VirtualBoundariesNumVer-1), then "clipRightPos" can be set as "VirtualBoundariesPosX[n]" (i.e., clipRightPos=VirtualBoundariesPosX[n]). If "VirtualBoundariesLoopfilterDisabled" is equal to 2 and (VirtualBoundariesPosX[n]-xCtb-x) is in the open space (0,5) for any n=0,...,(VirtualBoundariesNumVer-1), then "clipRightPos" can be set as "VirtualBoundariesPosX[n]" (i.e., clipRightPos=VirtualBoundariesPosX[n]).

[0115]

[0131] Otherwise, if "(CtbSizeY-x)" is less than 5 and one or more of the following conditions are true, then "clipRightPos" can be set as (xCtb+CtbSizeY), namely: (1) the right boundary of the current coding tree block is the right boundary of a tile and "loop_filter_across_tiless_enabled_flag" is equal to 0, (2) the right boundary of the current coding tree block is the right boundary of a slice and "loop_filter_across_slices_enabled_flag" is equal to 0, or (3) the right boundary of the current coding tree block is the right boundary of a subpicture and "loop_filter_across_subpic_enabled_flag[SubPicIdx]" is equal to 0.

[0116]

[0132] Compared to the VVC / H.266 standard (e.g., in VVC Draft 7), the variables 'clipTopLeftFlag' and 'clipBotRightFlag' can be determined as follows: if the coding tree block covering luma position (xCtb, yCtb) and the coding tree block covering luma position (xCtb-CtbSizeY, yCtb-CtbSizeY) belong to different slices and 'loop_filter_across_slices_enabled_flag' is equal to 0, 'clipTopLeftFlag' can be set to 1. If the coding tree block covering luma position (xCtb, yCtb) and the coding tree block covering luma position (xCtb+CtbSizeY, yCtb+CtbSizeY) belong to different slices and 'loop_filter_across_slices_enabled_flag' is equal to 0, 'clipBotRightFlag' can be set to 1.

[0117]

[0133] According to some embodiments of the present disclosure, the loop filter may include a sample adaptive offset (SAO) operation. When the SAO is partially disabled on a first side (e.g., left, right, top, or bottom) of a virtual boundary, if the SAO for a pixel on the first side requires a pixel on the second side (e.g., right, left, bottom, or top), the application of the SAO to the pixel on the first side may be skipped. Doing so makes the pixel on the second side unusable.

[0118]

[0134] According to some embodiments of the present disclosure, the CTB correction process involves calculating all sample positions (xS i ,yS j ) and (xY i ,yY j ) the following operations can be applied:

[0119]

[0135] If one or more of the following conditions are true, the variable "saoPicture[xS i ][yS j ]" can be unmodified, and the following conditions are: (1) the variable "SaoTypeIdx[cIdx][rx][ry]" is equal to 0, (2) "VirtualBoundariesLoopfilterDisabled" is equal to 1, and "xS j " equals ((VirtualBoundariesPosX[n] / scaleWidth)-1) for any n=0,...,(VirtualBoundariesNumVer-1), "SaoTypeIdx[cIdx][rx][ry]" equals 2, and the variable "SaoEoClass[cIdx][rx][ry]" is not equal to 1, (3) "VirtualBoundariesLoopfilterDisabled" equals 1, and "xS j " equals (VirtualBoundariesPosX[n] / scaleWidth) for any n=0,...,(VirtualBoundariesNumVer-1), "SaoTypeIdx[cIdx][rx][ry]" equals 2 and "SaoEoClass[cIdx][rx][ry]" does not equal 1, (4) "VirtualBoundariesLoopfilterDisabled" equals 1 and "yS j " equals ((VirtualBoundariesPosY[n] / scaleHeight)-1) for any n=0,...,(VirtualBoundariesNumHor-1), "SaoTypeIdx[cIdx][rx][ry]" equals 2, "SaoEoClass[cIdx][rx][ry]" does not equal 0, (5) "VirtualBoundariesLoopfilterDisabled" equals 1, and "yS j" equals (VirtualBoundariesPosY[n] / scaleHeight) for any n=0,...,(VirtualBoundariesNumHor-1), "SaoTypeIdx[cIdx][rx][ry]" equals 2, and "SaoEoClass[cIdx][rx][ry]" does not equal 0, (6) "VirtualBoundariesLoopfilterDisabled" equals 2, and "xS j " equals ((VirtualBoundariesPosX[n] / scaleWidth)-1) for any n=0,...,(VirtualBoundariesNumVer-1), "SaoTypeIdx[cIdx][rx][ry]" equals 2, and "SaoEoClass[cIdx][rx][ry]" does not equal 1, (7) "VirtualBoundariesLoopfilterDisabled" equals 3, and "xS j " equals (VirtualBoundariesPosX[n] / scaleWidth) for any n=0,...,(VirtualBoundariesNumVer-1), "SaoTypeIdx[cIdx][rx][ry]" equals 2, and "SaoEoClass[cIdx][rx][ry]" does not equal 1, (8) "VirtualBoundariesLoopfilterDisabled" equals 2, and "yS j " is equal to ((VirtualBoundariesPosY[n] / scaleHeight)-1) for any n=0,...,(VirtualBoundariesNumHor-1), "SaoTypeIdx[cIdx][rx][ry]" is equal to 2 and "SaoEoClass[cIdx][rx][ry]" is not equal to 0, or (9) "VirtualBoundariesLoopfilterDisabled" is equal to 3 and "yS j" is equal to (VirtualBoundariesPosY[n] / scaleHeight) for any n=0,...,(VirtualBoundariesNumHor-1), "SaoTypeIdx[cIdx][rx][ry]" is equal to 2, and "SaoEoClass[cIdx][rx][ry]" is equal to 0.

[0120]

[0136] In accordance with some embodiments of the present disclosure, the loop filter may include a deblocking filter. In some embodiments, when the deblocking filter is partially disabled on a first side (e.g., left side, right side, top side, or bottom side) of a virtual boundary, pixels on the first side may be skipped from being processed by the deblocking filter, and a second side (e.g., right side, left side, bottom side, or top side) of the virtual boundary may be processed by the deblocking filter. In some embodiments, when "VirtualBoundariesLoopfilterDisabled" is not false (e.g., has a value of 0), the deblocking filter may be completely disabled, and pixels on both sides of the virtual boundary may be skipped from being processed by the deblocking filter.

[0121]

[0137] 13-15 illustrate flowcharts of exemplary methods 1300-1500, according to some embodiments of the present disclosure. Methods 1300-1500 may be performed by at least one processor (e.g., processor 402 of FIG. 4 ) associated with a video encoder (e.g., an encoder described in connection with FIGS. 2A-2B ) or a video decoder (e.g., a decoder described in connection with FIGS. 3A-3B ). In some embodiments, methods 1300-1500 may be implemented as a computer program product (e.g., embodied by a computer-readable medium) that includes computer-executable instructions (e.g., program code) for execution by a computer (e.g., device 400 of FIG. 4 ). In some embodiments, methods 1300-1500 may be implemented as a hardware product (e.g., memory 404 of FIG. 4 ) that stores computer-executable instructions (e.g., program instructions in memory 404 of FIG. 4 ), which may be a separate or integral part of the computer.

[0122]

[0138] 13 shows a flowchart of an exemplary process 1300 for video processing according to some embodiments of the present disclosure. For example, the process 1300 may be performed by an encoder.

[0123]

[0139] In step 1302, in response to a processor (e.g., processor 402 of FIG. 4) receiving a video sequence (e.g., video sequence 202 of FIGS. 2A-2B), the processor may encode first flag data (e.g., “gdr_enabled_flag” shown and described with respect to FIGS. 10-12) in a parameter set (e.g., SPS) associated with the video sequence. The first flag data may indicate whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence.

[0124]

[0140] In step 1304, if the first flag data indicates that GDR is disabled for the video sequence (e.g., "gdr_enabled_flag" is false), the processor may encode a picture header associated with the picture in the video sequence to indicate that the picture is a non-GDR picture. As used herein, a GDR picture may refer to a picture that includes both clean and dirty regions. By way of example, a clean region may be any of clean regions 516, 520, 524, or 528 illustrated and described in FIG. 5, and a dirty region may be any of dirty regions 514, 518, 522, or 526 illustrated and described in FIG. 5. A non-GDR picture in this disclosure may refer to a picture that does not include a clean region or does not include a dirty region.

[0125]

[0141] In some embodiments, to encode the picture header, the processor may disable encoding of second flag data in the picture header (e.g., “gdr_pic_flag” shown and described with respect to FIGS. 10-12). The second flag data may indicate whether the picture is a GDR picture.

[0126]

[0142] In some embodiments, to encode the picture header, the processor may encode second flag data in the picture header (e.g., "gdr_pic_flag" shown and described with respect to Figures 10-12), where the second flag data indicates that the picture is a non-GDR picture (e.g., "gdr_pic_flag" is false).

[0127]

[0143] In step 1306, the processor may encode the non-GDR picture.

[0128]

[0144] According to some embodiments of the present disclosure, if the first flag data indicates that GDR is enabled for the video sequence (e.g., "gdr_enabled_flag" is true), the processor may enable encoding of second flag data in the picture header (e.g., "gdr_pic_flag" shown and described with respect to FIGS. 10-12), the second flag data indicating whether the picture is a GDR picture (e.g., "gdr_pic_flag" is true). The processor may then encode the picture.

[0129]

[0145] According to some embodiments of the present disclosure, when the first flag data indicates that GDR is enabled for the video sequence (e.g., “gdr_enabled_flag” is true) and the second flag data indicates that the picture is a GDR picture (e.g., “gdr_pic_flag” is true), the processor may use a virtual boundary to divide the picture into a first region (e.g., a clean region such as any of clean regions 516, 520, 524, or 528 illustrated and described in FIG. 5 ) and a second region (e.g., a dirty region such as any of dirty regions 514, 518, 522, or 526 illustrated and described in FIG. 5 ). For example, the first region may be a first side (e.g., left side, right side, top side, or bottom side) of the virtual boundary, and the second region may be a second side (e.g., right side, left side, bottom side, or top side) of the virtual boundary. The processor can then disable the loop filter (e.g., loop filter 232 shown and described with respect to FIGS. 2B and 3B) if the filtering of the first pixel uses information of the second pixel in the second region, or apply the loop filter for the first pixel using only information of the pixel in the first region. The processor can then apply the loop filter for the pixel in the second region using pixels in at least one of the first region or the second region.

[0130]

[0146] 14 shows a flowchart of another exemplary process 1400 for video processing according to some embodiments of the present disclosure. For example, the process 1400 may be performed by a decoder.

[0131]

[0147] In step 1402, in response to a processor (e.g., processor 402 of FIG. 4) receiving a video bitstream (e.g., video bitstream 228 of FIGS. 3A-3B), the processor may decode first flag data (e.g., “gdr_enabled_flag” shown and described with respect to FIGS. 10-12) in a parameter set (e.g., SPS) associated with a sequence of the video bitstream (e.g., a video sequence to be decoded). The first flag data may indicate whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence.

[0132]

[0148] In step 1404, if the first flag data indicates that GDR is disabled for the sequence (e.g., "gdr_enabled_flag" is false), the processor may decode a picture header associated with a picture in the sequence, the picture header indicating that the picture is a non-GDR picture.

[0133]

[0149] In some embodiments, to decode the picture header, the processor may disable decoding of second flag data in the picture header (e.g., “gdr_pic_flag” shown and described with respect to FIGS. 10-12) and determine that the picture is a non-GDR picture, and the second flag data may indicate whether the picture is a GDR picture. In some embodiments, to decode the picture header, the processor may decode second flag data in the picture header, and the second flag data may indicate that the picture is a non-GDR picture (e.g., “gdr_pic_flag” is false).

[0134]

[0150] In step 1406, the processor may decode the non-GDR picture.

[0135]

[0151] According to some embodiments of the present disclosure, if the first flag data indicates that GDR is enabled for the video sequence (e.g., "gdr_enabled_flag" is true), the processor may decode the second flag data in the picture header (e.g., "gdr_pic_flag" shown and described with respect to FIGS. 10-12) and determine whether the picture is a GDR picture (e.g., "gdr_pic_flag" is true) based on the second flag data. The processor may then decode the picture.

[0136]

[0152] According to some embodiments of the present disclosure, when the first flag data indicates that GDR is enabled for the video sequence and the second flag data indicates that the picture is a GDR picture, the processor may use a virtual boundary to divide the picture into a first region (e.g., a clean region such as any of clean regions 516, 520, 524, or 528 illustrated and described in FIG. 5 ) and a second region (e.g., a dirty region such as any of dirty regions 514, 518, 522, or 526 illustrated and described in FIG. 5 ). For example, the first region may be a first side (e.g., left side, right side, upper side, or lower side) of the virtual boundary, and the second region may be a second side (e.g., right side, left side, lower side, or upper side) of the virtual boundary. The processor can then disable the loop filter in the first region (e.g., loop filter 232 shown and described with respect to FIGS. 2B and 3B) or apply the loop filter for the first pixel using only information from the pixels in the first region. The processor can then apply the loop filter for the pixel in the second region using pixels in at least one of the first region or the second region.

[0137]

[0153] 15 shows a flowchart of yet another exemplary process 1500 for video processing according to some embodiments of the present disclosure. For example, the process 1500 may be performed by an encoder or a decoder.

[0138]

[0154] In step 1502, in response to a processor (e.g., processor 402 of FIG. 4) receiving a picture of a video, the processor determines whether the picture is a gradual decoding refresh (GDR) based on flag data associated with the picture. For example, the flag data can be at the sequence level (e.g., stored in an SPS) or at the picture level (e.g., stored in a picture header).

[0139]

[0155] In step 1504, based on a determination that the picture is a GDR picture, the processor may use the virtual boundary to determine a first region of the picture (e.g., a clean region such as any of clean regions 516, 520, 524, or 528 illustrated and described in FIG. 5) and a second region (e.g., a dirty region such as any of dirty regions 514, 518, 522, or 526 illustrated and described in FIG. 5). For example, the first region may include a left region or a top region, and the second region may include a right region or a bottom region.

[0140]

[0156] In step 1506, the processor may disable a loop filter (e.g., loop filter 232 shown and described with respect to FIGS. 2B and 3B) for a first pixel in the first region if the filtering of the first pixel uses information of a second pixel in the second region, or may apply a loop filter for the first pixel using only information of pixels in the first region. In some embodiments, the processor may disable at least one of a sample adaptive offset, a deblocking filter, or an adaptive loop filter in the first region, or apply at least one of a sample adaptive offset, a deblocking filter, or an adaptive loop filter using pixels in the first region.

[0141]

[0157] In step 1508, pixels in at least one of the first region or the second region may be used to apply a loop filter to pixels in the second region.

[0142]

[0158] According to some embodiments of the present disclosure, a processor may encode or decode flag data in at least one of a sequence parameter set (SPS), a picture parameter set (PPS), or a picture header. For example, as illustrated and described with respect to Figures 10-12, the processor may encode the flag data as a GDR indication flag "gdr_enabled_flag," a GDR indication flag "gdr_pic_flag," or both.

[0143]

[0159] In some embodiments, a non-transitory computer-readable storage medium containing instructions is also provided, which can be executed by a device (such as the encoders and decoders of the present disclosure) to perform the methods described above. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROMs and EPROMs, FLASH-EPROMs or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge, and networked versions thereof. A device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.

[0144]

[0160] The embodiments may be further described using the following clauses: 1. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by at least one processor of a device to cause the device to perform a method, the method comprising: In response to receiving the video sequence, encoding first flag data in a parameter set associated with the video sequence, the first flag data indicating whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence; encoding a picture header associated with a picture in the video sequence to indicate that the picture is a non-GDR picture if the first flag data indicates that GDR is disabled for the video sequence; and Encoding non-GDR pictures 1. A non-transitory computer-readable medium comprising: 2. Encoding the picture header disabling encoding second flag data in the picture header; 10. The non-transitory computer-readable medium of claim 1, wherein the second flag data indicates whether the picture is a GDR picture. 3. Encoding the picture header encoding second flag data in the picture header; 10. The non-transitory computer-readable medium of claim 1, wherein the second flag data indicates that the picture is a non-GDR picture. 4. The set of instructions executable by at least one processor of the device comprises: Enabling encoding second flag data in a picture header when the first flag data indicates that GDR is enabled for the video sequence, the second flag data indicating whether the picture is a GDR picture; and Encoding a Picture 2. The non-transitory computer-readable medium of claim 1, further causing the device to: 5. The set of instructions executable by at least one processor of the device comprises: dividing the picture into a first region and a second region using a virtual border if the first flag data indicates that GDR is enabled for the video sequence and the second flag data indicates that the picture is a GDR picture; Disabling the loop filter for the first pixel in the first region when filtering of the first pixel uses information of a second pixel in the second region, or applying the loop filter for the first pixel using only information of pixels in the first region; and applying a loop filter to pixels in the second region using pixels in at least one of the first region or the second region; The non-transitory computer-readable medium of any one of clauses 2 to 4 further causes the device to perform the following. 6. The non-transitory computer-readable medium of any one of clauses 2 to 5, wherein the second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. 7. The non-transitory computer-readable medium of any one of clauses 1 to 6, wherein the second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. 8. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by at least one processor of a device to cause the device to perform a method, the method comprising: In response to receiving the video bitstream, decoding first flag data in a parameter set associated with a sequence of the video bitstream, the first flag data indicating whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence; decoding a picture header associated with a picture in the sequence if the first flag data indicates that GDR is disabled for the sequence, the picture header indicating that the picture is a non-GDR picture; and Decoding non-GDR pictures 1. A non-transitory computer-readable medium comprising: 9. Decoding the picture header Disabling decoding second flag data in the picture header and determining that the picture is a non-GDR picture. 9. The non-transitory computer-readable medium of claim 8, comprising: 10. Decoding the picture header decoding second flag data in the picture header; 9. The non-transitory computer-readable medium of claim 8, wherein the second flag data indicates that the picture is a non-GDR picture. 11. A set of instructions executable by at least one processor of the device comprises: If the first flag data indicates that GDR is enabled for the video sequence, decoding second flag data in the picture header and determining whether the picture is a GDR picture based on the second flag data; and Decoding a Picture 9. The non-transitory computer-readable medium of claim 8, further causing the device to: 12. The set of instructions executable by at least one processor of the device comprises: dividing the picture into a first region and a second region using a virtual border if the first flag data indicates that GDR is enabled for the video sequence and the second flag data indicates that the picture is a GDR picture; Disabling the loop filter for the first pixel in the first region when filtering of the first pixel uses information of a second pixel in the second region, or applying the loop filter for the first pixel using only information of pixels in the first region; and applying a loop filter to pixels in the second region using pixels in at least one of the first region or the second region; The non-transitory computer-readable medium of any one of clauses 8 to 11, further causing the device to perform the following: 13. A non-transitory computer-readable medium described in any one of clauses 9 to 12, wherein the second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. 14. A non-transitory computer-readable medium described in any one of clauses 8 to 13, wherein the second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. 15. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by at least one processor of a device to cause the device to perform a method, the method comprising: In response to receiving a picture of the video, determining whether the picture is a gradual decoding refresh (GDR) picture based on flag data associated with the picture; determining a first region and a second region of the picture using the virtual boundary based on determining that the picture is a GDR picture; Disabling the loop filter for the first pixel in the first region when filtering of the first pixel uses information of a second pixel in the second region, or applying the loop filter for the first pixel using only information of pixels in the first region; and applying a loop filter to pixels in the second region using pixels in at least one of the first region or the second region; 1. A non-transitory computer-readable medium comprising: 16. Disabling a loop filter for a first pixel in a first region when filtering of the first pixel uses information of a second pixel in a second region, or applying a loop filter for the first pixel using only information of pixels in the first region, Disabling at least one of the sample adaptive offset, the deblocking filter, or the adaptive loop filter in the first region; or Applying at least one of a sample adaptive offset, a deblocking filter, or an adaptive loop filter using pixels in the first region. 16. The non-transitory computer-readable medium of clause 15, comprising: 17. A set of instructions executable by at least one processor of the device comprises: Encoding or decoding flag data in at least one of a sequence parameter set (SPS), a picture parameter set (PPS), or a picture header 17. The non-transitory computer-readable medium of clause 15 or 16, further causing the device to: 18. The non-transitory computer-readable medium of any one of clauses 15 to 17, wherein the first region includes a left region or an upper region, and the second region includes a right region or a lower region. 19. The non-transitory computer-readable medium of any one of clauses 15 to 18, wherein the flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. 20. An apparatus comprising: a memory configured to store a set of instructions; and one or more processors communicatively coupled to the memory, the one or more processors: In response to receiving the video sequence, encoding first flag data in a parameter set associated with the video sequence, the first flag data indicating whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence; encoding a picture header associated with a picture in the video sequence to indicate that the picture is a non-GDR picture if the first flag data indicates that GDR is disabled for the video sequence; and Encoding non-GDR pictures 10. An apparatus configured to execute a set of instructions to cause the apparatus to perform the following: 21. Encoding a picture header comprises: disabling encoding second flag data in the picture header; 21. The apparatus of claim 20, wherein the second flag data indicates whether the picture is a GDR picture. 22. Encoding a picture header comprises: encoding second flag data in the picture header; 21. The apparatus of clause 20, wherein the second flag data indicates that the picture is a non-GDR picture. 23. One or more processors may: Enabling encoding second flag data in a picture header when the first flag data indicates that GDR is enabled for the video sequence, the second flag data indicating whether the picture is a GDR picture; and Encoding a Picture 21. The apparatus of clause 20, further configured to execute a set of instructions to cause the apparatus to: 24. One or more processors may: dividing the picture into a first region and a second region using a virtual border if the first flag data indicates that GDR is enabled for the video sequence and the second flag data indicates that the picture is a GDR picture; Disabling the loop filter for the first pixel in the first region when filtering of the first pixel uses information of a second pixel in the second region, or applying the loop filter for the first pixel using only information of pixels in the first region; and applying a loop filter to pixels in the second region using pixels in at least one of the first region or the second region; 24. The apparatus of any one of clauses 21 to 23, further configured to execute a set of instructions to cause the apparatus to perform the following: 25. The device of any one of clauses 21 to 24, wherein the second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. 26. The device of any one of clauses 20 to 25, wherein the second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. 27. An apparatus comprising: a memory configured to store a set of instructions; and one or more processors communicatively coupled to the memory, the one or more processors: In response to receiving the video bitstream, decoding first flag data in a parameter set associated with a sequence of the video bitstream, the first flag data indicating whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence; decoding a picture header associated with a picture in the sequence if the first flag data indicates that GDR is disabled for the sequence, the picture header indicating that the picture is a non-GDR picture; and Decoding non-GDR pictures 10. An apparatus configured to execute a set of instructions to cause the apparatus to perform the following: 28. Decoding a picture header Disabling decoding the second flag data in the picture header and determining that the picture is a non-GDR picture. 28. The apparatus of clause 27, wherein the second flag data indicates whether the picture is a GDR picture. 29. Decoding a picture header decoding second flag data in the picture header; 28. The apparatus of clause 27, wherein the second flag data indicates that the picture is a non-GDR picture. 30. One or more processors may: If the first flag data indicates that GDR is enabled for the video sequence, decoding second flag data in the picture header and determining whether the picture is a GDR picture based on the second flag data; and Decoding a Picture 28. The apparatus of clause 27, further configured to execute a set of instructions to cause the apparatus to: 31. One or more processors may: dividing the picture into a first region and a second region using a virtual border if the first flag data indicates that GDR is enabled for the video sequence and the second flag data indicates that the picture is a GDR picture; Disabling the loop filter for the first pixel in the first region when filtering of the first pixel uses information of a second pixel in the second region, or applying the loop filter for the first pixel using only information of pixels in the first region; and applying a loop filter to pixels in the second region using pixels in at least one of the first region or the second region; 31. The apparatus of any one of clauses 28 to 30, further configured to execute a set of instructions to cause the apparatus to perform the 32. The device of any one of clauses 28 to 31, wherein the second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. 33. The device of any one of clauses 27 to 32, wherein the second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. 34. A device comprising: a memory configured to store a set of instructions; and one or more processors communicatively coupled to the memory, the one or more processors: In response to receiving a picture of the video, determining whether the picture is a gradual decoding refresh (GDR) picture based on flag data associated with the picture; determining a first region and a second region of the picture using the virtual boundary based on determining that the picture is a GDR picture; Disabling the loop filter for the first pixel in the first region when filtering of the first pixel uses information of a second pixel in the second region, or applying the loop filter for the first pixel using only information of pixels in the first region; and applying a loop filter to pixels in the second region using pixels in at least one of the first region or the second region; 10. An apparatus configured to execute a set of instructions to cause the apparatus to perform the following: 35. Disabling a loop filter for a first pixel in a first region when filtering of the first pixel uses information of a second pixel in a second region, or applying a loop filter for the first pixel using only information of pixels in the first region, Disabling at least one of a sample adaptive offset, a deblocking filter, or an adaptive loop filter in the first region or applying at least one of a sample adaptive offset, a deblocking filter, or an adaptive loop filter using pixels in the first region. Equipment as described in clause 34, including: 36. One or more processors may: Encoding or decoding flag data in at least one of a sequence parameter set (SPS), a picture parameter set (PPS), or a picture header 36. An apparatus according to clause 34 or 35, further configured to execute a set of instructions to cause the apparatus to: 37. A device according to any one of clauses 34 to 36, wherein the first region comprises a left region or an upper region, and the second region comprises a right region or a lower region. 38. The device of any one of clauses 34 to 37, wherein the flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. 39. In response to receiving the video sequence, encoding first flag data in a parameter set associated with the video sequence, the first flag data indicating whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence; encoding a picture header associated with a picture in the video sequence to indicate that the picture is a non-GDR picture if the first flag data indicates that GDR is disabled for the video sequence; and Encoding non-GDR pictures A method comprising: 40. Encoding a picture header disabling encoding second flag data in the picture header; 39. The method of claim 39, wherein the second flag data indicates whether the picture is a GDR picture. 41. Encoding a picture header comprises: encoding second flag data in the picture header; 39. The method of claim 39, wherein the second flag data indicates that the picture is a non-GDR picture. 42. Enabling encoding second flag data in a picture header when the first flag data indicates that GDR is enabled for the video sequence, the second flag data indicating whether the picture is a GDR picture; and Encoding a Picture 39. The method of claim 39, further comprising: 43. Dividing the picture into a first region and a second region using a virtual boundary when the first flag data indicates that GDR is enabled for the video sequence and the second flag data indicates that the picture is a GDR picture; Disabling the loop filter for the first pixel in the first region when filtering of the first pixel uses information of a second pixel in the second region, or applying the loop filter for the first pixel using only information of pixels in the first region; and applying a loop filter to pixels in the second region using pixels in at least one of the first region or the second region; 43. The method of any one of clauses 40 to 42, further comprising: 44. The method of any one of clauses 40 to 43, wherein the second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. 45. The method of any one of clauses 39 to 44, wherein the second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. 46. ​​In response to receiving the video bitstream, decoding first flag data in a parameter set associated with a sequence of the video bitstream, the first flag data indicating whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence; decoding a picture header associated with a picture in the sequence if the first flag data indicates that GDR is disabled for the sequence, the picture header indicating that the picture is a non-GDR picture; and Decoding non-GDR pictures A method comprising: 47. Decoding a picture header Disabling decoding the second flag data in the picture header and determining that the picture is a non-GDR picture. 47. The method of claim 46, wherein the second flag data indicates whether the picture is a GDR picture. 48. Decoding a picture header decoding second flag data in the picture header; 47. The method of claim 46, wherein the second flag data indicates that the picture is a non-GDR picture. 49. If the first flag data indicates that GDR is enabled for the video sequence, decoding second flag data in the picture header and determining whether the picture is a GDR picture based on the second flag data; and Decoding a Picture 47. The method of clause 46, further comprising: 50. Dividing the picture into a first region and a second region using a virtual boundary when the first flag data indicates that GDR is enabled for the video sequence and the second flag data indicates that the picture is a GDR picture; Disabling the loop filter for the first pixel in the first region when filtering of the first pixel uses information of a second pixel in the second region, or applying the loop filter for the first pixel using only information of pixels in the first region; and applying a loop filter to pixels in the second region using pixels in at least one of the first region or the second region; 50. The method of any one of clauses 47 to 49, further comprising: 51. The method of any one of clauses 47 to 50, wherein the second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. 52. The method of any one of clauses 46 to 51, wherein the second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. 53. In response to receiving a picture of the video, determining whether the picture is a gradual decoding refresh (GDR) picture based on flag data associated with the picture; determining a first region and a second region of the picture using the virtual boundary based on determining that the picture is a GDR picture; Disabling the loop filter for the first pixel in the first region when filtering of the first pixel uses information of a second pixel in the second region, or applying the loop filter for the first pixel using only information of pixels in the first region; and applying a loop filter to pixels in the second region using pixels in at least one of the first region or the second region; A method comprising: 54. Disabling a loop filter for a first pixel in a first region when filtering of the first pixel uses information of a second pixel in a second region, or applying a loop filter for the first pixel using only information of pixels in the first region, Disabling at least one of a sample adaptive offset, a deblocking filter, or an adaptive loop filter in the first region or applying at least one of a sample adaptive offset, a deblocking filter, or an adaptive loop filter using pixels in the first region. 54. The method of claim 53, including: 55. Encoding or decoding flag data in at least one of a sequence parameter set (SPS), a picture parameter set (PPS), or a picture header. 55. The method of clause 53 or 54, further comprising: 56. The method of any one of clauses 53 to 55, wherein the first region comprises a left region or an upper region, and the second region comprises a right region or a lower region. 57. The method of any one of clauses 53 to 56, wherein the flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture.

[0145]

[0161] It should be noted that relational terms herein, such as "first" and "second," are used merely to distinguish one entity or operation from another, and do not require or imply any actual relationship or order between those entities or operations. Furthermore, the words "comprise," "have," "contain," and "include," and other similar forms, are intended to be equivalent in meaning and open-ended in that the element or elements following any of these words are not meant to be an exclusive listing of such elements or elements, or to be limited to only the listed element or elements.

[0146]

[0162] As used herein, unless specifically stated otherwise, the term "or" includes all possible combinations unless impracticable. For example, if it is stated that a component can include A or B, then the component can include A or B, or A and B, unless specifically stated otherwise or impracticable. As a second example, if it is stated that a component can include A, B, or C, then the component can include A, B, or C, or A and B, or A and C, or B and C, or A and B and C, unless specifically stated otherwise or impracticable.

[0147]

[0163] It is understood that the above-described embodiments can be implemented by hardware or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-described computer-readable medium. The software, when executed by a processor, can perform the methods of the present disclosure. The computational units and other functional units described in the present disclosure can be implemented by hardware or software, or a combination of hardware and software. Those skilled in the art will also understand that multiple of the above-described modules / units can be combined into one module / unit, and that each of the above-described modules / units can be further divided into multiple sub-modules / sub-units.

[0148]

[0164] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications of the above-described embodiments may be made. Other embodiments may be apparent to those skilled in the art from consideration of the specification and practice of the disclosure disclosed herein. It is intended that the specification and examples be considered exemplary only, with the true scope and spirit of the disclosure being indicated by the appended claims. It is also intended that the sequences of steps depicted in the figures are for illustrative purposes only and are not intended to be limited to any particular sequence of steps. As such, one skilled in the art will recognize that these steps may be performed in different orders while implementing the same method.

[0149]

[0165] Illustrative embodiments have been disclosed in the drawings and herein. However, many variations and modifications to these embodiments may be made. Thus, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation.

Claims

1. 1. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by at least one processor of a device to cause the device to perform a method, the method comprising: responsive to receiving a video sequence, encoding first flag data in a parameter set associated with the video sequence, the first flag data indicating whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence; encoding a picture header associated with the picture in the video sequence to indicate that the picture is a non-GDR picture if the first flag data indicates that the GDR is disabled for the video sequence; and encoding the non-GDR picture; 1. A non-transitory computer-readable medium comprising:

2. encoding the picture header disabling encoding second flag data in the picture header; The non-transitory computer-readable medium of claim 1 , wherein the second flag data indicates whether the picture is a GDR picture.

3. 3. The non-transitory computer-readable medium of claim 2, wherein the second flag data has a value of "1" representing that the picture is a GDR picture or a value of "0" representing that the picture is a non-GDR picture.

4. encoding the picture header encoding second flag data in the picture header; The non-transitory computer-readable medium of claim 1 , wherein the second flag data indicates that the picture is the non-GDR picture.

5. The set of instructions executable by the at least one processor of the device comprises: enabling encoding second flag data in the picture header if the first flag data indicates that the GDR is enabled for the video sequence, the second flag data indicating whether the picture is a GDR picture; and encoding said picture The non-transitory computer-readable medium of claim 1 , further causing the device to perform:

6. The set of instructions executable by the at least one processor of the device comprises: dividing the picture into a first region and a second region using a virtual border if the first flag data indicates that GDR is enabled for the video sequence and the second flag data indicates that the picture is a GDR picture; Disabling a loop filter for a first pixel in the first region when filtering of the first pixel uses information of a second pixel in the second region, or applying the loop filter for the first pixel using only information of pixels in the first region; and applying the loop filter to pixels in the second region using pixels in at least one of the first region or the second region; The non-transitory computer-readable medium of claim 5 , further causing the device to perform:

7. 6. The non-transitory computer-readable medium of claim 5, wherein the second flag data has a value of "1" representing that the picture is a GDR picture or a value of "0" representing that the picture is a non-GDR picture.

8. 2. The non-transitory computer-readable medium of claim 1, wherein the first flag data has a value of "1" representing that the GDR is enabled for the video sequence or a value of "0" representing that the GDR is disabled for the video sequence.

9. 1. A non-transitory computer-readable medium storing a set of instructions, the set of instructions executable by at least one processor of a device to cause the device to perform a method, the method comprising: responsive to receiving a picture of a video, determining whether the picture is a gradual decoding refresh (GDR) picture based on flag data associated with the picture; determining a first region and a second region of the picture using a virtual boundary based on determining that the picture is the GDR picture; Disabling a loop filter for a first pixel in the first region when filtering of the first pixel uses information of a second pixel in the second region, or applying the loop filter for the first pixel using only information of pixels in the first region; and applying the loop filter to pixels in the second region using pixels in at least one of the first region or the second region; 1. A non-transitory computer-readable medium comprising:

10. Disabling the loop filter for the first pixel in the first region when filtering of the first pixel uses the information of the second pixel in the second region, or applying the loop filter for the first pixel using only the information of pixels in the first region, Disabling at least one of a sample adaptive offset, a deblocking filter, or an adaptive loop filter in the first region; or applying at least one of the sample adaptive offset, the deblocking filter, or the adaptive loop filter using the pixels in the first region; 10. The non-transitory computer-readable medium of claim 9, comprising:

11. The set of instructions executable by the at least one processor of the device comprises: encoding or decoding the flag data in at least one of a sequence parameter set (SPS), a picture parameter set (PPS), or a picture header; The non-transitory computer-readable medium of claim 9 , further causing the device to perform:

12. 10. The non-transitory computer-readable medium of claim 9, wherein the first region comprises a left region or a top region, and the second region comprises a right region or a bottom region.

13. 10. The non-transitory computer-readable medium of claim 9, wherein the flag data has a value of "1" representing that the picture is the GDR picture or a value of "0" representing that the picture is a non-GDR picture.

14. A device, a memory configured to store a set of instructions; one or more processors communicatively coupled to the memory, the one or more processors: responsive to receiving a video sequence, encoding first flag data in a parameter set associated with the video sequence, the first flag data indicating whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence; encoding a picture header associated with the picture in the video sequence to indicate that the picture is a non-GDR picture if the first flag data indicates that the GDR is disabled for the video sequence; and encoding the non-GDR picture; an apparatus configured to execute the set of instructions to cause the apparatus to perform

15. encoding the picture header disabling encoding second flag data in the picture header; The device of claim 14 , wherein the second flag data indicates whether the picture is a GDR picture.

16. 16. The device of claim 15, wherein the second flag data has a value of "1" representing that the picture is a GDR picture or a value of "0" representing that the picture is a non-GDR picture.

17. encoding the picture header encoding second flag data in the picture header; The device of claim 14 , wherein the second flag data indicates that the picture is the non-GDR picture.

18. The one or more processors: enabling encoding second flag data in the picture header if the first flag data indicates that the GDR is enabled for the video sequence, the second flag data indicating whether the picture is a GDR picture; and encoding said picture 15. The device of claim 14, further configured to execute the set of instructions to cause the device to:

19. The one or more processors: dividing the picture into a first region and a second region using a virtual border if the first flag data indicates that GDR is enabled for the video sequence and the second flag data indicates that the picture is a GDR picture; Disabling a loop filter for a first pixel in the first region when filtering of the first pixel uses information of a second pixel in the second region, or applying the loop filter for the first pixel using only information of pixels in the first region; and applying the loop filter to pixels in the second region using pixels in at least one of the first region or the second region; 15. The device of claim 14, further configured to execute the set of instructions to cause the device to:

20. The device of claim 14 , wherein the flag data has a value of “1” to indicate that the picture is the GDR picture or a value of “0” to indicate that the picture is a non-GDR picture.

Citation Information

Patent Citations

  • Incremental decoding refresh with temporal scalability support in video coding

    JP2016509404A

  • Method and system for performing progressive decoding refresh processing on a picture

    JP7701924B2