Method and system for performing progressive decoding refresh processing on a picture

GDR techniques and virtual boundary methods in the VVC/H.266 standard address latency and efficiency challenges in video encoding, enhancing encoding and decoding processes by optimizing loop filtering and random access in real-time video applications.

JP7701924B2Active Publication Date: 2025-07-02HFI INNOVATION INC

Patent Information

Application Number
JP2022532693
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-12-27
Filing Date
2020-12-01
Publication Date
2025-07-02
Estimated Expiration
2040-12-01

AI Technical Summary

Technical Problem

Existing video encoding standards face challenges in achieving low latency and efficient random access while maintaining high compression efficiency, particularly in applications requiring real-time video processing.

Method used

The implementation of Gradual Decoding Refresh (GDR) techniques and virtual boundary methods within the Versatile Video Coding (VVC/H.266) standard, allowing for flexible control of loop filtering operations across virtual boundaries to enhance encoding and decoding processes.

Benefits of technology

GDR techniques reduce latency and enable efficient random access by dispersing intra-coded regions, minimizing the size of hybrid-coded pictures, and optimizing loop filtering operations, thereby improving encoding and decoding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007701924000001
    Figure 0007701924000001
  • Figure 0007701924000002
    Figure 0007701924000002
  • Figure 0007701924000003
    Figure 0007701924000003
Patent Text Reader

Abstract

A method and apparatus for video processing includes, in response to receiving a video sequence, encoding first flag data in a parameter set associated with the video sequence, the first flag data indicating whether gradual decoding refresh (GDR) is enabled or disabled for the video sequence, encoding a picture header associated with a picture in the video sequence to indicate that the picture is a non-GDR picture if the first flag data indicates that GDR is disabled for the video sequence, and encoding the non-GDR picture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications

[0001] This disclosure claims priority to U.S. Provisional Patent Application No. 62 / 954,011, filed on December 27, 2019, the entire content of which is incorporated herein by reference.

[0002] Technical Field

[0002] This disclosure generally relates to video processing, and more particularly, to methods and systems for performing progressive decode refresh (GDR) processing on pictures.

Background Art

[0003] Background

[0003] Video is a set of static pictures (or "frames") that capture visual information. To reduce memory and transmission bandwidth, video can be compressed before storage or transmission and restored before display. The compression process is usually referred to as encoding, and the restoration process is usually referred to as decoding. Most commonly, there are various video encoding formats that use standardized video encoding techniques based on prediction, transformation, quantization, entropy encoding, and in - loop filtering. Video encoding standards such as the High - Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, which specify a particular video encoding format, have been developed by standardization organizations. As evolving video encoding techniques are successively adopted by video standards, the encoding efficiency of new video encoding standards becomes even higher.

Summary of the Invention

Means for Solving the Problems

[0004] Summary of the Disclosure

[0004] Embodiments of the present disclosure provide a method and apparatus for video processing. In one aspect, a non-transitory computer-readable medium is provided. The non-transitory computer-readable medium stores a set of instructions that are executable by at least one processor of the apparatus to cause the apparatus to perform the method. The method includes encoding first flag data in a parameter set associated with a video sequence in response to receiving the video sequence, where the first flag data indicates whether progressive decoding refresh (GDR) is enabled or disabled for the video sequence; encoding a picture header associated with a picture in the video sequence to indicate that the picture is a non-GDR picture when the first flag data indicates that GDR is disabled for the video sequence; and encoding non-GDR pictures.

[0005]

[0005] In another aspect, a non-transitory computer-readable medium is provided. The non-transitory computer-readable medium stores a set of instructions that are executable by at least one processor of the apparatus to cause the apparatus to perform the method. The method includes decoding first flag data in a parameter set associated with a sequence of a video bitstream in response to receiving the video bitstream, where the first flag data indicates whether progressive decoding refresh (GDR) is enabled or disabled for the video sequence; decoding a picture header associated with a picture in the sequence when the first flag data indicates that GDR is disabled for the sequence, where the picture header indicates that the picture is a non-GDR picture; and decoding non-GDR pictures.

[0006]

[0006] In yet another aspect, a non-transitory computer-readable medium is provided. The non-transitory computer-readable medium stores a set of instructions executable by at least one processor of the device to cause the device to perform a method. The method includes, in response to receiving a picture of a video, determining whether the picture is a Gradual Decoding Refresh (GDR) picture based on flag data associated with the picture, determining a first region and a second region of the picture using a virtual boundary based on the determination that the picture is a GDR picture, disabling a loop filter for a first pixel in the first region if filtering of the first pixel uses information of a second pixel in the second region or applying a loop filter for the first pixel using only information of pixels in the first region, and applying a loop filter for a pixel in the second region using pixels in at least one of the first region or the second region.

[0007]

[0007] In yet another aspect, a device is provided. The device includes a memory configured to store a set of instructions and one or more processors communicatively coupled to the memory. The one or more processors are configured to execute the set of instructions to cause the device to, in response to receiving a video sequence, encode first flag data in a set of parameters associated with the video sequence, the first flag data indicating whether Gradual Decoding Refresh (GDR) is enabled or disabled for the video sequence, encode a picture header associated with a picture in the video sequence to indicate that the picture is a non-GDR picture if the first flag data indicates that GDR is disabled for the video sequence, and encode the non-GDR picture.

[0008]

[0008] In yet another aspect, an apparatus is provided. The apparatus includes a memory configured to store a set of instructions and one or more processors communicatively coupled to the memory. The one or more processors are configured to execute the set of instructions to, in response to receiving a video bitstream, decode first flag data within a set of parameters associated with a sequence of the video bitstream, the first flag data indicating whether progressive decoding refresh (GDR) is enabled or disabled with respect to the video sequence, and, if the first flag data indicates that GDR is disabled with respect to the sequence, decode a picture header associated with a picture within the sequence, the picture header indicating that the picture is a non-GDR picture, and decode the non-GDR picture.

[0009]

[0009] In yet another aspect, an apparatus is provided. The apparatus includes a memory configured to store a set of instructions and one or more processors communicatively coupled to the memory. The one or more processors are configured to execute the set of instructions to, in response to receiving a picture of video, determine whether the picture is a progressive decoding refresh (GDR) picture based on flag data associated with the picture, determine, based on the determination that the picture is a GDR picture, a first region and a second region of the picture using a virtual boundary, disable a loop filter for a first pixel within the first region if filtering of the first pixel uses information of a second pixel within the second region, or apply a loop filter for the first pixel using only information of pixels within the first region, and apply a loop filter for a second pixel within the second region using pixels within at least one of the first region or the second region.

[0010]

[0010] In yet another aspect, a method is provided. The method includes, in response to receiving a video sequence, encoding first flag data within a parameter set associated with the video sequence, where the first flag data indicates whether progressive decoding refresh (GDR) is enabled or disabled for the video sequence, encoding a picture header associated with a picture within the video sequence to indicate that the picture is a non-GDR picture when the first flag data indicates that GDR is disabled for the video sequence, and encoding non-GDR pictures.

[0011]

[0011] In yet another aspect, a method is provided. The method includes, in response to receiving a video bitstream, decoding first flag data within a parameter set associated with a sequence of the video bitstream, where the first flag data indicates whether progressive decoding refresh (GDR) is enabled or disabled for the video sequence, decoding a picture header associated with a picture within the sequence, where the picture header indicates that the picture is a non-GDR picture when the first flag data indicates that GDR is disabled for the sequence, and decoding non-GDR pictures.

[0012]

[0012] In yet another aspect, a method is provided. The method includes, in response to receiving a picture of a video, determining whether the picture is a Gradual Decoding Refresh (GDR) picture based on flag data associated with the picture, determining, based on the determination that the picture is a GDR picture, a first region and a second region of the picture using a virtual boundary, disabling a loop filter for a first pixel in the first region if filtering of the first pixel uses information of a second pixel in the second region, or applying a loop filter for the first pixel using only information of pixels in the first region, and applying a loop filter for a pixel in the second region using pixels in at least one of the first region or the second region.

[0013] Brief Description of the Drawings

[0013] Embodiments and aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings. The various features shown in the drawings are not drawn to scale.

Brief Description of the Drawings

[0014]

Figure 1

[0014] It is a schematic diagram showing the structure of an exemplary video sequence according to some embodiments of the present disclosure.

Figure 2A

[0015] It shows a schematic diagram of an exemplary encoding process of a hybrid video encoding system according to an embodiment of the present disclosure.

Figure 2B

[0016] It shows a schematic diagram of another exemplary encoding process of a hybrid video encoding system according to an embodiment of the present disclosure.

Figure 3A

[0017] It shows a schematic diagram of an exemplary decoding process of a hybrid video encoding system according to an embodiment of the present disclosure.

Figure 3B

[0018] It shows a schematic diagram of another exemplary decoding process of a hybrid video encoding system according to an embodiment of the present disclosure.

Figure 4

[0019] A block diagram of an exemplary device for encoding or decoding video according to some embodiments of the present disclosure is shown.

Figure 5

[0020] A schematic diagram showing exemplary operations of Progressive Decoding Refresh (GDR) according to some embodiments of the present disclosure is shown.

Figure 6

[0021] Table 1 showing an exemplary syntax structure of a Sequence Parameter Set (SPS) enabling GDR according to some embodiments of the present disclosure is shown.

Figure 7

[0022] Table 2 showing an exemplary syntax structure of a Picture Header enabling GDR according to some embodiments of the present disclosure is shown.

Figure 8

[0023] Table 3 showing an exemplary syntax structure of an SPS enabling a virtual boundary according to some embodiments of the present disclosure is shown.

Figure 9

[0024] Table 4 showing an exemplary syntax structure of a Picture Header enabling a virtual boundary according to some embodiments of the present disclosure is shown.

Figure 10

[0025] Table 5 showing an exemplary syntax structure of a modified Picture Header according to some embodiments of the present disclosure is shown.

Figure 11

[0026] Table 6 showing an exemplary syntax structure of a modified SPS enabling a virtual boundary according to some embodiments of the present disclosure is shown.

Figure 12

[0027] Table 7 showing an exemplary syntax structure of a modified Picture Header enabling a virtual boundary according to some embodiments of the present disclosure is shown.

Figure 13

[0028] A flowchart of an exemplary process for video processing according to some embodiments of the present disclosure is shown.

Figure 14

[0029] A flowchart of another exemplary process for video processing according to some embodiments of the present disclosure is shown.

Figure 15

[0030] FIG. 1 shows a flowchart of yet another exemplary process for video processing according to some embodiments of the present disclosure. DETAILED DESCRIPTION

[0015] Detailed Description

[0031] Here, reference may be made in detail to exemplary embodiments illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which like numerals in different drawings represent the same or similar elements unless otherwise indicated. The implementations shown in the following description of the exemplary embodiments do not represent all implementations in accordance with the present invention. Rather, they are merely examples of devices and methods in accordance with aspects related to the present invention as recited in the appended claims. Specific aspects of the present disclosure are described in more detail below. Where terms and / or definitions incorporated by reference conflict, the terms and definitions provided herein shall prevail.

[0016]

[0032] The Joint Video Experts Team (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.

[0017]

[0033] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET is developing technologies beyond HEVC using the Joint Exploration Model (JEM) reference software. Since the coding technology is incorporated into JEM, JEM has achieved substantially higher coding performance than HEVC.

[0018]

[0034] The VVC standard has been recently developed and continues to incorporate more encoding techniques that bring better compression performance. VVC is based on the same hybrid video encoding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.

[0019]

[0035] Video is a set of static pictures (or "frames") arranged in time series to store visual information. A video capture device (e.g., a camera) can be used to capture and store those pictures in time series, and a video playback device (e.g., a TV, computer, smartphone, tablet computer, video player, or any end-user terminal with a display function) can be used to display such pictures in time series. Also, depending on the application, the video capture device can transmit the captured video in real time to a video playback device (e.g., a computer with a monitor) for supervision, holding a meeting, or live broadcast.

[0020]

[0036] To reduce the memory space and transmission bandwidth required by such applications, the video can be compressed before storage and transmission and restored before display. Compression and restoration can be implemented by software or special hardware executed by a processor (e.g., the processor of a general-purpose computer). The module for compression is generally referred to as an "encoder", and the module for restoration is generally referred to as a "decoder". The encoder and decoder can be collectively referred to as a "codec". The encoder and decoder can be implemented as any of various suitable hardware, software, or combinations thereof. For example, the hardware implementation of the encoder and decoder can include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. The software implementation of the encoder and decoder can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and restoration can be implemented by various algorithms or standards such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, or the like. Depending on the application, the codec can restore the video from a first encoding standard and recompress the restored video using a second encoding standard. In this case, the codec can be referred to as a "transcoder".

[0021]

[0037] The video encoding process can identify and retain the useful information that can be used to reconstruct the picture and ignore the information that is not important for reconstruction. If the ignored, unimportant information cannot be fully reconstructed, such an encoding process can be referred to as "irreversible". Otherwise, it can be referred to as "reversible". Most encoding processes are irreversible, which is a trade-off for reducing the required memory space and transmission bandwidth.

[0022]

[0038] Useful information of a picture being coded (referred to as the "current picture" or "target picture") includes changes with respect to a reference picture (e.g., a previously coded and reconstructed picture). Such changes can include changes in pixel position, brightness, or color, among which the change in position is the most important. The change in the position of a group of pixels representing an object can reflect the movement of the object between the reference picture and the target picture.

[0023]

[0039] A picture coded without referring to another picture (i.e., it is its own reference picture) is referred to as an "I picture". A picture coded using a previous picture as the reference picture is referred to as a "P picture". A picture coded using both a previous picture and a future picture as reference pictures (i.e., the reference is "bidirectional") is referred to as a "B picture".

[0024]

[0040] FIG. 1 shows the structure of an exemplary video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 can be a live video or a captured and archived video. The video 100 can be a real video, a computer-generated video (e.g., a computer game video), or a combination thereof (e.g., a real video with an augmented reality effect). The video sequence 100 can be input from a video capture device (e.g., a camera), a video archive containing previously captured videos (e.g., video files stored in a storage device), or a video supply interface (e.g., a video broadcast transceiver) for receiving videos from a video content provider.

[0025]

[0041] As shown in FIG. 1, video sequence 100 can include a series of pictures temporally arranged along a timeline, including pictures 102, 104, 106, and 108. Pictures 102 to 106 are consecutive, and there are additional pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I picture, and its reference picture is picture 102 itself. Picture 104 is a P picture, and its reference picture is picture 102, as indicated by the arrow. Picture 106 is a B picture, and its reference pictures are pictures 104 and 108, as indicated by the arrows. Depending on the embodiment, the reference picture of a picture (e.g., picture 104) may not be immediately before or after that picture. For example, the reference picture of picture 104 may be a picture before picture 102. The reference pictures of pictures 102 to 106 are merely examples, and it should be noted that the present disclosure does not limit the embodiments of the reference picture to the examples shown in FIG. 1.

[0026]

[0042] Typically, due to the computational complexity of the task, a video codec does not encode or decode an entire picture at once. Instead, the video codec can divide the picture into basic segments and encode or decode the picture segment by segment. Such a basic segment is referred to as a basic processing unit ("BPU") in the present disclosure. For example, structure 110 in FIG. 1 shows an exemplary structure of a picture (e.g., any of pictures 102-108) of video sequence 100. In structure 110, the picture is divided into 4×4 basic processing units, and their boundaries are shown as dashed lines. Depending on the embodiment, the basic processing unit may be referred to as a "macroblock" in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as a "coding tree unit" ("CTU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing unit can have a variable size or any arbitrary shape and size of pixels in the picture, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, etc. The size and shape of the basic processing unit can be selected based on a balance between coding efficiency and the level of detail to be maintained in the basic processing unit for the picture.

[0027]

[0043] The basic processing unit can be a logical unit that can include a group of different types of video data stored in a computer memory (e.g., within a video frame buffer). For example, the basic processing unit of a color picture can include a luma component (Y) representing achromatic luminance information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luma and chroma components can have the same size as the basic processing unit. The luma and chroma components may be referred to as "coding tree blocks" ("CTBs") in some video coding standards (e.g., H.265 / HEVC or H.266 / VVC). Any operation performed on the basic processing unit can be repeatedly performed on each of its luma and chroma components.

[0028]

[0044] Video encoding has multiple processing stages, examples of which are shown in FIGS. 2A-2B and FIGS. 3A-3B. At each stage, the size of the basic processing unit can still be too large for processing, and thus can be further divided into segments referred to herein as "basic processing subunits". Depending on the embodiment, the basic processing subunits can be referred to as "blocks" in some video encoding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as "coding units" ("CUs") in some other video encoding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunits can have a size that is the same as or smaller than the basic processing unit. Similar to the basic processing unit, the basic processing subunits are also logical units that can include groups of different types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in computer memory (e.g., within a video frame buffer). Any operation performed on a basic processing subunit can be repeatedly performed for each of its luma and chroma components. Note that such division can be carried out to further levels as required by the processing. Also note that different stages can divide the basic processing unit in different ways.

[0029]

[0045] For example, in the mode decision stage (an example of which is shown in FIG. 2B), the encoder can determine which prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for the basic processing unit, but the basic processing unit can be too large to make such a determination. The encoder can divide the basic processing unit into a plurality of basic processing subunits (e.g., CUs as in the case of H.265 / HEVC or H.266 / VVC) and determine the type of prediction for each individual basic processing subunit.

[0030]

[0046] As another example, in the prediction stage (an example of which is shown in FIGS. 2A - 2B), the coder can perform prediction operations at the level of a basic processing subunit (e.g., a CU). However, in some cases, the basic processing subunit may still be too large to process. The coder can further divide the basic processing subunit into smaller segments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), and prediction operations can be performed at that level.

[0031]

[0047] As another example, in the transform stage (an example of which is shown in FIGS. 2A - 2B), the coder can perform transform operations for a residual basic processing subunit (e.g., a CU). However, in some cases, the basic processing subunit may still be too large to process. The coder can further divide the basic processing subunit into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), and transform operations can be performed at that level. It should be noted that the division method of the same basic processing subunit can be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU can have different sizes and numbers.

[0032]

[0048] In the structure 110 of FIG. 1, the basic processing unit 112 is further divided into 3×3 basic processing subunits, and their boundaries are shown as dotted lines. Different basic processing units of the same picture can be divided into basic processing subunits in different ways.

[0033]

[0049] Depending on the implementation form, in order to bring parallel processing and error tolerance capabilities to video encoding and decoding, a picture can be divided into regions for processing. As a result, the encoding or decoding process does not need to depend on information from any other region of the picture with respect to the picture region. In other words, each region of the picture can be processed independently. By doing so, the codec can process different regions of the picture in parallel, thus increasing the encoding efficiency. Also, when the data of a region is damaged during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the damaged or lost data, thus bringing the ability of error tolerance. In some video encoding standards, a picture can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slice" and "tile". It should also be noted that different pictures of video sequence 100 can have different partitioning schemes for dividing the picture into regions.

[0034]

[0050] For example, in FIG. 1, structure 110 is divided into three regions 114, 116, and 118, and their boundaries are shown as solid lines inside structure 110. Region 114 includes four basic processing units. Each of regions 116 and 118 includes six basic processing units. It should be noted that the basic processing units, basic processing sub-units, and regions of structure 110 in FIG. 1 are merely examples, and the present disclosure does not limit its embodiments.

[0035]

[0051] FIG. 2A shows a schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure. For example, the encoding process 200A can be performed by an encoder. As shown in FIG. 2A, the encoder can encode a video sequence 202 into a video bitstream 228 according to process 200A. Similar to the video sequence 100 in FIG. 1, the video sequence 202 can include a set of pictures (referred to as “original pictures”) arranged in chronological order. Similar to the structure 110 in FIG. 1, each original picture of the video sequence 202 can be divided by the encoder into basic processing units, basic processing subunits, or regions for processing. Depending on the embodiment, the encoder can perform process 200A at the level of basic processing units for each original picture of the video sequence 202. For example, the encoder can perform process 200A in an iterative manner, in which case the encoder can encode a basic processing unit in one iteration of process 200A. Depending on the embodiment, the encoder can perform process 200A in parallel for regions (e.g., regions 114-118) of each original picture of the video sequence 202.

[0036]

[0052] In FIG. 2A, the coder can supply the basic processing unit of the original picture of the video sequence 202 (referred to as the "original BPU") to the prediction stage 204 and generate prediction data 206 and a prediction BPU 208. The coder can subtract the prediction BPU 208 from the original BPU to generate a residual BPU 210. The coder can supply the residual BPU 210 to the transformation stage 212 and the quantization stage 214 to generate quantized transform coefficients 216. The coder can supply the prediction data 206 and the quantized transform coefficients 216 to the binary coding stage 226 to generate a video bitstream 228. The components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the "forward path". During the process 200A, after the quantization stage 214, the coder can supply the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transformation stage 220 to generate a reconstructed residual BPU 222. The coder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 that is used in the prediction stage 204 for the next iteration of the process 200A. The components 218, 220, 222, and 224 of the process 200A may be referred to as the "reconstruction path". The reconstruction path may be used to ensure that both the coder and the decoder use the same reference data for prediction.

[0037]

[0053] The coder can iteratively perform the process 200A to encode each original BPU of the original picture (within the forward path) and generate a prediction reference 224 (within the reconstruction path) for encoding the next original BPU of the original picture. After encoding all the original BPUs of the original picture, the coder can proceed to encode the next picture within the video sequence 202.

[0038]

[0054] Referring to process 200A, the coder can receive video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" can refer to receiving, inputting, acquiring, obtaining, getting, reading, accessing, or any act by any means for inputting data.

[0039]

[0055] In prediction stage 204, in the current iteration, the coder can receive the original BPU and prediction criterion 224, perform a prediction operation, and generate prediction data 206 and prediction BPU 208. Prediction criterion 224 can be generated from the reconstruction path of a previous iteration of process 200A. The purpose of prediction stage 204 is to reduce information redundancy by extracting prediction data 206, and prediction data 206 can be used to reconstruct the original BPU as prediction BPU 208 from prediction data 206 and prediction criterion 224.

[0040]

[0056] Ideally, prediction BPU 208 can be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, prediction BPU 208 generally differs slightly from the original BPU. To record such a difference, after generating prediction BPU 208, the coder can subtract it from the original BPU to generate residual BPU 210. For example, the coder can subtract the pixel value (e.g., grayscale value or RGB value) of prediction BPU 208 from the corresponding pixel value of the original BPU. Each pixel of residual BPU 210 can have a residual value as a result of such subtraction between the corresponding pixels of the original BPU and prediction BPU 208. Compared to the original BPU, prediction data 206 and residual BPU 210 can have fewer bits, but they can be used to reconstruct the original BPU without significant quality degradation. Therefore, the original BPU is compressed.

[0041]

[0057] To further compress the residual BPU 210, in the transformation stage 212, the coder can reduce the spatial redundancy of the residual BPU 210 by decomposing it into a set of two-dimensional "basis patterns", where each basis pattern is associated with a "transformation coefficient". The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a frequency component of the change in the residual BPU 210 (e.g., the frequency of luminance change). None of the basis patterns can be reproduced from any combination (e.g., linear combination) of any other basis patterns. In other words, such a decomposition can decompose the change in the residual BPU 210 into the frequency domain. Such a decomposition is similar to the discrete Fourier transform of a function, where in this case, the basis patterns are similar to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the transformation coefficients are similar to the coefficients associated with the basis functions.

[0042]

[0058] Different transformation algorithms can use different basis patterns. For example, various transformation algorithms such as the discrete cosine transform, the discrete sine transform, or the like can be used in the transformation stage 212. The transformation in the transformation stage 212 is invertible. That is, the coder can recover the residual BPU 210 by means of the inverse operation of the transformation (referred to as "inverse transformation"). For example, to recover the pixels of the residual BPU 210, the inverse transformation can multiply the corresponding pixel values of the basis patterns by their respective associated coefficients and add up the products to generate a weighted sum. For a video coding standard, both the coder and the decoder can use the same transformation algorithm (and thus the same basis patterns). Therefore, the coder can record only the transformation coefficients, and the decoder can reconstruct the residual BPU 210 from the transformation coefficients without receiving the basis patterns from the coder. Compared with the residual BPU 210, the transformation coefficients can have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Therefore, the residual BPU 210 is further compressed.

[0043]

[0059] The coder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns can represent different change frequencies (e.g., luminance change frequency). Since the human eye is generally more adept at recognizing low-frequency changes, the coder can ignore the information of high-frequency changes without causing significant quality degradation in decoding. For example, in the quantization stage 214, the coder can generate the quantized transform coefficient 216 by dividing each transform coefficient by an integer value (referred to as the "quantization parameter") and rounding the quotient to the nearest integer. After such an operation, some of the transform coefficients of the high-frequency basis pattern can be converted to 0, and the transform coefficients of the low-frequency basis pattern can be converted to smaller integers. The coder can ignore the quantized transform coefficients 216 with a value of 0, thereby further compressing the transform coefficients. The quantization process is also invertible, in which case the quantized transform coefficients 216 can be reconstructed into transform coefficients in the inverse operation of quantization (referred to as "inverse quantization").

[0044]

[0060] Since the coder ignores the remainder of such division in the rounding operation, the quantization stage 214 can be non-invertible. Typically, the quantization stage 214 can contribute to the largest information loss in the process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To obtain different information loss levels, the coder can use different values of the quantization parameter or any other parameter of the quantization process.

[0045]

[0061] In the binary encoding stage 226, the encoder can encode the prediction data 206 and the quantized transform coefficients 216 using binary encoding techniques such as, for example, entropy encoding, variable length encoding, arithmetic encoding, Huffman encoding, context adaptive binary arithmetic encoding, or any other reversible or irreversible compression algorithm. Depending on the embodiment, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder can encode other information in the binary encoding stage 226, such as, for example, the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the type of transformation in the transformation stage 212, the parameters of the quantization process (e.g., quantization parameters), the encoder control parameters (e.g., bit rate control parameters), or the like. The encoder can generate a video bitstream 228 using the output data of the binary encoding stage 226. Depending on the embodiment, the video bitstream 228 can be further packetized for network transmission.

[0046]

[0062] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the encoder can perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In the inverse transformation stage 220, the encoder can generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 that will be used in the next iteration of process 200A.

[0047]

[0063] Note that other variations of process 200A can also be used to encode video sequence 202. Depending on the embodiment, the steps of process 200A can be performed in a different order by the encoder. Depending on the embodiment, one or more steps of process 200A can be combined into a single step. Depending on the embodiment, a single step of process 200A can be divided into multiple steps. For example, the transform step 212 and the quantization step 214 can be combined into a single step. Depending on the embodiment, process 200A can include additional steps. Depending on the embodiment, process 200A can omit one or more steps in FIG. 2A.

[0048]

[0064] FIG. 2B shows a schematic diagram of another exemplary encoding process 200B according to an embodiment of the present disclosure. Process 200B can be changed from process 200A. For example, process 200B can be used by an encoder compliant with a hybrid video encoding standard (e.g., the H.26x series). Compared with process 200A, the forward path of process 200B additionally includes a mode decision step 230, and the prediction step 204 is divided into a spatial prediction step 2042 and a temporal prediction step 2044. The reconstruction path of process 200B additionally includes a loop filter step 232 and a buffer 234.

[0049]

[0065] Generally, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra prediction") can use pixels from one or more already-encoded adjacent BPUs within the same picture to predict the target BPU. That is, the prediction reference 224 in spatial prediction can include adjacent BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter prediction") can use regions from one or more already-encoded pictures to predict the target BPU. That is, the prediction reference 224 in temporal prediction can include encoded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.

[0050]

[0066] Referring to process 200B, within the forward path, the coder performs prediction operations at spatial prediction stage 2042 and temporal prediction stage 2044. For example, at spatial prediction stage 2042, the coder can perform intra prediction. For the original BPU of the picture being encoded, prediction reference 224 can include one or more adjacent BPUs that are encoded (within the forward path) and reconstructed (within the reconstruction path) in the same picture. The coder can generate prediction BPU 208 by extrapolating the adjacent BPUs. The extrapolation technique can include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, or the like. In some embodiments, the coder can perform extrapolation at the pixel level, such as by extrapolating the value of each pixel of prediction BPU 208. The adjacent BPUs used for extrapolation can be located relative to the original BPU in various directions, such as the vertical direction (e.g., above the original BPU), the horizontal direction (e.g., to the left of the original BPU), the diagonal direction (e.g., bottom - left, bottom - right, top - left, or top - right of the original BPU), or any direction defined in the video coding standard being used. For intra prediction, prediction data 206 can include, for example, the location (e.g., coordinates) of the adjacent BPUs used, the size of the adjacent BPUs used, the parameters of the extrapolation, the direction of the adjacent BPUs used relative to the original BPU, or the like.

[0051]

[0067] As another example, in the temporal prediction stage 2044, the coder can perform inter prediction. For the original BPU of the target picture, the prediction reference 224 can include one or more pictures (referred to as "reference pictures") that are encoded (within the forward path) and reconstructed (within the reconstruction path). In some embodiments, the reference pictures can be encoded and reconstructed for each BPU. For example, the coder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all the reconstructed BPUs of the same picture have been generated, the coder can generate the reconstructed picture as a reference picture. The coder can perform an operation of "motion estimation" to search for a matching region within a range of reference pictures (referred to as a "search window"). The location of the search window within the reference picture can be determined based on the location of the original BPU of the target picture. For example, the search window can be centered at a location within the reference picture that has the same coordinates as the original BPU within the target picture and can be extended outward over a predetermined distance. When the coder identifies a region similar to the original BPU within the search window (e.g., by using a pixel recursive algorithm, a block matching algorithm, or the like), the coder can determine such a region as the matching region. The matching region can have dimensions that are different (e.g., smaller than, equal to, larger than, or of a different shape than) the original BPU. Since the reference picture and the target picture are temporally separated within a timeline (e.g., as shown in FIG. 1), the matching region can be considered to "move" to the location of the original BPU as time passes. The coder can record such a direction and distance of motion as a "motion vector". When multiple reference pictures are used (e.g., as picture 106 in FIG. 1), the coder can search for a matching region for each reference picture and determine its associated motion vector. In some embodiments, the coder can weight the pixel values of the matching region of each matching reference picture.

[0052]

[0068] Motion estimation can be used to identify various types of motion, such as translation, rotation, zooming, or the like. For inter prediction, the prediction data 206 can include, for example, the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference pictures, the weight associated with the reference pictures, or the like.

[0053]

[0069] To generate the prediction BPU 208, the coder can perform an operation of "motion compensation". Motion compensation can be used to reconstruct the prediction BPU 208 based on the prediction data 206 (e.g., motion vector) and the prediction reference 224. For example, the coder can move the matching region of the reference picture according to the motion vector so that the coder can predict the original BPU of the target picture. When multiple reference pictures are used (e.g., as picture 106 in FIG. 1), the coder can move the matching regions of the reference pictures according to their respective motion vectors and average the pixel values of the matching regions. In some embodiments, when the coder assigns weights to the pixel values of the matching regions of each matching reference picture, the coder can add the weighted sum of the pixel values to the moved matching region.

[0054]

[0070] In some embodiments, inter prediction can be unidirectional or bidirectional. Unidirectional inter prediction can use one or more reference pictures in the same temporal direction with respect to the target picture. For example, picture 104 in FIG. 1 is a unidirectional inter prediction picture where the reference picture (i.e., picture 102) precedes picture 104. Bidirectional inter prediction can use one or more reference pictures in both temporal directions with respect to the target picture. For example, picture 106 in FIG. 1 is a bidirectional inter prediction picture where the reference pictures (i.e., pictures 104 and 108) are in both temporal directions with respect to picture 104.

[0055]

[0071] Still referring to the forward path of process 200B, after the spatial prediction stage 2042 and the temporal prediction stage 2044, at the mode decision stage 230, the coder can select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of process 200B. For example, the coder can perform rate-distortion optimization techniques. In this technique, the coder can select a prediction mode to minimize the value of a cost function that depends on the bitrate of the candidate prediction modes and the distortion of the reconstructed reference pictures under the candidate prediction modes. Depending on the selected prediction mode, the coder can generate the corresponding prediction BPU 208 and prediction data 206.

[0056]

[0072] In the reconstruction path of process 200B, when the intra prediction mode is selected within the forward path, after generating a prediction reference 224 (e.g., the target BPU encoded and reconstructed in the target picture), the coder can directly supply the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for interpolation of the next BPU of the target picture). When the inter prediction mode is selected within the forward path, after generating a prediction reference 224 (e.g., the target picture where all BPUs are encoded and reconstructed), the coder can supply the prediction reference 224 to the loop filter stage 232, where the coder can apply a loop filter to the prediction reference 224 to reduce or eliminate the distortion (e.g., blocking artifacts) introduced by inter prediction. The coder can apply various loop filter techniques, such as deblocking, sample adaptive offset, adaptive loop filter, or the like, in the loop filter stage 232. The loop filtered reference picture can be stored in a buffer 234 (or "decoded picture buffer") for later use (e.g., for use as an inter prediction reference picture for future pictures of video sequence 202). The coder can store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the coder can encode loop filter parameters (e.g., loop filter strength) together with the quantized transform coefficients 216, prediction data 206, and other information in the binary encoding stage 226.

[0057]

[0073] FIG. 3A shows a schematic diagram of an exemplary decoding process 300A according to an embodiment of the present disclosure. Process 300A may be a restoration process corresponding to the compression process 200A in FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder can decode the video bitstream 228 into a video stream 304 according to process 300A. Video stream 304 may be very similar to video sequence 202. However, due to information loss in the compression and restoration processes (e.g., quantization stage 214 in FIGS. 2A-2B), generally, video stream 304 is not identical to video sequence 202. Similar to processes 200A and 200B in FIGS. 2A-2B, the decoder can perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, the decoder can perform process 300A in an iterative manner, in which case the decoder can decode the basic processing unit in one iteration of process 300A. In some embodiments, the decoder can perform process 300A in parallel for regions (e.g., regions 114-118) of each picture encoded in video bitstream 228.

[0058]

[0074] In FIG. 3A, the decoder can supply a portion of the video bitstream 228 associated with a basic processing unit of the encoded picture (referred to as the "encoded BPU") to the binary decoding stage 302. In the binary decoding stage 302, the decoder can decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder supplies the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220, and can generate a reconstructed residual BPU 222. The decoder supplies the prediction data 206 to the prediction stage 204 and can generate a prediction BPU 208. The decoder adds the reconstructed residual BPU 222 to the prediction BPU 208 and can generate a prediction reference 224. In some embodiments, the prediction reference 224 can be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder can supply the prediction reference 224 to the prediction stage 204 to perform the prediction operation in the next iteration of process 300A.

[0059]

[0075] The decoder can iteratively perform process 300A to decode each encoded BPU of the encoded picture and generate a prediction reference 224 for encoding the next encoded BPU of the encoded picture. After decoding all the encoded BPUs of the encoded picture, the decoder can output the picture to the video stream 304 for display and proceed to decode the next encoded picture in the video bitstream 228.

[0060]

[0076] In the binary decoding stage 302, the decoder can perform the inverse operation of the binary encoding technique (e.g., entropy encoding, variable-length encoding, arithmetic encoding, Huffman encoding, context-adaptive binary arithmetic encoding, or any other reversible compression algorithm) used by the encoder. Depending on the embodiment, in addition to the predicted data 206 and the quantized transform coefficients 216, the decoder can also decode other information in the binary decoding stage 302, such as, for example, the prediction mode, the parameters of the prediction operation, the type of transform, the parameters of the quantization process (e.g., quantization parameters), the encoder control parameters (e.g., bitrate control parameters), or the like. Depending on the embodiment, when the video bitstream 228 is transmitted in the form of packets through a network, the decoder can depacketize the video bitstream 228 before supplying it to the binary decoding stage 302.

[0061]

[0077] FIG. 3B shows a schematic diagram of another exemplary decoding process 300B according to an embodiment of the present disclosure. The process 300B can be changed from the process 300A. For example, the process 300B can be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with the process 300A, the process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filter stage 232 and a buffer 234.

[0062]

[0078] In process 300B, for an encoding basic processing unit (referred to as the "current BPU" or "target BPU") of an encoded picture being decoded (referred to as the "current picture" or "target picture"), the prediction data 206 decoded by the decoder from the binary decoding stage 302 can include various types of data depending on which prediction mode was used by the encoder to encode the target BPU. For example, if intra prediction was used by the encoder to encode the target BPU, the prediction data 206 can include an intra prediction, parameters of the intra prediction operation, or a prediction mode indicator (e.g., a flag value) indicating the like. The parameters of the intra prediction operation can include, for example, the location (e.g., coordinates) of one or more adjacent BPUs used as references, the size of the adjacent BPUs, extrapolation parameters, the direction of the adjacent BPUs with respect to the original BPU, or the like. As another example, if inter prediction was used by the encoder to encode the target BPU, the prediction data 206 can include an inter prediction, parameters of the inter prediction operation, or a prediction mode indicator (e.g., a flag value) indicating the like. The parameters of the inter prediction operation can include, for example, the number of reference pictures associated with the target BPU, the weights respectively associated with the reference pictures, the location (e.g., coordinates) of one or more matching regions in each reference picture, one or more motion vectors respectively associated with the matching regions, or the like.

[0063]

[0079] Based on the prediction mode indicator, the decoder can determine whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Details of performing such spatial or temporal prediction are described in FIG. 2B and will not be repeated here. After performing such spatial or temporal prediction, the decoder can generate a prediction BPU 208. As described in FIG. 3A, the decoder can add the prediction BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224.

[0064]

[0080] In process 300B, the decoder can supply the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing the prediction operation in the next iteration of process 300B. For example, when the target BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded target BPU), the decoder can directly supply the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the target picture). When the target BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture with all BPUs decoded), the encoder can supply the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder can apply the loop filter to the prediction reference 224 in the manner as described in FIG. 2B. The loop-filtered reference picture can be stored in buffer 234 (e.g., the decoded picture buffer in computer memory) for later use (e.g., for use as an inter prediction reference picture for future encoded pictures of video bitstream 228). The decoder can store one or more reference pictures in buffer 234 for use in the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter prediction is used to encode the target BPU, the prediction data can further include loop filter parameters (e.g., loop filter strength).

[0065]

[0081] FIG. 4 is a block diagram of an exemplary apparatus 400 for encoding or decoding video according to an embodiment of the present disclosure. As shown in FIG. 4, apparatus 400 can include a processor 402. When processor 402 executes the instructions described herein, apparatus 400 can become a special machine for video encoding or decoding. Processor 402 can be any kind of circuitry having the ability to manipulate or process information. For example, processor 402 can include a central processing unit (or “CPU”), a graphics processing unit (or “GPU”), a neural processing unit (“NPU”), a microcontroller unit (“MCU”), an optical processor, a programmable logic controller, a microcontroller, a microprocessor, a digital signal processor, an intellectual property (IP) core, a programmable logic array (PLA), a programmable array logic (PAL), a generic array logic (GAL), a complex programmable logic device (CPLD), a field programmable gate array (FPGA), a system on chip (SoC), an application specific integrated circuit (ASIC), or any number of any combination of the like. In some embodiments, processor 402 can be a set of processors grouped as a single logical component. For example, as shown in FIG. 4, processor 402 can include a plurality of processors, including processor 402a, processor 402b, and processor 402n.

[0066]

[0082] The machine 400 can also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, or the like). For example, as shown in FIG. 4, the stored data can include program instructions (e.g., program instructions for performing the steps in processes 200A, 200B, 300A, or 300B) as well as data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 can access (e.g., via bus 410) the program instructions and the data for processing, execute the program instructions, and perform operations or manipulations on the data for processing. The memory 404 can include a high-speed random access memory device or a non-volatile memory device. In some embodiments, the memory 404 can include any number of any combination of random access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (registered trademark) (CF) cards, or the like. The memory 404 can also be a group of memories grouped as a single logical component (not shown in FIG. 4).

[0067]

[0083] The bus 410 can be a communication device that transfers data between components inside the machine 400, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a universal serial bus port, a peripheral component interconnect express port), or the like.

[0068]

[0084] To facilitate explanation without creating ambiguity, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits" in this disclosure. The data processing circuits can be implemented entirely as hardware or as a combination of software, hardware, or firmware. Additionally, the data processing circuits can be a single stand-alone module or can be fully or partially integrated with any other component of the device 400.

[0069]

[0085] The device 400 can further include a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, or the like). Depending on the embodiment, the network interface 406 can include any number of any combination of a network interface controller (NIC), a radio frequency (RF) module, a transponder, a transceiver, a modem, a router, a gateway, a wired network adapter, a wireless network adapter, a Bluetooth® adapter, an infrared adapter, a near field communication ("NFC") adapter, a cellular network chip, or the like.

[0070]

[0086] Depending on the embodiment, optionally, the device 400 can further include a peripheral interface 408 for providing connection to one or more peripheral devices. As shown in FIG. 4, the peripheral devices can include, but are not limited to, a cursor control device (e.g., a mouse, a touchpad, or a touch screen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera or an input interface communicatively coupled to a video archive), or the like.

[0071]

[0087] Note that the video codec (e.g., the codec that performs processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules within device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instructions that can be loaded into memory 404. As another example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as special data processing circuits (e.g., FPGA, ASIC, NPU, or the like).

[0072]

[0088] In the quantization and inverse quantization functional blocks (e.g., quantization 214 and inverse quantization 218 in FIGS. 2A or 2B, inverse quantization 218 in FIGS. 3A or 3B), quantization parameters (QPs) are used to determine the amount of quantization (and inverse quantization) applied to the prediction residual. The initial QP value used for encoding a picture or slice can be signaled at a high level, for example, using the init_qp_minus26 syntax element within the picture parameter set (PPS) and the slice_qp_delta syntax element within the slice header. Further, the QP value can be adapted at a local level for each CU using the delta QP value transmitted with the subdivision of the quantization group.

[0073]

[0089] In some real-time applications (e.g., video conferencing or remote operation systems), system latency can be a critical issue that significantly affects the user experience and reliability of the system. For example, ITU-T G.114 stipulates that the latency tolerance for two-way audio-video communication is 150 milliseconds. In another example, virtual reality applications typically require an ultra-low latency of less than 20 milliseconds to prevent motion sickness caused by the timing deviation between head movement and the visual effects resulting from that movement.

[0074]

[0090] In a real-time video application system, the total latency includes the period from the time a frame is captured to the time the frame is displayed. That is, the total latency is the sum of the encoding time in the encoder, the transmission time in the transmission channel, the decoding time in the decoder, and the output delay in the decoder. Generally, the transmission time contributes the most to the total latency. The transmission time of an encoded picture is typically equal to the capacity of the coded picture buffer (CPB) divided by the bit rate of the video sequence.

[0075]

[0091] In the present disclosure, "random access" refers to the ability to start the decoding process at any random access point of a video sequence or stream and recover the correctly decoded pictures within the content. To support random access and prevent error propagation, intra-coded random access point (IRAP) pictures can be periodically inserted into the video sequence. However, at high coding efficiency, the size of an encoded I picture (e.g., an IRAP picture) is typically larger than the size of a P picture or a B picture. The larger size of the IRAP picture may result in a transmission delay higher than the average transmission delay. Therefore, periodically inserting IRAP pictures may not meet the requirements of low-latency video applications.

[0076]

[0092] In accordance with embodiments of the present disclosure, for low-latency encoding, a Gradual Decoding Refresh (GDR) technique, also referred to as a Progressive Intra Refresh (PIR) technique, can be used to reduce the latency caused by the insertion of IRAP pictures while enabling random access within a video sequence. GDR can gradually refresh pictures by dispersing intra-coded regions within non-intra-coded regions (e.g., B pictures or P pictures). By doing so, the sizes of hybrid-coded pictures can be made similar to each other, thereby reducing or minimizing the size of the CPB (e.g., to a value equal to the video sequence bitrate divided by the picture rate), and shortening the encoding time and decoding time within the total latency.

[0077]

[0093] As an example, FIG. 5 is a schematic diagram showing exemplary operations of Gradual Decoding Refresh (GDR) according to some embodiments of the present disclosure. FIG. 5 shows a GDR period 502 including a plurality of pictures (e.g., pictures 504, 506, 508, 510, and 512) within a video sequence (e.g., video sequence 100 of FIG. 1). The first picture within the GDR period 502 is called a GDR picture 504 which can be a random access picture, and the last picture within the GDR period 502 is called a recovery point picture 512. Each picture within the GDR period 502 includes an intra-coded region (represented by vertical boxes labeled "INTRA" within each picture of FIG. 5). Each intra-coded region can include various portions of the complete picture. As shown in FIG. 5, the intra-coded regions can gradually span the entire picture within the GDR period 502. In FIG. 5, the intra-coded regions are shown as rectangular slices, but it should be noted that these regions can be implemented in various shapes and sizes and are not limited by the examples described in the present disclosure.

[0078]

[0094] Pictures other than GDR picture 504 and recovery point picture 512 within GDR period 502 that are divided by the intra-coded region (e.g., any one of pictures 506 to 510) can include two regions: a "clean region" containing pixels that have already been refreshed and a "dirty region" containing pixels that are damaged due to transmission errors in previous pictures and still have not been refreshed (e.g., can be refreshed in subsequent pictures). The clean region of the current picture (e.g., picture 510) can include pixels that are reconstructed by referring to at least one of the clean region or the intra-coded region of the previous picture (e.g., pictures 508, 506, and GDR picture 504). The clean region of the current picture (e.g., picture 510) can include pixels that are reconstructed by referring to at least one of the dirty region, clean region, or intra-coded region of the previous picture (e.g., pictures 508, 506, and GDR picture 504).

[0079]

[0095] The principle of the GDR technique is to ensure that pixels within the clean region are reconstructed without using any information from any dirty region (e.g., the dirty region of the current picture or any previous picture). As an example, in Figure 5, the GDR picture 504 includes the dirty region 514. Picture 506 includes a clean region 516 that can be reconstructed using the intra-coded region of the GDR picture 504 as a reference, and a dirty region 518 that can be reconstructed using any part of the GDR picture 504 (e.g., at least one of the intra-coded region or the dirty region 514) as a reference. Picture 508 includes a clean region 520 that can be reconstructed using at least one of the intra-coded regions (e.g., the intra-coded region of picture 504 or 506) or the clean region (e.g., the clean region 516) of pictures 504 - 506 as a reference, and a dirty region 522 that can be reconstructed using at least one of a part of the GDR picture 504 (e.g., the intra-coded region or the dirty region 514), or a part of picture 506 (e.g., the clean region 516, the intra-coded region, or the dirty region 518) as a reference. Picture 510 includes a clean region 524 that can be reconstructed using at least one of the intra-coded regions (e.g., the intra-coded region of pictures 504, 506, or 508) or the clean region (e.g., the clean regions 516 or 520) of pictures 504 - 508 as a reference, and a dirty region 526 that can be reconstructed using at least one of a part of the GDR picture 504 (e.g., the intra-coded region or the dirty region 514), a part of picture 506 (e.g., the clean region 516, the intra-coded region, or the dirty region 518), or a part of picture 508 (e.g., the clean region 520, the intra-coded region, or the dirty region 522) as a reference.The recovery point picture 512 includes a clean area 528 that can be reconstructed by referring to at least one of a part of the GDR picture 504 (e.g., an intra-coded area or a dirty area 514), a part of the picture 506 (e.g., a clean area 516, an intra-coded area, or a dirty area 518), a part of the picture 508 (e.g., a clean area 520, an intra-coded area, or a dirty area 522), or a part of the picture 510 (e.g., a clean area 524, an intra-coded area, or a dirty area 526).

[0080]

[0096] As shown in the figure, all pixels of the recovery point picture 512 are refreshed. Decoding pictures in output order using the GDR technique after the recovery point picture 512 can be equivalent to decoding the pictures using the previous IRAP picture of the GDR picture 504 as if it existed, and the IRAP picture extends over all intra-coded areas of the pictures 504 to 512.

[0081]

[0097] As an example, FIG. 6 shows Table 1 illustrating the exemplary syntax structure of a sequence parameter set (SPS) enabling GDR according to some embodiments of the present disclosure. As shown in Table 1, a GDR sequence level enabling flag "gdr_enabled_flag" can be signaled within the SPS of the video sequence to indicate whether any GDR-compatible picture (e.g., any picture within the GDR period 502 of FIG. 5) exists within the video sequence. In some embodiments, in the VVC / H.266 standard, a true (e.g., equal to "1") "gdr_enabled_flag" can specify that a GDR-compatible picture exists within the coded layer video sequence (CLVS) referring to the SPS, and a false (e.g., equal to "0") "gdr_enabled_flag" can specify that no GDR-compatible picture exists within the CLVS.

[0082]

[0098] As an example, FIG. 7 shows Table 2, which illustrates an exemplary syntax structure of a picture header for enabling a GDR according to some embodiments of the present disclosure. As shown in Table 2, for a picture of a video sequence, in order to indicate whether the picture is GDR-compatible (e.g., any picture within the GDR period 502 in FIG. 5), a GDR picture-level enabling flag "gdr_pic_flag" can be signaled within the picture header of the picture. If the picture is GDR-compatible, a parameter "recovert_poc_cnt" can be signaled to specify the recovery point picture (e.g., the recovery point picture 512 in FIG. 5) in the output order. In some embodiments, in the VVC / H.266 standard, a true (e.g., equal to "1") "gdr_pic_flag" can specify that the picture related to the picture header is a GDR-compatible picture, and a false (e.g., equal to "0") "gdr_pic_flag" can specify that the picture related to the picture header is not a GDR-compatible picture.

[0083]

[0099] As an example, in the VVC / H.266 standard, when the "gdr_enabled_flag" shown in Table 1 is true and the parameter "PicOrderCntVal" of the current picture (not shown in FIG. 7) is greater than or equal to the sum of "PicOrderCntVal" and "recovery_poc_cnt" of the GDR-compatible picture (or pictures) related to the current picture, the current picture and subsequent pictures in the output order can be decoded as if they were decoded by starting the decoding process from an IRAP picture preceding the GDR-compatible picture (or pictures).

[0084]

[0100] According to embodiments of the present disclosure, virtual boundary techniques can be used to implement GDR (e.g., in the VVC / H.266 standard). In some applications (e.g., 360-degree video), the layout of a particular projection format can typically have multiple faces. When those projection formats include multiple faces, discontinuities can occur between two or more adjacent faces within a frame-packed picture, regardless of what kind of compact frame packing configuration is used. When in-loop filtering operations are performed across these discontinuities, seam artifacts of the faces may become visible in the reconstructed video after rendering.

[0085]

[0101] To reduce seam artifacts of the faces, in-loop filtering operations (e.g., deblocking filtering, sample adaptive offset filtering, or adaptive loop filtering) can be disabled across the discontinuities within a frame-packed picture, which can be referred to as virtual boundary techniques (e.g., the concept adopted by VVC draft 7). For example, the coder can set the discontinuous boundary as a virtual boundary and disable any loop filtering operations across the virtual boundary. By doing so, loop filtering across the discontinuities can be disabled.

[0086]

[0102] In the case of GDR, loop filtering operations should not be applied across the boundary between the clean region (e.g., the clean region 520 of picture 508 in FIG. 5) and the dirty region (e.g., the dirty region 522 of picture 508). The coder can set the boundary between the clean region and the dirty region as a virtual boundary and disable the loop filtering operations across the virtual boundary. By doing so, virtual boundaries can be used as a way to implement GDR.

[0087]

[0103] According to some embodiments, in the VVC / H.266 standard (e.g., within VVC Draft 7), virtual boundaries can be signaled in the SPS or picture header. As an example, FIG. 8 shows Table 3 which illustrates an exemplary syntax structure of the SPS enabling virtual boundaries according to some embodiments of the present disclosure. FIG. 9 shows Table 4 which illustrates an exemplary syntax structure of the picture header enabling virtual boundaries according to some embodiments of the present disclosure.

[0088]

[0104] As shown in Table 3, the sequence level virtual boundary presence flag "sps_virtual_boundaries_present_flag" can be signaled within the SPS. For example, a true "sps_virtual_boundaries_present_flag" (e.g., equal to "1") may specify that virtual boundary information is signaled within the SPS, and a false "sps_virtual_boundaries_present_flag" (e.g., equal to "0") may specify that virtual boundary information is not signaled within the SPS. If one or more virtual boundaries are signaled within the SPS, in-loop filtering operations across the virtual boundaries within a picture referring to the SPS can be disabled.

[0089]

[0105] As shown in Table 3, when the flag "sps_virtual_boundaries_present_flag" is true, the number of virtual boundaries (represented by the parameters "sps_num_ver_virtual_boundaries" and "sps_num_hor_virtual_boundaries" in Table 3) and their positions (represented by the arrays "sps_virtual_boundaries_pos_x" and "sps_virtual_boundaries_pos_y" in Table 3) can be signaled within the SPS. The parameters "sps_num_ver_virtual_boundaries" and "sps_num_hor_virtual_boundaries" can respectively specify the lengths of the arrays "sps_virtual_boundaries_pos_x" and "sps_virtual_boundaries_pos_y" within the SPS. In some embodiments, if "sps_num_ver_virtual_boundaries" (or "sps_num_hor_virtual_boundaries") does not exist within the SPS, its value can be inferred to be 0. The arrays "sps_virtual_boundaries_pos_x" and "sps_virtual_boundaries_pos_y" can respectively specify the positions of the i-th vertical or horizontal virtual boundary in units of luma samples divided by 8. For example, the value of "sps_virtual_boundaries_pos_x[i]" can be within the closed interval of 1 to Ceil(pic_width_in_luma_samples÷8)-1, where "Ceil" represents the ceiling function and "pic_width_in_luma_samples" is a parameter representing the width of the picture in units of luma samples. The value of "sps_virtual_boundaries_pos_y[i]" can be within the closed interval of 1 to Ceil(pic_height_in_luma_samples÷8)-1, where "pic_height_in_luma_samples" is a parameter representing the height of the picture in units of luma samples.

[0090]

[0106] Depending on the embodiment, when the flag "sps_virtual_boundaries_present_flag" is false (e.g., equal to "0"), as shown in Table 4, the picture level virtual boundary presence flag "ph_virtual_boundaries_present_flag" can be signaled within the picture header. For example, a true "ph_virtual_boundaries_present_flag" (e.g., equal to "1") can specify that virtual boundary information is signaled within the picture header, and a false "ph_virtual_boundaries_present_flag" (e.g., equal to "0") can specify that virtual boundary information is not signaled within the picture header. When one or more virtual boundaries are signaled within the picture header, the in-loop filtering operations over the virtual boundaries within the picture including the picture header can be disabled. Depending on the embodiment, when "ph_virtual_boundaries_present_flag" does not exist within the picture header, its value can be inferred to represent "false".

[0091]

[0107] As shown in Table 4, when the flag "ph_virtual_boundaries_present_flag" is true (e.g., equal to "1"), the number of virtual boundaries (represented by the parameters "ph_num_ver_virtual_boundaries" and "ph_num_hor_virtual_boundaries" in Table 4) and their positions (represented by the arrays "ph_virtual_boundaries_pos_x" and "ph_virtual_boundaries_pos_y" in Table 4) can be signaled within the picture header. The parameters "ph_num_ver_virtual_boundaries" and "ph_num_hor_virtual_boundaries" can respectively specify the lengths of the arrays "ph_virtual_boundaries_pos_x" and "ph_virtual_boundaries_pos_y" within the picture header. In some embodiments, if "ph_virtual_boundaries_pos_x" (or "ph_virtual_boundaries_pos_y") does not exist within the picture header, its value can be inferred to be 0. The arrays "ph_virtual_boundaries_pos_x" and "ph_virtual_boundaries_pos_y" can respectively specify the positions of the i-th vertical or horizontal virtual boundary in units of luma samples divided by 8. For example, the value of "ph_virtual_boundaries_pos_x[i]" can be within the closed interval of 1 to Ceil(pic_width_in_luma_samples÷8)-1, the value of "ph_virtual_boundaries_pos_y[i]" can be within the closed interval of 1 to Ceil(pic_height_in_luma_samples÷8)-1, and "pic_height_in_luma_samples" is a parameter representing the height of the picture in units of luma samples.

[0092]

[0108] According to some embodiments, in the VVC / H.266 standard (e.g., within VVC Draft 7), the variable "VirtualBoundariesDisabledFlag" can be defined as Equation (1). VirtualBoundariesDisabledFlag = sps_virtual_boundaries_present_flag || ph_virtual_boundaries_present_flag Equation (1)

[0093]

[0109] However, in the implementation form of GDR by using virtual boundaries, two problems may occur in the existing technical solutions. For example, as described above, in the existing technical solutions, regardless of the value of the sequence-level flag "gdr_enabled_flag", the picture-level flag "gdr_pic_flag" is always signaled within the picture header. That is, even if GDR is disabled for a sequence, the picture header of each picture in the sequence can still indicate whether the picture is GDR-compatible. Therefore, a contradiction may occur between the SPS level and the picture level. For example, a contradiction may occur when "gdr_enabled_flag" is false and "gdr_pic_flag" is true.

[0094]

[0110] As another example, when using virtual boundaries as the boundary between the clean area and the dirty area to implement GDR, in the existing technical solutions, the loop filtering operation across the virtual boundary is not applied. However, as a requirement of GDR, decoding pixels within the clean area cannot refer to pixels within the dirty area, but decoding pixels within the dirty area can refer to pixels within the clean area. In that case, completely disabling the loop filtering across the virtual boundary may impose overly strict restrictions, and such restrictions may degrade the encoding or decoding performance.

[0095]

[0111] To solve the above problems, the present disclosure provides a method, apparatus, and system for processing pictures. According to some embodiments of the present disclosure, in order to eliminate potential contradictions of the GDR indication flag at the SPS level and the picture level, the syntax structure of the picture header can be modified so that the picture level GDR indication flag can be signaled only when GDR is enabled at the sequence level.

[0096]

[0112] As an example, FIG. 10 shows Table 5 illustrating the exemplary syntax structure of a modified picture header according to some embodiments of the present disclosure. As shown in Table 5, the element 1002 (enclosed by the solid box) shows the syntax modification compared to Table 2 in FIG. 7. For example, a true (e.g., equal to "1") "gdr_pic_flag" may specify that the picture associated with the picture header is a GDR-capable picture, and a false (e.g., equal to "0") "gdr_pic_flag" may specify that the picture associated with the picture header is not a GDR-capable picture. In some embodiments, if "gdr_pic_flag" does not exist within the picture header, its value can be inferred to represent "false".

[0097]

[0113] In accordance with some embodiments of the present disclosure, to eliminate potential contradictions of the GDR indication flag at the SPS level and the picture level, the syntax structure of the picture header can be kept unchanged (e.g., as shown in Table 2 of FIG. 7), and the bitstream compliance requirements (e.g., bitstream compliance defined in the VVC / H.266 standard) can be implemented such that the picture-level GDR indication flag is not true (e.g., invalid or false) when the sequence-level GDR indication flag is not true (e.g., invalid or false). As used herein, the bitstream compliance requirements can refer to operations that can ensure that a bitstream subset related to the operation point conforms to a video coding standard (e.g., the VVC / H.266 standard). The "operation point" can refer to a first bitstream created from a second bitstream by a sub-bitstream extraction process, and in such an extraction process, for the network abstraction layer (NAL) units of the second bitstream, they can be removed if they do not belong to a target set determined by a list of target time identifiers and target layer identifiers. For example, the bitstream compliance requirements can be implemented such that the "gdr_pic_flag" is also set to false when the "gdr_enabled_flag" is false.

[0098]

[0114] According to some embodiments of the present disclosure, in order to enhance flexibility when disabling loop filtering operations across virtual boundaries, the syntax structures of the SPS and picture headers can be modified so that loop filtering operations across virtual boundaries can be partially disabled. By doing so, pixels on one side of the virtual boundary can be not filtered, while pixels on the opposite side of the virtual boundary can be filtered. For example, when the virtual boundary vertically divides the picture into a left side and a right side, the encoder or decoder can partially disable the loop filter on the right side so that pixels are not filtered there (e.g., the information of the left pixels is not used for the loop filtering of the right pixels), and enable the loop filter on the left side so that pixels are filtered there (e.g., the information of at least one of the left or right pixels can be used for loop filtering).

[0099]

[0115] As an example, FIG. 11 shows Table 6 which illustrates an exemplary syntax structure of an SPS with modifications to enable virtual boundaries according to some embodiments of the present disclosure. FIG. 12 shows Table 7 which illustrates an exemplary syntax structure of a picture header with modifications to enable virtual boundaries according to some embodiments of the present disclosure. As shown in the accompanying drawings of the present disclosure, the dashed boxes indicate that the content or elements enclosed therein are deleted or removed (indicated by a strikethrough). As shown in FIGS. 11-12, the sequence level GDR indication flag "sps_virtual_boundaries_present_flag" and the picture level GDR indication flag "ph_virtual_boundaries_present_flag" are respectively replaced by the GDR control parameters "sps_virtual_boundaries_loopfilter_disable" and "ph_virtual_boundaries_loopfilter_disable", which are extended to support loop filtering operations that are partially disabled at each of the sequence level and the picture level.

[0100]

[0116] If the direction of the GDR (e.g., from left to right, from right to left, from top to bottom, from bottom to top, or any combination thereof) is fixed across the entire sequence, the GDR control parameter (e.g., "sps_virtual_boundaries_loopfilter_disable") can be set within the SPS, thereby saving bits. If the direction of the GDR needs to be changed within the sequence, the GDR control parameter (e.g., "ph_virtual_boundaries_loopfilter_disable") can be set within the picture header, thereby providing higher flexibility for low-level control.

[0101]

[0117] According to some embodiments of the present disclosure, the GDR control parameters "sps_virtual_boundaries_loopfilter_disable" and "ph_virtual_boundaries_loopfilter_disable" can be configured to be a plurality of values (e.g., beyond the representation of "true" or "false") to represent various implementation methods.

[0102]

[0118] For example, "sps_virtual_boundaries_loopfilter_disable" being "0" may specify that the information on virtual boundaries is not signaled within the SPS. "sps_virtual_boundaries_loopfilter_disable" being "1" may specify that the information on virtual boundaries is signaled within the SPS and the in-loop filtering operation is disabled across the virtual boundaries. "sps_virtual_boundaries_loopfilter_disable" being "2" may specify that the information on virtual boundaries is signaled within the SPS and one of the following: (1) the in-loop filtering operation on the left side of the virtual boundary is disabled; (2) the in-loop filtering operation on the left side does not use the information of any pixel on the right side of the virtual boundary; (3) the in-loop filtering operation on the upper side of the virtual boundary is disabled; or (4) the in-loop filtering operation on the upper side does not use the information of any pixel on the lower side of the virtual boundary. "sps_virtual_boundaries_loopfilter_disable" being "3" may specify that the information on virtual boundaries is signaled within the SPS and one of the following: (1) the in-loop filtering operation on the right side of the virtual boundary is disabled; (2) the in-loop filtering operation on the right side does not use the information of any pixel on the left side of the virtual boundary; (3) the in-loop filtering operation on the lower side of the virtual boundary is disabled; or (4) the in-loop filtering operation on the lower side does not use the information of any pixel on the upper side of the virtual boundary.

[0103]

[0119] Similarly, in another example, "ph_virtual_boundaries_loopfilter_disable" being "0" may specify that the information of virtual boundaries is not signaled in the picture header. "ph_virtual_boundaries_loopfilter_disable" being "1" may specify that the information of virtual boundaries is signaled in the picture header and the in-loop filtering operation is disabled across the virtual boundaries. "ph_virtual_boundaries_loopfilter_disable" being "2" may specify that the information of virtual boundaries is signaled in the picture header and one of the following: (1) the in-loop filtering operation on the left side of the virtual boundary is disabled; (2) the left-side in-loop filtering operation does not use the information of any pixel on the right side of the virtual boundary; (3) the in-loop filtering operation on the upper side of the virtual boundary is disabled; or (4) the upper-side in-loop filtering operation does not use the information of any pixel on the lower side of the virtual boundary. "ph_virtual_boundaries_loopfilter_disable" being "3" may specify that the information of virtual boundaries is signaled in the picture header and one of the following: (1) the in-loop filtering operation on the right side of the virtual boundary is disabled; (2) the right-side in-loop filtering operation does not use the information of any pixel on the left side of the virtual boundary; (3) the in-loop filtering operation on the lower side of the virtual boundary is disabled; or (4) the lower-side in-loop filtering operation does not use the information of any pixel on the upper side of the virtual boundary. According to some embodiments, when "ph_virtual_boundaries_loopfilter_disable" does not exist in the picture header, its value can be inferred to be 0.

[0104]

[0120] According to some embodiments, the variable "VirtualBoundariesLoopfilterDisabled" can be defined as formula (2). VirtualBoundariesLoopfilterDisabled = sps_virtual_boundaries_loopfilter_disable ?sps_virtual_boundaries_loopfilter_disable: ph_virtual_boundaries_loopfilter_disable in Equation (2)

[0105]

[0121] According to some embodiments of the present disclosure, the loop filter can be an adaptive loop filter (ALF). When the ALF is partially disabled on the first side (e.g., left side, right side, upper side, or lower side), the pixels on the first side can be padded within the filtering, and the pixels on the second side (e.g., right side, left side, lower side, or upper side) are not used for filtering.

[0106]

[0122] In some embodiments, the boundary position of the ALF can be derived as described below. In the ALF boundary position derivation process, the variables "clipLeftPos", "clipRightPos", "clipTopPos", and "clipBottomPos" can be set to "-128".

[0107]

[0123] Compared with the VVC / H.266 standard (e.g., in VVC Draft 7), the variable "clipTopPos" can be determined as follows. When (y - (CtbSizeY - 4)) is greater than or equal to 0, the variable "clipTopPos" can be set to (yCtb + CtbSizeY - 4). When (y - (CtbSizeY - 4)) is negative, "VirtualBoundariesLoopfilterDisabled" is equal to 1, and (yCtb + y - VirtualBoundariesPosY[n]) is within the semi-open interval [1, 3) for any n = 0, 1, ..., (VirtualBoundariesNumHor - 1), "clipTopPos" can be set to "VirtualBoundariesPosY[n]" (i.e., clipTopPos = VirtualBoundariesPosY[n]). When (y - (CtbSizeY - 4)) is negative, "VirtualBoundariesLoopfilterDisabled" is equal to 3, and (yCtb + y - VirtualBoundariesPosY[n]) is within the semi-open interval [1, 3) for any n = 0, 1, ..., (VirtualBoundariesNumHor - 1), "clipTopPos" can be set to "VirtualBoundariesPosY[n]" (i.e., clipTopPos = VirtualBoundariesPosY[n]).

[0108]

[0124] If (y-(CtbSizeY-4)) is negative, y is less than 3, and one or more of the following conditions are true, then "clipTopPos" can be set to "yCtb". The following conditions are: (1) the upper boundary of the current coded tree block is the upper boundary of the tile and "loop_filter_across_tiles_enabled_flag" is equal to 0; (2) the upper boundary of the current coded tree block is the upper boundary of the slice and "loop_filter_across_slices_enabled_flag" is equal to 0; or (3) the upper boundary of the current coded tree block is the upper boundary of the subpicture and "loop_filter_across_subpic_enabled_flag[SubPicIdx]" is equal to 0.

[0109]

[0125] Compared with the VVC / H.266 standard (e.g., in VVC draft 7), the variable "clipBottomPos" can be determined as follows. If "VirtualBoundariesLoopfilterDisabled" is equal to 1, "VirtualBoundariesPosY[n]" is not equal to (pic_height_in_luma_samples-1) or 0, and (VirtualBoundariesPosY[n]-yCtb-y) is within the open interval (0,5) for any n = 0,...,(VirtualBoundariesNumHor-1), then "clipBottomPos" can be set to "VirtualBoundariesPosY[n]" (i.e., clipBottomPos = VirtualBoundariesPosY[n]).

[0110]

[0126] If 「VirtualBoundariesLoopfilterDisabled」 is equal to 2, 「VirtualBoundariesPosY[n]」 is not equal to (pic_height_in_luma_samples - 1) or 0, and (VirtualBoundariesPosY[n] - yCtb - y) is within the open interval (0, 5) for any n = 0,...,(VirtualBoundariesNumHor - 1), then 「clipBottomPos」 can be set to 「VirtualBoundariesPosY[n]」 (i.e., clipBottomPos = VirtualBoundariesPosY[n]).

[0111]

[0127] Otherwise, if (CtbSizeY - 4 - y) is within the open interval (0, 5), 「clipBottomPos」 can be set to 「yCtb + CtbSizeY - 4」. Otherwise, if (CtbSizeY - y) is less than 5 and one or more of the following conditions are true, 「clipBottomPos」 can be set to 「(yCtb + CtbSizeY)」, and the following conditions are: (1) the lower boundary of the current coded tree block is the lower boundary of the tile and 「loop_filter_across_tiles_enabled_flag」 is equal to 0; (2) the lower boundary of the current coded tree block is the lower boundary of the slice and 「loop_filter_across_slices_enabled_flag」 is equal to 0; or (3) the lower boundary of the current coded tree block is the lower boundary of the subpicture and 「loop_filter_across_subpic_enabled_flag[SubPicIdx]」 is equal to 0.

[0112]

[0128] Compared with the VVC / H.266 standard (e.g., in VVC draft 7), the variable "clipLeftPos" can be determined as follows. If "VirtualBoundariesLoopfilterDisabled" is equal to 1 and (xCtb + x - VirtualBoundariesPosX[n]) is within the semi-open interval [1, 3) for any n = 0,..., (VirtualBoundariesNumVer - 1), then "clipLeftPos" can be set to "VirtualBoundariesPosX[n]" (i.e., clipLeftPos = VirtualBoundariesPosX[n]). If "VirtualBoundariesLoopfilterDisabled" is equal to 3 and (xCtb + x - VirtualBoundariesPosX[n]) is within the semi-open interval [1, 3) for any n = 0,..., (VirtualBoundariesNumVer - 1), then "clipLeftPos" can be set to "VirtualBoundariesPosX[n]" (i.e., clipLeftPos = VirtualBoundariesPosX[n]).

[0113]

[0129] Otherwise, if x is less than 3 and one or more of the following conditions are true, then "clipLeftPos" can be set to "xCtb", and the following conditions are: (1) the left boundary of the current coding tree block is the left boundary of the tile and "loop_filter_across_tiles_enabled_flag" is equal to 0; (2) the left boundary of the current coding tree block is the left boundary of the slice and "loop_filter_across_slices_enabled_flag" is equal to 0; (3) the left boundary of the current coding tree block is the left boundary of the sub-picture and "loop_filter_across_subpic_enabled_flag[SubPicIdx]" is equal to 0.

[0114]

[0130] Compared with the VVC / H.266 standard (e.g., in VVC Draft 7), the variable "clipRightPos" can be determined as follows. If "VirtualBoundariesLoopfilterDisabled" is equal to 1 and "(VirtualBoundariesPosX[n] - xCtb - x)" is within the open interval (0, 5) for any n = 0,..., (VirtualBoundariesNumVer - 1), then "clipRightPos" can be set to "VirtualBoundariesPosX[n]" (i.e., clipRightPos = VirtualBoundariesPosX[n]). If "VirtualBoundariesLoopfilterDisabled" is equal to 2 and (VirtualBoundariesPosX[n] - xCtb - x) is within the open space (0, 5) for any n = 0,..., (VirtualBoundariesNumVer - 1), then "clipRightPos" can be set to "VirtualBoundariesPosX[n]" (i.e., clipRightPos = VirtualBoundariesPosX[n]).

[0115]

[0131] Otherwise, if "(CtbSizeY - x)" is less than 5 and one or more of the following conditions are true, then "clipRightPos" can be set to (xCtb + CtbSizeY), and the following conditions are: (1) the right boundary of the current coding tree block is the right boundary of the tile and "loop_filter_across_tiless_enabled_flag" is equal to 0; (2) the right boundary of the current coding tree block is the right boundary of the slice and "loop_filter_across_slices_enabled_flag" is equal to 0; or (3) the right boundary of the current coding tree block is the right boundary of the subpicture and "loop_filter_across_subpic_enabled_flag[SubPicIdx]" is equal to 0.

[0116]

[0132] Compared with the VVC / H.266 standard (e.g., in VVC Draft 7), the variables "clipTopLeftFlag" and "clipBotRightFlag" can be determined as follows. When the coding tree blocks including the luma position (xCtb, yCtb) and the coding tree blocks including the luma position (xCtb - CtbSizeY, yCtb - CtbSizeY) belong to different slices and "loop_filter_across_slices_enabled_flag" is equal to 0, "clipTopLeftFlag" can be set to 1. When the coding tree blocks including the luma position (xCtb, yCtb) and the coding tree blocks including the luma position (xCtb + CtbSizeY, yCtb + CtbSizeY) belong to different slices and "loop_filter_across_slices_enabled_flag" is equal to 0, "clipBotRightFlag" can be set to 1.

[0117]

[0133] According to some embodiments of the present disclosure, the loop filter may include sample adaptive offset (SAO) operations. When SAO is partially disabled on the first side (e.g., left, right, top, or bottom) of the virtual boundary, if the SAO for the pixels on the first side requires pixels on the second side (e.g., right, left, bottom, or top), the application of SAO for the pixels on the first side can be skipped. By doing so, the pixels on the second side become unavailable.

[0118]

[0134] According to some embodiments of the present disclosure, in the CTB correction process, for all sample positions (xS i , yS j ) and (xY i , yY j ) where i = 0,...,(nCtbSw - 1) and j = 0,...,(nCtbSh - 1), the following operations are applicable.

[0119]

[0135] If one or more of the following conditions are true, the variable "saoPicture[xS i [yS j " can be left unmodified, and the following conditions are: namely, (1) the variable "SaoTypeIdx[cIdx][rx][ry]" is equal to 0, (2) "VirtualBoundariesLoopfilterDisabled" is equal to 1, "xS j " is equal to ((VirtualBoundariesPosX[n] / scaleWidth)-1) for any n = 0,...,(VirtualBoundariesNumVer-1), "SaoTypeIdx[cIdx][rx][ry]" is equal to 2, and the variable "SaoEoClass[cIdx][rx][ry]" is not equal to 1, (3) "VirtualBoundariesLoopfilterDisabled" is equal to 1, "xS j " is equal to (VirtualBoundariesPosX[n] / scaleWidth) for any n = 0,...,(VirtualBoundariesNumVer-1), "SaoTypeIdx[cIdx][rx][ry]" is equal to 2, and "SaoEoClass[cIdx][rx][ry]" is not equal to 1, (4) "VirtualBoundariesLoopfilterDisabled" is equal to 1, "yS j " is equal to ((VirtualBoundariesPosY[n] / scaleHeight)-1) for any n = 0,...,(VirtualBoundariesNumHor-1), "SaoTypeIdx[cIdx][rx][ry]" is equal to 2, and "SaoEoClass[cIdx][rx][ry]" is not equal to 0, (5) "VirtualBoundariesLoopfilterDisabled" is equal to 1, "yS j" is equal to (VirtualBoundariesPosY[n] / scaleHeight) for any n = 0, ..., (VirtualBoundariesNumHor-1), "SaoTypeIdx[cIdx][rx][ry]" is equal to 2, and "SaoEoClass[cIdx][rx][ry]" is not equal to 0. (6) "VirtualBoundariesLoopfilterDisabled" is equal to 2, "xS j " is equal to ((VirtualBoundariesPosX[n] / scaleWidth)-1) for any n = 0, ..., (VirtualBoundariesNumVer-1), "SaoTypeIdx[cIdx][rx][ry]" is equal to 2, and "SaoEoClass[cIdx][rx][ry]" is not equal to 1. (7) "VirtualBoundariesLoopfilterDisabled" is equal to 3, "xS j " is equal to (VirtualBoundariesPosX[n] / scaleWidth) for any n = 0, ..., (VirtualBoundariesNumVer-1), "SaoTypeIdx[cIdx][rx][ry]" is equal to 2, and "SaoEoClass[cIdx][rx][ry]" is not equal to 1. (8) "VirtualBoundariesLoopfilterDisabled" is equal to 2, "yS j " is equal to ((VirtualBoundariesPosY[n] / scaleHeight)-1) for any n = 0, ..., (VirtualBoundariesNumHor-1), "SaoTypeIdx[cIdx][rx][ry]" is equal to 2, and "SaoEoClass[cIdx][rx][ry]" is not equal to 0, or (9) "VirtualBoundariesLoopfilterDisabled" is equal to 3, "yS jIt is equal to (VirtualBoundariesPosY[n] / scaleHeight) for any n = 0, ..., (VirtualBoundariesNumHor-1), "SaoTypeIdx[cIdx][rx][ry]" is equal to 2, and "SaoEoClass[cIdx][rx][ry]" is equal to 0.

[0120]

[0136] According to some embodiments of the present disclosure, the loop filter may include a deblocking filter. In some embodiments, when the deblocking filter is partially disabled on the first side of the virtual boundary (e.g., the left side, the right side, the upper side, or the lower side), the pixels on the first side can be skipped from being processed by the deblocking filter, and the second side of the virtual boundary (e.g., the right side, the left side, the lower side, or the upper side) can be processed by the deblocking filter. In some embodiments, if "VirtualBoundariesLoopfilterDisabled" is not false (e.g., having a value of 0), the deblocking filter can be completely disabled, and the pixels on both sides of the virtual boundary can be skipped from being processed by the deblocking filter.

[0121]

[0137] According to some embodiments of the present disclosure, FIGS. 13 to 15 show flowcharts of exemplary methods 1300 to 1500. Methods 1300 to 1500 may be executed by at least one processor (e.g., processor 402 of FIG. 4) associated with a video encoder (e.g., the encoder described in connection with FIGS. 2A to 2B) or a video decoder (e.g., the decoder described in connection with FIGS. 3A to 3B). According to some embodiments, methods 1300 to 1500 may be implemented as a computer program product including computer-executable instructions (e.g., program code) (e.g., embodied by a computer-readable medium) executed by a computer (e.g., device 400 of FIG. 4). According to some embodiments, methods 1300 to 1500 may be implemented as a hardware product (e.g., memory 404 of FIG. 4) storing computer-executable instructions (e.g., program instructions in memory 404 of FIG. 4), and the hardware product may be an independent part or an integral part of the computer.

[0122]

[0138] As an example, FIG. 13 shows a flowchart of an exemplary process 1300 for video processing according to some embodiments of the present disclosure. For example, process 1300 may be executed by an encoder.

[0123]

[0139] In step 1302, in response to a processor (e.g., processor 402 of FIG. 4) receiving a video sequence (e.g., video sequence 202 of FIGS. 2A to 2B), the processor may encode first flag data (e.g., "gdr_enabled_flag" illustrated and described with respect to FIGS. 10 to 12) within a parameter set (e.g., SPS) associated with the video sequence. The first flag data may indicate whether progressive decoding refresh (GDR) is enabled or disabled with respect to the video sequence.

[0124]

[0140] In step 1304, if the first flag data indicates that the GDR is disabled for the video sequence (for example, "gdr_enabled_flag" is false), the processor can encode the picture header associated with the pictures in the video sequence to indicate that the picture is a non-GDR picture. As used herein, a GDR picture can refer to a picture that includes both a clean area and a dirty area. By way of example, the clean area can be any of the clean areas 516, 520, 524, or 528 illustrated and described in FIG. 5, and the dirty area can be any of the dirty areas 514, 518, 522, or 526 illustrated and described in FIG. 5. A non-GDR picture of the present disclosure can refer to a picture that does not include a clean area or does not include a dirty area.

[0125]

[0141] In some embodiments, to encode the picture header, the processor can disable encoding the second flag data (for example, "gdr_pic_flag" illustrated and described with respect to FIGS. 10-12) within the picture header. The second flag data can indicate whether the picture is a GDR picture.

[0126]

[0142] In some embodiments, to encode the picture header, the processor can encode the second flag data (for example, "gdr_pic_flag" illustrated and described with respect to FIGS. 10-12) within the picture header, and the second flag data indicates that the picture is a non-GDR picture (for example, "gdr_pic_flag" is false).

[0127]

[0143] In step 1306, the processor can encode the non-GDR pictures.

[0128]

[0144] According to some embodiments of the present disclosure, when first flag data indicates that GDR is enabled for a video sequence (e.g., "gdr_enabled_flag" is true), the processor can be enabled to encode second flag data (e.g., "gdr_pic_flag" illustrated and described with respect to FIGS. 10-12) in the picture header, and the second flag data indicates whether the picture is a GDR picture (e.g., "gdr_pic_flag" is true). Then, the processor can encode the picture.

[0129]

[0145] According to some embodiments of the present disclosure, when first flag data indicates that GDR is enabled for a video sequence (e.g., "gdr_enabled_flag" is true) and second flag data indicates that the picture is a GDR picture (e.g., "gdr_pic_flag" is true), the processor can divide the picture into a first region (e.g., a clean region such as any of the clean regions 516, 520, 524, or 528 illustrated and described with respect to FIG. 5) and a second region (e.g., a dirty region such as any of the dirty regions 514, 518, 522, or 526 illustrated and described with respect to FIG. 5) using a virtual boundary. For example, the first region can be on a first side (e.g., left side, right side, upper side, or lower side) of the virtual boundary, and the second region can be on a second side (e.g., right side, left side, lower side, or upper side) of the virtual boundary. Then, when filtering of a first pixel uses information of a second pixel in the second region, the processor can disable (e.g., loop filter 232 illustrated and described with respect to FIGS. 2B and 3B) or apply a loop filter to the first pixel using only information of pixels in the first region. Thereafter, the processor can apply a loop filter to pixels in the second region using pixels in at least one of the first region or the second region.

[0130]

[0146] As an example, FIG. 14 shows a flowchart of another exemplary process 1400 for video processing according to some embodiments of the present disclosure. For example, process 1400 may be executed by a decoder.

[0131]

[0147] In step 1402, in response to a processor (e.g., processor 402 of FIG. 4) receiving a video bitstream (e.g., video bitstream 228 of FIGS. 3A - 3B), the processor can decode first flag data (e.g., the "gdr_enabled_flag" illustrated and described with respect to FIGS. 10 - 12) within a parameter set (e.g., SPS) related to a sequence of the video bitstream (e.g., the video sequence to be decoded). The first flag data can indicate whether progressive decoding refresh (GDR) is enabled or disabled for the video sequence.

[0132]

[0148] In step 1404, if the first flag data indicates that GDR is disabled for the sequence (e.g., the "gdr_enabled_flag" is false), the processor can decode a picture header related to a picture within the sequence, and the picture header indicates that the picture is a non - GDR picture.

[0133]

[0149] In some embodiments, to decode the picture header, the processor can disable decoding of second flag data (e.g., the "gdr_pic_flag" illustrated and described with respect to FIGS. 10 - 12) within the picture header and determine that the picture is a non - GDR picture, where the second flag data can indicate whether the picture is a GDR picture. In some embodiments, to decode the picture header, the processor can decode the second flag data within the picture header, and the second flag data can indicate that the picture is a non - GDR picture (e.g., the "gdr_pic_flag" is false).

[0134]

[0150] In step 1406, the processor can decode non-GDR pictures.

[0135]

[0151] According to some embodiments of the present disclosure, if first flag data indicates that GDR is enabled for a video sequence (e.g., "gdr_enabled_flag" is true), the processor can decode second flag data in the picture header (e.g., "gdr_pic_flag" illustrated and described with respect to FIGS. 10-12) to determine whether the picture is a GDR picture (e.g., whether "gdr_pic_flag" is true) based on the second flag data. Then, the processor can decode the picture.

[0136]

[0152] According to some embodiments of the present disclosure, if first flag data indicates that GDR is enabled for a video sequence and second flag data indicates that the picture is a GDR picture, the processor can use a virtual boundary to divide the picture into a first region (e.g., a clean region such as any of the clean regions 516, 520, 524, or 528 illustrated and described in FIG. 5) and a second region (e.g., a dirty region such as any of the dirty regions 514, 518, 522, or 526 illustrated and described in FIG. 5). For example, the first region can be on the first side (e.g., left side, right side, upper side, or lower side) of the virtual boundary, and the second region can be on the second side (e.g., right side, left side, lower side, or upper side) of the virtual boundary. Then, the processor can disable the loop filter in the first region (e.g., the loop filter 232 illustrated and described with respect to FIGS. 2B and 3B) or apply the loop filter to the first pixel using only the information of the pixels in the first region. Thereafter, the processor can apply the loop filter to the pixels in the second region using the pixels in at least one of the first region or the second region.

[0137]

[0153] As an example, FIG. 15 shows a flowchart of yet another exemplary process 1500 for video processing according to some embodiments of the present disclosure. For example, process 1500 may be performed by an encoder or a decoder.

[0138]

[0154] In step 1502, in response to the processor (e.g., processor 402 of FIG. 4) receiving a picture of video, the processor determines whether the picture is a Gradual Decoding Refresh (GDR) picture based on flag data associated with the picture. For example, the flag data can be at the sequence level (e.g., stored in the SPS) or at the picture level (e.g., stored in the picture header).

[0139]

[0155] In step 1504, based on the determination that the picture is a GDR picture, the processor can determine a first region (e.g., a clean region such as any of the clean regions 516, 520, 524, or 528 illustrated and described in FIG. 5) and a second region (e.g., a dirty region such as any of the dirty regions 514, 518, 522, or 526 illustrated and described in FIG. 5) of the picture using virtual boundaries. For example, the first region can include a left region or an upper region, and the second region can include a right region or a lower region.

[0140]

[0156] In step 1506, when the filtering of the first pixel uses the information of the second pixel in the second region, the processor can disable the loop filter (e.g., loop filter 232 illustrated and described with respect to FIGS. 2B and 3B) for the first pixel in the first region, or apply the loop filter for the first pixel using only the information of the pixels in the first region. According to some embodiments, the processor can disable at least one of the sample adaptive offset, deblocking filter, or adaptive loop filter in the first region, or apply at least one of the sample adaptive offset, deblocking filter, or adaptive loop filter using the pixels in the first region.

[0141]

[0157] In step 1508, the loop filter for the pixels in the second region can be applied using the pixels in at least one of the first region or the second region.

[0142]

[0158] According to some embodiments of the present disclosure, the processor can encode or decode flag data in at least one of the sequence parameter set (SPS), picture parameter set (PPS), or picture header. As illustrated and described with respect to FIGS. 10 - 12, for example, the processor can encode the flag data as the GDR indication flag "gdr_enabled_flag", the GDR indication flag "gdr_pic_flag", or both.

[0143]

[0159] According to some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by a device (such as the encoder and decoder of the present disclosure) to perform the above-described method. Examples of common forms of non-transitory media include, for example, floppy (registered trademark) disks, flexible disks, hard disks, solid state drives, magnetic tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media having a pattern of holes, RAM, PROM, and EPROM, FLASH (registered trademark)-EPROM or any other flash memory, NVRAM, caches, registers, any other memory chips or cartridges, and networked versions thereof. The device can include one or more processors (CPUs), an input / output interface, a network interface, and / or a memory.

[0144]

[0160] The embodiments can be further described using the following clauses: 1. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by at least one processor of a device to cause the device to perform a method, the method comprising: encoding first flag data in a parameter set associated with a video sequence in response to receiving the video sequence, the first flag data indicating whether progressive decoding refresh (GDR) is enabled or disabled for the video sequence; encoding a picture header associated with a picture in the video sequence to indicate that the picture is a non-GDR picture when the first flag data indicates that GDR is disabled for the video sequence; and encoding non-GDR pictures comprising the non-transitory computer-readable medium. 2. Encoding the picture header includes: disabling encoding of second flag data in the picture header; The second flag data is a non-transitory computer-readable medium as described in clause 1, indicating whether the picture is a GDR picture. 3. Encoding the picture header includes encoding the second flag data in the picture header, where the second flag data is a non-transitory computer-readable medium as described in clause 1, indicating that the picture is a non-GDR picture. 4. A set of instructions executable by at least one processor of the device enables encoding the second flag data in the picture header when the first flag data indicates that GDR is enabled for the video sequence, where the second flag data indicates whether the picture is a GDR picture, and further causes the device to encode the picture using a non-transitory computer-readable medium as described in clause 1. 5. A set of instructions executable by at least one processor of the device when the first flag data indicates that GDR is enabled for the video sequence and the second flag data indicates that the picture is a GDR picture, divides the picture into a first region and a second region using a virtual boundary, disables the loop filter for the first pixel in the first region or applies the loop filter for the first pixel using only the information of the pixels in the first region when the filtering of the first pixel uses the information of the second pixel in the second region, and further causes the device to apply a loop filter for the pixels in the second region using the pixels in at least one of the first region or the second region using a non-transitory computer-readable medium as described in any one of clauses 2 to 4. 6. The second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture, using a non-transitory computer-readable medium as described in any one of clauses 2 to 5. 7. The second flag data is a non-transitory computer-readable medium according to any one of clauses 1 to 6, having a value of "1" indicating that the picture is a GDR picture, or a value of "0" indicating that the picture is a non-GDR picture. 8. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by at least one processor of the device to cause the device to perform a method, the method comprising: decoding first flag data in a parameter set associated with a sequence of video bitstreams in response to receiving the video bitstream, the first flag data indicating whether progressive decoding refresh (GDR) is enabled or disabled for the video sequence; when the first flag data indicates that GDR is disabled for the sequence, decoding a picture header associated with a picture in the sequence, the picture header indicating that the picture is a non-GDR picture, and decoding the non-GDR picture The non-transitory computer-readable medium includes. 9. Decoding the picture header includes: disabling decoding of second flag data in the picture header and determining that the picture is a non-GDR picture, The second flag data is a non-transitory computer-readable medium according to clause 8, indicating whether the picture is a GDR picture. 10. Decoding the picture header includes: decoding second flag data in the picture header, The second flag data is a non-transitory computer-readable medium according to clause 8, indicating that the picture is a non-GDR picture. 11. The set of instructions executable by at least one processor of the device includes: If the first flag data indicates that the GDR is enabled for the video sequence, decrypt the second flag data in the picture header, and determine whether the picture is a GDR picture based on the second flag data, and decode the picture The non-transitory computer-readable medium according to clause 8, which further causes the device to perform. 12. A set of instructions executable by at least one processor of the device is If the first flag data indicates that the GDR is enabled for the video sequence and the second flag data indicates that the picture is a GDR picture, divide the picture into a first region and a second region using a virtual boundary, When the filtering of the first pixel uses the information of the second pixel in the second region, disable the loop filter for the first pixel in the first region, or apply the loop filter for the first pixel using only the information of the pixels in the first region, and Apply a loop filter to the pixels in the second region using the pixels in at least one of the first region or the second region The non-transitory computer-readable medium according to any one of clauses 8 to 11, which further causes the device to perform. 13. The second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. The non-transitory computer-readable medium according to any one of clauses 9 to 12. 14. The second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. The non-transitory computer-readable medium according to any one of clauses 8 to 13. 15. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by at least one processor of the device to cause the device to perform a method, the method being In response to receiving a picture of an image, determining whether the picture is a progressive decode refresh (GDR) picture based on flag data associated with the picture; Based on the determination that the picture is a GDR picture, determining a first region and a second region of the picture using a virtual boundary; When filtering of a first pixel uses information of a second pixel in a second region, disabling a loop filter for the first pixel in the first region, or applying a loop filter for the first pixel using only information of pixels in the first region, and Applying a loop filter for pixels in the second region using pixels in at least one of the first region or the second region A non-transitory computer-readable medium including the above. 16. When filtering of a first pixel uses information of a second pixel in a second region, disabling a loop filter for the first pixel in the first region, or applying a loop filter for the first pixel using only information of pixels in the first region is Disabling at least one of a sample adaptive offset, a deblocking filter, or an adaptive loop filter in the first region, or Applying at least one of a sample adaptive offset, a deblocking filter, or an adaptive loop filter using pixels in the first region The non-transitory computer-readable medium according to clause 15, including the above. 17. A set of instructions executable by at least one processor of a device is Encoding or decoding flag data in at least one of a sequence parameter set (SPS), a picture parameter set (PPS), or a picture header, to cause the device to further perform the above. The non-transitory computer-readable medium according to clause 15 or 16. 18. The non - transitory computer - readable medium according to any one of clauses 15 to 17, wherein the first region includes a left - hand region or an upper region, and the second region includes a right - hand region or a lower region. 19. The non - transitory computer - readable medium according to any one of clauses 15 to 18, wherein the flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non - GDR picture. 20. An apparatus, comprising a memory configured to store a set of instructions, and one or more processors communicatively coupled to the memory, the one or more processors being configured to, in response to receiving a video sequence, encode first flag data in a parameter set associated with the video sequence, the first flag data indicating whether progressive decoding refresh (GDR) is enabled or disabled for the video sequence, when the first flag data indicates that GDR is disabled for the video sequence, encode a picture header associated with a picture in the video sequence to indicate that the picture is a non - GDR picture, and encode non - GDR pictures by executing the set of instructions to cause the apparatus to perform. 21. Encoding the picture header includes disabling encoding of second flag data in the picture header, where the second flag data indicates whether the picture is a GDR picture, for the apparatus according to clause 20. 22. Encoding the picture header includes encoding second flag data in the picture header, where the second flag data indicates that the picture is a non - GDR picture, for the apparatus according to clause 20. 23. The one or more processors are When the first flag data indicates that GDR is enabled for the video sequence, enabling encoding the second flag data in the picture header, where the second flag data indicates whether the picture is a GDR picture, and encoding the picture The apparatus according to clause 20, further configured to execute a set of instructions to cause the apparatus to perform. 24. One or more processors When the first flag data indicates that GDR is enabled for the video sequence and the second flag data indicates that the picture is a GDR picture, dividing the picture into a first region and a second region using a virtual boundary, Disabling the loop filter for the first pixel in the first region or applying the loop filter for the first pixel using only the information of the pixels in the first region when the filtering of the first pixel uses the information of the second pixel in the second region, and Applying a loop filter for the pixels in the second region using the pixels in at least one of the first region or the second region The apparatus according to any one of clauses 21 to 23, further configured to execute a set of instructions to cause the apparatus to perform. 25. The second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture, for the apparatus according to any one of clauses 21 to 24. 26. The second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture, for the apparatus according to any one of clauses 20 to 25. 27. An apparatus A memory configured to store a set of instructions, and One or more processors communicatively coupled to the memory, where the one or more processors In response to receiving a video bitstream, decoding first flag data in a parameter set related to the sequence of the video bitstream, the first flag data indicating whether progressive decoding refresh (GDR) is enabled or disabled for the video sequence, When the first flag data indicates that GDR is disabled for the sequence, decoding a picture header related to a picture in the sequence, the picture header indicating that the picture is a non-GDR picture, and Decoding the non-GDR picture An apparatus configured to execute a set of instructions to cause the apparatus to perform the above. 28. Decoding the picture header includes Disabling decoding of second flag data in the picture header and determining that the picture is a non-GDR picture The apparatus according to clause 27, wherein the second flag data indicates whether the picture is a GDR picture. 29. Decoding the picture header includes Decoding second flag data in the picture header, The apparatus according to clause 27, wherein the second flag data indicates that the picture is a non-GDR picture. 30. One or more processors are further configured to When the first flag data indicates that GDR is enabled for the video sequence, decode second flag data in the picture header and determine whether the picture is a GDR picture based on the second flag data, and Decode the picture The apparatus according to clause 27, which is further configured to execute a set of instructions to cause the apparatus to perform the above. 31. One or more processors are When first flag data indicates that GDR is enabled for a video sequence and second flag data indicates that a picture is a GDR picture, dividing the picture into a first region and a second region using a virtual boundary, When filtering of a first pixel uses information of a second pixel within a second region, disabling a loop filter for the first pixel within the first region or applying a loop filter for the first pixel using only information of pixels within the first region, and Applying a loop filter for a pixel within the second region using pixels in at least one of the first region or the second region The apparatus according to any one of clauses 28 to 30, further configured to execute a set of instructions to cause the apparatus to perform the above. 32. The apparatus according to any one of clauses 28 to 31, wherein the second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. 33. The apparatus according to any one of clauses 27 to 32, wherein the second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture. 34. An apparatus, a memory configured to store a set of instructions, and one or more processors communicatively coupled to the memory, the one or more processors being configured to: in response to receiving a picture of a video, determine whether the picture is a gradual decoding refresh (GDR) picture based on flag data associated with the picture; based on a determination that the picture is a GDR picture, determine a first region and a second region of the picture using a virtual boundary; When filtering the first pixel uses information of a second pixel in a second region, disabling a loop filter for the first pixel in the first region, or applying a loop filter for the first pixel using only information of pixels in the first region, and applying a loop filter for a pixel in the second region using pixels in at least one of the first region or the second region A device configured to execute a set of instructions to cause the device to perform the above. 35. When filtering the first pixel uses information of a second pixel in a second region, disabling a loop filter for the first pixel in the first region, or applying a loop filter for the first pixel using only information of pixels in the first region is disabling at least one of a sample adaptive offset, a deblocking filter, or an adaptive loop filter in the first region, or applying at least one of a sample adaptive offset, a deblocking filter, or an adaptive loop filter using pixels in the first region The device according to clause 34, including the above. 36. One or more processors are further configured to execute a set of instructions to cause the device to encode or decode flag data in at least one of a sequence parameter set (SPS), a picture parameter set (PPS), or a picture header The device according to clause 34 or 35, further configured as above. 37. The device according to any one of clauses 34 to 36, wherein the first region includes a left region or an upper region, and the second region includes a right region or a lower region. 38. The device according to any one of clauses 34 to 37, wherein the flag data has a value of "1" indicating that the picture is a GDR picture, or a value of "0" indicating that the picture is a non-GDR picture. 39. In response to receiving a video sequence, encoding first flag data in a parameter set associated with the video sequence, the first flag data indicating whether progressive decoding refresh (GDR) is enabled or disabled for the video sequence, When the first flag data indicates that GDR is disabled for the video sequence, encoding a picture header associated with a picture in the video sequence to indicate that the picture is a non-GDR picture, and Encoding the non-GDR picture A method including. 40. Encoding the picture header Includes disabling encoding of second flag data in the picture header, The second flag data represents whether the picture is a GDR picture, the method according to clause 39. 41. Encoding the picture header Includes encoding second flag data in the picture header, The second flag data represents that the picture is a non-GDR picture, the method according to clause 39. 42. When the first flag data indicates that GDR is enabled for the video sequence, enabling encoding of second flag data in the picture header, the second flag data indicating whether the picture is a GDR picture, and Encoding the picture The method according to clause 39, further including. 43. When the first flag data indicates that GDR is enabled for the video sequence and the second flag data indicates that the picture is a GDR picture, dividing the picture into a first region and a second region using a virtual boundary, When filtering the first pixel uses information of a second pixel within a second region, disabling the loop filter for the first pixel within the first region, or applying the loop filter for the first pixel using only information of pixels within the first region, and applying a loop filter for a pixel within the second region using a pixel in at least one of the first region or the second region The method according to any one of clauses 40 to 42, further comprising. 44. The second flag data has a value of "1" indicating that the picture is a GDR picture, or a value of "0" indicating that the picture is a non-GDR picture, according to any one of clauses 40 to 43. The method described. 45. The second flag data has a value of "1" indicating that the picture is a GDR picture, or a value of "0" indicating that the picture is a non-GDR picture, according to any one of clauses 39 to 44. The method described. 46. In response to receiving a video bitstream, decoding first flag data in a parameter set related to the sequence of the video bitstream, where the first flag data indicates whether progressive decoding refresh (GDR) is enabled or disabled for the video sequence, When the first flag data indicates that GDR is disabled for the sequence, decoding a picture header related to a picture in the sequence, where the picture header indicates that the picture is a non-GDR picture, and Decoding non-GDR pictures A method including. 47. Decoding a picture header is Disabling decoding of second flag data in the picture header and determining that the picture is a non-GDR picture Including, the second flag data indicates whether the picture is a GDR picture, according to the method described in clause 46. 48. Decoding a picture header is including decrypting second flag data in a picture header, The second flag data is the method described in clause 46 indicating that the picture is a non-GDR picture. 49. If the first flag data indicates that GDR is enabled for a video sequence, decrypt the second flag data in the picture header, determine whether the picture is a GDR picture based on the second flag data, and decrypt the picture The method described in clause 46 further includes. 50. If the first flag data indicates that GDR is enabled for a video sequence and the second flag data indicates that the picture is a GDR picture, divide the picture into a first region and a second region using a virtual boundary, When the filtering of the first pixel uses the information of the second pixel in the second region, disable the loop filter for the first pixel in the first region or apply the loop filter for the first pixel using only the information of the pixels in the first region, and Apply a loop filter to the pixels in the second region using the pixels in at least one of the first region or the second region The method according to any one of clauses 47 to 49 further includes. 51. The second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture, according to the method described in any one of clauses 47 to 50. 52. The second flag data has a value of "1" indicating that the picture is a GDR picture or a value of "0" indicating that the picture is a non-GDR picture, according to the method described in any one of clauses 46 to 51. 53. In response to receiving a picture of a video, determine whether the picture is a progressive decoding refresh (GDR) picture based on the flag data related to the picture, Based on the determination that the picture is a GDR picture, determining a first region and a second region of the picture using a virtual boundary, When filtering the first pixel uses information of a second pixel in the second region, disabling the loop filter for the first pixel in the first region, or applying the loop filter for the first pixel using only information of pixels in the first region, and Applying a loop filter for a pixel in the second region using pixels in at least one of the first region or the second region A method including the above. 54. When filtering the first pixel uses information of a second pixel in the second region, disabling the loop filter for the first pixel in the first region, or applying the loop filter for the first pixel using only information of pixels in the first region is Disabling at least one of a sample adaptive offset, a deblocking filter, or an adaptive loop filter in the first region, or applying at least one of a sample adaptive offset, a deblocking filter, or an adaptive loop filter using pixels in the first region The method according to clause 53, including the above. 55. Further including encoding or decoding flag data in at least one of a sequence parameter set (SPS), a picture parameter set (PPS), or a picture header The method according to clause 53 or 54, further including the above. 56. The method according to any one of clauses 53 to 55, wherein the first region includes a left region or an upper region, and the second region includes a right region or a lower region. 57. The method according to any one of clauses 53 to 56, wherein the flag data has a value of "1" indicating that the picture is a GDR picture, or a value of "0" indicating that the picture is a non-GDR picture.

[0145]

[0161] Relative terms in this specification such as "first" and "second" are merely used to distinguish one entity or operation from another entity or operation, and it should be noted that no actual relationship or order between these entities or operations is required or implied. Further, the words "comprising", "having", "containing", "including" and other similar forms are of equal meaning, and elements or groups of elements following any of these words are not meant to be a limiting enumeration of such elements or groups of elements, or are not meant to be limited only to the enumerated elements or groups of elements, and are intended to be open-ended.

[0146]

[0162] As used in this specification, unless otherwise specifically stated, the term "or" includes all possible combinations except when it is infeasible. For example, if it is stated that a component can include A or B, then, unless otherwise specifically stated or infeasible, the component can include A or B or A and B. As a second example, if it is stated that a component can include A, B or C, then, unless otherwise specifically stated or infeasible, the component can include A, B or C, or A and B, A and C or B and C, or A and B and C.

[0147]

[0163] It is understood that the above-described embodiments can be implemented by hardware or software (program code) or a combination of hardware and software. When implemented by software, it can be stored in the above-described computer-readable medium. The software can perform the method of the present disclosure when executed by a processor. The computing units and other functional units described in the present disclosure can be implemented by hardware or software or a combination of hardware and software. Those skilled in the art can also understand that a plurality of the above-described modules / units can be combined into one module / unit, and each of the above-described modules / units can be further divided into a plurality of sub-modules / sub-units.

[0148]

[0164] In the above specification, embodiments have been described with reference to many specific details that may vary from implementation to implementation. Specific adaptations and modifications of the above embodiments may be made. From the consideration of this specification and the practice of the disclosure disclosed herein, other embodiments may become apparent to those skilled in the art. This specification and examples are intended to be considered only as examples, and the true scope and spirit of the disclosure are indicated by the appended claims. Also, the arrangement of steps shown in the figures is for illustrative purposes only and is not intended to be limited to any particular arrangement of steps. Therefore, those skilled in the art can understand that these steps can be performed in a different order while implementing the same method.

[0149]

[0165] In the drawings and this specification, exemplary embodiments have been disclosed. However, many variations and modifications to these embodiments can be made. Therefore, even if specific terms are employed, they are used only in a general descriptive sense and not for the purpose of limitation.

Claims

1. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by at least one processor of the device to cause the device to perform a method, the method comprising: encoding first flag data within a set of parameters associated with the video sequence in response to receiving the video sequence, the first flag data indicating whether progressive decoding refresh (GDR) is enabled or disabled for the video sequence; encoding a picture header associated with a picture within the video sequence; encoding the picture; wherein the picture header includes second flag data; the second flag data indicates whether the picture is a GDR picture, and when the first flag data indicates that GDR is disabled for the video sequence, the second flag data indicates that the picture is a non-GDR picture. A non-transitory computer-readable medium.

2. The non-transitory computer-readable medium of claim 1, wherein when the first flag data indicates that GDR is enabled for the video sequence, the second flag data has a first value indicating that the picture is a GDR picture or a second value indicating that the picture is a non-GDR picture.

3. The set of instructions executable by the at least one processor of the device further comprises: when the first flag data indicates that GDR is enabled for the video sequence and the second flag data indicates that the picture is a GDR picture, dividing the picture into a first region and a second region using a virtual boundary; disabling a loop filter for a first pixel in the first region if filtering of the first pixel in the first region uses information of a second pixel in the second region, or applying the loop filter for the first pixel using only information of pixels in the first region; and applying the loop filter for pixels in the second region using pixels in at least one of the first region or the second region. The non - transitory computer - readable medium according to claim 1, further causing the device to perform

4. The non - transitory computer - readable medium according to claim 1, wherein the first flag data has a first value indicating that the GDR is enabled for the video sequence or a second value indicating that the GDR is disabled for the video sequence.

5. A non - transitory computer - readable medium storing a set of instructions, the set of instructions being executable by at least one processor of the device to cause the device to perform a method, the method comprising: Determining whether a picture is a Gradual Decoding Refresh (GDR) picture based on flag data associated with the picture in response to receiving the picture of the video; Determining a first region and a second region of the picture using a virtual boundary based on the determination that the picture is the GDR picture; Disabling a loop filter for a first pixel in the first region if filtering of the first pixel in the first region uses information of a second pixel in the second region, or applying the loop filter for the first pixel using only information of pixels in the first region, and Applying the loop filter for pixels in the second region using pixels in at least one of the first region or the second region The non - transitory computer - readable medium comprising.

6. Disabling the loop filter for the first pixel in the first region if filtering of the first pixel uses the information of the second pixel in the second region, or applying the loop filter for the first pixel using only information of pixels in the first region, Comprises disabling at least one of a sample - adaptive offset, a de - blocking filter, or an adaptive loop filter in the first region, or Applying at least one of the sample - adaptive offset, the de - blocking filter, or the adaptive loop filter using the pixels in the first region The non - transitory computer - readable medium according to claim 5, comprising.

7. The set of instructions executable by the at least one processor of the device comprises Encoding or decoding the flag data in at least one of a sequence parameter set (SPS), a picture parameter set (PPS), or a picture header The non-transitory computer-readable medium according to claim 5, further causing the device to perform the above

8. The non-transitory computer-readable medium according to claim 5, wherein the first region includes a left region or an upper region, and the second region includes a right region or a lower region

9. The non-transitory computer-readable medium according to claim 5, wherein the flag data has a value of "1" indicating that the picture is the GDR picture, or a value of "0" indicating that the picture is a non-GDR picture

10. A device, comprising A memory configured to store a set of instructions, and One or more processors communicatively coupled to the memory, the one or more processors being configured to In response to receiving a video sequence, encoding first flag data in a parameter set related to the video sequence, the first flag data indicating whether progressive decoding refresh (GDR) is enabled or disabled for the video sequence, Encoding a picture header related to a picture in the video sequence, and Encoding the picture The device is configured to execute the set of instructions to perform the above, The picture header includes second flag data, The second flag data indicates whether the picture is a GDR picture. When the first flag data indicates that the GDR is disabled for the video sequence, the second flag data indicates that the picture is a non-GDR picture. A device

11. The device according to claim 10, wherein when the first flag data indicates that the GDR is enabled for the video sequence, the second flag data has a first value indicating that the picture is the GDR picture, or a second value indicating that the picture is the non-GDR picture

12. The one or more processors are configured to When the first flag data indicates that the GDR is enabled for the video sequence, and the second flag data indicates that the picture is a GDR picture, the picture is divided into a first region and a second region using a virtual boundary. When filtering of a first pixel in the first region uses information of a second pixel in the second region, disabling the loop filter for the first pixel, or applying the loop filter for the first pixel using only information of pixels in the first region, and applying the loop filter for pixels in the second region using pixels in at least one of the first region or the second region The apparatus according to claim 10, further configured to execute the set of instructions to cause the apparatus to perform the above.

13. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by at least one processor of the apparatus to cause the apparatus to perform a method, the method comprising: In response to receiving a video bitstream, decoding first flag data in a parameter set associated with a sequence of the video bitstream, the first flag data indicating whether progressive decoding refresh (GDR) is enabled or disabled for the sequence. Decoding a picture header associated with a picture in the sequence, and Decoding the picture including The decoded picture header includes second flag data. The second flag data indicates whether the picture is a GDR picture. When the first flag data indicates that the GDR is disabled for the sequence, the second flag data indicates that the picture is a non-GDR picture. A non-transitory computer-readable medium.

14. The non-transitory computer-readable medium according to claim 13, wherein when the first flag data indicates that the GDR is enabled for the sequence, the second flag data has a first value indicating that the picture is a GDR picture or a second value indicating that the picture is a non-GDR picture.

15. The set of instructions executable by the at least one processor of the machine comprises when the first flag data indicates that the GDR is enabled with respect to the sequence and the second flag data indicates that the picture is a GDR picture, dividing the picture into a first region and a second region using a virtual boundary, disabling a loop filter for a first pixel in the first region when filtering of the first pixel in the first region uses information of a second pixel in the second region, or applying the loop filter for the first pixel using only information of pixels in the first region, and applying the loop filter for pixels in the second region using pixels in at least one of the first region or the second region to further cause the machine to perform, the non-transitory computer-readable medium of claim 13. **Claim 16** A machine comprising a memory configured to store a set of instructions, and one or more processors communicatively coupled to the memory, the one or more processors being configured to decode first flag data within a set of parameters associated with a sequence of the video bitstream in response to receiving the video bitstream, the first flag data indicating whether progressive decoding refresh (GDR) is enabled or disabled with respect to the sequence, decode a picture header associated with a picture within the sequence, and decode the picture by executing the set of instructions to cause the machine to perform, the picture header to be decoded includes second flag data, the second flag data indicates whether the picture is a GDR picture, and when the first flag data indicates that the GDR is disabled with respect to the sequence, the second flag data indicates that the picture is a non-GDR picture. **Claim 17** When the first flag data indicates that the GDR is enabled for the sequence, the second flag data has a first value indicating that the picture is a GDR picture or a second value indicating that the picture is a non-GDR picture, the device according to claim 16.

18. The set of instructions executable by the one or more processors of the device is When the first flag data indicates that the GDR is enabled for the sequence and the second flag data indicates that the picture is a GDR picture, dividing the picture into a first region and a second region using a virtual boundary, When filtering of a first pixel in the first region uses information of a second pixel in the second region, disabling a loop filter for the first pixel or applying the loop filter for the first pixel using only information of pixels in the first region, and Applying the loop filter for pixels in the second region using pixels in at least one of the first region or the second region to further cause the device to perform, the device according to claim 16.

19. A method for storing a bitstream of a video sequence, the method comprising: Receiving a video sequence, Encoding one or more pictures of the video sequence, Generating a bitstream, and Storing the bitstream in a non-transitory computer-readable storage medium including, the encoding comprising: Encoding first flag data related to the video sequence, the first flag data indicating whether progressive decoding refresh (GDR) is enabled or disabled for the video sequence, Encoding second flag data related to a picture in the video sequence including, The second flag data indicates whether the picture is a GDR picture, and when the first flag data indicates that the GDR is disabled for the video sequence, the second flag data indicates that the picture is a non-GDR picture, the method.

Citation Information

Patent Citations

  • Encoders, decoders, and corresponding methods

    JP2022526088A

  • Image decoding method, image encoding method, image decoding device, image encoding device, and image encoding and decoding device

    WO2014006860A1

  • Gradual decoding refresh in video coding

    WO2020185962A1

Cited By

  • Methods and systems for performing gradual decoding refresh processing on pictures

    JP2025123466A