Method and apparatus for signaling video coding information
By using flags in the bitstream to control encoding modes at levels below the sequence level, the method addresses the challenge of efficient compression in advanced video coding standards, optimizing bandwidth usage and maintaining quality.
Patent Information
- Application Number
- JP2022516185
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2019-09-12
- Filing Date
- 2020-08-20
- Publication Date
- 2025-11-17
- Estimated Expiration
- 2040-08-20
AI Technical Summary
Existing video coding standards face challenges in achieving efficient compression while maintaining quality, particularly with the development of advanced standards like VVC/H.266, which require finer control over encoding modes to optimize bandwidth usage.
The method and apparatus provide control over encoding modes through flags in the bitstream, enabling or disabling modes at levels below the sequence level, allowing for more granular management of video encoding processes.
This approach enhances coding efficiency by allowing for more precise control over encoding processes, aligning with the goals of advanced standards like VVC/H.266, reducing bandwidth requirements without compromising quality.
Smart Images

Figure 0007771048000014 
Figure 0007771048000015 
Figure 0007771048000016
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This disclosure claims priority to U.S. Provisional Patent Application No. 62 / 899,169, filed September 12, 2019, the entirety of which is incorporated by reference herein. [Background technology]
[0002] background
[0002] A video is a set of static pictures (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, a video can be compressed before storage or transmission and decompressed before display. The compression process is typically referred to as encoding, and the decompression process is typically referred to as decoding. There are various video coding formats that use standardized video coding techniques, most commonly based on prediction, transform, quantization, entropy coding, and in-loop filtering. Standardization organizations have developed video coding standards that specify specific video coding formats, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard. As more and more advanced video coding techniques are adopted into video standards, the coding efficiency of new video coding standards becomes increasingly higher. Summary of the Invention [Means for solving the problem]
[0003] Disclosure Overview
[0003] Embodiments of the present disclosure provide a method and apparatus for controlling an encoding mode for video data. In an exemplary embodiment, the method includes receiving a bitstream of the video data, enabling or disabling an encoding mode for a video sequence based on a first flag in the bitstream, and determining, based on a second flag in the bitstream, control whether the encoding mode is enabled or disabled at a level below the sequence level.
[0004]
[0004] In another exemplary embodiment, a method includes receiving a bitstream of video data, enabling or disabling a first encoding mode for a video sequence based on a first flag in the bitstream, enabling or disabling a second encoding mode for the video sequence based on a second flag in the bitstream, and determining, based on a third flag in the bitstream, whether control of at least one of the first encoding mode or the second encoding mode is enabled at a level lower than the sequence level.
[0005]
[0005] In another exemplary embodiment, a method includes receiving a video sequence, a first flag, and a second flag; enabling or disabling an encoding mode for the video bitstream based on the first flag; and enabling or disabling control of the encoding mode at a level lower than the sequence level based on the second flag.
[0006]
[0006] In another exemplary embodiment, a method includes receiving a video sequence, a first flag, a second flag, and a third flag; enabling or disabling a first encoding mode for the video bitstream based on the first flag; enabling or disabling a second encoding mode for the video bitstream based on the second flag; and enabling or disabling control of at least one of the first encoding mode or the second encoding mode at a level lower than the sequence level based on the third flag.
[0007] In another exemplary embodiment, a non-transitory computer-readable medium stores a set of instructions executable by at least one processor of the device to cause the device to perform a method including receiving a bitstream of video data, enabling or disabling a coding mode for a video sequence based on a first flag in the bitstream, and determining, based on a second flag in the bitstream, whether control of the coding mode is enabled or disabled at a level below the sequence level.
[0008] In another exemplary embodiment, a non-transitory computer-readable medium stores a set of instructions executable by at least one processor of the device to cause the device to perform a method including receiving a bitstream of video data, enabling or disabling a first encoding mode for a video sequence based on a first flag in the bitstream, enabling or disabling a second encoding mode for the video sequence based on a second flag in the bitstream, and determining, based on a third flag in the bitstream, whether control of at least one of the first encoding mode or the second encoding mode is enabled at a level below the sequence level.
[0009] In another embodiment, an apparatus includes a memory configured to store a set of instructions and one or more processors communicatively coupled to the memory, the one or more processors configured to execute the set of instructions to cause the apparatus to receive a bitstream of video data, enable or disable a coding mode for a video sequence based on a first flag in the bitstream, and determine whether control of the coding mode is enabled or disabled at a level below the sequence level based on a second flag in the bitstream.
[0010] In another embodiment, an apparatus includes a memory configured to store a set of instructions and one or more processors communicatively coupled to the memory, wherein the one or more processors are configured to execute the set of instructions to cause the apparatus to receive a bitstream of video data, enable or disable a first encoding mode for a video sequence based on a first flag in the bitstream, enable or disable a second encoding mode for the video sequence based on a second flag in the bitstream, and determine, based on a third flag in the bitstream, whether control of at least one of the first encoding mode or the second encoding mode is enabled at a level below the sequence level.
[0011] BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings, in which the various features shown are not drawn to scale. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a schematic diagram illustrating the structure of an exemplary video sequence, according to some embodiments of the present disclosure. [Figure 2A]
[0013] FIG. 2A shows a schematic diagram of an exemplary encoding process of a hybrid video encoding system according to an embodiment of the present disclosure. [Figure 2B]
[0014] FIG. 2B shows a schematic diagram of another example encoding process of a hybrid video encoding system according to an embodiment of the present disclosure. [Figure 3A]
[0015] FIG. 3A shows a schematic diagram of an exemplary decoding process of a hybrid video coding system, according to an embodiment of the present disclosure. [Figure 3B]
[0016] FIG. 3B shows a schematic diagram of another example decoding process of a hybrid video coding system, according to an embodiment of the present disclosure. [Figure 4]
[0017] FIG. 4 illustrates a block diagram of an exemplary apparatus for encoding or decoding video, according to some embodiments of the present disclosure. [Figure 5]
[0018] FIG. 5 is a schematic diagram illustrating an example process of decoder side motion vector refinement (DMVR), according to some embodiments of the present disclosure. [Figure 6]
[0019] FIG. 6 is a schematic diagram illustrating an example DMVR discovery process according to some embodiments of the present disclosure. [Figure 7]
[0020] FIG. 7 is a schematic diagram illustrating an example pattern for DMVR integer luma sample search according to some embodiments of the present disclosure. [Figure 8]
[0021] FIG. 8 is a schematic diagram illustrating another example pattern for stages for integer sample offset search in a DMVR integer luma sample search, according to some embodiments of the present disclosure. [Figure 9]
[0022] FIG. 9 is a schematic diagram illustrating an example pattern for estimation of a DMVR parameter error surface, according to some embodiments of the present disclosure. [Figure 10]
[0023] FIG. 10 is a schematic diagram of an example of an extended coding unit (CU) region used in bi-directional optical flow (BDOF), according to some embodiments of the present disclosure. [Figure 11]
[0024] FIG. 11 is a schematic diagram of an example of sub-block-based affine motion and sample-based affine motion, according to some embodiments of the present disclosure. [Figure 12]
[0025] FIG. 12 shows Table 1 illustrating an example syntax structure of a Sequence Parameter Set (SPS) that implements control flags for DMVR and BDOF, according to some embodiments of the present disclosure. [Figure 13]
[0026] FIG. 13 shows Table 2 illustrating an example syntax structure of a slice header implementing control flags for DMVR and BDOF according to some embodiments of the present disclosure. [Figure 14A]
[0027] FIG. 14A shows Table 3A illustrating an example syntax structure of an SPS implementing slice-level control flags for DMVR, BDOF, and prediction refinement with optical flow (PROF) according to some embodiments of the present disclosure. [Figure 14B]
[0028] FIG. 14B shows Table 3B illustrating an example syntax structure of an SPS implementing picture-level control flags for DMVR, BDOF, and PROF, according to some embodiments of the present disclosure. [Figure 15A]
[0029] FIG. 15A shows Table 4A illustrating an example syntax structure of a slice header implementing control flags for DMVR, BDOF, and PROF, according to some embodiments of the present disclosure. [Figure 15B]
[0030] FIG. 15B shows Table 4B illustrating an example syntax structure of a picture header implementing control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. [Figure 16]
[0031] FIG. 16 shows Table 5 illustrating an example syntax structure of an SPS that implements separate sequence-level control flags for DMVR, BDOF, and PROF, according to some embodiments of the present disclosure. [Figure 17]
[0032] FIG. 17 shows Table 6 illustrating an example syntax structure of a slice header implementing joint control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. [Figure 18]
[0033] FIG. 18 shows Table 7 illustrating an example syntax structure of a slice header implementing separate control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. [Figure 19]
[0034] FIG. 19 shows Table 8 illustrating an example syntax structure of a slice header implementing hybrid control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. [Figure 20]
[0035] FIG. 20 shows Table 9 illustrating an example syntax structure of an SPS implementing hybrid sequence-level control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. [Figure 21]
[0036] FIG. 21 shows Table 10 illustrating another example syntax structure of a slice header implementing hybrid control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. [Figure 22]
[0037] FIG. 22 shows Table 11 illustrating another example syntax structure of a slice header implementing separate control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. [Figure 23]
[0038] FIG. 23 shows a flowchart of an exemplary process for controlling video decoding modes according to some embodiments of this disclosure. [Figure 24]
[0039] FIG. 24 shows a flowchart of another example process for controlling video decoding modes according to some embodiments of this disclosure. [Figure 25]
[0040] FIG. 25 illustrates a flowchart of an exemplary process for controlling a video coding mode according to some embodiments of the present disclosure. [Figure 26]
[0041] FIG. 26 shows a flowchart of another example process for controlling video coding modes according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0013] Detailed Description
[0042] Reference may be made in detail to the exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which like reference numerals in different drawings represent the same or similar elements unless otherwise indicated. The implementations set forth in the following description of exemplary embodiments do not represent all implementations in accordance with the present invention. Rather, they are merely examples of apparatus and methods in accordance with aspects related to the present invention as recited in the appended claims. Certain aspects of the present disclosure are described in more detail below. In the event of a conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall control.
[0014]
[0043] The ITU-T Video Coding Expert Group (ITU-T VCEG) and the ISO / IEC Moving Picture Expert Group (ISO / IEC MPEG) Joint Video Experts Team (JVET) are currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 while using half the bandwidth.
[0015]
[0044] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET is developing technology beyond HEVC using the JEM (joint exploration model) reference software. Because the coding technology was incorporated into JEM, JEM achieved substantially higher coding performance than HEVC.
[0016]
[0045] The VVC standard is a recent development and continues to incorporate more coding techniques that result in better compression performance. VVC is based on the same hybrid video coding system that has been used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.
[0017]
[0046] Video is a set of static pictures (or "frames") arranged in time sequence to store visual information. A video capture device (e.g., a camera) can be used to capture and store these pictures in time sequence, and a video playback device (e.g., a television, a computer, a smartphone, a tablet computer, a video player, or any end-user terminal with display capabilities) can be used to display such pictures in time sequence. In some applications, a video capture device can also transmit the captured video in real time to a video playback device (e.g., a computer with a monitor) for purposes such as supervision, conferencing, or live broadcasting.
[0018]
[0047] To reduce the storage space and transmission bandwidth required by such applications, video can be compressed before storage and transmission and decompressed before display. Compression and decompression can be performed by software executed by a processor (e.g., a processor in a general-purpose computer) or by specialized hardware. A module for compression is commonly referred to as an “encoder,” and a module for decompression is commonly referred to as a “decoder.” Collectively, the encoder and decoder can be referred to as a “codec.” The encoder and decoder can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of the encoder and decoder can include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of the encoder and decoder can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression may be performed by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, or the like. In some applications, a codec may decompress video from a first encoding standard and recompress the decompressed video using a second encoding standard. In this case, the codec may be referred to as a "transcoder."
[0019]
[0048] A video coding process can identify and retain useful information that can be used to reconstruct a picture and ignore information that is not important for reconstruction. If the ignored, unimportant information cannot be perfectly reconstructed, such a coding process may be called "lossy." Otherwise, it may be called "lossless." Most coding processes are lossy; this is a tradeoff to reduce the required storage space and transmission bandwidth.
[0020]
[0049] Useful information about the picture being coded (called the "current picture") includes changes relative to a reference picture (e.g., a previously coded and reconstructed picture). Such changes can include changes in pixel position, brightness, or color, of which position changes are the most important. Changes in the position of a group of pixels representing an object can reflect the movement of the object between the reference picture and the current picture.
[0021]
[0050] A picture that is coded without reference to another picture (i.e., it is its own reference picture) is called an "I-picture." A picture that is coded using a previous picture as a reference picture is called a "P-picture." A picture that is coded using both a previous picture and a future picture as a reference picture (i.e., the referencing is "bidirectional") is called a "B-picture."
[0022]
[0051] 1 illustrates the structure of an exemplary video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 can be live video or captured and archived video. The video 100 can be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). The video sequence 100 can be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., video files stored in a storage device), or a video supply interface (e.g., a video broadcast transceiver) for receiving video from a video content provider.
[0023]
[0052] As shown in FIG. 1, video sequence 100 may include a series of pictures arranged temporally along a timeline, including pictures 102, 104, 106, and 108. Pictures 102-106 are consecutive, with additional pictures between pictures 106 and 108. In FIG. 1, picture 102 is an I-picture, and its reference picture is picture 102 itself. Picture 104 is a P-picture, and its reference picture is picture 102, as indicated by the arrow. Picture 106 is a B-picture, and its reference pictures are pictures 104 and 108, as indicated by the arrows. In some embodiments, the reference picture of a picture (e.g., picture 104) need not immediately precede or follow that picture. For example, the reference picture of picture 104 can be a picture before picture 102. It should be noted that the reference pictures of pictures 102-106 are merely examples, and this disclosure does not limit the reference picture embodiments to the examples shown in FIG.
[0024]
[0053] Typically, video codecs do not encode or decode an entire picture at once due to the computational complexity of such a task. Rather, they may divide a picture into elementary segments and encode or decode the picture segment by segment. Such elementary segments are referred to as basic processing units ("BPUs") in this disclosure. For example, structure 110 in FIG. 1 shows an example structure of a picture (e.g., any of pictures 102-108) of video sequence 100. In structure 110, the picture is divided into 4x4 basic processing units, the boundaries of which are shown as dashed lines. In some embodiments, basic processing units may be referred to as "macroblocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) or as "coding tree units" ("CTUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing units can have variable sizes in pictures or any arbitrary shape and size of pixels, such as 128x128, 64x64, 32x32, 16x16, 4x8, 16x32, etc. The size and shape of the basic processing unit can be selected based on a balance between coding efficiency and the level of detail to be maintained in the basic processing unit for the picture.
[0025]
[0054] A basic processing unit may be a logical unit that can include groups of different types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color picture may include a luma component (Y) that represents colorless luminance information, one or more chroma components (e.g., Cb and Cr) that represent color information, and related syntax elements, where the luma and chroma components may have the same size of a basic processing unit. The luma and chroma components may be referred to as "coding tree blocks" (CTBs) in some video coding standards (e.g., H.265 / HEVC or H.266 / VVC). Any operation performed on a basic processing unit may be performed repeatedly on each of its luma and chroma components.
[0026]
[0055] Video coding has multiple computational stages, examples of which are shown in Figures 2A-2B and 3A-3B. At each stage, the size of the basic processing unit may still become too large for processing and therefore may be further divided into segments referred to as "basic processing subunits" in this disclosure. In some embodiments, the basic processing subunits may be referred to as "blocks" in some video coding standards (e.g., MPEG family, H.261, H.263, or H.264 / AVC) or as "coding units" ("CUs") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunits may have the same or smaller size than the basic processing units. Similar to the basic processing units, the basic processing subunits are also logical units that may contain groups of different types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on a basic processing sub-unit may be repeatedly performed on each of its luma and chroma components. Note that such division may be performed to further levels as needed for processing. Also note that different stages may use different schemes to divide the basic processing units.
[0027]
[0056] For example, in a mode decision stage (an example of which is shown in FIG. 2B ), an encoder can decide what prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for a basic processing unit, but the basic processing unit may be too large to make such a decision. The encoder can divide the basic processing unit into multiple basic processing sub-units (e.g., CUs, as in the case of H.265 / HEVC or H.266 / VVC) and decide the type of prediction for each individual basic processing sub-unit.
[0028]
[0057] As another example, in the prediction stage (an example of which is shown in FIGS. 2A-2B), the encoder may perform prediction operations at the level of basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunits may still be too large to process. The encoder may further divide the basic processing subunits into smaller segments (e.g., referred to as "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), at which level the prediction operations may be performed.
[0029]
[0058] As another example, in the transform stage (an example of which is shown in FIGS. 2A-2B), the encoder may perform transform operations for residual basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunits may still be too large to process. The encoder may further divide the basic processing subunits into smaller segments (e.g., referred to as "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), at which levels the transform operations may be performed. Note that the division scheme of the same basic processing subunit may be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU may have different sizes and numbers.
[0030]
[0059] 1, the fundamental processing units 112 are further divided into 3x3 fundamental processing sub-units, the boundaries of which are shown as dotted lines. Different fundamental processing units of the same picture may be divided into fundamental processing sub-units in different ways.
[0031]
[0060] In some implementations, to provide parallel processing and error resilience capabilities to video encoding and decoding, a picture can be divided into regions for processing, so that the encoding or decoding process does not rely on information about a picture region from any other region of the picture. In other words, each region of a picture can be processed independently. By doing so, the codec can process different regions of a picture in parallel, thus increasing coding efficiency. Also, when data for a region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same picture without relying on the corrupted or lost data, thus providing error resilience. Some video coding standards allow pictures to be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles." It should also be noted that different pictures in video sequence 100 can have different partitioning schemes for dividing the picture into regions.
[0032]
[0061] For example, in Figure 1, structure 110 is divided into three regions 114, 116, and 118, the boundaries of which are shown as solid lines within structure 110. Region 114 includes four basic processing units. Regions 116 and 118 each include six basic processing units. It should be noted that the basic processing units, basic processing subunits, and regions of structure 110 in Figure 1 are merely examples, and the present disclosure does not limit the embodiments thereof.
[0033]
[0062] FIG. 2A shows a schematic diagram of an exemplary encoding process 200A according to an embodiment of the present disclosure. For example, encoding process 200A may be performed by an encoder. As shown in FIG. 2A, the encoder may encode a video sequence 202 into a video bitstream 228 according to process 200A. Similar to video sequence 100 in FIG. 1, video sequence 202 may include a set of pictures (referred to as "original pictures") arranged in a temporal order. Similar to structure 110 in FIG. 1, each original picture in video sequence 202 may be divided by the encoder into basic processing units, basic processing sub-units, or regions for processing. In some embodiments, the encoder may perform process 200A at the level of basic processing units for each original picture in video sequence 202. For example, the encoder may perform process 200A in an iterative manner, and the encoder may encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder may perform process 200A in parallel for regions (eg, regions 114-118) of each original picture in video sequence 202.
[0034]
[0063] 2A , an encoder may provide a fundamental processing unit (referred to as an “original BPU”) of an original picture of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a predicted BPU 208. The encoder may subtract the predicted BPU 208 from the original BPU to generate a residual BPU 210. The encoder may provide the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may provide the prediction data 206 and the quantized transform coefficients 216 to a binary coding stage 226 to generate a video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a “forward path.” During process 200A, after quantization stage 214, the encoder may provide quantized transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate reconstructed residual BPU 222. The encoder may add reconstructed residual BPU 222 to prediction BPU 208 to generate prediction reference 224, which is used in prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as the "reconstruction path." The reconstruction path may be used to ensure that both the encoder and decoder use the same reference data for prediction.
[0035]
[0064] The encoder may iteratively perform process 200A to encode each original BPU of the original picture (in the forward path) and generate a prediction reference 224 for encoding the next original BPU of the original picture (in the reconstruction path). After encoding all the original BPUs of the original picture, the encoder may proceed to encode the next picture in the video sequence 202.
[0036]
[0065] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video capture device (e.g., a camera). As used herein, the term "receive" may refer to receiving, inputting, acquiring, obtaining, getting, reading, accessing, or any act in any manner to input data.
[0037]
[0066] In the prediction step 204, in the current iteration, the encoder receives an original BPU and a prediction reference 224, and may perform a prediction operation to generate predicted data 206 and a predicted BPU 208. The predicted reference 224 may be generated from a reconstruction path of a previous iteration of the process 200A. The purpose of the prediction step 204 is to reduce information redundancy by extracting predicted data 206, which can be used to reconstruct the original BPU as a predicted BPU 208 from the predicted data 206 and the prediction reference 224.
[0038]
[0067] Ideally, predicted BPU 208 can be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, predicted BPU 208 generally differs slightly from the original BPU. To record such differences, after generating predicted BPU 208, the encoder can subtract it from the original BPU to generate residual BPU 210. For example, the encoder can subtract pixel values (e.g., grayscale or RGB values) of predicted BPU 208 from corresponding pixel values of the original BPU. Each pixel of residual BPU 210 can have a residual value that is the result of such a subtraction between the corresponding pixel of the original BPU and predicted BPU 208. Compared to the original BPU, predicted data 206 and residual BPU 210 can have fewer bits, which can be used to reconstruct the original BPU without significant quality degradation. Therefore, the original BPU is compressed.
[0039]
[0068] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce spatial redundancy in the residual BPU 210 by decomposing it into a set of two-dimensional "basis patterns," each associated with a "transform coefficient." The basis patterns can have the same size (e.g., the size of the residual BPU 210). Each basis pattern can represent a change frequency (e.g., a frequency of luminance change) component of the residual BPU 210. No basis pattern can be reconstructed from any combination (e.g., a linear combination) of any other basis patterns. In other words, the decomposition can decompose the changes in the residual BPU 210 into the frequency domain. Such a decomposition is similar to a discrete Fourier transform of a function, where the basis patterns are similar to the basis functions (e.g., trigonometric functions) of the discrete Fourier transform, and the transform coefficients are similar to the coefficients associated with the basis functions.
[0040]
[0069] Different transform algorithms can use different basis patterns. For example, various transform algorithms, such as a discrete cosine transform, a discrete sine transform, or the like, can be used in transform stage 212. The transform in transform stage 212 is invertible. That is, an encoder can recover residual BPU 210 by inverting the transform (referred to as an "inverse transform"). For example, to recover pixels of residual BPU 210, the inverse transform can multiply the values of corresponding pixels in the basis pattern by their associated coefficients and add the products to generate a weighted sum. For video coding standards, both the encoder and decoder can use the same transform algorithm (and therefore the same basis pattern). Therefore, the encoder can record only the transform coefficients, and the decoder can reconstruct residual BPU 210 from the transform coefficients without receiving the basis pattern from the encoder. Compared to residual BPU 210, the transform coefficients can have fewer bits, which can be used to reconstruct residual BPU 210 without significant quality degradation. Therefore, the residual BPU 210 is further compressed.
[0041]
[0070] The encoder can further compress the transform coefficients in the quantization stage 214. In the transform process, different basis patterns can represent different change frequencies (e.g., luminance change frequencies). Because the human eye is generally better at perceiving low-frequency changes, the encoder can ignore high-frequency change information without significant quality degradation during decoding. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (referred to as a "quantization parameter") and rounding the quotient to its nearest integer. After such an operation, some transform coefficients of high-frequency basis patterns can be converted to zero, and transform coefficients of low-frequency basis patterns can be converted to smaller integers. The encoder can ignore zero-valued quantized transform coefficients 216, thereby further compressing the transform coefficients. The quantization process can also be inverted, in which case the quantized transform coefficients 216 can be reconstructed into transform coefficients in the inverse operation of quantization (referred to as "dequantization").
[0042]
[0071] Because the encoder ignores the remainder of such a division in a rounding operation, the quantization stage 214 may be lossy. Typically, the quantization stage 214 may contribute the greatest information loss in the process 200A. The greater the information loss, the fewer bits the quantized transform coefficients 216 may require. To achieve different levels of information loss, the encoder may use different values of the quantization parameter or any other parameter of the quantization process.
[0043]
[0072] In binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as, for example, entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in binary encoding stage 226, such as, for example, a prediction mode used in prediction stage 204, parameters of the prediction operation, the type of transform in transform stage 212, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), or the like. The encoder may generate a video bitstream 228 using the output data of binary encoding stage 226. In some embodiments, the video bitstream 228 may be further packetized for network transmission.
[0044]
[0073] Referring to the reconstruction path of process 200A, in an inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In an inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to a predicted BPU 208 to generate a predicted reference 224 to be used in the next iteration of process 200A.
[0045]
[0074] It should be noted that other variations of process 200A may be used to encode video sequence 202. In some embodiments, the stages of process 200A may be performed in a different order by the encoder. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be split into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit one or more stages in FIG. 2A.
[0046]
[0075] 2B shows a schematic diagram of another exemplary encoding process 200B according to an embodiment of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B additionally includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.
[0047]
[0076] Generally, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-picture prediction or "intra-prediction") can use pixels from one or more already-encoded neighboring BPUs within the same picture to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of a picture. Temporal prediction (e.g., inter-picture prediction or "inter-prediction") can use regions from one or more already-encoded pictures to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include encoded pictures. Temporal prediction can reduce the inherent temporal redundancy of a picture.
[0048]
[0077] Referring to process 200B, within the forward path, the encoder performs prediction operations in a spatial prediction step 2042 and a temporal prediction step 2044. For example, in the spatial prediction step 2042, the encoder may perform intra prediction. For an original BPU of a picture being encoded, the prediction reference 224 may include one or more neighboring BPUs within the same picture that were coded (in the forward path) and reconstructed (in the reconstruction path). The encoder may generate the predicted BPU 208 by extrapolating the neighboring BPUs. Extrapolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, or the like. In some embodiments, the encoder may perform extrapolation at the pixel level, such as by extrapolating, for each pixel of the predicted BPU 208, the value of the corresponding pixel. The neighboring BPUs used for extrapolation can be located relative to the original BPU from various directions, such as vertically (e.g., above the original BPU), horizontally (e.g., to the left of the original BPU), diagonally (e.g., below-left, below-right, above-left, or above-right of the original BPU), or any direction defined in the video coding standard used. For intra prediction, the prediction data 206 may include, for example, the locations (e.g., coordinates) of the neighboring BPUs used, the sizes of the neighboring BPUs used, parameters of the extrapolation, the orientations of the neighboring BPUs used relative to the original BPU, or the like.
[0049]
[0078] As another example, in the temporal prediction stage 2044, the encoder may perform inter-prediction. For an original BPU of the current picture, the prediction reference 224 may include one or more pictures (referred to as "reference pictures") that have been coded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference pictures may be coded and reconstructed for each BPU. For example, the encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs of the same picture have been generated, the encoder may generate the reconstructed picture as the reference picture. The encoder may perform a "motion estimation" operation to search for a matching region within a range (referred to as a "search window") of the reference picture. The location of the search window in the reference picture may be determined based on the location of the original BPU of the current picture. For example, the search window may be centered in the reference picture at a location having the same coordinates as the original BPU in the current picture and may extend outward over a predetermined distance. When the encoder identifies a region similar to the original BPU within the search window (e.g., by using a pel-recursive algorithm, a block matching algorithm, or the like), the encoder can determine such a region as a matching region. The matching region can have different dimensions than the original BPU (e.g., smaller than, equal to, larger than, or a different shape than the original BPU). Because the reference picture and the current picture are temporally separated in a timeline (e.g., as shown in FIG. 1), the matching region can be considered to "move" toward the location of the original BPU over time. The encoder can record the direction and distance of such movement as a "motion vector." When multiple reference pictures are used (e.g., as picture 106 in FIG. 1), the encoder can search for a matching region for each reference picture and determine the associated motion vector. In some embodiments, the encoder can assign weights to the pixel values of the matching region in each matching reference picture.
[0050]
[0079] Motion estimation can be used to identify various types of motion, such as, for example, translation, rotation, zooming, or the like. For inter prediction, prediction data 206 can include, for example, the location (e.g., coordinates) of the matching region, a motion vector associated with the matching region, the number of reference pictures, weights associated with the reference pictures, or the like.
[0051]
[0080] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., a motion vector) and the prediction reference 224. For example, the encoder may shift the matching region of a reference picture according to the motion vector, thereby enabling the encoder to predict the original BPU of the current picture. When multiple reference pictures are used (e.g., as picture 106 in FIG. 1), the encoder may shift the matching region of the reference picture according to each motion vector and average the pixel values of the matching region. In some embodiments, if the encoder weights the pixel values of the matching region of each matching reference picture, the encoder may add a weighted sum of pixel values to the shifted matching region.
[0052]
[0081] Depending on the embodiment, inter-prediction can be unidirectional or bidirectional. Unidirectional inter-prediction can use one or more reference pictures in the same temporal direction relative to the current picture. For example, picture 104 in FIG. 1 is a unidirectional inter-predicted picture in which a reference picture (i.e., picture 102) precedes picture 104. Bidirectional inter-prediction can use one or more reference pictures in both temporal directions relative to the current picture. For example, picture 106 in FIG. 1 is a bidirectional inter-predicted picture in which reference pictures (i.e., pictures 104 and 108) are in both temporal directions relative to picture 104.
[0053]
[0082] Still referring to the forward path of process 200B, after spatial prediction step 2042 and temporal prediction step 2044, in mode decision step 230, the encoder may select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of process 200B. For example, the encoder may perform a rate-distortion optimization technique. In this technique, the encoder may select a prediction mode to minimize the value of a cost function that depends on the bitrate of the candidate prediction mode and the distortion of the reconstructed reference picture under the candidate prediction mode. Depending on the selected prediction mode, the encoder may generate a corresponding predicted BPU 208 and predicted data 206.
[0054]
[0083] Within the reconstruction path of process 200B, if an intra-prediction mode is selected within the forward path, after generating the prediction reference 224 (e.g., the current BPU coded and reconstructed in the current picture), the encoder can directly provide the prediction reference 224 to spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If an inter-prediction mode is selected within the forward path, after generating the prediction reference 224 (e.g., the current picture coded and reconstructed in all BPUs), the encoder can provide the prediction reference 224 to loop filter stage 232, where the encoder can apply a loop filter to the prediction reference 224 to reduce or eliminate distortions (e.g., blocking artifacts) introduced by inter prediction. The encoder can apply various loop filter techniques within loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filter, or the like. The loop-filtered reference picture may be stored in a buffer 234 (or "decoded picture buffer") for later use (e.g., to be used as an inter-prediction reference picture for a future picture in the video sequence 202). The encoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) along with the quantized transform coefficients 216, the prediction data 206, and other information in the binary encoding stage 226.
[0055]
[0084] FIG. 3A shows a schematic diagram of an exemplary decoding process 300A according to an embodiment of the present disclosure. Process 300A may be a decompression process corresponding to compression process 200A in FIG. 2A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. A decoder may follow process 300A to decode video bitstream 228 into video stream 304. Video stream 304 may be similar to video sequence 202. However, due to information loss in the compression and decompression processes (e.g., quantization stage 214 in FIGS. 2A-2B), video stream 304 is generally not identical to video sequence 202. Similar to processes 200A and 200B in FIGS. 2A-2B, a decoder may perform process 300A at the level of a basic processing unit (BPU) for each picture encoded in video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder may decode a basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for regions (e.g., regions 114-118) of each picture encoded in video bitstream 228.
[0056]
[0085] In FIG. 3A , a decoder may provide a portion of a video bitstream 228 associated with a basic processing unit (referred to as a “coding BPU”) of a coded picture to a binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may provide the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may provide the prediction data 206 to a prediction stage 204 to generate a prediction BPU 208. The decoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in computer memory). The decoder may provide the prediction reference 224 to the prediction stage 204 for performing prediction operations in the next iteration of process 300A.
[0057]
[0086] The decoder may perform process 300A iteratively to decode each coded BPU of a coded picture and generate a predictive reference 224 for encoding the next coded BPU of the coded picture. After decoding all coded BPUs of a coded picture, the decoder may output the picture to video stream 304 for display and proceed to decode the next coded picture in video bitstream 228.
[0058]
[0087] In binary decoding step 302, the decoder may perform the inverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to prediction data 206 and quantized transform coefficients 216, the decoder may decode other information in binary decoding step 302, such as, for example, a prediction mode, parameters of the prediction operation, type of transform, parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bitrate control parameters), or the like. In some embodiments, if video bitstream 228 is transmitted in packets over a network, the decoder may depacketize video bitstream 228 before providing it to binary decoding step 302.
[0059]
[0088] 3B shows a schematic diagram of another exemplary decoding process 300B according to an embodiment of the present disclosure. Process 300B may be modified from process 300A. For example, process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared to process 300A, process 300B additionally divides prediction stage 204 into spatial prediction stage 2042 and temporal prediction stage 2044, and additionally includes loop filter stage 232 and buffer 234.
[0060]
[0089] In process 300B, prediction data 206 decoded by the decoder from binary decoding stage 302 for a coding basic processing unit (referred to as a "current BPU") of a coding picture being decoded (referred to as a "current picture") may include various types of data, depending on what prediction mode was used by the encoder to encode the current BPU. For example, if intra prediction was used by the encoder to encode the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, or the like. Parameters of the intra prediction operation may include, for example, the location (e.g., coordinates) of one or more neighboring BPUs used as references, the size of the neighboring BPUs, parameters of extrapolation, the orientation of the neighboring BPUs relative to the original BPU, or the like. As another example, if inter prediction was used by the encoder to encode the current BPU, prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, or the like. Parameters for the inter prediction operation may include, for example, the number of reference pictures associated with the current BPU, weights associated with each of the reference pictures, locations (e.g., coordinates) of one or more matching regions within each reference picture, one or more motion vectors associated with each of the matching regions, or the like.
[0061]
[0090] Based on the prediction mode indicator, the decoder may determine whether to perform spatial prediction (e.g., intra prediction) in spatial prediction step 2042 or temporal prediction (e.g., inter prediction) in temporal prediction step 2044. Details of performing such spatial or temporal prediction are described in FIG. 2B and will not be repeated below. After performing such spatial or temporal prediction, the decoder may generate a predicted BPU 208. The decoder may add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a predicted reference 224, as described in FIG. 3A.
[0062]
[0091] In process 300B, the decoder may provide the prediction reference 224 to spatial prediction stage 2042 or temporal prediction stage 2044 to perform a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may provide the prediction reference 224 directly to spatial prediction stage 2042 for later use (e.g., for extrapolation of the next BPU of the current picture). If the current BPU is decoded using inter prediction in temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference picture from which all BPUs are decoded), the encoder may provide the prediction reference 224 to loop filter stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may apply a loop filter to the prediction reference 224 in the manner described in FIG. 2B . The loop-filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer in computer memory) for later use (e.g., to be used as an inter-prediction reference picture for future coded pictures of the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-prediction was used to encode the current BPU, the prediction data may further include parameters of the loop filter (e.g., loop filter strength).
[0063]
[0092] 4 is a block diagram of an exemplary device 400 for encoding or decoding video, according to an embodiment of the present disclosure. As shown in FIG. 4, the device 400 may include a processor 402. When the processor 402 executes the instructions described herein, the device 400 may become a dedicated machine for video encoding or decoding. The processor 402 may be any type of circuitry capable of manipulating or processing information. For example, processor 402 may include any number and combination of a central processing unit (or "CPU"), a graphics processing unit (or "GPU"), a neural processing unit ("NPU"), a microcontroller unit ("MCU"), an optical processor, a programmable logic controller, a microcontroller, a microprocessor, a digital signal processor, an intellectual property (IP) core, a programmable logic array (PLA), a programmable array logic (PAL), a generic array logic (GAL), a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a system on chip (SoC), an application-specific integrated circuit (ASIC), or the like. In some embodiments, processor 402 may also be a set of processors grouped as a single logical entity. For example, as shown in Figure 4, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0064]
[0093] The device 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, or the like). For example, as shown in Figure 4, the stored data may include program instructions (e.g., program instructions for performing steps in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access (e.g., via bus 410) the program instructions and data for processing, execute the program instructions, and perform operations or manipulations on the data for processing. The memory 404 may include a high-speed random-access storage device or a non-volatile storage device. In some embodiments, memory 404 may include any number or combination of random-access memory (RAM), read-only memory (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, security digital (SD) cards, memory sticks, compact flash (CF) cards, or the like. Memory 404 may also be a group of memories (not shown in FIG. 4) grouped as a single logical entity.
[0065]
[0094] Bus 410 can be a communication device that transfers data between components internal to apparatus 400, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Express port), or the like.
[0066]
[0095] For ease of explanation and without ambiguity, the processor 402 and other data processing circuitry will be collectively referred to in this disclosure as "data processing circuitry." The data processing circuitry may be implemented entirely as hardware or as a combination of software, hardware, or firmware. In addition, the data processing circuitry may be a single, stand-alone module or may be fully or partially combined with any other component of the device 400.
[0067]
[0096] Device 400 may further include a network interface 406 for providing wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, or the like). In some embodiments, network interface 406 may include any number or combination of a network interface controller (NIC), a radio frequency (RF) module, a transponder, a transceiver, a modem, a router, a gateway, a wired network adapter, a wireless network adapter, a Bluetooth adapter, an infrared adapter, a near-field communication ("NFC") adapter, a cellular network chip, or the like.
[0068]
[0097] In some embodiments, apparatus 400 may optionally further include a peripheral interface 408 for providing connection to one or more peripheral devices. As shown in Figure 4, the peripheral devices may include, but are not limited to, a cursor control device (e.g., a mouse, touchpad, or touchscreen), a keyboard, a display (e.g., a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (e.g., a camera, or an input interface communicatively coupled to a video archive), or the like.
[0069]
[0098] It should be noted that a video codec (e.g., a codec performing process 200A, 200B, 300A, or 300B) may be implemented as any combination of software or hardware modules within device 400. For example, some or all of the stages of process 200A, 200B, 300A, or 300B may be implemented as one or more software modules of device 400, such as program instructions that may be loaded into memory 404. As another example, some or all of the stages of process 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of device 400, such as specialized data processing circuitry (e.g., FPGA, ASIC, NPU, or the like).
[0070]
[0099] In the quantization and inverse quantization functional blocks (e.g., quantization 214 and inverse quantization 218 in FIG. 2A or 2B, inverse quantization 218 in FIG. 3A or 3B), a quantization parameter (QP) is used to determine the amount of quantization (and inverse quantization) applied to the prediction residual. The initial QP value used for coding a picture or slice can be signaled at a high level, for example, using the init_qp_minus26 syntax element in the Picture Parameter Set (PPS) and the slice_qp_delta syntax element in the slice header. Furthermore, the QP value can be adapted at a local level per CU using delta QP values sent at the granularity of the quantization group.
[0071]
[0100] To increase the accuracy of merge-mode motion vectors (MVs), bilateral-matching (BM)-based decoder-side motion vector refinement (DMVR) is adopted in Versatile Video Coding (VVC) Draft 6. In bilateral prediction operations, refined MVs are searched around an initial MV within reference picture lists L0 and L1. BM-based DMVR calculates distortion between two candidate blocks within reference picture lists L0 and L1. FIG. 5 shows an example process 500 of decoder-side motion vector refinement (DMVR) according to some embodiments of the present disclosure. FIG. 5 shows a current picture 502, a first reference picture 504 in a first reference picture list L0, and a second reference picture 506 in a second reference picture list L1. A first initial MV 508 is directed from a current block 510 in the current picture 502 to a first initial reference block 512 in the first reference picture 504. The second initial MV 514 points from the current block 510 to a second initial reference block 516 in the second reference picture 506. The process 500 performs BM-based DMVR and can determine a first candidate reference block 518 in the first reference picture 504, a second candidate reference block 520 in the second reference picture 506, a first candidate MV 522 connecting the current block 510 and the first candidate reference block 518, and a second candidate MV 524 connecting the current block 510 and the second candidate reference block 516. As shown in FIG. 5 , the first candidate MV 522 and the second candidate MV 524 are close to the first initial MV 508 and the second initial MV 514, respectively. In some embodiments, the process 500 can calculate the sum of absolute differences (SAD) between the first initial reference block 518 and the second initial reference block 520 based on each MV candidate (e.g., the first candidate MV 522 or the second candidate MV 524) around the initial MV (e.g., the first initial MV 508 or the second initial MV 514).The first candidate MV 522 and the second candidate MV 524 with the lowest SAD can be the refinement MV to be used to generate the bi-predicted signal.
[0072]
[0101] In some embodiments, as described in VVC Draft 6, DMVR is applied to a CU that satisfies all of the following conditions: (1) the merge mode is at the CU level with bi-predictive MVs; (2) the block is predicted using bi-predictive MVs with equal weights, e.g., by not applying bi-prediction with weighted averaging (BWA) to the block; (3) one reference picture is in the past (e.g., the first reference picture 504) and another reference picture is in the future (e.g., the second reference picture 506) with respect to the current picture (e.g., the current picture 502); (4) the distances (e.g., the difference in picture order counts (POCs)) from both reference pictures to the current picture are the same; and (5) the block has at least 128 luma samples and the width and height of the block are both at least 8 luma samples.
[0073]
[0102] The refined MV derived by the DMVR process (e.g., process 500) can be used to generate inter-predicted samples and for temporal motion vector prediction for encoding future pictures. The original MV (e.g., the first initial MV 508 or the second initial MV 514) can be used in the deblocking process and in spatial MV prediction for encoding future CUs within the current picture (e.g., the current picture 502).
[0074]
[0103] 5, the first MV offset 526 represents a refinement offset between the first initial MV 508 and the first candidate MV 522, and the second MV offset 528 represents a refinement offset between the second initial MV 514 and the second candidate MV 524. The first MV offset 526 and the second MV offset 528 can have the same magnitude and opposite direction. In some embodiments, the search point can surround the initial MV (e.g., the first initial MV 508 and the second initial MV 514), and the MV offsets (e.g., the first MV offset 526 and the second MV offset 528) can follow an MV difference mirroring rule. For example, the point checked by the DMVR can be represented by a candidate MV pair, MV0 (e.g., the first candidate MV 522) and MV1 (e.g., the second candidate MV 524), and can be based on equations (1) and (2): MV0'=MV0+MV_offset Equation (1) MV1'=MV1-MV_offset formula (2)
[0075]
[0104] In equations (1) and (2), MV0′ (e.g., the first initial MV 508) and MV1′ (e.g., the second candidate MV 514) represent a pair of initial MVs, and MV_offset represents a refinement offset (e.g., the first MV offset 526 or the second MV offset 528) between the initial MV (e.g., MV0′ or MV1′) and the refinement MV (e.g., MV0 or MV1). Note that MV_offset is a vector having motion displacements (e.g., in the X and Y dimensions). In some embodiments, as described in VVC Draft 5, the refinement search range (e.g., the search range of the DMVR) can be two integer luma samples from the initial MV (e.g., the first initial MV 508 and the second initial MV 514) in both the horizontal and vertical dimensions.
[0076]
[0105] FIG. 6 illustrates an exemplary DMVR search process 600 according to some embodiments of the present disclosure. In some embodiments, process 600 may be performed by a codec (e.g., the encoder in FIGS. 2A-2B or the decoder in FIGS. 3A-3B). For example, the codec may be implemented as one or more software or hardware components of an apparatus for encoding or transcoding a video sequence (e.g., apparatus 400 in FIG. 4). In some embodiments, process 600 may serve as an example of a search process for a DMVR as described in VVC Draft 4. As shown in FIG. 6, process 600 includes a stage 602 for integer sample offset search and a stage 604 for fractional sample refinement. To reduce search complexity, in some embodiments, a fast search method with an early termination mechanism may be applied in stage 602. For example, rather than using a 25-point full search, a two-iteration search scheme may be applied in stage 602 to reduce the SAD checkpoint.
[0077]
[0106] 6, step 602 may be followed by step 604. To reduce computational complexity, in some embodiments, the fractional sample refinement in step 604 may be derived using a parametric error surface equation instead of performing an additional search requiring SAD comparisons. Step 604 may be conditionally invoked based on the output of step 602.
[0078]
[0107] FIG. 7 illustrates an exemplary pattern 700 for a DMVR integer luma sample search according to some embodiments of the present disclosure. The DMVR integer luma sample search can determine the point with the smallest SAD among the searched samples. For example, the DMVR integer luma sample search can be implemented as process 600 in FIG. 6, including a stage for an integer sample offset search (e.g., stage 602) and a stage for a fractional sample refinement search (e.g., stage 604), where each stage can be performed in at least one iteration. In some embodiments, up to six SADs can be checked in the first iteration of the DMVR integer luma sample search. Using FIG. 7 as an example, the SADs of five points 702-710 (represented as black blocks) can be compared in the first iteration, with point 702 serving as the center point for the search. If the center point (i.e., point 702) has the smallest SAD, the DMVR integer sample stage can be terminated. Otherwise, another point 712 (represented as a shaded block) can be checked, as determined by the SAD distribution of points 704-710. In a second iteration of the DMVR integer luma sample search, the point with the smallest SAD among points 704-712 can be selected as the new center point for the search. In some embodiments, the second iteration can be performed in the same way as the first iteration. In some embodiments, the SAD calculated in the first iteration can be reused in the second iteration, and thus only the SAD of additional points may need to be further calculated.
[0079]
[0108] In some embodiments, as described in VVC Draft 6, the two-iteration search illustrated in FIG. 7 can be eliminated. Then, in the stage for integer sample offset search (e.g., stage 602 in FIG. 6 ), all SADs of 25 points can be calculated in a single iteration. FIG. 8 shows another exemplary pattern 800 for the stage for integer sample offset search in DMVR integer luma sample search according to some embodiments of the present disclosure. For example, the stage for integer sample offset search can be stage 602 in FIG. 6 . FIG. 8 shows an initial MV 802 and 25 points for which SADs can be calculated simultaneously. In some embodiments, the SAD of the initial MV 802 can be reduced (e.g., by a factor of four) to flavor the initial MV 802. In some embodiments, the position with the smallest SAD can be further refined in the stage for fractional sample refinement (e.g., stage 604 in FIG. 6 ). The stage for fractional sample refinement can be conditionally invoked based on the position with the smallest SAD. 8, if the position with the smallest SAD is one of the nine points around the initial MV 802 (as represented by box 804), a stage for fractional sample refinement can be invoked to determine a refined MV as the output of the DMVR integer luma sample search. If the position with the smallest SAD is not one of the nine points around the initial MV 802, the position with the smallest SAD can be directly used as the output of the DMVR integer luma sample search.
[0080]
[0109] 9 is a schematic diagram illustrating an example pattern 900 for DMVR parameter error surface estimation, according to some embodiments of the present disclosure. FIG. 8 shows an initial MV 902 and 25 points. The initial MV 902 is connected to a center point 904 with the smallest SAD. In parameter error surface-based sub-pixel offset estimation, as shown in FIG. 9, the sum of absolute difference (SAD) cost of the center point 904 and the SAD costs of four neighboring points 906-912 around the center point 904 can be used to fit a two-dimensional parabolic error surface equation. For example, the two-dimensional parabolic error surface equation can be based on Equation (3): E(x,y)=((A(xx min ) 2 +B(yy min ) 2 +)>>mvShift)+E(0,0) Equation (3)
[0081]
[0110] In equation (3), (x min ,y min ) corresponds to the fractional position with the smallest SAD cost, E(x,y) corresponds to the SAD costs of the center point 904 and the four adjacent points 906-912, mvShift can be set to 4 as in VVC (in VVC, the MV precision is 1 / 16 pel), and A and B can be determined based on equations (4) and (5), respectively:
number
[0082]
[0111] By solving equations (3) to (5) using the SAD cost values of the five search points (i.e., points 904 to 912), (x min ,y min ) can be determined:
number
[0083]
[0112] In some embodiments, all SAD cost values are positive, with the smallest value being E(0,0), so x min and y min The value of x may be automatically constrained to be between -8 and 8 (e.g., with a precision of 1 / 16 sample), which corresponds to a half-pel offset with MV precision of 1 / 16 pel in VVC. min ,y min ) can be added to the integer distance refinement MV, allowing the refinement MV to have sub-pel accuracy.
[0084]
[0113] VVC includes a bidirectional optical flow (BDOF) tool. As the name suggests, BDOF mode is based on the concept of optical flow, which assumes smooth object motion. BDOF was previously called BIO and is also included in the Joint Video Exploration Model (JEM) software. Compared to BIO in JEM, BDOF in VVC is a simpler version that requires much less computation, especially in terms of the number of multiplications and the size of the multipliers.
[0085]
[0114] BDOF can be used to refine the bidirectional prediction signal of a CU at the 4x4 sub-block level. In some embodiments, BDOF is applied to a CU that meets the following conditions: (1) the height of the CU is not 4 and the size of the CU is not 4x8, (2) the CU is not coded using an affine mode or an advanced temporal motion vector prediction (ATMVP) merge mode, and (3) the CU is coded using a "true" bidirectional prediction mode in which one of two reference pictures (e.g., the first reference picture 504 in FIG. 5) is before the current picture (e.g., the current picture 502 in FIG. 5) in display order, and the other (e.g., the second reference picture 506 in FIG. 5) is after the current picture in display order. In some embodiments, BDOF can be applied to the luma component.
[0086]
[0115] In some embodiments, when BDOF is used to refine the bidirectionally predicted signal of a CU at the 4x4 sub-block level, motion refinement (v) is performed by minimizing the difference between the predicted samples in the two reference picture lists L0 and L1 for each 4x4 sub-block. x ,v y ) can be calculated. (v x ,v y ) can then be used to adjust the bi-predicted sample values within the 4x4 sub-block.
[0087]
[0116] In some embodiments, the following steps are applied in the BDOF process: First, the horizontal and vertical gradients of the two predicted signals, k=0,1,
number
number
[0088]
[0117] In equations (8) and (9), I (k) (i,j) is the sample value at coordinate (i,j) of the predicted signal in list k, where k=0,1, and shift1 is calculated based on the luma bit depth ("bitDepth") as equation (10): shift1=max(2,14-bitDepth) Equation (10)
[0089]
[0118] Next, the auto- and cross-correlations of the gradients S1, S2, S3, S5, and S6 can be determined based on equations (11)-(15):
number
[0090]
[0119] Regarding equations (11) to (15), ψx (i,j), ψ y The values of (i,j) and θ(i,j) can be determined based on equations (16) to (18):
number
[0091]
[0120] In equations (11) to (18), Ω is a 6 × 6 window around a 4 × 4 sub-block, and n a and n b The values of are set to satisfy equations (19) and (20), respectively: n a =min(5,bitDepth-7) Equation (19) n b =min(8,bitDepth-4) Equation (20)
[0092]
[0121] Next, we use equations (21) and (22) to refine the motion (v x ,v y ) is derived:
number
[0093]
[0122] In equations (21) and (22),
number
number
number
number
number
[0094]
[0123] Based on the motion refinement and gradients, the following adjustment b(x,y) can be determined for each sample in the 4×4 sub-block using equation (23):
number
[0095]
[0124] Finally, the BDOF samples of the CU can be determined based on adjusting the bi-prediction samples according to equation (24): pred BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+o offset )≫shift expression (24)
[0096]
[0125] In some embodiments, the values in equations (8) to (23) can be selected so that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit width of the intermediate parameters in the BDOF process can be kept within 32 bits.
[0097]
[0126] To derive the gradient value, we select some predicted samples I in list k (k=0,1) outside the current CU boundary. (k)(i,j) can be generated. Figure 10 is a schematic diagram of an example of an extended coding unit (CU) region 1000 used in BDOF according to some embodiments of the present disclosure. As shown in Figure 10, a 4x4 block 1002 (surrounded by a solid black line) used in BDOF is surrounded by one extended row or column (represented by a dashed black line) around the boundary of the block 1002, forming a surrounding region 1004. To control the computational complexity of generating prediction samples outside the boundary, prediction samples within the extended area 1006 (represented by a white frame) can be generated by directly taking reference samples at nearby integer positions without using interpolation (using a round-down operation on coordinates), and a normal 8-tap motion compensation interpolation filter can be used to generate prediction samples within the CU 1008 (represented by a gray frame). These extended sample values can only be used in gradient calculation. For the remaining steps in the BDOF process, any samples and gradient values outside the boundary of the CU 1008 can be padded (or repeated) from their nearest neighbors if needed.
[0098]
[0127] At the JVET meeting, a coding tool called optical flow prediction refinement (PROF) was adopted. PROF improves the accuracy of affine motion compensation prediction by refining subblock-based affine motion compensation prediction using optical flow. Affine motion model parameters can be used to derive motion vectors for each sample position within a CU. However, due to the complexity and high memory access bandwidth required to generate affine motion compensation predictions for each sample, affine prediction in VVC uses a subblock-based affine motion compensation method in which a CU is divided into 4x4 subblocks and each subblock is assigned a motion vector derived from the control point motion vectors of the affine CU. Subblock-based affine motion compensation is a trade-off between coding efficiency, complexity, and memory access bandwidth. It loses some prediction accuracy due to the use of subblock-based prediction instead of theoretical sample-based motion compensation prediction.
[0099]
[0128] To achieve finer granularity of affine motion compensation, in some embodiments, PROF can be applied after regular sub-block-based affine motion compensation. Sample-based refinement can be derived based on the optical flow equation, such as Equation (25): ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) Equation (25)
[0100]
[0129] In equation (25), g x (i,j) and g y (i,j) is the spatial gradient at sample position (i,j). Δv is the motion offset from the sub-block-based motion vector to the sample-based motion vector derived from the affine model parameters.
[0101]
[0130] 11 is a schematic diagram of an example of sub-block-based translational motion and sample-based affine motion according to some embodiments of the present disclosure. As shown in FIG. 11, V(i,j) is a theoretical motion vector for sample position (i,j) derived using an affine model, and V SB is a subblock-based motion vector, and ΔV(i,j) (represented as a dotted arrow) is the relationship between V(i,j) and V SB This is the difference between
[0102]
[0131] The prediction refinement ΔI(i,j) can then be added to the sub-block prediction I(i,j). The final prediction I′ can be generated based on equation (26): I'(i,j)=I(i,j)+ΔI(i,j) Equation (26)
[0103]
[0132] According to embodiments of the present disclosure, both DMVR and BDOF can have control flags at two levels in the syntax structure. The first control flag can be transmitted in the sequence parameter set (SPS) at the sequence level, and the second control flag can be transmitted in the slice header at the slice level. FIG. 12 shows Table 1 illustrating an example syntax structure of a sequence parameter set (SPS) implementing control flags for DMVR and BDOF according to some embodiments of the present disclosure. As shown in Table 1 of FIG. 12, sps_bdof_enabled_flag and sps_dmvr_enabled_flag are control flags for BDOF and DMVR at the sequence level, respectively, transmitted in the SPS. When sps_bdof_enabled_flag or sps_dmvr_enabled_flag is false, BDOF or DMVR can be disabled within the entire video sequence that references this SPS. When sps_bdof_enabled_flag and sps_dmvr_enabled_flag are true, BDOF or DMVR can be enabled for the current video sequence. In this case, another flag sps_bdof_dmvr_slice_present_flag can be further signaled to indicate whether slice-level control of BDOF and DMVR is enabled.
[0104]
[0133] 13 shows Table 2 illustrating an example syntax structure of a slice header implementing control flags for DMVR and BDOF according to some embodiments of the present disclosure. As shown in Table 2, when sps_bdof_dmvr_slice_present_flag is true as set in Table 1, slice_disable_bdof_dmvr_flag can be included in the slice header to signal whether BDOF and DMVR are disabled for the current slice.
[0105]
[0134] 12-13 show a two-level control mechanism for DMVR and BDOF. By using such a mechanism, an encoder (e.g., an encoder implementing process 200A or 200B in FIGS. 2A-2B) can use a slice-level flag slice_disable_bdof_dmvr_flag to switch DMVR and BDOF on or off for individual slices. Such slice-level adaptation can have two advantages: (1) when at least one of DMVR or BDOF is not useful for the current slice, switching it off can improve coding performance; and (2) DMVR and BDOF have relatively high computational complexity, and therefore, switching them off can reduce the coding and decoding complexity of the current slice.
[0106]
[0135] In embodiments of the present disclosure, control flags for PROF can also be used at both the sequence level and the slice level. In some embodiments, three separate flags can be included in the SPS to signal, respectively, to indicate whether DMVR, BDOF, and PROF are enabled. If any of them is enabled, a corresponding lower-level control enable flag can be signaled to indicate whether the enabled tool is controlled at a lower level. The lower level can be the slice level or the picture level. If slice-level or picture-level control is enabled, then a slice-level or picture-level disable flag can be included in each slice or picture header to signal whether the enabled tool is disabled for the current slice or picture.
[0107]
[0136] In accordance with an embodiment of the present disclosure, Figure 14A shows Table 3A illustrating an example syntax structure of a sequence parameter set (SPS) that implements slice-level control flags for DMVR, BDOF, and PROF, according to some embodiments of the present disclosure. Figure 14B shows Table 3B illustrating an example syntax structure of an SPS that implements picture-level control flags for DMVR, BDOF, and PROF, according to some embodiments of the present disclosure. As shown in Tables 3A-3B with italic highlighting, sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag are flags that are signaled in the SPS to indicate whether BDOF, DMVR, and PROF, respectively, are enabled for the video sequence. If BDOF, DMVR, or PROF is enabled, sps_bdof_slice_present_flag, sps_dmvr_slice_present_flag, or sps_affine_prof_slice_present_flag may be further signaled to indicate whether slice-level control of BDOF, DMVR, and PROF is enabled, respectively, as shown in Table 3 A. If BDOF, DMVR, or PROF is enabled, sps_bdof_picture_present_flag, sps_dmvr_picture_present_flag, or sps_affine_prof_picture_present_flag may be further signaled to indicate whether picture-level control of BDOF, DMVR, and PROF is enabled, respectively, as shown in Table 3B.
[0108]
[0137]
[0071] Figure 15A shows Table 4A illustrating an example syntax structure of a slice header implementing control flags for DMVR, BDOF, and PROF, according to some embodiments of the present disclosure. As shown in Table 4A with italicized highlighting, if any of sps_bdof_slice_present_flag, sps_dmvr_slice_present_flag, or sps_affine_prof_slice_present_flag as set in Table 3A is true, then slice_disable_bdof_flag, slice_disable_dmvr_flag, or slice_disable_affine_prof_flag may be signaled to indicate whether BDOF, DMVR, or PROF is disabled for the current slice, respectively. Figure 15B shows Table 4B illustrating an example syntax structure of a picture header implementing control flags for DMVR, BDOF, and PROF, according to some embodiments of the present disclosure. As shown in Table 4B with italic highlighting, if any of sps_bdof_picture_present_flag, sps_dmvr_picture_present_flag, or sps_affine_prof_picture_present_flag as set in Table 3B is true, then ph_disable_bdof_flag, ph_disable_dmvr_flag, or ph_disable_affine_prof_flag may be signaled to indicate whether BDOF, DMVR, or PROF is disabled for the current picture, respectively.
[0109]
[0138] In some embodiments, DMVR, BDOF, and PROF may have three separate sequence-level enable flags but share the same slice-level control enable flag. For example, one slice-level disable flag may be signaled for DMVR, BDOF, and PROF. In another example, three slice-level disable flags may be signaled separately for DMVR, BDOF, and PROF. As another example, two slice-level disable flags may be signaled for DMVR, BDOF, and PROF. It should be noted that various syntax structures may be implemented for control of DMVR, BDOF, and PROF at the sequence level and at levels lower than the sequence level (referred to herein as "lower levels," such as the slice level or picture level), not limited to the examples described herein.
[0110]
[0139] 16 shows Table 5 illustrating an example syntax structure of a sequence parameter set (SPS) that implements separate sequence-level control flags for DMVR, BDOF, and PROF, according to some embodiments of the present disclosure. As shown in Table 5 with italic highlighting, three separate flags, sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag, are signaled in the SPS to indicate whether DMVR, BDOF, and PROF are enabled, respectively. If at least one of DMVR, BDOF, or PROF is enabled, a slice control enable flag, sps_bdof_dmvr_affine_prof_slice_present_flag, can be signaled to indicate whether at least one of DMVR, BDOF, or PROF enabled at the sequence level is controlled at a lower level.
[0111]
[0140] 17 shows Table 6 illustrating an example syntax structure of a slice header implementing joint control flags for DMVR, BDOF, and PROF, according to some embodiments of the present disclosure. As shown in Table 6 with italic highlighting, if slice-level control is enabled as described in Table 5 (e.g., sps_bdof_dmvr_affine_prof_slice_present_flag is true), then a slice-level disable flag slice_disable_bdof_dmvr_affine_prof_flag may be included in each slice header to signal whether at least one of DMVR, BDOF, or PROF enabled at the sequence level is disabled for the current slice. In the illustrated syntax structure of Table 6, if multiple DMVRs, BDOFs, and PROFs are enabled at the sequence level, the enabled ones can be jointly controlled at the slice level if slice-level control is enabled.
[0112]
[0141] 18 shows Table 7 illustrating an example syntax structure of a slice header implementing separate control flags for DMVR, BDOF, and PROF, according to some embodiments of the present disclosure. As shown in Table 7 with italic highlighting, if slice-level control is enabled as described in Table 5 (e.g., sps_bdof_dmvr_affine_prof_slice_present_flag is true), then for enabled ones of DMVR, BDOF, and PROF at the sequence level, a slice-level disable flag may be signaled to indicate whether the enabled one is disabled for the current slice. For example, if sps_bdof_enabled_flag and sps_bdof_dmvr_affine_prof_slice_present_flag in Table 5 are set as true, slice_disable_bdof_flag in Table 7 may be signaled to indicate whether BDOF is disabled for the current slice. As another example, if sps_dmvr_enabled_flag and sps_bdof_dmvr_affine_prof_slice_present_flag in Table 5 are set as true, then slice_disable_dmvr_flag in Table 7 can be signaled to indicate whether DMVR is disabled for the current slice. In yet another example, if sps_affine_prof_enabled_flag and sps_bdof_dmvr_affine_prof_slice_present_flag in Table 5 are set as true, then slice_disable_affine_prof_flag in Table 7 can be signaled to indicate whether PROF is disabled for the current slice. In Table 7, each of DMVR, BDOF, and PROF can be controlled separately at the slice level if slice-level control is enabled.
[0113]
[0142] Considering the fact that both BDOF and PROF use optical flow to refine inter-predictors, in some embodiments, BDOF and PROF may share the same slice-level control flag, and DMVR may use a separate slice-level control flag. Figure 19 shows Table 8, which illustrates an example syntax structure of a slice header implementing hybrid control flags for DMVR, BDOF, and PROF, according to some embodiments of the present disclosure. In the syntax structure of Table 8, BDOF and PROF may share the same slice-level disable flag, and DMVR may use a separate slice-level disable flag. As shown in Table 8 with italic highlighting, when slice-level control is enabled in Table 5 (e.g., sps_bdof_dmvr_affine_prof_slice_present_flag is true), two slice-level disable flags may be signaled to indicate whether DMVR, BDOF, and PROF are disabled for the current slice. For example, if at least one of sps_bdof_enabled_flag and sps_affine_prof_enabled_flag in Table 5 is set as true and sps_bdof_dmvr_affine_prof_slice_present_flag in Table 5 is set as true, then slice_disable_bdof_affine_prof_flag in Table 8 may be signaled to indicate whether at least one of BDOF or PROF enabled at the sequence level is disabled for the current slice. In another example, if sps_dmvr_enabled_flag and sps_bdof_dmvr_affine_prof_slice_present_flag in Table 5 are set as true, then slice_disable_dmvr_flag in Table 8 may be signaled to indicate whether DMVR is disabled for the current slice.In Table 8, BDOF and PROF are jointly controlled at the slice level, and DMVR is controlled separately from BDOF and PROF when slice-level control is enabled. Figure 20 shows Table 9, which illustrates an example syntax structure of a sequence parameter set (SPS) that implements hybrid sequence-level control flags for DMVR, BDOF, and PROF, according to some embodiments of the present disclosure. In the syntax structure of Table 9, three slice-level disable flags can be signaled separately for DMVR, BDOF, and PROF. As shown in Table 9 with italic highlighting, three separate flags, sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag, can be included in the SPS and signaled to indicate whether DMVR, BDOF, or PROF is enabled, respectively. If sps_dmvr_enabled_flag is true, the slice level control enable flag sps_dmvr_slice_present_flag may be signaled to indicate whether the DMVR is controlled at the slice level. If at least one of sps_bdof_enabled_flag or sps_affine_prof_enabled_flag is true, the slice level control enable flag sps_bdof_affine_prof_slice_present_flag may be signaled to indicate whether at least one of BDOF or PROF is controlled at the slice level.
[0114]
[0143] 21 shows Table 10 illustrating another example syntax structure of a slice header implementing hybrid control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As shown in Table 10 with italicized highlighting, if sps_dmvr_slice_present_flag in Table 9 is set as true, then the slice-level disable flag slice_disable_dmvr_flag can be signaled to indicate whether DMVR is disabled for the current slice. If sps_bdof_affine_prof_slice_present_flag in Table 9 is set as true, then the slice-level disable flag slice_disable_bdof_affine_prof_flag can be signaled to indicate whether at least one of BDOF or PROF enabled at the sequence level (as described in Table 8) can be disabled for the current slice. In the syntax structure of Table 9, BDOF and PROF can be jointly controlled at the slice level, and DMVR can be separately controlled at the slice level if slice level control is enabled.
[0115]
[0144] 22 shows Table 11 illustrating another example syntax structure of a slice header implementing separate control flags for DMVR, BDOF, and PROF, according to some embodiments of the present disclosure. As shown in Table 11 with italicized highlighting, if sps_dmvr_enabled_flag in Table 9 is set as true, then the slice-level disable flag slice_disable_dmvr_flag can be signaled to indicate whether DMVR is disabled for the current slice. If sps_bdof_enabled_flag and sps_bdof_affine_prof_slice_present_flag in Table 9 are set as true, then the slice-level disable flag slice_disable_bdof_flag can be signaled to indicate whether BDOF is disabled for the current slice. If sps_affine_prof_enabled_flag and sps_bdof_affine_prof_slice_present_flag in Table 9 are set as true, the slice-level disable flag slice_disable_affine_prof_flag can be signaled to indicate whether PROF is disabled for the current slice. In the syntax structure of Table 11, each of DMVR, BDOF, and PROF is controlled separately at the slice level if slice-level control is enabled.
[0116]
[0145] 23-26 show flowcharts of example processes 2300-2600 for controlling video coding modes according to some embodiments of the present disclosure. In some embodiments, processes 2300-2600 may be performed by a codec (e.g., the encoder in FIGS. 2A-2B or the decoder in FIGS. 3A-3B). For example, the codec may be implemented as one or more software or hardware components of an apparatus (e.g., apparatus 400) for controlling coding modes for encoding or decoding a video sequence.
[0117]
[0146] 23 shows a flowchart of an example process 2300 for controlling a video decoding mode according to some embodiments of the present disclosure. In step 2302, a codec (e.g., a decoder in FIGS. 3A-3B) may receive a bitstream of video data (e.g., video bitstream 228 in processes 300A or 300B in FIGS. 3A-3B).
[0118]
[0147] In step 2304, the codec may enable or disable a coding mode for a video sequence (e.g., video stream 304 in process 300A or 300B in FIGS. 3A-3B) based on a first flag in the bitstream. For example, the coding mode may be at least one of a bidirectional optical flow (BDOF) mode, a prediction refinement by optical flow (PROF) mode, or a decoder-side motion vector refinement (DMVR) mode. In some embodiments, the codec may detect a first flag in a sequence parameter set (SPS) of the video sequence. For example, the first flag may be a flag sps_bdof_enabled_flag, a flag sps_dmvr_enabled_flag, or a flag sps_affine_prof_enabled_flag as described in FIG. 12, 14A-14B, 16, or 20.
[0119]
[0148] In step 2306, the codec may determine whether coding mode control is enabled or disabled at a level below the sequence level based on a second flag in the bitstream. Levels below the sequence level may include the slice level or the picture level. In some embodiments, the codec may detect a second flag in the SPS of the video sequence in response to a coding mode being enabled for the video sequence. For example, the second flag may be the flag sps_bdof_dmvr_slice_present_flag, the flag sps_bdof_slice_present_flag, the flag sps_dmvr_slice_present_flag, the flag sps_affine_prof_slice_present_flag, the flag sps_bdof_picture_present_flag, the flag sps_dmvr_picture_present_flag, the flag sps_affine_prof_picture_present_flag, the flag sps_bdof_affine_prof_slice_present_flag, or the flag sps_bdof_dmvr_affine_prof_slice_present_flag as described in Figure 12, Figures 14A to 14B, Figure 16, or Figure 20.
[0120]
[0149] In some embodiments, after step 2306, in response to the coding mode control being enabled at a level lower than the sequence level, the codec can enable or disable the coding mode for the target lower-level region based on a third flag in the bitstream. The target lower-level region can be a target slice or a target picture. If the lower level is the slice level, in some embodiments, the codec can detect the third flag in a slice header of the target slice. If the lower level is the picture level, in some embodiments, the codec can detect the third flag in a picture header of the target picture. For example, the third flag can be the flag slice_disable_bdof_dmvr_flag, the flag slice_disable_bdof_flag, the flag slice_disable_dmvr_flag, the flag slice_disable_affine_prof_flag, the flag ph_disable_bdof_flag, the flag ph_disable_dmvr_flag, the flag ph_disable_affine_prof_flag, the flag slice_disable_bdof_dmvr_affine_prof_flag, or the flag slice_disable_bdof_affine_prof_flag as described in Figures 13, 15A to 15B, 17 to 19, or 21 to 22.
[0121]
[0150] 24 shows a flowchart of another example process 2400 for controlling a video decoding mode according to some embodiments of the present disclosure. In step 2402, a codec (e.g., a decoder in FIGS. 3A-3B) may receive a bitstream of video data (e.g., video bitstream 228 in processes 300A or 300B in FIGS. 3A-3B).
[0122]
[0151] In step 2404, the codec may enable or disable a first coding mode for a video sequence (e.g., video stream 304 in process 300A or 300B in FIGS. 3A-3B) based on a first flag in the bitstream. In step 2406, the codec may enable or disable a second coding mode for a video sequence based on a second flag in the bitstream. The first and second coding modes may be two different coding modes that may be selected from a bidirectional optical flow (BDOF) mode, a prediction refinement based on optical flow (PROF) mode, and a decoder-side motion vector refinement (DMVR) mode. For example, the first coding mode and the second coding mode may be a bidirectional optical flow (BDOF) mode and a prediction refinement based on optical flow (PROF) mode, respectively.
[0123]
[0152] In some embodiments, the codec may detect first and second flags in a sequence parameter set (SPS) of a video sequence. For example, the first and second flags may be selected from the flags sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag as described in Figure 12, Figures 14A-14B, Figure 16, or Figure 20. As another example, if the first and second encoding modes are the BDOF mode and the PROF mode, respectively, the first and second flags may be the flags sps_bdof_enabled_flag and sps_affine_prof_enabled_flag, respectively, as described in Figure 12, Figures 14A-14B, Figure 16, or Figure 20.
[0124]
[0153] In step 2408, the codec may determine, based on a third flag in the bitstream, whether control of at least one of the first coding mode or the second coding mode is enabled at a level below the sequence level. The level below the sequence level may include the slice level or the picture level. In some embodiments, the codec may detect a third flag in the SPS of the video sequence in response to at least one of the first coding mode or the second coding mode being enabled for the video sequence. For example, the third flag may be the flag sps_bdof_dmvr_slice_present_flag, the flag sps_bdof_slice_present_flag, the flag sps_dmvr_slice_present_flag, the flag sps_affine_prof_slice_present_flag, the flag sps_bdof_picture_present_flag, the flag sps_dmvr_picture_present_flag, the flag sps_affine_prof_picture_present_flag, the flag sps_bdof_affine_prof_slice_present_flag, or the flag sps_bdof_dmvr_affine_prof_slice_present_flag as described in Figure 12, Figures 14A to 14B, Figure 16, or Figure 20.
[0125]
[0154] In some embodiments, after step 2408, the codec can enable or disable the first coding mode (e.g., BDOF) for the target lower-level region based on a fourth flag (e.g., slice_disable_bdof_flag as described in FIG. 22) in the bitstream in response to the first flag (e.g., sps_bdof_enabled_flag as described in FIG. 20) indicating that the first coding mode is enabled for the video sequence and the third flag (e.g., sps_bdof_affine_prof_slice_present_flag as described in FIG. 20) indicating that control of at least one of the first coding mode or the second coding mode (e.g., PROF) is enabled at a level lower than the sequence level. The target lower-level region can be a target slice or a target picture. If the target lower-level region is a target slice, in some embodiments, the codec can detect the fourth flag in the slice header of the target slice. If the target lower level is a target picture, in some embodiments the codec can detect a fourth flag in the picture header of the target picture. For example, the fourth flag can be the flag slice_disable_bdof_dmvr_flag, the flag slice_disable_bdof_flag, the flag slice_disable_dmvr_flag, the flag slice_disable_affine_prof_flag, the flag ph_disable_bdof_flag, the flag ph_disable_dmvr_flag, the flag ph_disable_affine_prof_flag, the flag slice_disable_bdof_dmvr_affine_prof_flag, or the flag slice_disable_bdof_affine_prof_flag as described in Figures 13, 15A-15B, 17-19, or 21-22.
[0126]
[0155] In some embodiments, after step 2408, the codec may enable or disable both the first coding mode (e.g., BDOF) and the second coding mode (e.g., PROF) for the target lower-level region based on a fourth flag in the bitstream (e.g., slice_disable_bdof_affine_prof_flag as described in FIG. 21) in response to control of at least one of the first coding mode or the second coding mode being enabled at a level lower than the sequence level. For example, a third flag (e.g., sps_bdof_affine_prof_slice_present_flag as described in FIG. 20) may indicate that at least one of the first coding mode or the second coding mode is enabled at a lower level (e.g., slice level).
[0127]
[0156] In some embodiments, after enabling or disabling both the first coding mode and the second coding mode for a target lower-level region (e.g., a target slice or a target picture) based on a fourth flag in the bitstream, the codec may further enable or disable a third coding mode for a video sequence based on a second flag in the bitstream, and determine whether control of the third coding mode is enabled at a level lower than the sequence level based on a fifth flag in the bitstream. For example, the first, second, and third coding modes may be a BDOF mode, a PROF mode, and a DMVR mode, respectively. In this example, the fourth flag may be the flag slice_disable_bdof_affine_prof_flag as described in FIG. 21, the second flag may be the flag sps_dmvr_enabled_flag as described in FIG. 20, and the fifth flag may be the flag sps_dmvr_slice_present_flag as described in FIG. 20.
[0128]
[0157] In some embodiments, in response to the third coding mode being enabled at a lower level (e.g., slice level or picture level), the codec may further enable or disable the third coding mode for a target lower-level region (e.g., target slice or target picture) based on a sixth flag in the bitstream. For example, when the first, second, and third coding modes can be BDOF mode, PROF mode, and DMVR mode, respectively, the sixth flag may be slice_disable_dmvr_flag as described in Figure 21.
[0129]
[0158] 25 shows a flowchart of an example process 2500 for controlling a video encoding mode according to some embodiments of the present disclosure. In step 2502, a codec (e.g., the encoder in FIGS. 2A-2B) may receive a video sequence (e.g., the video sequence 202 in processes 200A or 200B in FIGS. 2A-2B), a first flag, and a second flag. For example, the first flag may be the flag sps_bdof_enabled_flag, the flag sps_dmvr_enabled_flag, or the flag sps_affine_prof_enabled_flag as described in FIG. 12, 14A-14B, 16, or 20. As another example, the second flag may be the flag sps_bdof_dmvr_slice_present_flag, the flag sps_bdof_slice_present_flag, the flag sps_dmvr_slice_present_flag, the flag sps_affine_prof_slice_present_flag, the flag sps_bdof_picture_present_flag, the flag sps_dmvr_picture_present_flag, the flag sps_affine_prof_picture_present_flag, the flag sps_bdof_affine_prof_slice_present_flag, or the flag sps_bdof_dmvr_affine_prof_slice_present_flag as described in Figure 12, Figures 14A to 14B, Figure 16, or Figure 20.
[0130]
[0159] In step 2504, the codec may enable or disable a coding mode for a video bitstream (e.g., video bitstream 228 in process 200A or 200B in FIGS. 2A-2B) based on a first flag in the bitstream. The coding mode may be at least one of a bidirectional optical flow (BDOF) mode, a prediction refinement by optical flow (PROF) mode, or a decoder-side motion vector refinement (DMVR) mode.
[0131]
[0160] In step 2506, the codec can enable or disable coding mode control at a level below the sequence level based on the second flag, which can include the slice level or the picture level.
[0132]
[0161] 26 shows a flowchart of another example process 2600 for controlling a video encoding mode according to some embodiments of the present disclosure. In step 2602, a codec (e.g., the encoder in FIGS. 2A-2B) may receive a video sequence (e.g., the video sequence 202 in processes 200A or 200B in FIGS. 2A-2B), a first flag, a second flag, and a third flag. For example, the first and second flags may be selected from the flags sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag as described in FIG. 12, 14A-14B, 16, or 20. As another example, the third flag may be the flag sps_bdof_dmvr_slice_present_flag, the flag sps_bdof_slice_present_flag, the flag sps_dmvr_slice_present_flag, the flag sps_affine_prof_slice_present_flag, the flag sps_bdof_picture_present_flag, the flag sps_dmvr_picture_present_flag, the flag sps_affine_prof_picture_present_flag, the flag sps_bdof_affine_prof_slice_present_flag, or the flag sps_bdof_dmvr_affine_prof_slice_present_flag as described in Figure 12, Figures 14A to 14B, Figure 16, or Figure 20.
[0133]
[0162] In step 2604, the codec may enable or disable a first encoding mode for the video bitstream (e.g., video bitstream 228 in process 200A or 200B in FIGS. 2A-2B) based on the first flag. In step 2606, the codec may enable or disable a second encoding mode for the video bitstream based on the second flag. The first and second encoding modes may be two different encoding modes that may be selected from a bidirectional optical flow (BDOF) mode, a prediction refinement by optical flow (PROF) mode, and a decoder-side motion vector refinement (DMVR) mode. For example, the first encoding mode and the second encoding mode may be a bidirectional optical flow (BDOF) mode and a prediction refinement by optical flow (PROF) mode, respectively.
[0134]
[0163] In step 2608, the codec may enable or disable control of at least one of the first encoding mode or the second encoding mode at a level below the sequence level based on the third flag. The level below the sequence level may include the slice level or the picture level.
[0135]
[0164] Some embodiments also provide a non-transitory computer-readable storage medium containing instructions that can be executed by a device (such as the encoders and decoders of the present disclosure) to perform the above-described methods. Common forms of non-transitory media include, for example, a floppy disk, a flexible disk, a hard disk, a solid-state drive, a magnetic tape, or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROM, and EPROM, FLASH-EPROM or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge, and networked versions thereof. A device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.
[0136]
[0165] The embodiments may be further described using the following clauses: 1. A computer-implemented method comprising: receiving a bitstream of video data; enabling or disabling a coding mode for the video sequence based on a first flag in the bitstream; determining whether coding mode control is enabled or disabled at a level below the sequence level based on a second flag in the bitstream; 20. A computer-implemented method comprising: 2. The computer-implemented method of clause 1, wherein the level below the sequence level includes the slice level or the picture level. 3. The computer-implemented method of clause 1 or 2, further comprising enabling or disabling the coding mode for the target lower-level region based on a third flag in the bitstream in response to coding mode control being enabled at a level lower than the sequence level. 4. Detecting a third flag in the slice header of the target slice, where the target slice is a target lower-level region; or detecting a third flag in a picture header of the target picture, wherein the target picture is a target lower-level region; 4. The computer-implemented method of claim 3, further comprising: 5.Encoding mode is Bidirectional Optical Flow (BDOF) mode, Optical Flow Prediction Refinement (PROF) mode, or Decoder-side Motion Vector Refinement (DMVR) mode, 5. The computer-implemented method of any one of clauses 1 to 4, wherein at least one of 6. The computer-implemented method of any one of clauses 1 to 5, further comprising detecting a first flag in a sequence parameter set (SPS) of the video sequence. 7. A computer-implemented method described in any one of clauses 1 to 6, further comprising detecting a second flag in the SPS of the video sequence in response to the encoding mode being enabled for the video sequence. 8. A computer-implemented method comprising: receiving a bitstream of video data; enabling or disabling a first coding mode for the video sequence based on a first flag in the bitstream; enabling or disabling a second coding mode for the video sequence based on a second flag in the bitstream; determining whether control of at least one of the first coding mode or the second coding mode is enabled at a level below the sequence level based on a third flag in the bitstream; 20. A computer-implemented method comprising: 9. The computer-implemented method of clause 8, wherein the level below the sequence level includes the slice level or the picture level. 10. The computer-implemented method described in clause 8 or 9, further comprising enabling or disabling both the first coding mode and the second coding mode for the target lower-level region based on a fourth flag in the bitstream in response to control of at least one of the first coding mode or the second coding mode being enabled at a level lower than the sequence level. 11. Detecting a fourth flag in the slice header of the target slice, where the target slice is a target lower-level region; or detecting a fourth flag in a picture header of the target picture, wherein the target picture is a target lower-level region; 11. The computer-implemented method of clause 10, further comprising: 12. Enabling or disabling a third coding mode for the video sequence based on a second flag in the bitstream; determining whether control of the third coding mode is enabled at a level below the sequence level based on a fifth flag in the bitstream; 12. The computer-implemented method of clause 10 or 11, further comprising: 13. The computer-implemented method of clause 12, further comprising enabling or disabling the third coding mode for the target lower-level region based on a sixth flag in the bitstream in response to the third coding mode being enabled at a level lower than the sequence level. 14. The computer-implemented method of clause 12 or 13, wherein the first, second, and third encoding modes are a bidirectional optical flow (BDOF) mode, a prediction refinement by optical flow (PROF) mode, and a decoder-side motion vector refinement (DMVR) mode, respectively. 15. A computer-implemented method described in any one of clauses 8 to 14, further comprising enabling or disabling the first coding mode for a target lower-level region based on a fourth flag in the bitstream in response to the first flag indicating that the first coding mode is enabled for the video sequence and the third flag indicating that control of at least one of the first coding mode or the second coding mode is enabled at a level lower than the sequence level. 16. The first and second encoding modes are: Bidirectional Optical Flow (BDOF) mode, Optical Flow Prediction Refinement (PROF) mode, and Decoder-side Motion Vector Refinement (DMVR) mode, 16. The computer-implemented method of any one of clauses 8 to 15, wherein the two different encoding modes are selected from: 17. The computer-implemented method of any one of clauses 8 to 16, wherein the first encoding mode and the second encoding mode are a bidirectional optical flow (BDOF) mode and a prediction refinement by optical flow (PROF) mode, respectively. 18. The computer-implemented method of any one of clauses 8-17, further comprising detecting first and second flags in a sequence parameter set (SPS) of the video sequence. 19. The computer-implemented method of any one of clauses 8 to 18, further comprising detecting a third flag in the SPS of the video sequence in response to at least one of the first encoding mode or the second encoding mode being enabled for the video sequence. 20. Receiving a video sequence, a first flag, and a second flag; enabling or disabling a coding mode for the video bitstream based on a first flag; enabling or disabling control of the coding mode at a level below the sequence level based on a second flag; 20. A computer-implemented method comprising: 21. The computer-implemented method of clause 20, wherein the level below the sequence level includes the slice level or the picture level. 22. Receiving a third flag; enabling or disabling the coding mode for the target lower-level region based on a third flag in response to controlling whether the coding mode is enabled or disabled at a level lower than the sequence level; 22. The computer-implemented method of clause 20 or 21, further comprising: 23. Storing a third flag in the slice header of the target slice, where the target slice is a target lower-level region; or storing a third flag in a picture header of the target picture, wherein the target picture is a target lower-level region; 23. The computer-implemented method of claim 22, further comprising: 24.Encoding mode is Bidirectional Optical Flow (BDOF) mode, Optical Flow Prediction Refinement (PROF) mode, or Decoder-side Motion Vector Refinement (DMVR) mode, 24. The computer-implemented method of any one of clauses 20 to 23, wherein at least one of 25. The computer-implemented method of any one of clauses 20-24, further comprising storing the first flag in a sequence parameter set (SPS) of the video bitstream. 26. The computer-implemented method of any one of clauses 20-25, further comprising storing a second flag within the SPS of the video bitstream in response to enabling or disabling the encoding mode for the video bitstream. 27. A computer-implemented method comprising: receiving a video sequence, a first flag, a second flag, and a third flag; enabling or disabling a first coding mode for the video bitstream based on the first flag; enabling or disabling a second coding mode for the video bitstream based on the second flag; enabling or disabling control of at least one of the first encoding mode or the second encoding mode at a level lower than the sequence level based on the third flag; 20. A computer-implemented method comprising: 28. The computer-implemented method of clause 27, wherein the level below the sequence level includes the slice level or the picture level. 29. Receiving a fourth flag; enabling or disabling the first coding mode for the target lower-level region based on a fourth flag in response to enabling control of the first coding mode for the video bitstream based on the first flag and enabling control of at least one of the first coding mode or the second coding mode at a level lower than the sequence level based on the third flag; 29. The computer-implemented method of clause 27 or 28, further comprising: 30. Receiving a fourth flag; The computer-implemented method of clause 27 or 28, further comprising enabling or disabling both the first encoding mode and the second encoding mode for the target lower-level region based on a fourth flag in response to enabling or disabling control of at least one of the first encoding mode or the second encoding mode at a level lower than the sequence level. 31. Storing a fourth flag in the slice header of the target slice, where the target slice is a target lower-level region; or storing a fourth flag in a picture header of the target picture, wherein the target picture is a target lower-level region; 31. The computer-implemented method of clause 29 or 30, further comprising: 32. Receiving a fifth flag; enabling or disabling a third coding mode for the video bitstream based on the second flag; enabling or disabling control of the third coding mode at a level below the sequence level based on a fifth flag; 32. The computer-implemented method of claim 31, further comprising: 33. Receiving a sixth flag; enabling or disabling the third coding mode for the target lower-level region based on a sixth flag in response to enabling or disabling control of the third coding mode at a level lower than the sequence level; 33. The computer-implemented method of clause 31 or 32, further comprising: 34. The computer-implemented method of any one of clauses 27 to 33, wherein the first, second, and third encoding modes are a bidirectional optical flow (BDOF) mode, a prediction refinement by optical flow (PROF) mode, and a decoder-side motion vector refinement (DMVR) mode, respectively. 35. The first and second encoding modes are: Bidirectional Optical Flow (BDOF) mode, Optical Flow Prediction Refinement (PROF) mode, and Decoder-side Motion Vector Refinement (DMVR) mode, 34. The computer-implemented method of any one of clauses 27 to 33, wherein the two different encoding modes are selected from: 36. The computer-implemented method of any one of clauses 27 to 35, wherein the first encoding mode and the second encoding mode are a bidirectional optical flow (BDOF) mode and a prediction refinement by optical flow (PROF) mode, respectively. 37. The computer-implemented method of any one of clauses 27-36, further comprising storing the first and second flags in a sequence parameter set (SPS) of the video bitstream. 38. The computer-implemented method of any one of clauses 27 to 37, further comprising storing a third flag within the SPS of the video sequence in response to enabling at least one of the first encoding mode or the second encoding mode for the video bitstream. 39. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by at least one processor of a device to cause the device to perform a method, the method comprising: receiving a bitstream of video data; enabling or disabling a coding mode for the video sequence based on a first flag in the bitstream; determining whether coding mode control is enabled or disabled at a level below the sequence level based on a second flag in the bitstream; 1. A non-transitory computer-readable medium comprising: 40. The non-transitory computer-readable medium of clause 39, wherein the level below the sequence level includes the slice level or the picture level. 41. A set of instructions executable by at least one processor of the device includes: 41. The non-transitory computer-readable medium of clause 39 or 40, further comprising, in response to control of the encoding mode being enabled at a level lower than the sequence level, enabling or disabling the encoding mode for the target lower-level region based on a third flag in the bitstream. 42. A set of instructions executable by at least one processor of the device includes: Detecting a third flag in the slice header of the target slice, where the target slice is a target lower-level region; or detecting a third flag in a picture header of the target picture, wherein the target picture is a target lower-level region; 42. The non-transitory computer-readable medium of claim 41, further comprising: 43.Encoding mode is Bidirectional Optical Flow (BDOF) mode, Optical Flow Prediction Refinement (PROF) mode, or Decoder-side Motion Vector Refinement (DMVR) mode, 43. The non-transitory computer-readable medium of any one of clauses 39 to 42, wherein at least one of 44. A set of instructions executable by at least one processor of the device includes: 44. The non-transitory computer-readable medium of any one of clauses 39 to 43, further comprising detecting a first flag in a sequence parameter set (SPS) of the video sequence. 45. A set of instructions executable by at least one processor of the device includes: A non-transitory computer-readable medium as described in any one of clauses 39 to 44, further comprising detecting a second flag in the SPS of the video sequence in response to the encoding mode being enabled for the video sequence. 46. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by at least one processor of a device to cause the device to perform a method, the method comprising: receiving a bitstream of video data; enabling or disabling a first coding mode for the video sequence based on a first flag in the bitstream; enabling or disabling a second coding mode for the video sequence based on a second flag in the bitstream; determining whether control of at least one of the first coding mode or the second coding mode is enabled at a level below the sequence level based on a third flag in the bitstream; 1. A non-transitory computer-readable medium comprising: 47. The non-transitory computer-readable medium of clause 46, wherein the level below the sequence level includes the slice level or the picture level. 48. A set of instructions executable by at least one processor of the device includes: 48. The non-transitory computer-readable medium of clause 46 or 47, further comprising enabling or disabling both the first coding mode and the second coding mode for the target lower-level region based on a fourth flag in the bitstream in response to control of at least one of the first coding mode or the second coding mode being enabled at a level lower than the sequence level. 49. A set of instructions executable by at least one processor of the device includes: Detecting a fourth flag in the slice header of the target slice, where the target slice is a target lower-level region; or detecting a fourth flag in a picture header of the target picture, wherein the target picture is a target lower-level region; 49. The non-transitory computer-readable medium of claim 48, further comprising: 50. A set of instructions executable by at least one processor of the device includes: enabling or disabling a third coding mode for the video sequence based on a second flag in the bitstream; determining whether control of the third coding mode is enabled at a level below the sequence level based on a fifth flag in the bitstream; 49. The non-transitory computer-readable medium of claim 48 or 49, further comprising: 51. A set of instructions executable by at least one processor of the device includes: The non-transitory computer-readable medium of clause 50, further comprising enabling or disabling the third coding mode for the target lower-level region based on a sixth flag in the bitstream in response to the third coding mode being enabled at a level lower than the sequence level. 52. The non-transitory computer-readable medium of clause 50 or 51, wherein the first, second, and third encoding modes are a bidirectional optical flow (BDOF) mode, a prediction refinement by optical flow (PROF) mode, and a decoder-side motion vector refinement (DMVR) mode, respectively. 53. A set of instructions executable by at least one processor of the device includes: A non-transitory computer-readable medium as described in any one of clauses 46 to 52, further comprising, in response to a first flag indicating that a first encoding mode is enabled for a video sequence and a third flag indicating that control of at least one of the first encoding mode or the second encoding mode is enabled at a level lower than the sequence level, enabling or disabling the first encoding mode for a target lower-level region based on a fourth flag in the bitstream. 54. The first and second encoding modes are: Bidirectional Optical Flow (BDOF) mode, Optical Flow Prediction Refinement (PROF) mode, and Decoder-side Motion Vector Refinement (DMVR) mode, 54. The non-transitory computer-readable medium of any one of clauses 46 to 53, wherein the encoding modes are two different encoding modes selected from: 55. A non-transitory computer-readable medium described in any one of clauses 46 to 54, wherein the first encoding mode and the second encoding mode are a bidirectional optical flow (BDOF) mode and a prediction refinement by optical flow (PROF) mode, respectively. 56. A set of instructions executable by at least one processor of the device includes: 56. The non-transitory computer-readable medium of any one of clauses 46 to 55, further comprising detecting first and second flags in a sequence parameter set (SPS) of the video sequence. 57. A set of instructions executable by at least one processor of the device includes: A non-transitory computer-readable medium as described in any one of clauses 46 to 56, further comprising detecting a third flag in the SPS of the video sequence in response to at least one of the first encoding mode or the second encoding mode being enabled for the video sequence. 58. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by at least one processor of a device to cause the device to perform a method, the method comprising: receiving a video sequence, a first flag, and a second flag; enabling or disabling a coding mode for the video bitstream based on a first flag; enabling or disabling control of the coding mode at a level below the sequence level based on a second flag; 1. A non-transitory computer-readable medium comprising: 59. The non-transitory computer-readable medium of clause 58, wherein the level below the sequence level includes the slice level or the picture level. 60. A set of instructions executable by at least one processor of the device includes: receiving a third flag; enabling or disabling the coding mode for the target lower-level region based on a third flag in response to controlling enabling or disabling of the coding mode at a level lower than the sequence level; 59. The non-transitory computer-readable medium of clause 58 or 59, further comprising: 61. A set of instructions executable by at least one processor of the device includes: storing a third flag in the slice header of the target slice, where the target slice is a target lower-level region; or storing a third flag in a picture header of the target picture, wherein the target picture is a target lower-level region; 61. The non-transitory computer-readable medium of clause 60, further causing 62.Encoding mode is Bidirectional Optical Flow (BDOF) mode, Optical Flow Prediction Refinement (PROF) mode, or Decoder-side Motion Vector Refinement (DMVR) mode, 62. The non-transitory computer-readable medium of any one of clauses 58 to 61, wherein at least one of the following is true: 63. A set of instructions executable by at least one processor of the device includes: 63. The non-transitory computer-readable medium of any one of clauses 58-62, further comprising storing the first flag within a sequence parameter set (SPS) of the video bitstream. 64. A set of instructions executable by at least one processor of the device includes: A non-transitory computer-readable medium as described in any one of clauses 58 to 63, further comprising storing a second flag within the SPS of the video bitstream in response to enabling or disabling the encoding mode for the video bitstream. 65. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by at least one processor of a device to cause the device to perform a method, the method comprising: receiving a video sequence, a first flag, a second flag, and a third flag; enabling or disabling a first coding mode for the video bitstream based on the first flag; enabling or disabling a second coding mode for the video bitstream based on the second flag; enabling or disabling control of at least one of the first encoding mode or the second encoding mode at a level lower than the sequence level based on the third flag; 1. A non-transitory computer-readable medium comprising: 66. The non-transitory computer-readable medium of clause 65, wherein the level below the sequence level includes the slice level or the picture level. 67. A set of instructions executable by at least one processor of the device includes: receiving a fourth flag; enabling or disabling the first coding mode for the target lower-level region based on a fourth flag in response to enabling control of the first coding mode for the video bitstream based on the first flag and enabling control of at least one of the first coding mode or the second coding mode at a level lower than the sequence level based on the third flag; 67. The non-transitory computer-readable medium of clause 65 or 66, further causing the computer to: 68. A set of instructions executable by at least one processor of the device includes: receiving a fourth flag; enabling or disabling both the first encoding mode and the second encoding mode for the target lower-level region based on a fourth flag in response to enabling or disabling control of at least one of the first encoding mode or the second encoding mode at a level lower than the sequence level; 67. The non-transitory computer-readable medium of clause 65 or 66, further causing the computer to: 69. A set of instructions executable by at least one processor of the device includes: storing a fourth flag in the slice header of the target slice, where the target slice is a target lower-level region; or storing a fourth flag in a picture header of the target picture, wherein the target picture is a target lower-level region; 69. The non-transitory computer-readable medium of clause 67 or 68, further causing 70. A set of instructions executable by at least one processor of the device includes: receiving a fifth flag; enabling or disabling a third coding mode for the video bitstream based on the second flag; enabling or disabling control of the third coding mode at a level below the sequence level based on a fifth flag; 69. The non-transitory computer-readable medium of claim 69, further comprising: 71. A set of instructions executable by at least one processor of the device includes: receiving a sixth flag; enabling or disabling the third coding mode for the target lower-level region based on a sixth flag in response to enabling or disabling control of the third coding mode at a level lower than the sequence level; 71. The non-transitory computer-readable medium of clause 69 or 70, further causing 72. A non-transitory computer-readable medium according to any one of clauses 65 to 71, wherein the first, second, and third encoding modes are a bidirectional optical flow (BDOF) mode, a prediction refinement by optical flow (PROF) mode, and a decoder-side motion vector refinement (DMVR) mode, respectively. 73. The first and second encoding modes are: Bidirectional Optical Flow (BDOF) mode, Optical Flow Prediction Refinement (PROF) mode, and Decoder-side Motion Vector Refinement (DMVR) mode, 73. The non-transitory computer-readable medium of any one of clauses 65 to 72, wherein the encoding modes are two different encoding modes selected from: 74. A non-transitory computer-readable medium described in any one of clauses 65 to 73, wherein the first encoding mode and the second encoding mode are a bidirectional optical flow (BDOF) mode and a prediction refinement by optical flow (PROF) mode, respectively. 75. A set of instructions executable by at least one processor of the device includes: 75. The non-transitory computer-readable medium of any one of clauses 65 to 74, further comprising storing the first and second flags in a sequence parameter set (SPS) of a video bitstream. 76. A set of instructions executable by at least one processor of the device includes: A non-transitory computer-readable medium as described in any one of clauses 65 to 75, further comprising storing a third flag within an SPS of the video sequence in response to enabling at least one of the first encoding mode or the second encoding mode for the video bitstream. 77. An apparatus comprising: a memory configured to store a set of instructions; one or more processors communicatively coupled to the memory, the one or more processors executing the set of instructions to cause the apparatus to: receiving a bitstream of video data; enabling or disabling a coding mode for the video sequence based on a first flag in the bitstream; determining whether coding mode control is enabled or disabled at a level below the sequence level based on a second flag in the bitstream; An apparatus configured to cause 78. The apparatus of clause 77, wherein the level below the sequence level includes the slice level or the picture level. 79. One or more processors execute a set of instructions to cause a device to: enabling or disabling the coding mode for the target lower-level region based on a third flag in the bitstream in response to the coding mode control being enabled at a level lower than the sequence level; 79. The apparatus of clause 77 or 78, further configured to: 80. One or more processors execute a set of instructions to cause a device to: Detecting a third flag in the slice header of the target slice, where the target slice is a target lower-level region; or detecting a third flag in a picture header of the target picture, wherein the target picture is a target lower-level region; 80. The apparatus of clause 79, further configured to: 81.Encoding mode is Bidirectional Optical Flow (BDOF) mode, Optical Flow Prediction Refinement (PROF) mode, or Decoder-side Motion Vector Refinement (DMVR) mode, 81. The device according to any one of clauses 77 to 80, wherein the device is at least one of: 82. One or more processors execute a set of instructions to cause a device to: Detecting a first flag in a sequence parameter set (SPS) of a video sequence 82. The apparatus of any one of clauses 77 to 81, further configured to: 83. One or more processors execute a set of instructions to cause a device to: Detecting a second flag in the SPS of the video sequence in response to the coding mode being enabled for the video sequence. 83. The apparatus of any one of clauses 77 to 82, further configured to: 84. An apparatus comprising: a memory configured to store a set of instructions; one or more processors communicatively coupled to the memory, the one or more processors executing the set of instructions to cause the apparatus to: receiving a bitstream of video data; enabling or disabling a first coding mode for the video sequence based on a first flag in the bitstream; enabling or disabling a second coding mode for the video sequence based on a second flag in the bitstream; determining whether control of at least one of the first coding mode or the second coding mode is enabled at a level below the sequence level based on a third flag in the bitstream; An apparatus configured to cause 85. The apparatus of clause 84, wherein the level below the sequence level includes the slice level or the picture level. 86. One or more processors execute a set of instructions to cause a device to: and enabling or disabling both the first coding mode and the second coding mode for the target lower-level region based on a fourth flag in the bitstream in response to control of at least one of the first coding mode or the second coding mode being enabled at a level lower than the sequence level. 86. The apparatus of clause 84 or 85, further configured to: 87. One or more processors execute a set of instructions to cause a device to: Detecting a fourth flag in the slice header of the target slice, where the target slice is a target lower-level region; or detecting a fourth flag in a picture header of the target picture, wherein the target picture is a target lower-level region; 87. The apparatus of clause 86, further configured to: 88. One or more processors execute a set of instructions to cause a device to: enabling or disabling a third coding mode for the video sequence based on a second flag in the bitstream; determining whether control of the third coding mode is enabled at a level below the sequence level based on a fifth flag in the bitstream; 88. The apparatus of clause 86 or 87, further configured to: 89. One or more processors execute a set of instructions to cause a device to: and enabling or disabling the third coding mode for the target lower-level region based on a sixth flag in the bitstream in response to the third coding mode being enabled at a level lower than the sequence level. 89. The apparatus of clause 88, further configured to: 90. The apparatus of clause 88 or 89, wherein the first, second, and third encoding modes are a bidirectional optical flow (BDOF) mode, a prediction refinement by optical flow (PROF) mode, and a decoder-side motion vector refinement (DMVR) mode, respectively. 91. One or more processors execute a set of instructions to cause a device to: enabling or disabling the first coding mode for the target lower-level region based on a fourth flag in the bitstream in response to the first flag indicating that the first coding mode is enabled for the video sequence and the third flag indicating that control of at least one of the first coding mode or the second coding mode is enabled at a level lower than the sequence level; 91. The apparatus of any one of clauses 84 to 90, further configured to: 92. The first and second encoding modes are: Bidirectional Optical Flow (BDOF) mode, Optical Flow Prediction Refinement (PROF) mode, and Decoder-side Motion Vector Refinement (DMVR) mode, 92. The apparatus of any one of clauses 84 to 91, wherein the encoding modes are two different modes selected from: 93. The device of any one of clauses 84 to 92, wherein the first encoding mode and the second encoding mode are a bidirectional optical flow (BDOF) mode and a prediction refinement by optical flow (PROF) mode, respectively. 94. One or more processors execute a set of instructions to cause the device to: Detecting first and second flags in a sequence parameter set (SPS) of a video sequence 94. The apparatus of any one of clauses 84 to 93, further configured to: 95. One or more processors execute a set of instructions to cause the device to: Detecting a third flag in the SPS of the video sequence in response to at least one of the first encoding mode or the second encoding mode being enabled for the video sequence. 95. The apparatus of any one of clauses 84 to 94, further configured to: 96. An apparatus comprising: a memory configured to store a set of instructions; one or more processors communicatively coupled to the memory, the one or more processors executing the set of instructions to cause the apparatus to: receiving a video sequence, a first flag, and a second flag; enabling or disabling a coding mode for the video bitstream based on a first flag; enabling or disabling control of the coding mode at a level below the sequence level based on a second flag; An apparatus configured to cause 97. The apparatus of clause 96, wherein the level below the sequence level includes the slice level or the picture level. 98. One or more processors execute a set of instructions to cause the device to: receiving a third flag; enabling or disabling the coding mode for the target lower-level region based on a third flag in response to controlling enabling or disabling of the coding mode at a level lower than the sequence level; 98. The apparatus of clause 96 or 97, further configured to: 99. One or more processors execute a set of instructions to cause the device to: storing a third flag in the slice header of the target slice, where the target slice is a target lower-level region; or storing a third flag in a picture header of the target picture, wherein the target picture is a target lower-level region; 99. The apparatus of clause 98, further configured to: 100.Encoding mode is Bidirectional Optical Flow (BDOF) mode, Optical Flow Prediction Refinement (PROF) mode, or Decoder-side Motion Vector Refinement (DMVR) mode, 99. The device of any one of clauses 96 to 99, wherein the device is at least one of: 101. One or more processors execute a set of instructions to cause a device to: Storing the first flag in a sequence parameter set (SPS) of the video bitstream 101. The apparatus of any one of clauses 96 to 100, further configured to: 102. One or more processors execute a set of instructions to cause a device to: storing a second flag in the SPS of the video bitstream in response to enabling or disabling the coding mode for the video bitstream; 102. The apparatus of any one of clauses 96 to 101, further configured to: 103. An apparatus comprising: a memory configured to store a set of instructions; one or more processors communicatively coupled to the memory, the one or more processors executing the set of instructions to cause the apparatus to: receiving a video sequence, a first flag, a second flag, and a third flag; enabling or disabling a first coding mode for the video bitstream based on the first flag; enabling or disabling a second coding mode for the video bitstream based on the second flag; enabling or disabling control of at least one of the first encoding mode or the second encoding mode at a level lower than the sequence level based on the third flag; An apparatus configured to cause 104. The apparatus of clause 102, wherein the level below the sequence level includes the slice level or the picture level. 105. One or more processors execute a set of instructions to cause the device to: receiving a fourth flag; enabling or disabling the first coding mode for the target lower-level region based on a fourth flag in response to enabling control of the first coding mode for the video bitstream based on the first flag and enabling control of at least one of the first coding mode or the second coding mode at a level lower than the sequence level based on the third flag; 105. The apparatus of clause 103 or 104, further configured to: 106. One or more processors execute a set of instructions to cause a device to: receiving a fourth flag; enabling or disabling both the first encoding mode and the second encoding mode for the target lower-level region based on a fourth flag in response to enabling or disabling control of at least one of the first encoding mode or the second encoding mode at a level lower than the sequence level; 105. The apparatus of clause 103 or 104, further configured to: 107. One or more processors execute a set of instructions to cause a device to: storing a fourth flag in the slice header of the target slice, where the target slice is a target lower-level region; or storing a fourth flag in a picture header of the target picture, wherein the target picture is a target lower-level region; 107. The apparatus of clause 105 or 106, further configured to: 108. One or more processors execute a set of instructions to cause a device to: receiving a fifth flag; enabling or disabling a third coding mode for the video bitstream based on the second flag; enabling or disabling control of the third coding mode at a level below the sequence level based on a fifth flag; 108. The apparatus of clause 107, further configured to: 109. One or more processors execute a set of instructions to cause a device to: receiving a sixth flag; enabling or disabling the third coding mode for the target lower-level region based on a sixth flag in response to enabling or disabling control of the third coding mode at a level lower than the sequence level; 109. The apparatus of clause 107 or 108, further configured to: 110. The apparatus of any one of clauses 103 to 109, wherein the first, second, and third encoding modes are a bidirectional optical flow (BDOF) mode, a prediction refinement by optical flow (PROF) mode, and a decoder-side motion vector refinement (DMVR) mode, respectively. 111. The first and second encoding modes are: Bidirectional Optical Flow (BDOF) mode, Optical Flow Prediction Refinement (PROF) mode, and Decoder-side Motion Vector Refinement (DMVR) mode, 111. The apparatus according to any one of clauses 103 to 110, wherein the two different encoding modes are selected from: 112. The device of any one of clauses 103 to 110, wherein the first encoding mode and the second encoding mode are a bidirectional optical flow (BDOF) mode and a prediction refinement by optical flow (PROF) mode, respectively. 113. One or more processors execute a set of instructions to cause a device to: Storing the first and second flags in a sequence parameter set (SPS) of the video bitstream 113. The apparatus of any one of clauses 103 to 112, further configured to: 114. One or more processors execute a set of instructions to cause a device to: storing a third flag in the SPS of the video sequence in response to enabling at least one of the first encoding mode or the second encoding mode for the video bitstream; 114. The apparatus of any one of clauses 103 to 113, further configured to:
[0137]
[0166] It should be noted that relational terms herein, such as "first" and "second," are used merely to distinguish one entity or operation from another, and do not require or imply any actual relationship or order among those entities or operations. Furthermore, the words "comprising," "having," "containing," and "including," and other similar forms, are intended to be equivalent in meaning and to be open-ended in that the element or elements following any of these words are not meant to be an exclusive list of such elements or elements, or to be limited to only the listed element or elements.
[0138]
[0167] As used herein, unless specifically stated otherwise, the term "or" includes all possible combinations unless impracticable. For example, if it is stated that a component can include A or B, then, unless specifically stated otherwise or impracticable, the component can include A, or B, or A and B. As a second example, if it is stated that a component can include A, B, or C, then, unless specifically stated otherwise or impracticable, the component can include A, B, or C, or A and B, A and C, or B and C, or A and B and C.
[0139]
[0168] It is understood that the above-described embodiments can be implemented by hardware, or software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-described computer-readable medium. The software, when executed by a processor, can perform the methods of the present disclosure. The computational units and other functional units described in the present disclosure can be implemented by hardware, or software, or a combination of hardware and software. Those skilled in the art will also understand that multiple of the above-described modules / units can be combined into one module / unit, and that each of the above-described modules / units can be further divided into multiple sub-modules / sub-units.
[0140]
[0169] In the foregoing specification, embodiments have been described with reference to numerous specific details that may vary from implementation to implementation. Certain adaptations and modifications of the above-described embodiments may be made. Other embodiments may be apparent to those skilled in the art from consideration of the specification and practice of the invention disclosed herein. It is intended that the specification and examples be considered as examples only, with the true scope and spirit of the invention being indicated by the appended claims. It is also intended that the sequences of steps depicted in the figures are for illustrative purposes only and are not intended to be limited to any particular sequence of steps. Thus, one skilled in the art will recognize that these steps may be performed in different orders while implementing the same method.
[0141]
[0170] In the drawings and specification, illustrative embodiments have been disclosed. However, many variations and modifications to these embodiments may be made. Accordingly, although specific terms are employed, they are used in a generic and descriptive sense only and not for purposes of limitation.
Claims
1. 1. A computer-implemented decoding method comprising: receiving a bitstream associated with a video sequence; decoding a first plurality of flags in a sequence parameter set (SPS) of the bitstream; determining whether a second plurality of flags is present in the bitstream based on the values of the first plurality of flags, respectively; the second plurality of flags respectively indicating whether a plurality of coding modes are enabled or disabled at a first syntax level below the sequence level; the plurality of coding modes include a prediction refinement by optical flow (PROF) mode, a bidirectional optical flow (BDOF) mode, and a decoder-side motion vector refinement (DMVR) mode; one or more flags of the second plurality of flags are present in the bitstream, and the bitstream is decoded based on the values of the one or more flags of the second plurality of flags. A computer-implemented decoding method.
2. the second plurality of flags: a first flag associated with the PROF mode; a second flag associated with the BDOF mode; and a third flag associated with the DMVR mode; The method comprises: when the first flag is present in the bitstream, enabling or disabling the PROF mode at the first syntax level based on the value of the first flag; when the second flag is present in the bitstream, enabling or disabling the BDOF mode at the first syntax level based on the value of the second flag; When the third flag is present in the bitstream, enabling or disabling the DMVR mode at the first syntax level based on the value of the third flag; 10. The computer-implemented decoding method of claim 1, further comprising:
3. decoding a third plurality of flags in the SPS; and enabling or disabling the plurality of encoding modes at the first syntax level based on values of the third plurality of flags, respectively; 10. The computer-implemented decoding method of claim 1, further comprising:
4. 2. The computer-implemented decoding method of claim 1, wherein the first syntax level is a slice level or a picture level.
5. the first plurality of flags includes a first flag, the second plurality of flags includes a second flag associated with the first flag, and the method further comprises: determining, in response to the first flag having a first value, that the second flag is present in a slice header or a picture header of the bitstream; 10. The computer-implemented decoding method of claim 1, further comprising:
6. 6. The computer-implemented decoding method of claim 5, wherein the first value is one.
7. 1. A computer-implemented coding method comprising: encoding a first plurality of flags in a sequence parameter set (SPS) of a bitstream associated with a video sequence; determining whether to signal a second plurality of flags in the bitstream based on the first plurality of flags, respectively; the second plurality of flags respectively indicating whether a plurality of coding modes are enabled or disabled at a first syntax level below the sequence level; the plurality of coding modes include a prediction refinement by optical flow (PROF) mode, a bidirectional optical flow (BDOF) mode, and a decoder-side motion vector refinement (DMVR) mode; when it is determined to signal one or more flags of the second plurality of flags present in the bitstream, the bitstream is encoded based on values of the one or more flags of the second plurality of flags. Computer-implemented coding method.
8. the second plurality of flags: a first flag associated with the PROF mode; a second flag associated with the BDOF mode; and a third flag associated with the DMVR mode; The method comprises: enabling or disabling the PROF mode at the first syntax level when it is determined to signal the first flag; enabling or disabling the BDOF mode at the first syntax level when it is determined to signal the second flag; enabling or disabling the DMVR mode at the first syntax level when it is determined to signal the third flag; 8. The computer-implemented coding method of claim 7, further comprising:
9. 8. The computer-implemented encoding method of claim 7, further comprising: encoding a third plurality of flags in the SPS, the third plurality of flags indicating whether the plurality of encoding modes are enabled or disabled for the video sequence, respectively.
10. 8. The computer-implemented coding method of claim 7, wherein the first syntax level is a slice level or a picture level.
11. the first plurality of flags includes a first flag, the second plurality of flags includes a second flag associated with the first flag, and the method further comprises: encoding the second flag in a slice header or a picture header of the bitstream in response to the first flag having a first value.
8. The computer-implemented coding method of claim 7, further comprising:
12. 12. The computer-implemented coding method of claim 11, wherein the first value is one.
13. 1. A method for storing a bitstream associated with a video, comprising: encoding a first plurality of flags in a sequence parameter set (SPS) of a bitstream associated with a video sequence; determining whether to signal a second plurality of flags in the bitstream based on the first plurality of flags, respectively; the second plurality of flags respectively indicating whether a plurality of coding modes are enabled or disabled at a first syntax level below the sequence level; the plurality of coding modes include a prediction refinement by optical flow (PROF) mode, a bidirectional optical flow (BDOF) mode, and a decoder-side motion vector refinement (DMVR) mode; determining, when it is determined to signal one or more flags of the second plurality of flags present in the bitstream, that the bitstream is encoded based on values of the one or more flags of the second plurality of flags; generating a bitstream including at least one of the first plurality of flags or the second plurality of flags; storing the bitstream on a non-transitory computer readable medium; A method comprising:
14. the second plurality of flags: a first flag associated with the PROF mode; a second flag associated with the BDOF mode; and a third flag associated with the DMVR mode; The method comprises: enabling or disabling the PROF mode at the first syntax level when it is determined to signal the first flag; enabling or disabling the BDOF mode at the first syntax level when it is determined to signal the second flag; enabling or disabling the DMVR mode at the first syntax level when it is determined to signal the third flag; 14. The method of claim 13, further comprising:
15. 14. The method of claim 13, further comprising encoding a third plurality of flags in the SPS, the third plurality of flags indicating whether the plurality of coding modes are enabled or disabled for the video sequence, respectively.
16. The method of claim 13 , wherein the first syntax level is a slice level or a picture level.
17. the first plurality of flags includes a first flag, the second plurality of flags includes a second flag associated with the first flag, and the method further comprises: encoding the second flag in a slice header or a picture header of the bitstream in response to the first flag having a first value.
14. The method of claim 13, further comprising:
18. 18. The method of claim 17, wherein the first value is one.
Citation Information
Patent Citations
Method and apparatus for prediction refinement using optical flow for affine coded blocks
JP2022527701A
PROF method, computing device, computer-readable storage medium, and program
JP2022536208A