Computer-implemented encoding and decoding method and non-transitory computer readable storage medium
By introducing sequence level and control flags below sequence level into the video bitstream, dynamically control the encoding mode, the technical problems of coding efficiency and quality improvement in VVC/H.266 are solved, and more efficient video encoding is achieved.
Patent Information
- Application Number
- CN202510428513.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2019-09-12
- Filing Date
- 2020-08-20
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2040-08-20
AI Technical Summary
In the efficient video encoding standards such as VVC/H.266, existing video encoding technologies have not fully utilized control flags below the sequence level to optimize the activation or disabling of encoding mode, resulting in the potential improvement of encoding efficiency and quality not being fully utilized.
Enable or disable encoding modes such as bidirectional optical flow (BDOF), optical flow prediction correction (PROF), and decoder side motion vector correction (DMVR) dynamically enable or disable encoding efficiency and quality by introducing sequence-level and lower-sequence control flags into the video bitstream.
It realizes higher encoding efficiency and quality under the VVC/H.266 standard, and through flexible encoding mode control, the calculation complexity is reduced and the performance of video encoding is improved.
Smart Images

Figure CN120281923A_ABST
Abstract
Description
Cross - Reference to Related Applications
[0001] This disclosure claims priority to U.S. Provisional Application No. 62 / 899,169, filed on September 12, 2019, the entire content of which is incorporated herein by reference. Background Art
[0002] Video is a set of static images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and then decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats using standardized video coding techniques, and the most common ones are based on prediction, transformation, quantization, entropy coding, and loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, specify specific video coding formats and are formulated by standardization organizations. As more and more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards is also getting higher and higher. Summary of the Invention
[0003] Embodiments of the present invention provide a method and apparatus for controlling an encoding mode for video data. In one example embodiment, a method includes: receiving a bitstream of video data; enabling or disabling an encoding mode for a video sequence based on a first flag in the bitstream; and determining, based on a second flag in the bitstream, to enable or disable control of the encoding mode at a level lower than the sequence level.
[0004] In another example embodiment, a method includes: receiving a bitstream of video data; enabling or disabling a first encoding mode for a video sequence based on a first flag in the bitstream; enabling or disabling a second encoding mode for the video sequence based on a second flag in the bitstream; and determining, based on a third flag in the bitstream, whether to enable control of at least one of the first encoding mode or the second encoding mode at a level lower than the sequence level.
[0005] In another example embodiment, a method includes: receiving a video sequence, a first flag, and a second flag; enabling or disabling an encoding mode for a video bitstream based on the first flag; and enabling or disabling control of the encoding mode at a level lower than the sequence level based on the second flag.
[0006] In another example embodiment, a method includes: receiving a video sequence, a first flag, a second flag, and a third flag; enabling or disabling a first coding mode for a video bitstream based on the first flag; enabling or disabling a second coding mode for the video bitstream based on the second flag; and enabling or disabling control of at least one of the first coding mode or the second coding mode at a level lower than the sequence level based on the third flag.
[0007] In another example embodiment, a non-transitory computer-readable medium stores a set of instructions executable by at least one processor of a device to cause the device to perform a method that includes: receiving a video data bitstream; enabling or disabling a coding mode for a video sequence based on a first flag in the bitstream; and determining whether to enable or disable control of the coding mode at a level lower than the sequence level based on a second flag in the bitstream.
[0008] In another example embodiment, a non-transitory computer-readable medium stores a set of instructions executable by at least one processor of a device to cause the device to perform a method that includes: receiving a video data bitstream; enabling or disabling a first coding mode for a video sequence based on a first flag in the bitstream; enabling or disabling a second coding mode for the video sequence based on a second flag in the bitstream; and determining whether to enable control of at least one of the first coding mode or the second coding mode at a level lower than the sequence level based on a third flag in the bitstream.
[0009] In another embodiment, a device includes a memory configured to store a set of instructions and one or more processors communicatively coupled to the memory, the one or more processors configured to execute the set of instructions to cause the device to: receive a bitstream of video data; enable or disable a coding mode for a video sequence based on a first flag in the bitstream; and determine whether to enable or disable control of the coding mode at a level lower than the sequence level based on a second flag in the bitstream.
[0010] In another embodiment, a device includes a memory configured to store a set of instructions and one or more processors communicatively coupled to the memory, the one or more processors configured to execute the set of instructions to cause the device to: receive a bitstream of video data; enable or disable a first coding mode for a video sequence based on a first flag in the bitstream; enable or disable a second coding mode for the video sequence based on a second flag in the bitstream; and determine whether to enable control of at least one of the first coding mode or the second coding mode at a level lower than the sequence level based on a third flag in the bitstream. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Embodiments and aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings. The various features shown in the drawings are not drawn to scale.
[0012] Figure 1 It is a schematic diagram showing the structure of an example video sequence according to some embodiments of the present disclosure.
[0013] Figure 2A It is a schematic diagram showing an example encoding process of a hybrid video coding system consistent with embodiments of the present disclosure.
[0014] Figure 2B It is a schematic diagram showing another example encoding process of a hybrid video coding system consistent with embodiments of the present disclosure.
[0015] Figure 3A It is a schematic diagram showing an example decoding process of a hybrid video coding system consistent with embodiments of the present disclosure.
[0016] Figure 3B It is a schematic diagram showing another example decoding process of a hybrid video coding system consistent with embodiments of the present disclosure.
[0017] Figure 4 It is a block diagram showing an example apparatus for encoding or decoding video according to some embodiments of the present disclosure.
[0018] Figure 5 It is a schematic diagram showing an example process of decoder-side motion vector refinement (DMVR) according to some embodiments of the present disclosure.
[0019] Figure 6 It is a schematic diagram showing an example DMVR search process according to some embodiments of the present disclosure.
[0020] Figure 7 It is a schematic diagram illustrating an example pattern for DMVR integer luminance sample search according to some embodiments of the present disclosure.
[0021] Figure 8 It is a schematic diagram showing another example pattern of the stage of integer sample offset search in DMVR integer luminance sample search according to some embodiments of the present disclosure.
[0022] Figure 9 It is a schematic diagram illustrating an example pattern for estimating the DMVR parameter error surface according to some embodiments of the present disclosure.
[0023] Figure 10 It is a schematic diagram of an example of an extended coding unit (CU) region used in bidirectional optical flow (BDOF) according to some embodiments of the present disclosure.
[0024] Figure 11Schematic diagram of examples of sub-block based affine motion and sample based affine motion according to some embodiments of the present disclosure.
[0025] Figure 12 Table 1 is shown, which shows an example syntax structure of a sequence parameter set (SPS) for implementing control flags for DMVR and BDOF according to some embodiments of the present disclosure.
[0026] Figure 13 Table 2 is shown, which shows an example syntax structure of a slice header for implementing control flags for DMVR and BDOF according to some embodiments of the present disclosure.
[0027] Figure 14A Table 3A is shown, which shows an example syntax structure of an SPS for implementing slice level control flags for DMVR, BDOF, and optical flow prediction correction (PROF) according to some embodiments of the present disclosure.
[0028] Figure 14B Table 3B is shown, which shows an example syntax structure of an SPS for implementing picture level control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.
[0029] Figure 15A Table 4A is shown, which shows an example syntax structure of a slice header for implementing control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.
[0030] Figure 15B Table 4B is shown, which shows an example syntax structure of a picture header for implementing control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.
[0031] Figure 16 Table 5 is shown, which shows an example syntax structure of an SPS for implementing individual sequence level control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.
[0032] Figure 17 Table 6 is shown, which shows an example syntax structure of a slice head for implementing joint control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.
[0033] Figure 18 Table 7 is shown, which shows an example syntax structure of a slice head for implementing individual control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.
[0034] Figure 19 Table 8 is shown, and Table 8 shows an example syntax structure of a slice head that implements hybrid control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.
[0035] Figure 20 Table 9 is shown, and Table 9 shows an example syntax structure of an SPS that implements hybrid sequence-level control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.
[0036] Figure 21 Table 10 is shown, and Table 10 shows another example syntax structure of a slice head that implements hybrid control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.
[0037] Figure 22 Table 11 is shown, and Table 11 shows another example syntax structure of a slice head that implements individual control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.
[0038] Figure 23 A flowchart of an example process for controlling a video decoding mode according to some embodiments of the present disclosure is shown.
[0039] Figure 24 A flowchart of another example process for controlling a video decoding mode according to some embodiments of the present disclosure is shown.
[0040] Figure 25 A flowchart of an example process for controlling a video encoding mode according to some embodiments of the present disclosure is shown.
[0041] Figure 26 A flowchart of another example process for controlling a video encoding mode according to some embodiments of the present disclosure is shown. Detailed Description
[0042] Reference may now be made in detail to example embodiments, which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings unless otherwise noted, in which like numbers in different drawings represent the same or similar elements. The embodiments set forth in the following description of the example embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of apparatus and methods consistent with aspects of the present invention as recited in the appended claims. Specific aspects of the present disclosure are described in more detail below. If there is a conflict with the terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.
[0043] The Joint Video Exploration Team (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, the goal of VVC is to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.
[0044] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET has been using the Joint Exploration Model (JEM) reference software to explore technologies beyond HEVC. As coding techniques are incorporated into JEM, JEM has achieved higher coding performance than HEVC.
[0045] The VVC standard has been recently developed and continues to incorporate more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.
[0046] Video is a set of static images (or "frames") arranged in chronological order to store visual information. Video capture devices (e.g., cameras) can be used to capture and store these images in chronological order, and video playback devices (e.g., TVs, computers, smartphones, tablets, video players, or any end-user terminal with a display function) can be used to display such images in chronological order. Additionally, in some applications, a video capture device can transmit the captured video in real time to a video playback device (e.g., a computer with a monitor), such as for surveillance, conferencing, or live streaming.
[0047] To reduce the storage space and transmission bandwidth required for such applications, the video can be compressed before storage and transmission and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., the processor of a general-purpose computer) or dedicated hardware. The module for compression is generally called an "encoder", and the module for decompression is generally called a "decoder". The encoder and decoder can be collectively referred to as a "codec". The encoder and decoder can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, the hardware implementation of the encoder and decoder can include circuitry such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. The software implementation of the encoder and decoder can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be implemented by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, the H.26x series, etc. In some applications, the codec can decompress the video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec can be called a "transcoder".
[0048] The video encoding process can identify and retain useful information that can be used to reconstruct the image while ignoring unimportant information for reconstruction. If the ignored, unimportant information cannot be fully reconstructed, such an encoding process can be called "lossy". Otherwise, it can be called "lossless". Most encoding processes are lossy, which is a trade-off made to reduce the required storage space and transmission bandwidth.
[0049] The useful information of the encoded image (referred to as the "current image") includes the changes relative to a reference image (e.g., a previously encoded and reconstructed image). Such changes can include changes in the position of pixels, changes in brightness, or changes in color, with the most concerned being the changes in position. The change in the position of a group of pixels representing an object can reflect the movement of the object between the reference image and the current image.
[0050] An image encoded without referring to another image (i.e., it is its own reference image) is called an "I-image". An image encoded using a previous image as the reference image is called a "P-image". An image encoded using a previous image and a future image as the reference image (i.e., the reference is "bidirectional") is called a "B-image".
[0051] Figure 1Illustrates the structure of an example video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 can be a live video or a video that has been captured and archived. The video 100 can be a real video, a computer-generated video (e.g., a computer game video), or a combination thereof (e.g., a real video with augmented reality effects). The video sequence 100 can be input from a video capture device (e.g., a camera), a video archive containing previously captured videos (e.g., a video file stored in a storage device), or a provision interface that receives videos from a video content provider (e.g., a video broadcast transceiver).
[0052] As Figure 1 shown, the video sequence 100 can include a series of images arranged along a time axis over time, including images 102, 104, 106, and 108. Images 102-106 are consecutive, and there are more images between images 106 and 108. In Figure 1 this example, image 102 is an I image, and its reference image is image 102 itself. Image 104 is a P image, and its reference image is image 102, as indicated by the arrow. Image 106 is a B image, and its reference images are images 104 and 108, as indicated by the arrows. In some embodiments, the reference image of an image (e.g., image 104) may not be immediately before or after the image. For example, the reference image of image 104 can be an image before image 102. It should be noted that the reference images of images 102-106 are only examples, and the present disclosure does not limit the embodiments of the reference images to the examples shown in Figure 1 this example.
[0053] Generally, due to the computational complexity of the encoding and decoding tasks, video codecs do not encode or decode an entire image at once. Instead, they can divide an image into basic segments and encode or decode the image segment by segment. Such basic segments are referred to as basic processing units ("BPUs") in the present disclosure. For example, Figure 1The structure 110 therein shows an example structure of an image (e.g., any one of images 102-108) of the video sequence 100. In structure 110, the image is divided into 4×4 basic processing units, the boundaries of which are shown as dashed lines. In some embodiments, the basic processing unit may be referred to as a "macroblock" in some video coding standards (e.g., the MPEG series, H.261, H.263, or H.264 / AVC), or as a "coding tree unit" ("CTU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing units in the image can have different sizes, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or pixels of any shape and size. The size and shape of the basic processing unit can be selected for the image based on a balance between coding efficiency and the level of detail to be retained in the basic processing unit.
[0054] The basic processing unit can be a logical unit that can include a set of different types of video data stored in a computer memory (e.g., in a video frame buffer). For example, the basic processing unit of a color image can include a luminance component (Y) representing achromatic luminance information, one or more chrominance components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luminance and chrominance components can have basic processing units of the same size. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luminance and chrominance components can be referred to as "coding tree blocks" ("CTB"). Any operation performed on the basic processing unit can be repeated for each of its luminance and chrominance components.
[0055] Video coding has multiple operation stages, examples of which are in Figures 2A - 2B and Figures 3A - 3BShown in. For each stage, the size of the basic processing unit may still be too large to process, so it can be further divided into segments, which are referred to as "basic processing subunits" in this disclosure. In some embodiments, the basic processing subunit may be referred to as a "block" in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC), or as a "coding unit" ("CU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunit may have the same size as the basic processing unit or a smaller size than the basic processing subunit. Similar to the basic processing unit, the basic processing subunit is also a logical unit, which may include a set of different types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in a computer memory (e.g., in a video frame buffer). Any operation performed on the basic processing subunit can be repeated for each of its luminance and chrominance components. It should be noted that this division can be performed to a further level according to the processing needs. It should also be noted that different stages can use different schemes to divide the basic processing unit.
[0056] For example, in the mode decision stage (an example of which is shown in Figure 2B ), the encoder can decide what prediction mode (e.g., intra-picture prediction or inter-picture prediction) to use for the basic processing unit, and the basic processing unit may be too large to make this decision. The encoder can split the basic processing unit into multiple basic processing subunits (e.g., CUs in H.265 / HEVC or H.266 / VVC) and determine the prediction type for each individual basic processing subunit.
[0057] For another example, in the prediction stage (an example of which is shown in Figures 2A - 2B ), the encoder can perform prediction operations at the basic processing subunit (e.g., CU) level. However, in some cases, the basic processing subunit may still be too large to process. The encoder can further split the basic processing subunit into smaller segments (e.g., referred to as "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), and the prediction operations can be performed at this level.
[0058] For another example, in the transform stage (an example of which is shown in Figures 2A - 2BAs shown in []. The encoder can perform a transformation operation on a residual basic processing unit (e.g., a CU). However, in some cases, the basic processing unit may still be too large to process. The encoder can further divide the basic processing unit into smaller segments (e.g., called "transformation blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), and the transformation operation can be performed at this level. It should be noted that the partitioning scheme of the same basic processing unit can be different in the prediction stage and the transformation stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transformation blocks of the same CU can have different sizes and numbers.
[0059] In Figure 1 the structure 110 of [], the basic processing unit 112 is further divided into 3×3 basic processing subunits, and their boundaries are represented by dotted lines. Different basic processing units of the same image can be divided into basic processing subunits with different schemes.
[0060] In some embodiments, to provide parallel processing and fault tolerance capabilities for video encoding and decoding, an image can be divided into multiple regions for processing, such that for one region of the image, the encoding or decoding process does not have to depend on information from any other region of the image. In other words, each region of the image can be processed independently. By doing so, the codec can process different regions of the image in parallel, thereby improving the encoding efficiency. Additionally, when the data of one region is damaged or lost during network transmission during processing, the codec can correctly encode or decode other regions of the same image without relying on the damaged or lost data, thereby providing fault tolerance. In some video coding standards, an image can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles". It should also be noted that different images of the video sequence 100 can have different partitioning schemes for dividing the image into multiple regions.
[0061] For example, in Figure 1 [], the structure 110 is divided into three regions 114, 116, and 118, and their boundaries are shown as solid lines inside the structure 110. Region 114 includes four basic processing units. Each of regions 116 and 118 includes six basic processing units. It should be noted that Figure 1 the basic processing units, basic processing subunits, and regions of the structure 110 in [] are only examples, and the present disclosure does not limit its embodiments.
[0062] Figure 2A illustrates a schematic diagram of an example encoding process 200A consistent with an embodiment of the present disclosure. For example, the encoding process 200A can be performed by an encoder. AsFigure 2A As shown, the encoder may encode the video sequence 202 into a video bitstream 228 according to process 200A. Similar to Figure 1 the video sequence 100 in , the video sequence 202 may include a set of images arranged in chronological order (referred to as "original images"). Similar to Figure 1 the structure 110 in , each original image of the video sequence 202 may be divided by the encoder into basic processing units, basic processing subunits, or regions for processing. In some embodiments, the encoder may execute process 200A at the basic processing unit level for each original image of the video sequence 202. For example, the encoder may execute process 200A iteratively, where the encoder may encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder may execute process 200A in parallel for regions (e.g., regions 114 - 118) of each original image of the video sequence 202.
[0063] In Figure 2A the encoder may provide the basic processing unit of the original image of the video sequence 202 (referred to as "original BPU") to the prediction stage 204 to generate prediction data 206 and a prediction BPU 208. The encoder may subtract the prediction BPU 208 from the original BPU to generate a residual BPU 210. The encoder may provide the residual BPU 210 to the transform stage 212 and the quantization stage 214 to generate quantized transform coefficients 216. The encoder may provide the prediction data 206 and the quantized transform coefficients 216 to the binary coding stage 226 to generate the video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as the "forward path". During process 200A, after the quantization stage 214, the encoder may provide the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224, which is used in the prediction stage 204 for the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A may be referred to as the "reconstruction path". The reconstruction path may be used to ensure that both the encoder and the decoder use the same reference data for prediction.
[0064] The encoder may iteratively execute process 200A to encode each original BPU (in the forward path) of the original image and generate a prediction reference 224 (in the reconstruction path) for the next original BPU for encoding the original image. After encoding all the original BPUs of the original image, the encoder may continue to encode the next image in the video sequence 202.
[0065] Referring to process 200A, the encoder may receive video sequence 202 generated by a video capture device (e.g., a camera). The term "receive" as used herein may refer to any action of receiving, inputting, obtaining, retrieving, acquiring, reading, accessing, or in any way inputting data.
[0066] In the prediction stage 204, in the current iteration, the encoder may receive the original BPU and prediction reference 224 and perform a prediction operation to generate prediction data 206 and a predicted BPU 208. The prediction reference 224 may be generated from the reconstruction path of a previous iteration of process 200A. The purpose of the prediction stage 204 is to reduce information redundancy by extracting the prediction data 206, which can be used to reconstruct the original BPU into the predicted BPU 208 from the prediction data 206 and the prediction reference 224.
[0067] Ideally, the predicted BPU 208 may be the same as the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is typically slightly different from the original BPU. To record this difference, after generating the predicted BPU 208, the encoder may subtract it from the original BPU to generate a residual BPU 210. For example, the encoder may subtract the value of the corresponding pixel of the predicted BPU 208 (e.g., grayscale value or RGB value) from the pixel value of the original BPU 208. Each pixel of the residual BPU 210 may have a residual value, which is the result of subtracting the corresponding pixels of the original BPU and the predicted BPU 208. Compared with the original BPU, the prediction data 206 and the residual BPU 210 may have fewer bits, but they can be used to reconstruct the original BPU without significantly degrading the quality. Thus, the original BPU is compressed.
[0068] To further compress the residual BPU 210, in the transform stage 212, the encoder may reduce its spatial redundancy by decomposing the residual BPU 210 into a set of two-dimensional "basis patterns", each basis pattern associated with a "transformation coefficient". The basis patterns may have the same size (e.g., the size of the residual BPU 210). Each basis pattern may represent a frequency component of the variation of the residual BPU 210 (e.g., the frequency of brightness variation). No basis pattern can be reproduced from any combination (e.g., linear combination) of any other basis patterns. In other words, the decomposition may decompose the variation of the residual BPU 210 into the frequency domain. This decomposition is similar to the discrete Fourier transform of a function, where the basis patterns are similar to the basis functions of the discrete Fourier transform (e.g., trigonometric functions), and the transformation coefficients are similar to the coefficients associated with the basis functions.
[0069] Different transformation algorithms can use different basic patterns. Various transformation algorithms can be used in the transformation stage 212, such as discrete cosine transform, discrete sine transform, etc. The transformation in the transformation stage 212 is reversible. That is, the encoder can recover the residual BPU 210 through the inverse operation of the transformation (referred to as "inverse transformation"). For example, to recover the pixels of the residual BPU 210, the inverse transformation can be to multiply the values of the corresponding pixels of the basic pattern by their respective correlation coefficients and sum the products to produce a weighted sum. For video coding standards, both the encoder and the decoder can use the same transformation algorithm (and thus the same basic pattern). Therefore, the encoder can record only the transformation coefficients, and the decoder can reconstruct the residual BPU 210 from these coefficients without receiving the basic pattern from the encoder. Compared with the residual BPU 210, the transformation coefficients can have fewer bits, but they can be used to reconstruct the residual BPU 210 without significant quality degradation. Therefore, the residual BPU 210 is further compressed.
[0070] The encoder can further compress the transformation coefficients in the quantization stage 214. During the transformation process, different basic patterns can represent different change frequencies (e.g., luminance change frequencies). Since the human eye is generally better at recognizing low-frequency changes, the encoder can ignore the information of high-frequency changes without causing a significant decrease in the decoding quality. For example, in the quantization stage 214, the encoder can generate the quantized transformation coefficients 216 by dividing each transformation coefficient by an integer value (referred to as "quantization parameter") and rounding the quotient to its nearest integer. After such an operation, some transformation coefficients of the high-frequency basic pattern can be converted to zero, while the transformation coefficients of the low-frequency basic pattern can be converted to smaller integers. The encoder can ignore the zero-valued quantized transformation coefficients 216, further compressing the transformation coefficients through this operation. The quantization process is also reversible, where the quantized transformation coefficients 216 can be reconstructed as transformation coefficients in the inverse operation of quantization (referred to as "inverse quantization").
[0071] Because the encoder ignores the remainder of this division in the rounding operation, the quantization stage 214 may be lossy. Generally, the quantization stage 214 causes the most information loss in the process 200A. The greater the information loss, the fewer bits the quantized transformation coefficients 216 may require. To obtain different degrees of information loss, the encoder can use different quantization parameter values or any other parameters of the quantization process.
[0072] In the binary coding stage 226, the encoder may use binary coding techniques to code the predicted data 206 and the quantized transform coefficients 216, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context - adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the predicted data 206 and the quantized transform coefficients 216, the encoder may code other information in the binary coding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the type of transform in the transform stage 212, the parameters of the quantization process (e.g., quantization parameter), the encoder control parameters (e.g., bit - rate control parameter), etc. The encoder may use the output data of the binary coding stage 226 to generate the video bitstream 228. In some embodiments, the video bitstream 228 may be further packetized for network transmission.
[0073] Referring to the reconstruction path of process 200A, in the inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In the inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224 that will be used in the next iteration of process 200A.
[0074] It should be noted that other variations of process 200A may be adopted in the encoded video sequence 202. In some embodiments, the stages of process 200A may be executed by the encoder in a different order. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, the transform stage 212 and the quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may omit Figure 2A one or more of the
[0075] Figure 2B FIG. illustrates a schematic diagram of another example encoding process 200B consistent with embodiments of the present disclosure. Process 200B may be modified from process 200A. For example, process 200B may be used by an encoder compliant with a hybrid video coding standard (e.g., H.26x series). Compared with process 200A, the forward path of process 200B additionally includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.
[0076] Generally speaking, prediction techniques can be classified into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-image prediction or "intra-frame prediction") can use pixels from one or more encoded adjacent BPUs in the same image to predict the current BPU. That is to say, the prediction reference 224 in spatial prediction can include adjacent BPUs. Spatial prediction can reduce the spatial redundancy inherent in the image. Temporal prediction (e.g., inter-image prediction or "inter-frame prediction") can use regions from one or more encoded images to predict the current BPU. That is to say, the prediction reference 224 in temporal prediction can include encoded images. Temporal prediction can reduce the temporal redundancy inherent in the image.
[0077] Referring to process 200B, in the forward path, the encoder performs prediction operations in the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder can perform intra-frame prediction. For the original BPU of the encoded image, the prediction reference 224 can include one or more adjacent BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same image. The encoder can generate the predicted BPU 208 by interpolating the adjacent BPUs. The interpolation techniques can include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder can perform interpolation at the pixel level, for example, by interpolating the values of the corresponding pixels for each pixel used to predict BPU 208. The adjacent BPUs used for interpolation can be located in various directions relative to the original BPU, such as in the vertical direction (e.g., at the top of the original BPU), the horizontal direction (e.g., to the left of the original BPU), the diagonal direction (e.g., bottom left, bottom right, top left, or top right of the original BPU), or any direction defined in the video coding standard being used. For intra-frame prediction, the prediction data 206 can include, for example, the positions (e.g., coordinates) of the adjacent BPUs used, the sizes of the adjacent BPUs used, the parameters of the interpolation, the direction of the adjacent BPUs used relative to the original BPU, etc.
[0078] In another example, during the temporal prediction stage 2044, the encoder may perform inter-frame prediction. For the original BPU of the current image, the prediction reference 224 may include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images may be encoded and reconstructed on a per-BPU basis. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all the reconstructed BPUs of the same image are generated, the encoder may generate a reconstructed image as a reference image. The encoder may perform a "motion estimation" operation to search for a matching region within the range of the reference image (referred to as the "search window"). The position of the search window in the reference image may be determined based on the position of the original BPU in the current image. For example, the search window may be centered at the position in the reference image that has the same coordinates as the original BPU in the current image and may extend a predetermined distance. When the encoder identifies (e.g., by using a pixel recursive algorithm, a block matching algorithm, etc.) a region in the search window that is similar to the original BPU, the encoder may determine such a region as the matching region. The matching region may have a different size (e.g., smaller than, equal to, larger than the original BPU or have a different shape from the original BPU). Since the reference image and the current image are temporally separated on the time axis (e.g., as Figure 1 shown), the matching region can be considered to "move" over time to the position of the original BPU. The encoder may record the direction and distance of this motion as a "motion vector". When multiple reference images are used (e.g., as in Figure 1 image 106), the encoder may search for the matching region and determine its associated motion vector for each reference image. In some embodiments, the encoder may assign weights to the pixel values of the matching regions of the respective matching reference images.
[0079] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, the prediction data 206 may include, for example, the position (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference images, the weights associated with the reference images, etc.
[0080] To generate the prediction BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation can be used to reconstruct the prediction BPU 208 based on the prediction data 206 (e.g., the motion vector) and the prediction reference 224. For example, the encoder may move the matching region of the reference image according to the motion vector, where the encoder may predict the original BPU of the current image. When multiple reference images are used (e.g., as in Figure 1In the image 106), the encoder can move the matching region of the reference image according to the respective motion vectors and the average pixel value of the matching region. In some embodiments, if the encoder has assigned weights to the pixel values of the matching regions of the respective matching reference images, the encoder can sum the weighted sums of the pixel values of the moved matching regions.
[0081] In some embodiments, the inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference images in the same temporal direction relative to the current image. For example, Figure 1 The image 104 in is an unidirectional inter-frame prediction image, where the reference image (i.e., image 102) is before image 104. Bidirectional inter-frame prediction can use one or more reference images in two temporal directions relative to the current image. For example, Figure 1 The image 106 in is a bidirectional inter-frame prediction image, where the reference images (i.e., images 104 and 108) are in two temporal directions relative to image 104.
[0082] Still referring to the forward path of process 200B, after the spatial prediction stage 2042 and the temporal prediction stage 2044, at the mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra-frame prediction or inter-frame prediction) for the current iteration of process 200B. For example, the encoder can perform rate-distortion optimization techniques, where the encoder can select a prediction mode to minimize the value of a cost function based on the bit rate of the candidate prediction modes and the distortion of the reconstructed reference image under the candidate prediction modes. According to the selected prediction mode, the encoder can generate the corresponding prediction BPU 208 and prediction data 206.
[0083] In the reconstruction path of process 200B, if an intra prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current image), the encoder can directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for interpolation of the next BPU of the current image). If an inter prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current image in which all BPUs have been encoded and reconstructed), the encoder can provide the prediction reference 224 to the loop filter stage 232, where the encoder can apply loop filtering to the prediction reference 224 to reduce or eliminate the distortion (e.g., block artifacts) introduced by inter prediction. The encoder can apply various loop filtering techniques at the loop filter stage 232, such as deblocking, sample adaptive offset, adaptive loop filtering, etc. The reference image for loop filtering can be stored in the buffer 234 (or "decoded picture buffer") for later use (e.g., as an inter prediction reference image for future pictures of the video sequence 202). The encoder can store one or more reference images in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the encoder can encode the parameters of loop filtering (e.g., loop filter strength), as well as the quantized transform coefficients 216, prediction data 206, and other information, in the binary coding stage 226.
[0084] Figure 3A FIG. illustrates a schematic diagram of an example decoding process 300A consistent with embodiments of the present disclosure. Process 300A can be a decompression process corresponding to Figure 2A the compression process 200A in. In some embodiments, process 300A can be similar to the reconstruction path of process 200A. The decoder can decode the video bitstream 228 into a video stream 304 according to process 300A. The video stream 304 can be very similar to the video sequence 202. However, due to information loss during the compression and decompression processes (e.g., Figures 2A - 2B the quantization stage 214 in), generally, the video stream 304 is not the same as the video sequence 202. Similar to Figures 2A - 2B processes 200A and 200B in, the decoder can perform process 300A on each image encoded in the video bitstream 228 at the basic processing unit (BPU) level. For example, the decoder can perform process 300A in an iterative manner, where the decoder can decode one basic processing unit in one iteration of process 300A. In some embodiments, the decoder can perform process 300A in parallel on regions (e.g., regions 114 - 118) of each image encoded in the video bitstream 228.
[0085] In Figure 3AIn [the figure], the decoder may provide a portion of the video bitstream 228 associated with a basic processing unit of the encoded image (referred to as an "encoded BPU") to the binary decoding stage 302. At the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may provide the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may provide the prediction data 206 to the prediction stage 204 to generate a predicted BPU 208. The decoder may add the reconstructed residual BPU 222 to the predicted BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded image buffer in a computer memory). The decoder may provide the prediction reference 224 to the prediction stage 204 to perform a prediction operation in the next iteration of process 300A.
[0086] The decoder may iteratively execute process 300A to decode each encoded BPU of the encoded image and generate a prediction reference 224 for decoding the next encoded BPU of the encoded image. After decoding all the encoded BPUs of the encoded image, the decoder may output the image to the video stream 304 for display and continue to decode the next encoded image in the video bitstream 228.
[0087] At the binary decoding stage 302, the decoder may perform the inverse operations of the binary coding techniques used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder may decode other information at the binary decoding stage 302, such as a prediction mode, parameters of a prediction operation, a transform type, a quantization parameter process (e.g., quantization parameter), encoder control parameters (e.g., bit rate control parameters), etc. In some embodiments, if the video bitstream 228 is transmitted in packets over a network, the decoder may unpack the video bitstream 228 before feeding it to the binary decoding stage 302.
[0088] Figure 3B A schematic diagram of another example decoding process 300B consistent with embodiments of the present disclosure is shown. Process 300B may be modified from process 300A. For example, process 300B may be used by a decoder compliant with a hybrid video coding standard (e.g., the H.26x series). Compared with process 300A, process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filtering stage 232 and a buffer 234.
[0089] In process 300B, for the encoded basic processing unit (referred to as "current BPU") of the decoded encoded image (referred to as "current image"), the prediction data 206 decoded by the decoder in the binary decoding stage 302 can include various types of data, depending on what prediction mode the encoder uses to encode the current BPU. For example, if the encoder uses intra prediction to encode the current BPU, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating intra prediction, parameters of the intra prediction operation, and so on. The parameters of the intra prediction operation can include, for example, the positions (e.g., coordinates) of one or more adjacent BPUs used as references, the sizes of the adjacent BPUs, interpolation parameters, the orientations of the adjacent BPUs relative to the original BPU, and so on. For another example, if the encoder uses inter prediction to encode the current BPU, the prediction data 206 can include a prediction mode indicator (e.g., a flag value) indicating inter prediction, parameters of the inter prediction operation, and so on. The parameters of the inter prediction operation can include, for example, the number of reference images associated with the current BPU, the weights respectively associated with the reference images, the positions (e.g., coordinates) of one or more matching regions in each reference image, one or more motion vectors associated with each matching region, and so on.
[0090] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. The details of performing this spatial prediction or temporal prediction are described in Figure 2B and will not be repeated hereinafter. After performing such spatial prediction or temporal prediction, the decoder can generate a predicted BPU 208. The decoder can add the predicted BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as Figure 3A described.
[0091] In process 300B, the decoder can provide the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing prediction operations in the next iteration of process 300B. For example, if the current BPU is decoded using intra prediction in the spatial prediction stage 2042, after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder can directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current image). If the current BPU is decoded using inter prediction in the temporal prediction stage 2044, after generating the prediction reference 224 (e.g., the reference image in which all BPUs have been decoded), the encoder can provide the prediction reference 224 to the loop filter stage 232 to reduce or eliminate distortions (e.g., blocking artifacts). The decoder can Figure 2BApply loop filtering to the prediction reference 224 in the manner described. The loop-filtered reference image can be stored in a buffer 234 (e.g., a decoded picture buffer in a computer memory) for later use (e.g., as an inter-prediction reference image for future coded pictures of a video bitstream 228). The decoder can store one or more reference images in the buffer 234 for use during the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-prediction is used to encode the current BPU, the prediction data can further include parameters of the loop filter (e.g., loop filter strength).
[0092] Figure 4 is a block diagram of an example apparatus 400 for encoding or decoding video that is consistent with embodiments of the present disclosure. As Figure 4 shown, the apparatus 400 can include a processor 402. When the processor 402 executes the instructions described herein, the apparatus 400 can become a special-purpose machine for video encoding or decoding. The processor 402 can be any type of circuit capable of manipulating or processing information. For example, the processor 402 can include any combination of any number of central processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), generic array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems on a chip (SoCs), application specific integrated circuits (ASICs), etc. In some embodiments, the processor 402 can also be a group of processors grouped as a single logical component. For example, as Figure 4 shown, the processor 402 can include multiple processors, including processor 402a, processor 402b, and processor 402n.
[0093] The device 400 can also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as shown. As Figure 4As shown, the stored data may include program instructions (e.g., program instructions for implementing the stages in processes 200A, 200B, 300A, or 300B) and data to be processed (e.g., video sequence 202, video bitstream 228, or video stream 304). The processor 402 may access the program instructions and data for processing (e.g., via bus 410) and execute the program instructions to operate on or manipulate the data for processing. The memory 404 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, the memory 404 may include any number of random access memories (RAMs), read-only memories (ROMs), optical discs, magnetic disks, hard disk drives, solid state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, etc. The memory 404 may also be a set of memories grouped as a single logical component ( Figure 4 not shown in
[0094] The bus 410 may be a communication device for transferring data between components inside the device 400, such as an internal bus (e.g., a CPU-memory bus), an external bus (e.g., a universal serial bus port, a peripheral component interconnect express port), etc.
[0095] For ease of explanation without causing ambiguity, in the present disclosure, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits". The data processing circuits may be implemented entirely in hardware, or as a combination of software, hardware, or firmware. Additionally, the data processing circuits may be a single stand-alone module, or may be fully or partially incorporated into any other component of the device 400.
[0096] The device 400 may also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, repeaters, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication ("NFC") adapters, cellular network chips, etc.
[0097] In some embodiments, optionally, the device 400 may further include a peripheral interface 408 to provide connections to one or more peripheral devices. As Figure 4 shown, the peripheral devices may include, but are not limited to, cursor control devices (such as a mouse, a touchpad, or a touchscreen), a keyboard, a display (such as a cathode ray tube display, a liquid crystal display, or a light-emitting diode display), a video input device (e.g., a camera or an input interface communicatively coupled to a video archive), etc.
[0098] Note that a video codec (e.g., the codec performing processes 200A, 200B, 300A, or 300B) may be implemented as any combination of any software or hardware modules in apparatus 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more software modules of apparatus 400, such as program instructions that may be loaded into memory 404. As another example, some or all stages of processes 200A, 200B, 300A, or 300B may be implemented as one or more hardware modules of apparatus 400, such as dedicated data processing circuitry (e.g., FPGA, ASIC, NPU, etc.).
[0099] In quantization and inverse quantization functional blocks (e.g., Figure 2A or Figure 2B the quantization 214 and inverse quantization 218 of Figure 3A or Figure 3B the inverse quantization 218 of), a quantization parameter (QP) is used to determine the amount of quantization (and inverse quantization) applied to a prediction residual. An initial QP value for encoding an image or a slice may be signaled at a high level, e.g., using the init_qp_minus26 syntax element in a picture parameter set (PPS) and using the slice_qp_delta syntax element in a slice header. Additionally, incremental QP values sent in the granularity of quantization groups may be used to adjust the QP value for each CU at a local level.
[0100] To improve the accuracy of motion vectors (MVs) in the merge mode, a decoding-side motion vector refinement (DMVR) based on bilateral matching (BM) is adopted in Versatile Video Coding (VVC) draft 6. In the bilateral operation, a refined MV is searched around an initial MV in reference picture list L0 and reference picture list L1. The BM-based DMVR calculates the distortion between two candidate blocks in reference picture list L0 and list L1. Figure 5 FIG. illustrates an example process 500 of decoding-side motion vector refinement (DMVR) according to some embodiments of the present disclosure. Figure 5Shows the current image 502, the first reference image 504 in the first reference image list L0, and the second reference image 506 in the second reference image list L1, the first initial MV 508 pointing from the current block 510 in the current image 502 to the first initial reference block 512 in the first reference image 504, and the second initial MV 514 pointing from the current block 510 to the second initial reference block 516 in the second reference image 506. The process 500 may perform BM-based DMVR to determine the first candidate reference block 518 in the first reference image 504, the second candidate reference block 520 in the second reference image 506, the first candidate MV 522 connecting the current block 510 and the first candidate reference block 518, and the second candidate MV 524 connecting the current block 510 and the second candidate reference block 516. As Figure 5 shown, the first candidate MV 522 and the second candidate MV 524 are respectively close to the first initial MV 508 and the second initial MV 514. In some embodiments, the process 500 may calculate the sum of absolute differences (SAD) between the first initial reference block 518 and the second initial reference block 520 based on each MV candidate (e.g., the first candidate MV 522 or the second candidate MV 524) around the initial MV (e.g., the first initial MV 508 or the second initial MV 514). The one with the lowest SAD among the first candidate MV 522 and the second candidate MV 524 may become the corrected MV for generating the bi-prediction signal.
[0101] In some embodiments, as described in the VVC draft 6, DMVR is applied to a CU that satisfies all of the following conditions: (1) The merge mode is at the CU level with bi-prediction MVs; (2) Bi-prediction MVs with equal weights are used to predict the block (e.g., by not applying bi-prediction with weighted average (BWA) to the block); (3) Relative to the current image (e.g., the current image 502), one reference image is in the past (e.g., the first reference image 504) and the other reference image is in the future (e.g., the second reference image 506); (4) The distances (e.g., picture order count (POC) differences) of the two reference images from the current image are the same; (5) The block has at least 128 luminance samples, and both the width and height of the block are at least 8 luminance samples.
[0102] The corrected MV obtained by the DMVR process (e.g., the process 500) can be used to generate inter-prediction samples and temporal motion vector prediction for future image coding. The initial MV (e.g., the first initial MV 508 or the second initial MV 514) can be used for the deblocking process and spatial MV prediction for future CU coding within the current image (e.g., the current image 502).
[0103] As Figure 5As shown, the first MV offset 526 represents the correction offset between the first initial MV 508 and the first candidate MV 522, and the second MV offset 528 represents the correction offset between the second initial MV 514 and the second candidate MV 524. The first MV offset 526 and the second MV offset 528 may have the same magnitude and opposite directions. In some embodiments, the search points may be around the initial MVs (e.g., the first initial MV 508 and the second initial MV 514), and the MV offsets (e.g., the first MV offset 526 and the second MV offset 528) may follow the MV difference mirroring rule. For example, the points examined by the DMVR may be represented by a pair of candidate MVs MV0 (e.g., the first candidate MV 522) and MV1 (e.g., the second candidate MV 524) based on equations (1) and (2): MV0′ = MV0 + MV_offset Equation (1) MV1′ = MV1 - MV_offset Equation (2)
[0104] In equations (1) and (2), MV0' (e.g., the first initial MV 508) and MV1' (e.g., the second candidate MV 514) represent a pair of initial MVs, and MV_offset represents the correction offset (e.g., the first MV offset 526 or the second MV offset 52) between the initial MV (e.g., MV0' or MV1') and the corrected MV (e.g., MV0 or MV1). Note that MV_offset is a vector with motion displacement (e.g., in the X and Y dimensions). In some embodiments, as described in VVC draft 5, the corrected search range (e.g., the search range of the DMVR) may be two integer luminance samples of the initial MVs (e.g., the first initial MV 508 and the second initial MV 514) from the horizontal and vertical directions.
[0105] Figure 6 FIG. illustrates an example DMVR search process 600 according to some embodiments of the present disclosure. In some embodiments, the process 600 may be performed by a codec (e.g., Figures 2A - 2B the encoder in Figures 3A - 3B or the decoder in Figure 4 . For example, the codec may be implemented as one or more software or hardware components of a device (e.g., the device 400 in Figure 6As shown, process 600 includes a stage 602 for integer sample offset search and a stage 604 for fractional sample correction. To reduce the search complexity, in some embodiments, a fast search method with an early termination mechanism is applied in stage 602. For example, a 2-iteration search scheme can be applied in stage 602 instead of using a 25-point full search to reduce the SAD checkpoints.
[0106] As Figure 6 shown, stage 604 can be after stage 602. To save computational complexity, in some embodiments, the fractional sample correction in stage 604 can be derived using a parametric error surface equation instead of performing an additional search involving SAD comparison. Stage 604 can be conditionally invoked based on the output of stage 602.
[0107] Figure 7 FIG. illustrates an example pattern 700 for DMVR integer luminance sample search according to some embodiments of the present disclosure. The DMVR integer luminance sample search can determine the point with the minimum SAD among the search samples. For example, the DMVR integer luminance sample search can be implemented as Figure 6 process 600 in, including a stage for integer sample offset search (e.g., stage 602) and a stage for fractional sample correction search (e.g., stage 604), where at least one iteration can be performed in each stage. In some embodiments, in the first iteration of the DMVR integer luminance sample search, at most 6 SADs can be checked. Taking Figure 7 as an example, in the first iteration, the SADs of five points 702-710 (represented as black blocks) can be compared, where point 702 can be used as the center point of the search. If the SAD of the center point (i.e., point 702) is the minimum, the integer sampling stage of DMVR can be terminated. Otherwise, another point 712 (represented as a shaded block) determined by the SAD distribution of points 704-710 can be checked. In the second iteration of the DMVR integer luminance sample search, the point with the minimum SAD among points 704-712 can be selected as the new search center point. In some embodiments, the second iteration can be performed in the same way as the first iteration. In some embodiments, the SADs calculated in the first iteration can be reused in the second iteration, so only the SADs of additional points need to be further calculated.
[0108] In some embodiments, as described in VVC draft 6, the 2-iteration search as described in Figure 7 can be removed. Then, in the integer sample offset search stage (e.g., Figure 6 stage 602 in), the SADs of all 25 points can be calculated in one iteration. Figure 8Illustrated is an example pattern 800 for the stage of integer sample offset search in DMVR integer luminance sample search according to some embodiments of the present disclosure. For example, the stage of integer sample offset search can be Figure 6 stage 602 in Figure 8 Shows the initial MV 802 and 25 points, and their SADs can be calculated together. In some embodiments, the SAD of the initial MV 802 can be reduced (e.g., reduced by a quarter) to adjust the initial MV 802. In some embodiments, it can be further corrected at a stage for partial sample correction (e.g., in Figure 6 stage 604 in Figure 8 the position with the minimum SAD). The fractional sample correction stage can be conditionally called according to the position with the minimum SAD. For example, as
[0109] Figure 9 shown, if the position with the minimum SAD is one of the nine points around the initial MV 802 (as shown in box 804), the stage for fractional sample correction can be called to determine the corrected MV as the output of the DMVR integer luminance sample search. If the position with the minimum SAD is not any of the nine points around the initial MV 802, the position with the minimum SAD can be directly used as the output of the DMVR integer luminance sample search. Figure 8 Shows the initial MV 902 and 25 points. The initial MV 902 is connected to the central point 904 with the minimum SAD. In the sub-pixel offset estimation based on the parameter error surface, as Figure 9 shown, the sum of the absolute difference (SAD) cost of the central point 904 and the SAD costs of four adjacent points 906 - 912 around the central point 904 can be used to fit a two-dimensional parabolic error surface equation. For example, the two-dimensional parabolic error surface equation can be based on Equation (3). E(x,y) = ((A(x - x mim ) 2 + B(y - y min ) 2 +) >> mvShift) + E(0,0) Equation (3)
[0110] In Equation (3), (x min , y min ) corresponds to the fractional position with the lowest SAD cost, E(x, y) corresponds to the SAD costs of the central point 904 and the four adjacent points 906 - 912, mvShift can be set to 4 as in VVC (in VVC, the MV accuracy is 1 / 16 pixel), and A and B can be determined according to Equations (4) and (5) respectively:
[0111] By solving equations (3) to (5) using the SAD cost values of five search points (i.e., points 904 - 912), (x min , y min ) can be determined according to equations (6) and (7).
[0112] In some embodiments, the values of x min and y min can be automatically limited between -8 and 8 (e.g., in 1 / 16 sample precision), because all SAD cost values are positive and the minimum value is E(0, 0), which corresponds to a half-pixel offset with 1 / 16 pixel MV precision in VVC. The calculated fractions (x min , y min ) can be added to the integer distance corrected MV to make the corrected MV have sub-pixel precision.
[0113] The Bidirectional Optical Flow (BDOF) tool is included in VVC. As the name implies, the BDOF mode is based on the optical flow concept, assuming that the motion of objects is smooth. BDOF, previously called BIO, is also included in the Joint Video Exploration Model (JEM) software. Compared with BIO in JEM, BDOF in VVC is a simpler version, especially in terms of the number of multiplications and the size of the multipliers, requiring much less computational effort.
[0114] BDOF can be used to correct the bidirectional prediction signal of a CU at the 4×4 sub-block level. In some embodiments, BDOF is applied to a CU that satisfies the following conditions: (1) the height of the CU is not 4 and the size of the CU is not 4×8; (2) the CU is not encoded using the affine mode or the Advanced Temporal Motion Vector Prediction (ATMVP) merge mode; (3) the CU is encoded using the "true" bidirectional prediction mode, where one of the two reference images (e.g., Figure 5 the first reference image 504) is before the current image (e.g., Figure 5 the current image 502) in the display order, and the other (e.g., Figure 5 the second reference image 506) is after the current image in the display order. In some embodiments, BDOF can be applied to the luminance component.
[0115] In some embodiments, when BDOF is used to correct the bidirectional prediction signal of a CU at the 4×4 sub-block level, for each 4×4 sub-block, the motion correction (v x , v y ) can be calculated by minimizing the difference between the predicted samples in the two reference image lists L0 and L1. (v x , vy ) It can then be used to adjust the bidirectional prediction sample values in the 4×4 sub-block.
[0116] In some embodiments, the following steps are applied during the BDOF process. First, the horizontal and vertical gradients of the two prediction signals at levels k = 0, 1 can be determined based on calculating the difference between two adjacent samples and as shown in equations (8) and (9):
[0117] In equations (8) and (9), I (k) (i,j) is the sample value of the prediction signal in list k at coordinates (i,j) when k = 0, 1. shift1 is calculated based on the luminance bit depth (“bitDepth”), as in equation (10): shift1 = max(2, 14 - bitDepth) Equation (10)
[0118] Then, the auto-correlation and cross-correlation of the gradients S1, S2, S3, S5, and S6 can be determined according to equations (11) to (15):
[0119] For equations (11) to (15), ψ x (i,j), ψ y (i,j), θ(i,j) values θ(i,j) = (I (1) (i,j) >> n b ) - (I (0) (i,j) >> n b ) Equation (18)
[0120] In equations (11) to (18), Ω is the 6×6 window around the 4×4 sub-block, n a and n b values are set by equations (19) and (20) respectively: n a = min(5, bitDepth - 7) Equation (19) n b = min(8, bitDepth - 4) Equation (20)
[0121] Then, the motion correction (v x , v y ) is obtained using equations (21) and (22) based on the cross-correlation and auto-correlation terms:
[0122] In equations (21) and (22), th′ BIO = 2 13-BD represents the floor function, and
[0123] Based on motion correction and gradients, the following adjustment b(x, y) for each sample in the 4×4 sub-block can be determined using equation (23):
[0124] Finally, the BDOF samples of the CU can be determined by adjusting the bi-directional prediction samples according to equation (24): pred BDoF (x, y) = (I (0) (x, y)+I (1) (x, y)+b(x, y)+o offset ) >> shift Equation (24)
[0125] In some embodiments, the values in equations (8) to (23) can be selected such that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit-width of the intermediate parameters in the BDOF process can be kept within 32 bits.
[0126] To derive the gradient values, some prediction samples I (k) (i, j) outside the current CU boundary in the list k (k = 0, 1) can be generated. Figure 10 is a schematic diagram of an example of an extended coding unit (CU) region 1000 used in BDOF according to some embodiments of the present disclosure. As Figure 10 shown, the 4×4 block 1002 (enclosed by the solid black line) used in BDOF is surrounded by an extended row or column around the boundary of the block 1002 (represented by the dashed black line), forming an enclosed region 1004. To control the computational complexity of generating prediction samples outside the boundary, prediction samples within the extended region 1006 (represented by the white box) can be generated by directly obtaining reference samples at nearby integer positions (using the floor operation on the coordinates) without interpolation, and prediction samples within the CU 1008 (represented by the gray box) can be generated using ordinary 8-tap motion compensation interpolation filtering. These extended sample values can only be used for gradient calculation. For the remaining steps of the BDOF process, if any samples and gradient values outside the boundary of the CU 1008 are needed, they can be filled (or repeated) from their nearest neighbors.
[0127] At the JVET conference, an encoding tool called Prediction Refinement based on Optical Flow (PROF) was adopted. PROF improves the accuracy of affine motion compensation prediction by using optical flow to correct sub-block-based affine motion compensation prediction. The affine motion model parameters can be used to derive the motion vector for each sample position in a CU. However, due to the high complexity and memory access bandwidth of generating per-sample affine motion compensation prediction, affine prediction in VVC uses a sub-block-based affine motion compensation method, where a CU is divided into 4×4 sub-blocks, each assigned an MV derived from the control point MVs of the affine CU. Sub-block-based affine motion compensation is a trade-off among coding efficiency, complexity, and memory access bandwidth. Due to its sub-block-based prediction rather than per-sample motion compensation prediction, it loses some prediction accuracy.
[0128] To achieve a finer affine motion compensation granularity, in some embodiments, PROF can be applied after conventional sub-block-based affine motion compensation. The sample-based refinement can be obtained based on the optical flow equation, such as Equation (25): ΔI(i,j) = g x (i,j) * Δv x (i,j) + g y (i,j) * Δv y (i,j) Equation (25)
[0129] In Equation (25), g x (i,j) and g y (i,j) are the spatial gradients at the sample position (i,j), and Δv is the motion offset from the sub-block-based motion vector to the sample-based motion vector derived from the affine model parameters.
[0130] Figure 11 is a schematic diagram of an example of sub-block-based translational motion and sample-based affine motion according to some embodiments of the present disclosure. As Figure 11 shown, V(i,j) is the theoretical motion vector at the sample position (i,j) derived using the affine model, V SB is the sub-block-based motion vector, and ΔV(i,j) (represented by the dashed arrow) is the difference between V(i,j) and V SB .
[0131] Then, the prediction refinement ΔI(i,j) can be added to the sub-block prediction I(i,j). The final prediction I' can be generated based on Equation (26): I′(i,j) = I(i,j) + ΔI(i,j) Equation (26)
[0132] Consistent with the disclosed embodiments, both DMVR and BDOF can have control flags at two levels in the syntax structure. The first control flag can be sent in the sequence parameter set (SPS) at the sequence level, and the second control flag can be sent in the slice header at the slice level. Figure 12 Table 1 is shown, and Table 1 shows an example syntax structure of a sequence parameter set (SPS) that implements control flags for DMVR and BDOF according to some embodiments of the present disclosure. As Figure 12 shown in Table 1 of, sps_bdof_enabled_flag and sps_dmvr_enabled_flag are control flags for BDOF and DMVR, respectively, at the sequence level sent in the SPS. When sps_bdof_enabled_flag or sps_dmvr_enabled_flag is false, BDOF or DMVR can be disabled throughout the video sequence that references this SPS. When sps_bdof_enabled_flag and sps_dmvr_enabled_flag are true, BDOF or DMVR can be enabled for the current video sequence. In this case, another flag sps_bdof_dmvr_slice_present_flag can be further signaled to indicate whether slice-level control of BDOF and DMVR is enabled.
[0133] Figure 13 Table 2 is shown, and Table 2 shows an example syntax structure of a slice header that implements control flags for DMVR and BDOF according to some embodiments of the present disclosure. As shown in Table 2, when sps_bdof_dmvr_slice_present_flag set in Table 1 is true, slice_disable_bdof_dmvr_flag can signal in the slice header whether the current slice disables BDOF and DMVR.
[0134] Figures 12 - 13 shows a two-level control mechanism for DMVR and BDOF. By using this mechanism, an encoder (e.g., an encoder that implements Figures 2A - 2B process 200A or 200B in) can use the slice-level flag slice_disable_bdof_dmvr_flag to turn on or off DMVR and BDOF for each slice. This slice-level adaptability can have two benefits: (1) when at least one of DMVR or BDOF is useless for the current slice, turning it (or them) off can improve encoding performance; (2) both the computational complexity of DMVR and BDOF is relatively high, and turning them off can reduce the encoding and decoding complexity of the current slice.
[0135] In the disclosed embodiments, the control flags for PROF can also be used at the sequence level and slice level. In some embodiments, three separate flags can be signaled in the SPS to indicate whether DMVR, BDOF, and PROF are enabled, respectively. If any of them is enabled, a corresponding lower-level control enable flag can be signaled to indicate whether the enabled tool is controlled at the lower level. The lower level can be the slice level or the picture level. If slice-level or picture-level control is enabled, in each slice header or picture header, a slice-level or picture-level disable flag can be signaled to indicate that the enabled tool for the current slice or picture is disabled.
[0136] Consistent with the disclosed embodiments, Figure 14A Table 3A is shown, and Table 3A shows an example syntax structure of a sequence parameter set (SPS) that implements slice-level control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. Figure 14B Table 3B is shown, and Table 3B shows an example syntax structure of an SPS that implements picture-level control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As shown in Tables 3A - 3B, the emphasis is in italics, and sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag are the flags signaled in the SPS to indicate whether BDOF, DMVR, and PROF are enabled for the video sequence, respectively. If BDOF, DMVR, or PROF is enabled, as shown in Table 3A, sps_bdof_slice_present_flag, sps_dmvr_slice_present_flag, or sps_affine_prof_slice_present_flag can be further signaled to indicate whether slice-level control of BDOF, DMVR, and PROF is enabled, respectively. If BDOF, DMVR, or PROF is enabled, as shown in Table 3B, sps_bdof_picture_present_flag, sps_dmvr_picture_present_flag, or sps_affine_prof_picture_present_flag can be further signaled, respectively, to indicate whether picture-level control of BDOF, DMVR, and PROF is enabled.
[0137] Figure 15AFigure 4A is illustrated, and Figure 4A shows an example syntax structure of a slice header that implements control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As highlighted in italics in Table 4A, if any of sps_bdof_slice_present_flag, sps_dmvr_slice_present_flag, or sps_affine_prof_slice_present_flag in Table 3A is set to true, then slice_disable_bdof_flag, slice_disable_dmvr_flag, or slice_disable_affine_prof_flag can be signaled respectively to indicate whether BDOF, DMVR, or PROF is disabled for the current slice. Figure 15B Figure 4B is shown, and Figure 4B shows an example syntax structure of a picture header that implements control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As highlighted in italics in Table 4B, if any of sps_bdof_picture_present_flag, sps_dmvr_picture_present_flag, or sps_affine_prof_picture_present_flag in Table 3B is set to true, then ph_disable_bdof_flag, ph_disable_dmvr_flag, or ph_disable_affine_prof_flag can be signaled respectively to indicate whether BDOD, DMVR, or PROF is disabled for the current picture.
[0138] In some embodiments, DMVR, BDOF, and PROF may have three separate sequence-level enable flags but share the same slice-level control enable flag. For example, a single slice-level disable flag can be sent for DMVR, BDOF, and PROF. In another example, three slice-level disable flags can be signaled separately for DMVR, BDOF, and PROF. As another example, two slice-level disable flags can be sent for DMVR, BDOF, and PROF. It should be noted that the control of DMVR, BDOF, and PROF can implement various syntax structures at the sequence level and levels below the sequence level (hereinafter simply referred to as the "lower levels", such as the slice level or the picture level), and it is not limited to the examples described herein.
[0139] Figure 16FIG. 5 illustrates an example syntax structure of a sequence parameter set (SPS) that implements separate sequence-level control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As shown in Table 5, highlighted in italics, three independent flags, sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag, are signaled in the SPS respectively to indicate whether DMVR, BDOF, and PROF are enabled. If at least one of DMVR, BDOF, or PROF is enabled, the slice control enable flag sps_bdof_dmvr_affine_prof_slice_present_flag can be signaled to indicate whether at least one of DMVR, BDOF, or PROF enabled at the sequence level is controlled at a lower level.
[0140] Figure 17 FIG. 6 illustrates an example syntax structure of a slice header that implements joint control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As shown highlighted in italics in Table 6, if slice-level control is enabled as described in Table 5 (e.g., sps_bdof_dmvr_affine_prof_slice_present_flag is true), then in each slice header, the slice-level disable flag slice_disable_bdof_dmvr_affine_prof_flag can be signaled to indicate whether at least one of DMVR, BDOF, or PROF enabled at the sequence level is disabled for the current slice. In the syntax structure shown in Table 6, if multiple DMVR, BDOF, and PROF are enabled at the sequence level, then in the case where slice-level control is enabled, the multiple enables can be jointly controlled at the slice level.
[0141] Figure 18FIG. 7 illustrates an example syntax structure of a slice header that implements separate control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As highlighted in italics in Table 7, if the slice-level control described in Table 5 is enabled (e.g., sps_bdof_dmvr_affine_prof_slice_present_flag is true), then for one of DMVR, BDOF, and PROF enabled at the sequence level, a slice-level disable flag can be signaled to indicate whether the aforementioned enabled one is disabled for the current slice. For example, if sps_bdof_enabled_flag and sps_bdof_dmvr_affine_prof_slice_present_flag in Table 5 are set to true, then slice_disable_bdof_flag in Table 7 can be signaled to indicate whether BDOF is disabled for the current slice. As another example, if sps_dmvr_enabled_flag and sps_bdof_dmvr_affine_prof_slice_present_flag in Table 5 are set to true, then slice_disable_dmvr_flag in Table 7 can be signaled to indicate whether DMVR is disabled for the current slice. In yet another example, if sps_affine_prof_enabled_flag and sps_bdof_dmvr_affine_prof_slice_present_flag in Table 5 are set to true, then slice_disable_affine_prof_flag in Table 7 can be signaled to indicate whether PROF is disabled for the current slice. In Table 7, if slice-level control is enabled, each of DMVR, BDOF, and PROF can be controlled separately at the slice level
[0142] Considering the fact that both BDOF and PROF use optical flow to correct the inter-frame predictor, in some embodiments, BDOF and PROF can share the same slice-level control flag, and DMVR can use a separate slice-level control flag. Figure 19Table 8 is shown, which shows an example syntax structure of a slice header that implements hybrid control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. In the syntax structure of Table 8, BDOF and PROF can share the same slice-level disable flag, and DMVR can use another slice-level disable flag. As highlighted in italics in Table 8, if slice-level control is enabled in Table 5 (e.g., sps_bdof_dmvr_affine_prof_slice_present_flag is true), then two slice-level disable flags can be signaled to indicate whether DMVR, BDOF, and PROF are disabled for the current slice. For example, if at least one of sps_bdof_enabled_flag and sps_affine_prof_enabled_flag in Table 5 is set to true, and sps_bdof_dmvr_affine_prof_slice_present_flag in Table 5 is set to true, then slice_disable_bdof_affine_prof_flag in Table 8 can be signaled to indicate whether at least one of BDOF or PROF enabled at the sequence level is disabled for the current slice. In another example, if sps_dmvr_enabled_flag and sps_bdof_dmvr_affine_prof_slice_present_flag in Table 5 are set to true, then slice_disable_dmvr_flag in Table 8 can be signaled to indicate whether DMVR is disabled for the current slice. In Table 8, if slice-level control is enabled, BDOF and PROF are jointly controlled at the slice level, and DMVR is controlled separately from BDOF and PROF. Figure 20Table 9 is shown, which shows an example syntax structure of a sequence parameter set (SPS) that implements hybrid sequence-level control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. In the syntax structure of Table 9, three slice-level disable flags can be issued for DMVR, BDOF, and PROF, respectively. As shown in Table 9, highlighted in italics, three separate flags, sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag, can be signaled in the SPS to indicate whether DMVR, BDOF, or PROF is enabled, respectively. If sps_dmvr_enabled_flag is true, the slice-level control enable flag sps_dmvr_slice_present_flag can be signaled to indicate whether DMVR is controlled at the slice level. If at least one of sps_bdof_enabled_flag or sps_affine_prof_enabled_flag is true, the slice-level control enable flag sps_bdof_affine_prof_slice_present_flag can be signaled to indicate whether at least one of BDOF or PROF is controlled at the slice level.
[0143] Figure 21 Table 10 is shown, which shows another example syntax structure of a slice header that implements hybrid control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As highlighted in italics in Table 10, if sps_dmvr_slice_present_flag in Table 9 is set to true, the slice-level disable flag slice_disable_dmvr_flag can be signaled to indicate whether DMVR is disabled for the current slice. If sps_bdof_affine_prof_slice_present_flag in Table 9 is set to true, the slice-level disable flag slice_disable_bdof_affine_prof_flag can be sent to indicate whether at least one of BDOF or PROF enabled at the sequence level (as described in Table 8) can be disabled for the current slice. In the syntax structure of Table 9, BDOF and PROF can be jointly controlled at the slice level, and if slice-level control is enabled, DMVR can be controlled individually at the slice level.
[0144] Figure 22Table 11 is shown, and Table 11 shows another example syntax structure of a slice header that implements separate control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As highlighted in italics in Table 11, if the sps_dmvr_enabled_flag in Table 9 is set to true, the slice-level disable flag slice_disable_dmvr_flag can be signaled to indicate whether DMVR is disabled for the current slice. If the sps_bdof_enabled_flag and sps_bdof_affine_prof_slice_present_flag in Table 9 are set to true, the slice-level disable flag slice_disable_bdof_flag can be signaled to indicate whether BDOF is disabled for the current slice. If the sps_affine_prof_enabled_flag and sps_bdof_affine_prof_slice_present_flag in Table 9 are set to true, the slice-level disable flag slice_disable_affine_prof_flag can be signaled to indicate whether PROF is disabled for the current slice. In the syntax structure of Table 11, if slice-level control is enabled, each of DMVR, BDOF, and PROF is controlled separately at the slice level.
[0145] Figures 23 - 26 A flowchart of an example process 2300-2600 for controlling a video coding / decoding mode according to some embodiments of the present disclosure is shown. In some embodiments, the process 2300-2600 may be performed by a codec (e.g., Figures 2A - 2B the encoder in Figures 3A - 3B or
[0146] Figure 23 the decoder in Figures 3A - 3B ). For example, the codec may be implemented as one or more software or hardware components of a device (e.g., device 400) for controlling the coding / decoding mode for encoding or decoding a video sequence. Figures 3A - 3B ).
[0147] Figures 3A - 3B In step 2304, the codec may enable or disable a video sequence based on a first flag in the bitstream (e.g.,the encoding mode of the video stream 304 in process 300A or 300B. For example, the encoding mode can be at least one of a bi-directional optical flow (BDOF) mode, an optical flow prediction refinement (PROF) mode, or a decoder-side motion vector refinement (DMVR) mode. In some embodiments, the codec can detect a first flag in the sequence parameter set (SPS) of the video sequence. For example, the first flag can be a flag sps_bdof_enabled_flag, a flag sps_dmvr_enabled_flag, or a flag sps_affine_prof_enabled_flag as described in Figure 12 , 14A -14B, 16 or 20.
[0148] In step 2306, the codec can determine whether to enable or disable control of the encoding mode at a level lower than the sequence level based on a second flag in the bitstream. The level lower than the sequence level can include a slice level or a picture level. In some embodiments, the codec can detect a second flag in the SPS of the video sequence in response to enabling the encoding mode for the video sequence. For example, the second flag can be a flag sps_bdof_dmvr_slice_present_flag, a flag sps_bdof_slice_present_flag, a flag sps_dmvr_slice_present_flag, a flag sps_affine_prof_slice_present_flag, a flag sps_bdof_picture_present_flag, a flag sps_dmvr_picture_present_flag, a flag sps_affine_prof_picture_present_flag, a flag sps_bdof_affine_prof_slice_present_flag, or a flag sps_bdof_dmvr_affine_prof_slice_present_flag as described in Figure 12 , 14A -14B, 16 or 20.
[0149] In some embodiments, after step 2306, in response to controlling the enabling of an encoding mode at a level lower than the sequence level, the codec may enable or disable the encoding mode for a target low-level region based on a third flag in the bitstream. The target low-level region may be a target slice or a target image. If the low level is the slice level, in some embodiments, the codec may detect the third flag in the slice header of the target slice. If the low level is the image level, in some embodiments, the codec may detect the third flag in the image header of the target image. For example, the third flag may be a flag such as Figure 13 , 15A -15B, 17-19 or 21-22 described flag slice_disable_bdof_dmvr_flag, flag slice_disable_bdof_flag, flag slice_disable_dmvr_flag, flag slice_disable_affine_prof_flag, flag ph_disable_bdof_flag, flag ph_disable_dmvr_flag, flag ph_disable_affine_prof_flag, flag slice_disable_bdof_dmvr_affine_prof_flag, or flag slice_disable_bdof_affine_prof_flag.
[0150] Figure 24 FIG. 2400 is a flowchart showing another example process for controlling a video decoding mode according to some embodiments of the present disclosure. At step 2402, a codec (e.g., Figures 3A - 3B the decoder in) may receive a bitstream of video data (e.g., Figures 3A - 3B the video bitstream 228 in process 300A or 300B in).
[0151] At step 2404, the codec may enable or disable a first encoding mode of a video sequence (e.g., Figures 3A - 3B the video stream 304 in process 300A or 300B in) based on a first flag in the bitstream. At step 2406, the codec may enable or disable a second encoding mode of the video sequence based on a second flag in the bitstream. The first and second encoding modes may be two different encoding modes, which may be selected from a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode, and a decoder-side motion vector correction (DMVR) mode. For example, the first encoding mode and the second encoding mode may be the bidirectional optical flow (BDOF) mode and the optical flow prediction correction (PROF) mode, respectively.
[0152] In some embodiments, the codec may detect first and second flags in the sequence parameter set (SPS) of a video sequence. For example, the first and second flags may be selected from the flags sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag described in Figure 12 , 14A -14B, 16, or 20. As another example, if the first coding mode and the second coding mode are the BDOF mode and the PROF mode, respectively, the first flag and the second flag may be the flags sps_bdof_enabled_flag and sps_affine_prof_enabled_flag described in Figure 12 , 14A -14B, 16, or 20.
[0153] In step 2408, the codec may determine whether to enable control of at least one of the first coding mode or the second coding mode at a level lower than the sequence level based on a third flag in the bitstream. The level lower than the sequence level may include the slice level or the picture level. In some embodiments, the codec may detect the third flag in the SPS of the video sequence in response to enabling at least one of the first coding mode or the second coding mode for the video sequence. For example, the third flag may be the flag sps_bdof_dmvr_slice_present_flag, sps_bdof_slice_present_flag, sps_dmvr_slice_present_flag, sps_affine__slice_present_flag, sps_bdof_picture_present_flag, sps_dmvr_picture_present_flag, sps_affine_prof_picture_present_flag, sps_bdof_affine_prof_slice_present_flag, or sps_bdof_dmvr_affine_prof_slice_present_flag described in FIGS. 12, 14A-14B, 16, or 20.
[0154] In some embodiments, after step 2408, in response to a first flag indicating that the first coding mode is enabled for the video sequence (e.g., as in Figure 20(the sps_bdof_enabled_flag described in) and a third flag (such as Figure 20 (the sps_bdof_affine_prof_slice_present_flag described in) indicating that at least one of the first coding mode or the second coding mode is enabled at a level lower than the sequence level (e.g., PROF), the codec may enable or disable the first coding mode (e.g., BDOF) for the target lower-level region based on a fourth flag in the bitstream (e.g., as Figure 22 (the slice_disable_bdof_flag described in). The target lower-level region may be a target slice or a target picture. If the target lower level is a target slice, in some embodiments, the codec may detect the fourth flag in the slice header of the target slice. If the target lower level is a target picture, in some embodiments, the codec may detect the fourth flag in the picture header of the target picture. For example, the fourth flag may be a flag such as Figure 13 , 15A -15B, 17-19 or 21-22 described flag slice_disable_bdof_dmvr_flag, flag slice_disable_bdof_flag, flag slice_disable_dmvr_flag, flag slice_disable_affine_prof_flag, flag ph_disable_bdof_flag, flag ph_disable_dmvr_flag, flag ph_disable_affine_prof_flag, flag slice_disable_bdof_dmvr_affine_prof_flag, or flag slice_disable_bdof_affine_prof_flag.
[0155] In some embodiments, after step 2408, in response to the control of enabling at least one of the first coding mode or the second coding mode at a level lower than the sequence level, the codec may enable or disable the first coding mode (e.g., BDOF) and the second coding mode (e.g., PROF) for the target lower-level region based on a fourth flag in the bitstream (e.g., Figure 21 (the slice_disable_bdof_affine_prof_flag shown). For example, the third flag (e.g., as Figure 20 (the sps_bdof_affine_prof_slice_present_flag described in) may indicate that at least one of the first coding mode or the second coding mode is enabled at a lower level (e.g., slice level).
[0156] In some embodiments, after enabling or disabling a first coding mode and a second coding mode of a target lower-level region (e.g., a target slice or a target image) based on a fourth flag in a bitstream, a codec may further enable or disable a third coding mode of a video sequence according to a second flag in the bitstream, and determine whether to enable control of the third coding mode at a level lower than the sequence level based on a fifth flag in the bitstream. For example, the first coding mode, the second coding mode, and the third coding mode may be a BDOF mode, a PROF mode, and a DMVR mode, respectively. In this example, the fourth flag may be the flag slice_disable_bdof_affine_prof_flag as described in Figure 21 , the second flag may be the flag sps_dmvr_enabled_flag as described in Figure 20 , and the fifth flag may be the flag sps_dmvr_slice_present_flag as described in Figure 20 .
[0157] In some embodiments, in response to enabling the third coding mode at a lower level (e.g., the slice level or the image level), the codec may further enable or disable the third coding mode of the target lower-level region (e.g., the target slice or the target image) based on a sixth flag in the bitstream. For example, when the first, second, and third coding modes may be a BDOF mode, a PROF mode, and a DMVR mode, respectively, the sixth flag may be the slice_disable_dmvr_flag as described in Figure 21 .
[0158] Figure 25 FIG. 2500 is a flowchart showing an example process for controlling a video coding mode according to some embodiments of the present disclosure. At step 2502, a codec (e.g., the encoder in Figures 2A - 2B ) may receive a video sequence (e.g., the video sequence 202 in process 200A or 200B in Figures 2A - 2B ), a first flag, and a second flag. For example, the first flag may be the flag sps_bdof_enabled_flag, the flag sps_dmvr_enabled_flag, or the flag sps_affine_prof_enabled_flag as described in Figure 12 , 14A -14B, 16, or 20. As another example, the second flag may be as in Figure 12 , 14A-14B, 16, or 20 of the flags sps_bdof_dmvr_slice_present_flag, sps_bdof_slice_present_flag, sps_dmvr_slice_present_flag, sps_affine_prof_slice_present_flag, sps_bdof_picture_present_flag, sps_dmvr_picture_present_flag, sps_affine_prof_picture_present_flag, sps_bdof_affine_prof_slice_present_flag, or sps_bdof_dmvr_affine_prof_slice_present_flag.
[0159] In step 2504, the codec may enable or disable the encoding mode of a video bitstream (e.g., the video bitstream 228 in process 200A or 200B in Figures 2A - 2B ) based on a first flag in the bitstream. For example, the encoding mode may be at least one of a bi-directional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode, or a decoder-side motion vector correction (DMVR) mode.
[0160] In step 2506, the codec may enable or disable control of the encoding mode at a level below the sequence level based on a second flag. The level below the sequence level may include the slice level or the picture level.
[0161] Figure 26 FIG. shows a flowchart of another example process 2600 for controlling a video encoding mode according to some embodiments of the present disclosure. In step 2602, a codec (e.g., Figures 2A - 2B the encoder in Figures 2A - 2B ) may receive a video sequence (e.g., the video sequence 202 in process 200A or 200B in Figure 12 , 14A ), a first flag, a second flag, and a third flag. For example, the first flag and the second flag may be selected from the flags sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag as described in Figure 12 , 14AThe flags sps_bdof_dmvr_slice_present_flag, sps_bdof_slice_present_flag, sps_dmvr_slice_present_flag, sps_affine_prof_slice_present_flag, sps_bdof_picture_present_flag, sps_dmvr_picture_present_flag, sps_affine_prof_picture_present_flag, sps_bdof_affine_prof_slice_present_flag, or sps_bdof_dmvr_affine_prof_slice_present_flag as described in -14B, 16, or 20.
[0162] In step 2604, the codec may enable or disable a first coding mode for a video bitstream (e.g., the video bitstream 228 in process 200A or 200B in Figures 2A - 2B ) based on a first flag. In step 2606, the codec may enable or disable a second coding mode for the video bitstream based on a second flag. The first coding mode and the second coding mode may be two different coding modes, which may be selected from a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode, and a decoder-side motion vector correction (DMVR) mode. For example, the first coding mode and the second coding mode may be the bidirectional optical flow (BDOF) mode and the optical flow prediction correction (PROF) mode, respectively.
[0163] In step 2608, the codec may enable or disable control of at least one of the first coding mode or the second coding mode at a level lower than the sequence level. The level lower than the sequence level may include the slice level or the picture level.
[0164] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by a device (such as the disclosed encoder and decoder) to perform the above method. Common forms of non-transitory media include, for example, floppy disks, hard disks, solid state drives, magnetic tapes, or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a hole pattern, RAM, PROM, and EPROM, FLASH-EPROM, or any other flash memory, NVRAM, caches, registers, any other storage chip or cartridge memory, and their network versions. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memories.
[0165] The embodiments can be further described using the following clauses: 1. A computer-implemented method, comprising: Receiving a bitstream of video data; Enabling or disabling an encoding mode for a video sequence based on a first flag in the bitstream; and Determining whether to enable or disable control of the encoding mode at a level lower than the sequence level based on a second flag in the bitstream. 2. The computer-implemented method according to clause 1, wherein the level lower than the sequence level includes a slice level or a picture level. 3. The computer-implemented method according to any one of clauses 1-2, further comprising: In response to enabling control of the encoding mode at the level lower than the sequence level, enabling or disabling the encoding mode for a target low-level region based on a third flag in the bitstream. 4. The computer-implemented method according to clause 3, further comprising: Detecting the third flag in the slice header of a target slice, where the target slice is the target low-level region; or Detecting the third flag in the picture header of a target picture, where the target picture is the target low-level region. 5. The computer-implemented method according to any one of clauses 1-4, wherein the encoding mode is at least one of the following: Bidirectional optical flow (BDOF) mode; Optical flow prediction correction (PROF) mode; or Decoder-side motion vector correction (DMVR) mode. 6. The computer-implemented method according to any one of clauses 1-5, further comprising: Detect a first flag in the sequence parameter set (SPS) of the video sequence. 7. The computer-implemented method according to any one of clauses 1-6, further comprising: In response to enabling the encoding mode for the video sequence, detect the second flag in the SPS of the video sequence. 8. A computer-implemented method, comprising: Receive a bitstream of video data; Based on a first flag in the bitstream, enable or disable a first encoding mode for the video sequence; Based on a second flag in the bitstream, enable or disable a second encoding mode for the video sequence; and Based on a third flag in the bitstream, determine whether to enable control of at least one of the first encoding mode or the second encoding mode at a level lower than the sequence level. 9. The computer-implemented method according to clause 8, wherein the level lower than the sequence level includes the slice level or the picture level. 10. The computer-implemented method according to any one of clauses 8-9, further comprising: In response to enabling control of at least one of the first encoding mode or the second encoding mode at the level lower than the sequence level, based on a fourth flag in the bitstream, enable or disable the first encoding mode and the second encoding mode for a target low-level region. 11. The computer-implemented method according to clause 10, further comprising: Detect a fourth flag in the slice header of a target slice, where the target slice is the target low-level region; or Detect a fourth flag in the picture header of a target picture, where the target picture is the target low-level region. 12. The computer-implemented method according to any one of clauses 10-11, further comprising: Based on the second flag in the bitstream, enable or disable a third encoding mode for the video sequence; and Based on a fifth flag in the bitstream, determine whether to enable control of the third encoding mode at the level lower than the sequence level. 13. The computer-implemented method according to clause 12, further comprising: In response to enabling the third encoding mode at the level lower than the sequence level, based on a sixth flag in the bitstream, enable or disable the third encoding mode for the target low-level region. 14. The computer-implemented method according to any one of clauses 12 and 13, wherein the first coding mode, the second coding mode, and the third coding mode are a bi-directional optical flow (BDOF) mode, a predicted optical flow refinement (PROF) mode, and a decoder-side motion vector refinement (DMVR) mode, respectively. 15. The computer-implemented method according to any one of clauses 8-14, further comprising: In response to enabling a first flag indicating a first coding mode for a video sequence and a third flag indicating enabling control of at least one of the first coding mode or the second coding mode at a level lower than the sequence level, enabling or disabling the first coding mode for a target lower-level region based on a fourth flag in the bitstream. 16. The computer-implemented method according to any one of clauses 8-15, wherein the first coding mode and the second coding mode are two different coding modes selected from the following: Bi-directional optical flow (BDOF) mode; Predicted optical flow refinement (PROF) mode; and Decoder-side motion vector refinement (DMVR) mode. 17. The computer-implemented method according to any one of clauses 8-16, wherein the first coding mode and the second coding mode are a bi-directional optical flow (BDOF) mode and a predicted optical flow refinement (PROF) mode, respectively. 18. The computer-implemented method according to any one of clauses 8-17, further comprising: Detecting a first flag and a second flag in a sequence parameter set (SPS) of a video sequence. 19. The computer-implemented method according to any one of clauses 8-18, further comprising: In response to enabling at least one of the first coding mode or the second coding mode for the video sequence, detecting a third flag in the SPS of the video sequence. 20. A computer-implemented method, comprising: Receiving a video sequence, a first flag, and a second flag; Enabling or disabling a coding mode for a video bitstream based on the first flag; and Enabling or disabling control of the coding mode at a level lower than the sequence level based on the second flag. 21. The computer-implemented method according to clause 20, wherein the level lower than the sequence level includes a slice level or an image level. 22. The computer-implemented method according to any one of clauses 20-21, further comprising: Receiving a third flag; and In response to enabling or disabling control of the encoding mode at a level lower than the sequence level, enable or disable the encoding mode for a target low-level region based on the third flag. 23. The computer-implemented method according to clause 22, further comprising: Store the third flag in the slice header of a target slice, where the target slice is the target low-level region; or Store the third flag in the picture header of a target picture, where the target picture is the target low-level region. 24. The computer-implemented method according to any one of clauses 20-23, wherein the encoding mode is at least one of the following: Bidirectional optical flow (BDOF) mode; Optical flow prediction refinement (PROF) mode; or Decoder-side motion vector refinement (DMVR) mode. 25. The computer-implemented method according to any one of clauses 20-24, further comprising: Store the first flag in the sequence parameter set (SPS) of the video bitstream. 26. The computer-implemented method according to any one of clauses 20-25, further comprising: In response to enabling or disabling the encoding mode of the video bitstream, store the second flag in the SPS of the video bitstream. 27. A computer-implemented method, comprising: Receiving a video sequence, a first flag, a second flag, and a third flag; enabling or disabling a first encoding mode for a video bitstream based on the first flag; Enabling or disabling a second encoding mode for the video bitstream based on the second flag; and Based on the third flag, enabling or disabling control of at least one of the first encoding mode or the second encoding mode at a level lower than the sequence level. 28. The computer-implemented method according to clause 27, wherein the level lower than the sequence level includes the slice level or the picture level. 29. The computer-implemented method according to any one of clauses 27-28, further comprising: Receiving a fourth flag; and In response to enabling control of the first coding mode for the video bitstream based on the first flag and enabling control of at least one of the first coding mode or the second coding mode at a level lower than the sequence level based on the third flag, enable or disable the first coding mode for a target low-level region according to the fourth flag. 30. The computer-implemented method according to any one of clauses 27-28, further comprising: Receiving a fourth flag; and In response to enabling or disabling control of at least one of the first coding mode or the second coding mode at a level lower than the sequence level, enable or disable the first coding mode and the second coding mode for a target low-level region based on the fourth flag. 31. The computer-implemented method according to any one of clauses 29-30, further comprising: Storing the fourth flag in the slice header of a target slice, where the target slice is the target low-level region; or Storing the fourth flag in the picture header of a target picture, where the target picture is the target low-level region. 32. The computer-implemented method according to clause 31, further comprising: Receiving a fifth flag; Enabling or disabling a third coding mode for the video bitstream based on the second flag; and Enabling or disabling control of the third coding mode at a level lower than the sequence level based on the fifth flag. 33. The computer-implemented method according to any one of clauses 31 and 32, further comprising: Receiving a sixth flag; and In response to enabling or disabling control of the third coding mode at a level lower than the sequence level, enable or disable the third coding mode for the target low-level region based on the sixth flag. 34. The computer-implemented method according to any one of clauses 27-33, wherein the first coding mode, the second coding mode, and the third coding mode are respectively a bi-directional optical flow (BDOF) mode, a predictive optical flow refinement (PROF) mode, and a decoder-side motion vector refinement (DMVR) mode. 35. The computer-implemented method according to any one of clauses 27-33, wherein the first coding mode and the second coding mode are two different coding modes selected from: Bi-directional optical flow (BDOF) mode; Predictive optical flow refinement (PROF) mode; and Decoder-side Motion Vector Refinement (DMVR) mode. 36. The computer-implemented method according to any one of clauses 27-35, wherein the first coding mode and the second coding mode are respectively a Bi-Directional Optical Flow (BDOF) mode and a Prediction Refinement of Optical Flow (PROF) mode. 37. The computer-implemented method according to any one of clauses 27-36, further comprising: Storing the first flag and the second flag in a Sequence Parameter Set (SPS) of the video bitstream. 38. The computer-implemented method according to any one of clauses 27-37, further comprising: In response to enabling at least one of the first coding mode or the second coding mode for the video bitstream, storing the third flag in the SPS of the video sequence. 39. A non-transitory computer-readable medium storing a set of instructions executable by at least one processor of a device to cause the device to perform a method comprising: Receiving a bitstream of video data; Enabling or disabling a coding mode for a video sequence based on a first flag in the bitstream; and Determining to enable or disable control of the coding mode at a level below the sequence level based on a second flag in the bitstream. 40. The non-transitory computer-readable medium according to clause 39, wherein the level below the sequence level includes a slice level or a picture level. 41. The non-transitory computer-readable medium according to any one of clauses 39-40, wherein the set of instructions executable by the at least one processor of the device causes the device to further perform: In response to enabling control of the coding mode at the level below the sequence level, enabling or disabling a coding mode for a target low-level region based on a third flag in the bitstream. 42. The non-transitory computer-readable medium according to clause 41, wherein the set of instructions executable by the at least one processor of the device causes the device to further perform: Detecting the third flag in a slice header of a target slice, wherein the target slice is the target low-level region; or Detecting the third flag in a picture header of a target picture, wherein the target picture is the target low-level region. 43. The non-transitory computer-readable medium according to any one of clauses 39-42, wherein the coding mode is at least one of the following: Bi-Directional Optical Flow (BDOF) mode; Optical flow prediction correction (PROF) mode; or Decoder-side motion vector correction (DMVR) mode. 44. The non-transitory computer-readable medium according to any one of clauses 39-43, wherein the instruction set executable by the at least one processor of the device causes the device to further perform: Detect the first flag in the sequence parameter set (SPS) of the video sequence. 45. The non-transitory computer-readable medium according to any one of clauses 39-44, wherein the instruction set executable by the at least one processor of the device causes the device to further perform: In response to enabling the encoding mode for the video sequence, detect the second flag in the SPS of the video sequence. 46. A non-transitory computer-readable medium storing an instruction set executable by at least one processor of a device to cause the device to perform a method, the method comprising: Receive a bitstream of video data; Based on a first flag in the bitstream, enable or disable a first encoding mode for a video sequence; Based on a second flag in the bitstream, enable or disable a second encoding mode for the video sequence; and Based on a third flag in the bitstream, determine whether to enable control of at least one of the first encoding mode or the second encoding mode at a level lower than the sequence level. 47. The non-transitory computer-readable medium according to clause 46, wherein the level lower than the sequence level includes a slice level or a picture level. 48. The non-transitory computer-readable medium according to any one of clauses 46-47, wherein the instruction set executable by the at least one processor of the device causes the device to further perform: In response to enabling control of at least one of the first encoding mode or the second encoding mode at the level lower than the sequence level, based on a fourth flag in the bitstream, enable or disable the first encoding mode and the second encoding mode for a target low-level region. 49. The non-transitory computer-readable medium according to clause 48, wherein the instruction set executable by the at least one processor of the device causes the device to further perform: Detect the fourth flag in the slice header of a target slice, wherein the target slice is the target low-level region; or Detect the fourth flag in the image header of the target image, where the target image is the target low-level region. 50. The non-transitory computer-readable medium according to any one of clauses 48 - 49, wherein the set of instructions executable by the at least one processor of the device causes the device to further perform: Enable or disable a third coding mode for the video sequence based on the second flag in the bitstream; and Determine whether to enable control of the third coding mode at a level lower than the sequence level based on a fifth flag in the bitstream. 51. The non-transitory computer-readable medium according to clause 50, wherein the set of instructions executable by the at least one processor of the device causes the device to further perform: In response to enabling the third coding mode at a level lower than the sequence level, enable or disable the third coding mode for the target low-level region based on a sixth flag in the bitstream. 52. The non-transitory computer-readable medium according to any one of clauses 50 and 51, wherein the first coding mode, the second coding mode, and the third coding mode are respectively a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode, and a decoder-side motion vector correction (DMVR) mode. 53. The non-transitory computer-readable medium according to any one of clauses 46 - 52, wherein the set of instructions executable by the at least one processor of the device causes the device to further perform: In response to the first flag indicating enabling the first coding mode for the video sequence and the third flag indicating enabling control of at least one of the first coding mode or the second coding mode at a level lower than the sequence level, enable or disable the first coding mode for a target lower-level region based on a fourth flag in the bitstream. 54. The non-transitory computer-readable medium according to any one of clauses 46 - 53, wherein the first coding mode and the second coding mode are two different coding modes selected from: Bidirectional optical flow (BDOF) mode; Optical flow prediction correction (PROF) mode; and Decoder-side motion vector correction (DMVR) mode. 55. The non-transitory computer-readable medium according to any one of clauses 46 - 54, wherein the first coding mode and the second coding mode are respectively a bidirectional optical flow (BDOF) mode and an optical flow prediction correction (PROF) mode. 56. The non-transitory computer-readable medium according to any one of clauses 46-55, wherein the instruction set executable by the at least one processor of the device causes the device to further perform: Detect a first flag and a second flag in a sequence parameter set (SPS) of the video sequence. 57. The non-transitory computer-readable medium according to any one of clauses 46-56, wherein the instruction set executable by the at least one processor of the device causes the device to further perform: In response to enabling at least one of the first coding mode or the second coding mode for the video sequence, detect the third flag in the SPS of the video sequence. 58. A non-transitory computer-readable medium storing an instruction set executable by at least one processor of a device to cause the device to perform a method, the method comprising: Receive a video sequence, a first flag, and a second flag; Enable or disable a coding mode for a video bitstream based on the first flag; and Enable or disable control of the coding mode at a level lower than the sequence level based on the second flag. 59. The non-transitory computer-readable medium according to clause 58, wherein the level lower than the sequence level includes a slice level or a picture level. 60. The non-transitory computer-readable medium according to any one of clauses 58-59, wherein the instruction set executable by at least one processor of the device causes the device to further perform: Receive a third flag; and In response to enabling or disabling control of the coding mode at the level lower than the sequence level, enable or disable a coding mode for a target low-level region based on the third flag. 61. The non-transitory computer-readable medium according to clause 60, wherein the instruction set executable by the at least one processor of the device causes the device to further perform: Store the third flag in a slice header of a target slice, where the target slice is the target low-level region; or Store the third flag in a picture header of a target picture, where the target picture is the target low-level region. 62. The non-transitory computer-readable medium according to any one of clauses 58-61, wherein the coding mode is at least one of the following: Bidirectional optical flow (BDOF) mode; Optical flow prediction correction (PROF) mode; or Decoder-side Motion Vector Refinement (DMVR) mode. 63. The non-transitory computer-readable medium according to any one of clauses 58 - 62, wherein the set of instructions executable by the at least one processor of the device causes the device to further perform: Store the first flag in the Sequence Parameter Set (SPS) of the video bitstream. 64. The non-transitory computer-readable medium according to any one of clauses 58 - 63, wherein the set of instructions executable by the at least one processor of the device causes the device to further perform: In response to enabling or disabling an encoding mode for the video bitstream, store the second flag in the SPS of the video bitstream. 65. A non-transitory computer-readable medium storing a set of instructions executable by at least one processor of a device to cause the device to perform a method, the method comprising: Receiving a video sequence, a first flag, a second flag, and a third flag; enabling or disabling a first encoding mode for the video bitstream based on the first flag; Enabling or disabling a second encoding mode for the video bitstream based on the second flag; and Based on the third flag, enabling or disabling control of at least one of the first encoding mode or the second encoding mode at a level lower than the sequence level. 66. The non-transitory computer-readable medium according to clause 65, wherein the level lower than the sequence level includes the slice level or the picture level. 67. The non-transitory computer-readable medium according to any one of clauses 65 - 66, wherein the set of instructions executable by the at least one processor of the device causes the device to further perform: Receiving a fourth flag; and In response to enabling control of the first encoding mode for the video bitstream based on the first flag and enabling control of at least one of the first encoding mode or the second encoding mode at the level lower than the sequence level based on the third flag, enabling or disabling the first encoding mode for a target low-level region based on the fourth flag. 68. The non-transitory computer-readable medium according to any one of clauses 65 - 66, wherein the set of instructions executable by the at least one processor of the device causes the device to further perform: Receiving a fourth flag; and In response to enabling or disabling control of at least one of the first coding mode or the second coding mode at the level below the sequence level, enable or disable the first coding mode and the second coding mode for a target low-level region based on the fourth flag. 69. The non-transitory computer-readable medium according to any one of clauses 67 - 68, wherein the instruction set executable by the at least one processor of the device causes the device to further perform: Store the fourth flag in the slice header of a target slice, where the target slice is the target low-level region; or Store the fourth flag in the image header of a target image, where the target image is the target low-level region. 70. The non-transitory computer-readable medium according to clause 69, wherein the instruction set executable by the at least one processor of the device causes the device to further perform: Receive a fifth flag; Enable or disable a third coding mode for the video bitstream based on the second flag; and Enable or disable control of the third coding mode at the level below the sequence level based on the fifth flag. 71. The non-transitory computer-readable medium according to any one of clauses 69 and 70, wherein the instruction set executable by the at least one processor of the device causes the device to further perform: Receive a sixth flag; and In response to enabling or disabling control of the third coding mode at the level below the sequence level, enable or disable the third coding mode for the target low-level region based on the sixth flag. 72. The non-transitory computer-readable medium according to any one of clauses 65 - 71, wherein the first coding mode, the second coding mode, and the third coding mode are respectively a bidirectional optical flow (BDOF) mode, a predictive optical flow correction (PROF) mode, and a decoder-side motion vector refinement (DMVR) mode. 73. The non-transitory computer-readable medium according to any one of clauses 65 - 72, wherein the first coding mode and the second coding mode are two different coding modes selected from: Bidirectional optical flow (BDOF) mode; Predictive optical flow correction (PROF) mode; and Decoder-side motion vector refinement (DMVR) mode. 74. The non-transitory computer-readable medium according to any one of clauses 65-73, wherein the first coding mode and the second coding mode are a bi-directional optical flow (BDOF) mode and a predicted optical flow refinement (PROF) mode, respectively. 75. The non-transitory computer-readable medium according to any one of clauses 65-74, wherein the instruction set executable by the at least one processor of the device further causes the device to: Store the first flag and the second flag in a sequence parameter set (SPS) of the video bitstream. 76. The non-transitory computer-readable medium according to any one of clauses 65-75, wherein the instruction set executable by the at least one processor of the device further causes the device to: In response to enabling at least one of the first coding mode or the second coding mode for the video bitstream, store the third flag in the SPS of the video sequence. 77. A device, comprising: A memory configured to store an instruction set; and One or more processors communicatively coupled to the memory and configured to execute the instruction set to cause the device to: Receive a bitstream of video data; Enable or disable a coding mode for a video sequence based on a first flag in the bitstream; and Based on a second flag in the bitstream, determine to enable or disable control of the coding mode at a level lower than the sequence level. 78. The device according to clause 77, wherein the level lower than the sequence level includes a slice level or a picture level. 79. The device according to any one of clauses 77-78, wherein the one or more processors are further configured to execute the instruction set to cause the device to: In response to enabling control of the coding mode at the level lower than the sequence level, enable or disable a coding mode for a target low-level region based on a third flag in the bitstream. 80. The device according to clause 79, wherein the one or more processors are further configured to execute the instruction set to cause the device to: Detect the third flag in a slice header of a target slice, wherein the target slice is the target low-level region; or Detect the third flag in a picture header of a target picture, wherein the target picture is the target low-level region. 81. The apparatus according to any one of clauses 77 - 80, wherein the coding mode is at least one of the following: Bidirectional optical flow (BDOF) mode; Optical flow prediction correction (PROF) mode; or Decoder - side motion vector correction (DMVR) mode. 82. The apparatus according to any one of clauses 77 - 81, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: Detect the first flag in the sequence parameter set (SPS) of the video sequence. 83. The apparatus according to any one of clauses 77 - 82, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: In response to enabling the coding mode for the video sequence, detect the second flag in the SPS of the video sequence. 84. An apparatus, comprising: A memory configured to store an instruction set; and One or more processors communicatively coupled to the memory and configured to execute the instruction set to cause the apparatus to: Receive a bitstream of video data; Enable or disable a first coding mode for a video sequence based on a first flag in the bitstream; Enable or disable a second coding mode for the video sequence based on a second flag in the bitstream; and Based on a third flag in the bitstream, determine whether to enable control of at least one of the first coding mode or the second coding mode at a level lower than the sequence level. 85. The apparatus according to clause 84, wherein the level lower than the sequence level includes a slice level or a picture level. 86. The apparatus according to any one of clauses 84 - 85, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: In response to enabling control of at least one of the first coding mode or the second coding mode at the level lower than the sequence level, enable or disable the first coding mode and the second coding mode for a target low - level region based on a fourth flag in the bitstream. 87. The apparatus according to clause 86, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: Detect the fourth flag in the slice header of the target slice, where the target slice is the target low-level region; or Detect the fourth flag in the picture header of the target picture, where the target picture is the target low-level region. 88. The apparatus according to any one of clauses 86-87, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: Based on the second flag in the bitstream, enable or disable a third coding mode for a video sequence; and Based on a fifth flag in the bitstream, determine whether to enable control of the third coding mode at a level lower than the sequence level. 89. The apparatus according to clause 88, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: In response to enabling the third coding mode at a level lower than the sequence level, based on a sixth flag in the bitstream, enable or disable the third coding mode for the target low-level region. 90. The apparatus according to any one of clauses 88 and 89, wherein the first coding mode, the second coding mode, and the third coding mode are respectively a bidirectional optical flow (BDOF) mode, a predictive optical flow refinement (PROF) mode, and a decoder-side motion vector refinement (DMVR) mode. 91. The apparatus according to any one of clauses 84-90, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: In response to the first flag indicating enabling the first coding mode for the video sequence and the third flag indicating enabling control of at least one of the first coding mode or the second coding mode at a level lower than the sequence level, based on a fourth flag in the bitstream, enable or disable the first coding mode for a target lower-level region. 92. The apparatus according to any one of clauses 84-91, wherein the first coding mode and the second coding mode are two different coding modes selected from: Bidirectional optical flow (BDOF) mode; Predictive optical flow refinement (PROF) mode; and Decoder-side motion vector refinement (DMVR) mode. 93. The apparatus according to any one of clauses 84-92, wherein the first coding mode and the second coding mode are respectively a bidirectional optical flow (BDOF) mode and a predictive optical flow refinement (PROF) mode. 94. The apparatus according to any one of clauses 84 - 93, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: Detect a first flag and a second flag in a sequence parameter set (SPS) of the video sequence. 95. The apparatus according to any one of clauses 84 - 94, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: In response to enabling at least one of a first encoding mode or a second encoding mode for the video sequence, detect the third flag in the SPS of the video sequence. 96. An apparatus, comprising: A memory configured to store an instruction set; and One or more processors communicatively coupled to the memory and configured to execute the instruction set to cause the apparatus to: Receive a video sequence, a first flag, and a second flag; Enable or disable an encoding mode for a video bitstream based on the first flag; and Enable or disable control of the encoding mode at a level lower than the sequence level based on the second flag. 97. The apparatus according to clause 96, wherein the level lower than the sequence level includes a slice level or a picture level. 98. The apparatus according to any one of clauses 96 - 97, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: Receive a third flag; and In response to enabling or disabling control of the encoding mode at the level lower than the sequence level, enable or disable an encoding mode for a target low - level region based on the third flag. 99. The apparatus according to clause 98, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: Store the third flag in a slice header of a target slice, where the target slice is the target low - level region; or Store the third flag in a picture header of a target picture, where the target picture is the target low - level region. 100. The apparatus according to any one of clauses 96 - 99, wherein the encoding mode is at least one of the following: Bidirectional optical flow (BDOF) mode; Optical flow prediction correction (PROF) mode; or Decoder-side Motion Vector Refinement (DMVR) mode. 101. The apparatus according to any one of clauses 96 - 100, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: Store the first flag in a Sequence Parameter Set (SPS) of the video bitstream. 102. The apparatus according to any one of clauses 96 - 101, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: In response to enabling or disabling an encoding mode for the video bitstream, store the second flag in the SPS of the video bitstream. 103. An apparatus, comprising: A memory configured to store an instruction set; and One or more processors communicatively coupled to the memory and configured to execute the instruction set to cause the apparatus to: Receive a video sequence, a first flag, a second flag, and a third flag; Enable or disable a first encoding mode for the video bitstream based on the first flag; Enable or disable a second encoding mode for the video bitstream based on the second flag; and Based on the third flag, enable or disable control of at least one of the first encoding mode or the second encoding mode at a level lower than the sequence level. 104. The apparatus according to clause 102, wherein the level lower than the sequence level includes a slice level or a picture level. 105. The apparatus according to any one of clauses 103 - 104, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: Receive a fourth flag; and In response to enabling control of the first encoding mode for the video bitstream based on the first flag and enabling control of at least one of the first encoding mode or the second encoding mode at the level lower than the sequence level based on the third flag, enable or disable the first encoding mode for a target low-level region based on the fourth flag. 106. The apparatus according to any one of clauses 103 - 104, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: Receive a fourth flag; and In response to enabling or disabling control of at least one of the first coding mode or the second coding mode at the level lower than the sequence level, enable or disable the first coding mode and the second coding mode for a target low-level region based on the fourth flag. 107. The apparatus according to any one of clauses 105 - 106, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: Store the fourth flag in a slice header of a target slice, where the target slice is the target low-level region; or Store the fourth flag in an image header of a target image, where the target image is the target low-level region. 108. The apparatus according to clause 107, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: Receive a fifth flag; Enable or disable a third coding mode for the video bitstream based on the second flag; and Enable or disable control of the third coding mode at the level lower than the sequence level based on the fifth flag. 109. The apparatus according to any one of clauses 107 and 108, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: Receive a sixth flag; and In response to enabling or disabling control of the third coding mode at the level lower than the sequence level, enable or disable the third coding mode for the target low-level region based on the sixth flag. 110. The apparatus according to any one of clauses 103 - 109, wherein the first coding mode, the second coding mode, and the third coding mode are a bidirectional optical flow (BDOF) mode, a predictive optical flow refinement (PROF) mode, and a decoder-side motion vector refinement (DMVR) mode, respectively. 111. The apparatus according to any one of clauses 103 - 110, wherein the first coding mode and the second coding mode are two different coding modes selected from: Bidirectional optical flow (BDOF) mode; Predictive optical flow refinement (PROF) mode; and Decoder-side motion vector refinement (DMVR) mode. 112. The apparatus according to any one of clauses 103 - 110, wherein the first coding mode and the second coding mode are a bidirectional optical flow (BDOF) mode and a predictive optical flow refinement (PROF) mode, respectively. 113. The apparatus according to any one of clauses 103 - 112, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: Store the first flag and the second flag in a sequence parameter set (SPS) of the video bitstream. 114. The apparatus according to any one of clauses 103 - 113, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to: In response to enabling at least one of the first coding mode or the second coding mode for the video bitstream, store the third flag in the SPS of the video sequence.
[0166] It should be noted that relational terms such as "first" and "second" herein are only used to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. In addition, words such as "comprising", "having", "including", and "containing" and other similar forms are intended to represent the same meaning and are open-ended, because one or more items after any of these words are not intended to exhaustively list such items or items, or be limited to the listed items or items.
[0167] As used herein, unless otherwise expressly stated, the term "or" encompasses all possible combinations, unless infeasible. For example, if it is stated that a component may include A or B, then unless otherwise expressly stated or infeasible, the component may include A, or B, or A and B. As a second example, if it is stated that a component may include A, B, or C, then unless otherwise expressly stated or infeasible, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.
[0168] It can be understood that the above embodiments can be implemented by hardware, software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above computer-readable medium. The software can execute the disclosed method when executed by a processor. The computing units and other functional units described in the present invention can be implemented by hardware, or by software, or by a combination of hardware and software. Those of ordinary skill in the art can also understand that multiple of the above modules / units can be combined into one module / unit, and each of the above modules / units can be further divided into multiple sub-modules / sub-units.
[0169] In the foregoing specification, the embodiments have been described with reference to numerous specific details which may vary with implementation. Certain modifications and changes may be made to the described embodiments. Other embodiments will be apparent to those skilled in the art in view of the specification and practice of the invention disclosed herein. The specification and examples are to be considered as illustrative only, and the true scope and spirit of the invention are indicated by the appended claims. The order of steps shown in the figures is also intended for illustrative purposes only and is not intended to be limiting to any particular order of steps. Accordingly, those skilled in the art will appreciate that these steps may be performed in a different order when implementing the same method.
[0170] In the drawings and the specification, exemplary embodiments have been disclosed. However, many variations and modifications can be made to these embodiments. Accordingly, although specific terms are used, they are for general and descriptive purposes only and not for purposes of limitation.
Claims
1. A computer-implemented decoding method, comprising: Receiving a bitstream associated with a video sequence; Decoding a first flag and a second flag in a sequence parameter set (SPS) of the bitstream, wherein the first flag indicates enabling or disabling at least one of multiple coding modes at a sequence level; the second flag indicates whether a coding mode enabled at the sequence level is controlled at a lower level; Based on a value of the second flag, determining whether a third flag exists in the bitstream; wherein the third flag indicates whether the coding mode enabled at the sequence level is enabled or disabled at a lower level lower than the sequence level; Multiple ones of the coding modes include an optical flow prediction correction mode; When the third flag exists in the bitstream, decoding the bitstream based on the value of the third flag.
2. The computer-implemented decoding method according to claim 1, wherein the lower level lower than the sequence level is a slice level or a picture level.
3. The computer-implemented decoding method according to claim 1, the method further comprising: In response to the second flag having a first value, determining that the third flag exists in a slice header or a picture header of the bitstream.
4. The computer-implemented decoding method according to claim 3, wherein the first value is 1.
5. A computer-implemented encoding method, comprising: Encoding a first flag and a second flag in a sequence parameter set (SPS) of a bitstream associated with a video sequence, wherein the first flag indicates enabling or disabling at least one of multiple coding modes at a sequence level; the second flag indicates whether a coding mode enabled at the sequence level is controlled at a lower level; Based on a value of the second flag, determining whether to signal a third flag in the bitstream; The third flag indicates whether the coding mode enabled at the sequence level is enabled or disabled at a lower level lower than the sequence level; Multiple ones of the coding modes include an optical flow prediction correction mode; When it is determined to signal the third flag in the bitstream, encoding the bitstream based on the value of the third flag.
6. The computer-implemented encoding method according to claim 5, wherein, The lower level lower than the sequence level is a slice level or a picture level.
7. The computer-implemented encoding method according to claim 5, the method further comprising: In response to the second flag having a first value, encoding based on the third flag in a slice header or a picture header of the bitstream.
8. The computer-implemented encoding method according to claim 7, wherein the first value is 1.
9. A non-transitory computer-readable storage medium storing a bitstream associated with a video sequence, the bitstream being generated by a method executable by a video processing device, the method comprising: Encoding a first flag and a second flag in a sequence parameter set (SPS) of a bitstream associated with a video sequence, wherein the first flag indicates enabling or disabling at least one of multiple coding modes at a sequence level; the second flag indicates whether a coding mode enabled at the sequence level is controlled at a lower level; Based on a value of the second flag, determining whether to signal a third flag in the bitstream; The third flag indicates whether an encoding mode enabled at the sequence level is enabled or disabled at a lower level below the sequence level; The plurality of encoding modes includes an optical flow prediction correction mode; When it is determined to signal the third flag in the bitstream, the bitstream is encoded based on the value of the third flag.
10. The non-transitory computer-readable storage medium according to claim 9, wherein the lower level below the sequence level is a slice level or a picture level.
11. The non-transitory computer-readable storage medium according to claim 9, the method further comprising: In response to the second flag having a first value, encoding in a slice header or a picture header of the bitstream based on the third flag.
12. The non-transitory computer-readable storage medium according to claim 11, wherein, The first value is 1.
Citation Information
Patent Citations
Features of intra block copy prediction mode for video and image coding and decoding
CN105765974A
Method and apparatus for intra prediction coding with boundary filtering control
CN105981385A
Methods of palette coding with inter-prediction in video coding
CN107409227A
Method and System for Lossless Coding Mode in Video Coding
US20130077696A1
Bypass bins for reference index coding in video coding
US20130272377A1