Computer-implemented coding method and non-transitory computer-readable storage medium

By receiving video data bitstreams and controlling the encoding mode of video sequences based on multiple flags, the problem of improving encoding efficiency in existing technologies is solved, and the compression efficiency and subjective quality of video encoding are improved under fine control below the sequence level.

CN120281923BActive Publication Date: 2026-07-21HFI INNOVATION INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HFI INNOVATION INC
Filing Date
2020-08-20
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing video coding standards still have room for improvement in coding efficiency among high-efficiency video coding technologies, especially when using half the bandwidth to achieve the same subjective quality as HEVC/H.265, it is difficult to flexibly control the coding mode to optimize compression efficiency.

Method used

By receiving video data bitstreams and controlling the encoding mode of video sequences based on multiple flags, different levels of encoding modes can be enabled or disabled, including sequence-level and sub-sequence-level control, dynamically adjusting the encoding process to optimize compression efficiency.

Benefits of technology

It achieves improved compression efficiency and subjective quality of video encoding under fine control below the sequence level, adapting to the needs of different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281923B_ABST
    Figure CN120281923B_ABST
Patent Text Reader

Abstract

The present disclosure provides methods and apparatuses for controlling coding modes for video data. The methods and apparatuses include receiving a bitstream of video data; enabling or disabling a coding mode for a video sequence based on a first flag in the bitstream; and determining whether to enable or disable control of the coding mode at a level lower than a sequence level based on a second flag in the bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications

[0002] This disclosure claims priority to U.S. Provisional Application No. 62 / 899,169, filed September 12, 2019, the entire contents of which are incorporated herein by reference. Background Technology

[0003] Video is a set of still images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and then decompressed before display. The compression process is usually called encoding, and the decompression process is usually called decoding. There are various video coding formats that use standardized video coding techniques, the most common being prediction, transform, quantization, entropy coding, and loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard and the Universal Video Coding (VVC / H.266) standard (AVS standard), specify concrete video coding formats and are developed by standardization organizations. As more and more advanced video coding technologies are adopted in video standards, the coding efficiency of new video coding standards is also increasing. Summary of the Invention

[0004] This invention provides a method and apparatus for controlling encoding modes for video data. In one example embodiment, a method includes: receiving a bitstream of video data; enabling or disabling an encoding mode for a video sequence based on a first flag in the bitstream; and determining, based on a second flag in the bitstream, to enable or disable control of the encoding mode at a level below the sequence level.

[0005] In another example embodiment, a method includes: receiving a bitstream of video data; enabling or disabling a first encoding mode for a video sequence based on a first flag in the bitstream; enabling or disabling a second encoding mode for the video sequence based on a second flag in the bitstream; and determining, based on a third flag in the bitstream, whether to enable control over at least one of the first encoding mode or the second encoding mode at a level below the sequence level.

[0006] In another example embodiment, a method includes: receiving a video sequence, a first flag, and a second flag; enabling or disabling an encoding mode for a video bitstream based on the first flag; and enabling or disabling control of the encoding mode at a level below the sequence level based on the second flag.

[0007] In another example embodiment, a method includes: receiving a video sequence, a first flag, a second flag, and a third flag; enabling or disabling a first encoding mode for a video bitstream based on the first flag; enabling or disabling a second encoding mode for the video bitstream based on the second flag; and enabling or disabling control over at least one of the first encoding mode or the second encoding mode at a level below the sequence level based on the third flag.

[0008] In another example embodiment, a non-transitory computer-readable medium stores a set of instructions executable by at least one processor of a device to cause the device to perform a method comprising: receiving a video data bitstream; enabling or disabling an encoding mode for a video sequence based on a first flag in the bitstream; and determining, based on a second flag in the bitstream, whether to enable or disable control of the encoding mode at a level below the sequence level.

[0009] In another example embodiment, a non-transitory computer-readable medium stores a set of instructions executable by at least one processor of a device to cause the device to perform a method comprising: receiving a video data bitstream; enabling or disabling a first encoding mode for a video sequence based on a first flag in the bitstream; enabling or disabling a second encoding mode for the video sequence based on a second flag in the bitstream; and determining, based on a third flag in the bitstream, whether to enable control over at least one of the first encoding mode or the second encoding mode at a level below the sequence level.

[0010] In another embodiment, an apparatus includes a memory configured to store an instruction set and one or more processors communicatively coupled to the memory, the one or more processors being configured to execute the instruction set to cause the apparatus to: receive a bitstream of video data; enable or disable an encoding mode for a video sequence based on a first flag in the bitstream; and determine whether to enable or disable control of the encoding mode at a level below the sequence level based on a second flag in the bitstream.

[0011] In another embodiment, an apparatus includes a memory configured to store an instruction set and one or more processors communicatively coupled to the memory, the one or more processors being configured to execute the instruction set to cause the apparatus to: receive a bitstream of video data; enable or disable a first encoding mode for a video sequence based on a first flag in the bitstream; enable or disable a second encoding mode for a video sequence based on a second flag in the bitstream; and determine, based on a third flag in the bitstream, whether to enable control over at least one of the first encoding mode or the second encoding mode at a level below the sequence level. Attached Figure Description

[0012] Embodiments and aspects of this disclosure are illustrated in the following detailed description and accompanying drawings. The various features shown in the figures are not drawn to scale.

[0013] Figure 1 This is a schematic diagram illustrating the structure of an example video sequence according to some embodiments of the present disclosure.

[0014] Figure 2A A schematic diagram of an example encoding process for a hybrid video encoding system consistent with embodiments of this disclosure is shown.

[0015] Figure 2B A schematic diagram of another example encoding process of a hybrid video coding system consistent with embodiments of this disclosure is shown.

[0016] Figure 3A A schematic diagram of an example decoding process for a hybrid video coding system consistent with embodiments of this disclosure is shown.

[0017] Figure 3B A schematic diagram of another example decoding process of a hybrid video coding system consistent with embodiments of this disclosure is shown.

[0018] Figure 4 A block diagram of an example apparatus for encoding or decoding video according to some embodiments of the present disclosure is shown.

[0019] Figure 5 This is a schematic diagram illustrating an example process of decoder-side motion vector refinement (DMVR) according to some embodiments of the present disclosure.

[0020] Figure 6 This is a schematic diagram illustrating an example DMVR search process according to some embodiments of the present disclosure.

[0021] Figure 7 This is a schematic diagram illustrating an example mode for DMVR integer luminance sample search according to some embodiments of the present disclosure.

[0022] Figure 8 This is a schematic diagram illustrating another example mode of the integer sample offset search stage in integer luminance sample search for DMVR according to some embodiments of the present disclosure.

[0023] Figure 9 This is a schematic diagram illustrating an example mode for estimating the error surface of DMVR parameters according to some embodiments of the present disclosure.

[0024] Figure 10 This is a schematic diagram of an example of an extended coding unit (CU) region used in bidirectional optical flow (BDOF) according to some embodiments of the present disclosure.

[0025] Figure 11This is a schematic diagram illustrating examples of sub-block-based affine motion and sample-based affine motion according to some embodiments of this disclosure.

[0026] Figure 12 Table 1 is shown, illustrating example syntax structures for the Sequence Parameter Set (SPS) of control flags for DMVR and BDOF according to some embodiments of the present disclosure.

[0027] Figure 13 Table 2 is shown, illustrating example syntax structures for slice headers of control flags for DMVR and BDOF according to some embodiments of this disclosure.

[0028] Figure 14A Table 3A is shown, illustrating example syntax structures of SPS for implementing slice-level control flags for DMVR, BDOF, and optical flow prediction correction (PROF) according to some embodiments of the present disclosure.

[0029] Figure 14B Table 3B is shown, illustrating example syntax structures for SPS implementations of image-level control flags for DMVR, BDOF, and PROF according to some embodiments of this disclosure.

[0030] Figure 15A Table 4A is shown, which illustrates example syntax structures for the control flags of the silce header for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.

[0031] Figure 15B Table 4B is shown, illustrating example syntax structures for image headers used for control flags of DMVR, BDOF, and PROF according to some embodiments of this disclosure.

[0032] Figure 16 Table 5 is shown, illustrating example syntax structures for SPS implementations of individual sequence-level control flags for DMVR, BDOF, and PROF according to some embodiments of this disclosure.

[0033] Figure 17 Table 6 is shown, illustrating example syntax structures for slice heads implementing joint control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.

[0034] Figure 18 Table 7 is shown, illustrating example syntax structures for slice heads implementing individual control flags for DMVR, BDOF, and PROF according to some embodiments of this disclosure.

[0035] Figure 19 Table 8 is shown, illustrating example syntax structures for slice heads implementing mixed control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.

[0036] Figure 20 Table 9 is shown, illustrating example syntax structures for implementing SPS for mixed sequence-level control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.

[0037] Figure 21 Table 10 is shown, which illustrates another example syntax structure for a slice head of mixed control flags for DMVR, BDOF, and PROF, according to some embodiments of the present disclosure.

[0038] Figure 22 Table 11 is shown, which illustrates another example syntax structure for the slice head of individual control flags for DMVR, BDOF, and PROF, according to some embodiments of this disclosure.

[0039] Figure 23 A flowchart illustrating an example process for controlling a video decoding mode according to some embodiments of the present disclosure is shown.

[0040] Figure 24 A flowchart of another example process for controlling a video decoding mode according to some embodiments of the present disclosure is shown.

[0041] Figure 25 A flowchart illustrating an example process for controlling a video encoding mode according to some embodiments of the present disclosure is shown.

[0042] Figure 26 A flowchart is shown as another example process for controlling a video encoding mode according to some embodiments of the present disclosure. Detailed Implementation

[0043] Reference can now be made to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, wherein, unless otherwise stated, the same numerals in different drawings denote the same or similar elements. The embodiments set forth in the following description of the exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with aspects of the invention as described in the appended claims. Specific aspects of this disclosure are described below in more detail. In the event of any conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.

[0044] The Joint Video Experts Group (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Universal Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, VVC aims to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.

[0045] To achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET has been developing techniques beyond HEVC using the Joint Exploratory Model (JEM) reference software. With the incorporation of coding techniques into JEM, JEM achieves higher coding performance than HEVC.

[0046] The VVC standard is a recent development and continues to incorporate more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.

[0047] Video is a set of still images (or "frames") arranged chronologically to store visual information. Video capture devices (e.g., cameras) can be used to capture and store these images chronologically, and video playback devices (e.g., televisions, computers, smartphones, tablets, video players, or any end-user terminal with a display capability) can be used to display such images chronologically. Furthermore, in some applications, video capture devices can transmit captured video in real time to video playback devices (e.g., computers with monitors), such as for surveillance, conferencing, or live streaming.

[0048] To reduce the storage space and transmission bandwidth required for such applications, video can be compressed before storage and transmission, and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., a processor in a general-purpose computer) or dedicated hardware. The module used for compression is typically called an "encoder," and the module used for decompression is typically called a "decoder." Encoders and decoders can be collectively referred to as a "codec." Encoders and decoders can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, hardware implementations of encoders and decoders can include circuits such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, or any combination thereof. Software implementations of encoders and decoders can include program code, computer-executable instructions, firmware, or any suitable computer-implemented algorithm or process embedded in a computer-readable medium. Video compression and decompression can be implemented using various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, etc. In some applications, a codec can decompress video from a first encoding standard and recompress the decompressed video using a second encoding standard; in this case, the codec can be called a "transcoder."

[0049] Video encoding processes can identify and retain useful information that can be used to reconstruct an image, while ignoring unimportant information that would otherwise be lost in the reconstruction. If the ignored, unimportant information cannot be fully reconstructed, this encoding process can be called "lossy." Otherwise, it can be called "lossless." Most encoding processes are lossy, a trade-off made to reduce required storage space and transmission bandwidth.

[0050] Useful information about the encoded image (referred to as the "current image") includes changes relative to a reference image (e.g., a previously encoded and reconstructed image). These changes can include variations in pixel position, brightness, or color, with positional changes being of primary concern. The positional changes of a set of pixels representing an object can reflect the object's movement between the reference and current images.

[0051] An image encoded without referencing another image (i.e., it is its own reference image) is called an "I-image". An image encoded using a previous image as a reference image is called a "P-image". An image encoded using both a previous image and a future image as reference images (i.e., the reference is "bidirectional") is called a "B-image".

[0052] Figure 1The illustration shows the structure of an example video sequence 100 according to some embodiments of the present disclosure. The video sequence 100 may be live video or video that has been captured and archived. The video 100 may be real video, computer-generated video (e.g., computer game video), or a combination thereof (e.g., real video with augmented reality effects). The video sequence 100 may be input from a video capture device (e.g., a camera), a video archive containing previously captured video (e.g., a video file stored on a storage device), or a providing interface (e.g., a video broadcast transceiver) receiving video from a video content provider.

[0053] like Figure 1 As shown, video sequence 100 may include a series of images arranged along a time axis, including images 102, 104, 106, and 108. Images 102-106 are consecutive, and there are more images between images 106 and 108. Figure 1 In this diagram, image 102 is an I-image, and its reference image is image 102 itself. Image 104 is a P-image, and its reference image is image 102, as indicated by the arrow. Image 106 is a B-image, and its reference images are images 104 and 108, as indicated by the arrow. In some embodiments, the reference image of an image (e.g., image 104) may not immediately precede or follow the image. For example, the reference image of image 104 may be an image preceding image 102. It should be noted that the reference images of images 102-106 are merely examples, and this disclosure does not limit the embodiments of the reference images to... Figure 1 The example shown.

[0054] Typically, due to the computational complexity of encoding and decoding tasks, video codecs do not encode or decode the entire image at once. Instead, they can segment the image into basic segments and encode or decode the image segment by segment. Such basic segments are referred to in this disclosure as basic processing units (“BPUs”). For example, Figure 1Structure 110 illustrates an example structure of an image (e.g., any one of images 102-108) from video sequence 100. In structure 110, the image is divided into 4×4 basic processing units, whose boundaries are shown as dashed lines. In some embodiments, the basic processing unit may be referred to as a “macroblock” in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC), or as a “coding tree unit” (“CTU”) in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing units in the image can have different sizes, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or pixels of any shape and size. The size and shape of the basic processing units for the image can be selected based on a balance between coding efficiency and the level of detail to be preserved in the basic processing units.

[0055] A basic processing unit can be a logical unit that may include a set of different types of video data stored in computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color image may include a luminance component (Y) representing luminance information, one or more chrominance components (e.g., Cb and Cr) representing color information, and associated syntax elements, where the luminance and chrominance components may have basic processing units of the same size. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), the luminance and chrominance components may be referred to as “code tree blocks” (“CTBs”). Any operation performed on a basic processing unit can be repeated for each of its luminance and chrominance components.

[0056] Video encoding involves multiple operational stages, examples of which are shown in... Figure 2A-2B and Figures 3A-3BAs shown in the diagram. For each stage, the size of the basic processing unit may still be too large to process, and therefore it can be further divided into segments, referred to herein as "basic processing subunits". In some embodiments, the basic processing subunit may be referred to as a "block" in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC), or as a "coding unit" ("CU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunit may have the same size as the basic processing unit or a smaller size. Similar to the basic processing unit, the basic processing subunit is also a logical unit that may include a set of different types of video data (e.g., Y, Cb, Cr, and associated syntax elements) stored in computer memory (e.g., in a video frame buffer). Any operation performed on the basic processing subunit may be repeated for each of its luma and chroma components. It should be noted that this division can be performed to a further level as needed for processing. It should also be noted that different stages may use different schemes to divide the basic processing units.

[0057] For example, in the pattern decision-making stage (examples of which are in...) Figure 2B As shown in the diagram, the encoder can decide which prediction mode to use for the basic processing unit (e.g., intra-image prediction or inter-image prediction), even if the basic processing unit is too large to make this decision. The encoder can break down the basic processing unit into multiple basic processing subunits (e.g., CUs in H.265 / HEVC or H.266 / VVC) and determine the prediction type for each individual basic processing subunit.

[0058] For example, in the prediction phase (examples are in...) Figure 2A-2B As shown in the diagram, the encoder can perform prediction operations at the level of a basic processing subunit (e.g., CU). However, in some cases, the basic processing subunit may still be too large to handle. The encoder can further break down the basic processing subunit into smaller segments (e.g., referred to as "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), at which prediction operations can be performed.

[0059] For another example, in the transformation phase (the example of which is in...) Figure 2A-2B(As shown in the diagram). The encoder can perform transform operations on residual basic processing subunits (e.g., CUs). However, in some cases, the basic processing subunits may still be too large to process. The encoder can further divide the basic processing subunits into smaller segments (e.g., referred to as "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), at which point the transform operation can be performed. It should be noted that the partitioning scheme of the same basic processing subunit can differ between the prediction and transform phases. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU can have different sizes and numbers.

[0060] exist Figure 1 In structure 110, the basic processing unit 112 is further divided into 3×3 basic processing sub-units, the boundaries of which are indicated by dashed lines. Different basic processing units of the same image can be divided into basic processing sub-units with different schemes.

[0061] In some implementations, to provide parallel processing and fault tolerance for video encoding and decoding, an image can be divided into multiple regions for processing. This allows the encoding or decoding process for one region of the image to be independent of information from any other region of the image. In other words, each region of the image can be processed independently. By doing so, the codec can process different regions of the image in parallel, thereby improving encoding efficiency. Furthermore, when data in one region is corrupted during processing or lost during network transmission, the codec can correctly encode or decode other regions of the same image without relying on the corrupted or lost data, thus providing fault tolerance. In some video coding standards, images can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles." It should also be noted that different images in the video sequence 100 can have different partitioning schemes for dividing the image into multiple regions.

[0062] For example, in Figure 1 In the diagram, structure 110 is divided into three regions 114, 116, and 118, whose boundaries are shown as solid lines within structure 110. Region 114 includes four basic processing units. Each of regions 116 and 118 includes six basic processing units. It is important to note that... Figure 1 The basic processing unit, basic processing subunit, and region of structure 110 are merely examples, and this disclosure does not limit its embodiments.

[0063] Figure 2A A schematic diagram of an example encoding process 200A consistent with embodiments of this disclosure is illustrated. For example, the encoding process 200A may be performed by an encoder. Figure 2A As shown, the encoder can encode the video sequence 202 into a video bitstream 228 according to process 200A. Similar to... Figure 1 Video sequence 100 and video sequence 202 may include a set of images arranged in chronological order (referred to as "original images"). Similar to... Figure 1 In the structure 110, each raw image of the video sequence 202 can be divided into basic processing units, basic processing sub-units, or regions by the encoder for processing. In some embodiments, the encoder can perform process 200A at the basic processing unit level for each raw image of the video sequence 202. For example, the encoder can perform process 200A iteratively, wherein the encoder can encode a basic processing unit in one iteration of process 200A. In some embodiments, the encoder can perform process 200A in parallel for a region (e.g., region 114-118) of each raw image of the video sequence 202.

[0064] exist Figure 2A In this process, the encoder can provide the basic processing unit (referred to as the "raw BPU") of the original image of video sequence 202 to prediction stage 204 to generate prediction data 206 and prediction BPU 208. The encoder can subtract prediction BPU 208 from the raw BPU to generate residual BPU 210. The encoder can provide residual BPU 210 to transform stage 212 and quantization stage 214 to generate quantization transform coefficients 216. The encoder can provide prediction data 206 and quantization transform coefficients 216 to binary encoding stage 226 to generate video bitstream 228. Components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 can be referred to as the "forward path". During process 200A, after quantization stage 214, the encoder can provide quantization transform coefficients 216 to inverse quantization stage 218 and inverse transform stage 220 to generate reconstructed residual BPU 222. The encoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224, which is used in the prediction stage 204 of the next iteration of process 200A. Components 218, 220, 222, and 224 of process 200A can be referred to as the "reconstruction path". The reconstruction path can be used to ensure that both the encoder and decoder use the same reference data for prediction.

[0065] The encoder can iteratively execute process 200A to encode each raw BPU (in the forward path) of the original image and generate a prediction reference 224 for the next raw BPU (in the reconstruction path) for encoding the original image. After encoding all raw BPUs of the original image, the encoder can continue to encode the next image in the video sequence 202.

[0066] Referring to process 200A, the encoder may receive a video sequence 202 generated by a video acquisition device (e.g., a camera). The term "receive" as used herein may refer to any action of receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or otherwise inputting data.

[0067] In prediction phase 204, during the current iteration, the encoder can receive the original BPU and prediction reference 224, and perform prediction operations to generate prediction data 206 and prediction BPU 208. Prediction reference 224 can be generated from the reconstruction path of a previous iteration of process 200A. The purpose of prediction phase 204 is to reduce information redundancy by extracting prediction data 206, which can be used to reconstruct the original BPU into prediction BPU 208 from prediction data 206 and prediction reference 224.

[0068] Ideally, the predicted BPU 208 should be identical to the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 will typically differ slightly from the original BPU. To record this difference, after generating the predicted BPU 208, the encoder can subtract it from the original BPU to generate the residual BPU 210. For example, the encoder can subtract the value of the corresponding pixel in the predicted BPU 208 (e.g., grayscale or RGB value) from the pixel values ​​of the original BPU 208. Each pixel in the residual BPU 210 can have a residual value, which is the result of subtracting the corresponding pixel from the original BPU and the predicted BPU 208. Compared to the original BPU, the predicted data 206 and the residual BPU 210 can have fewer bits, but they can be used to reconstruct the original BPU without significantly degrading quality. Therefore, the original BPU is compressed.

[0069] To further compress the residual BPU 210, in the transform phase 212, the encoder can reduce its spatial redundancy by decomposing the residual BPU 210 into a set of two-dimensional “fundamental patterns,” each fundamental pattern being associated with a “transform coefficient.” The fundamental patterns can have the same size (e.g., the size of the residual BPU 210). Each fundamental pattern can represent a frequency component of the residual BPU 210 (e.g., the frequency of brightness variation). No fundamental pattern can be reproduced from any combination of any other fundamental patterns (e.g., a linear combination). In other words, the decomposition decomposes the variation of the residual BPU 210 into the frequency domain. This decomposition is analogous to the discrete Fourier transform of a function, where the fundamental patterns are analogous to the fundamental functions of the discrete Fourier transform (e.g., trigonometric functions), and the transform coefficients are analogous to the coefficients associated with the fundamental functions.

[0070] Different transform algorithms can use different base modes. Various transform algorithms, such as discrete cosine transform, discrete sine transform, etc., can be used in the transform stage 212. The transform in the transform stage 212 is reversible. That is, the encoder can recover the residual BPU 210 through the inverse operation of the transform (called the "inverse transform"). For example, to recover the pixels of the residual BPU 210, the inverse transform can be to multiply the values ​​of the corresponding pixels of the base mode by their respective correlation coefficients and sum the products to produce a weighted sum. For video coding standards, both the encoder and decoder can use the same transform algorithm (and therefore the same base mode). Therefore, the encoder can only record the transform coefficients, from which the decoder can reconstruct the residual BPU 210 without receiving the base mode from the encoder. Compared to the residual BPU 210, the transform coefficients can have fewer bits, but they can be used to reconstruct the residual BPU 210 without significantly degrading the quality. Therefore, the residual BPU 210 is further compressed.

[0071] The encoder can further compress the transform coefficients in the quantization stage 214. During the transform process, different fundamental modes can represent different frequencies of change (e.g., brightness change frequency). Since the human eye is generally better at recognizing low-frequency changes, the encoder can ignore information about high-frequency changes without causing a significant degradation in decoding quality. For example, in the quantization stage 214, the encoder can generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called the "quantization parameter") and rounding the quotient to its nearest integer. After this operation, some transform coefficients of high-frequency fundamental modes can be converted to zero, while transform coefficients of low-frequency fundamental modes can be converted to smaller integers. The encoder can ignore zero-value quantized transform coefficients 216, further compressing the transform coefficients through this operation. The quantization process is also reversible, where the quantized transform coefficients 216 can be reconstructed into transform coefficients in the inverse operation of quantization (called "inverse quantization").

[0072] Because the encoder ignores the remainder of such division during rounding operations, quantization stage 214 can be lossy. Typically, quantization stage 214 causes the greatest information loss in process 200A. The greater the information loss, the fewer bits the quantization transform coefficients 216 may require. To obtain different levels of information loss, the encoder can use different quantization parameter values ​​or any other parameter of the quantization process.

[0073] In the binary encoding stage 226, the encoder can encode the prediction data 206 and the quantized transform coefficients 216 using binary encoding techniques, such as entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder can encode other information in the binary encoding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the transform type in the transform stage 212, the parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bit rate control parameters), etc. The encoder can use the output data of the binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 can be further packaged for network transmission.

[0074] Following the reconstruction path of process 200A, in the inverse quantization stage 218, the encoder can perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. In the inverse transform stage 220, the encoder can generate a reconstruction residual BPU 222 based on the reconstructed transform coefficients. The encoder can add the reconstruction residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 that will be used in the next iteration of process 200A.

[0075] It should be noted that other variations of process 200A may be employed in the encoded video sequence 202. In some embodiments, the stages of process 200A may be performed by the encoder in different orders. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may be omitted. Figure 2A One or more stages in the process.

[0076] Figure 2B A schematic diagram of another example encoding process 200B consistent with embodiments of the present disclosure is illustrated. Process 200B can be modified from process 200A. For example, process 200B can be used by an encoder conforming to a hybrid video coding standard (e.g., H.26x series). Compared to process 200A, the forward path of process 200B additionally includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filtering stage 232 and a buffer 234.

[0077] Generally, prediction techniques can be categorized into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-image prediction or "intra-frame prediction") uses pixels from one or more coded neighboring BPUs in the same image to predict the current BPU. That is, the prediction reference 224 in spatial prediction can include neighboring BPUs. Spatial prediction can reduce the inherent spatial redundancy of an image. Temporal prediction (e.g., inter-image prediction or "inter-frame prediction") uses regions from one or more coded images to predict the current BPU. That is, the prediction reference 224 in temporal prediction can include the coded image. Temporal prediction can reduce the inherent temporal redundancy of an image.

[0078] In reference process 200B, during the forward path, the encoder performs prediction operations in spatial prediction stage 2042 and temporal prediction stage 2044. For example, in spatial prediction stage 2042, the encoder may perform intra-frame prediction. For the original BPU of the encoded image, prediction reference 224 may include one or more adjacent BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same image. The encoder can generate a predicted BPU 208 by interpolating adjacent BPUs. Interpolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder may perform interpolation at the pixel level, for example by interpolating the value of the corresponding pixel used for predicting BPU 208 for each pixel. The adjacent BPUs used for interpolation may be located in various directions relative to the original BPU, such as in the vertical direction (e.g., at the top of the original BPU), the horizontal direction (e.g., to the left of the original BPU), the diagonal direction (e.g., the lower left, lower right, upper left, or upper right of the original BPU), or any direction defined in the video coding standard used. For intra-frame prediction, prediction data 206 may include, for example, the location (e.g., coordinates) of the neighboring BPUs used, the size of the neighboring BPUs used, the interpolation parameters, and the neighboring BPUs used relative to the original orientation BPU.

[0079] In another example, during the temporal prediction phase 2044, the encoder can perform inter-frame prediction. For the original BPU of the current image, the prediction reference 224 can include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images can be encoded and reconstructed on a BPU-by-BPU basis. For example, the encoder can add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs of the same image are generated, the encoder can generate the reconstructed image as the reference image. The encoder can perform a "motion estimation" operation to search for matching regions within the range of the reference image (referred to as a "search window"). The position of the search window in the reference image can be determined based on the position of the original BPU in the current image. For example, the search window can be centered at a position in the reference image that has the same coordinates as the original BPU in the current image and can extend a predetermined distance. When the encoder identifies (e.g., by using a pixel recursive algorithm, block matching algorithm, etc.) a region in the search window that is similar to the original BPU, the encoder can determine such a region as a matching region. The matching region can have a different size than the original BPU (e.g., smaller than, equal to, larger than, or with a different shape than the original BPU). This is because the reference image and the current image are temporally separated on the time axis (e.g., as...). Figure 1 As shown in the image, the matching region can be considered to have "moved" to the original BPU's location over time. The encoder can record the direction and distance of this movement as a "motion vector." When using multiple reference images (e.g., such as...), Figure 1 In image 106, the encoder can search for matching regions and determine the associated motion vector for each reference image. In some embodiments, the encoder can assign weights to the pixel values ​​of the matching regions of each matching reference image.

[0080] Motion estimation can be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference images, the weights associated with the reference images, etc.

[0081] To generate the predicted BPU 208, the encoder can perform a "motion compensation" operation. Motion compensation can be used to reconstruct the predicted BPU 208 based on the predicted data 206 (e.g., motion vectors) and the predicted reference 224. For example, the encoder can move a matching region of the reference image according to the motion vectors, where the encoder can predict the original BPU of the current image. When using multiple reference images (e.g., such as...), Figure 1In image 106), the encoder can move the matching region of the reference image based on the respective motion vectors and the average pixel value of the matching region. In some embodiments, if the encoder has already assigned weights to the pixel values ​​of the matching regions of the respective matching reference images, the encoder can add the weighted sum of the pixel values ​​of the moved matching regions.

[0082] In some embodiments, inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference images in the same temporal direction relative to the current image. For example, Figure 1 Image 104 in the image is a one-way inter-frame prediction image, where the reference image (i.e., image 102) precedes image 104. Two-way inter-frame prediction can use one or more reference images in two temporal directions relative to the current image. For example, Figure 1 Image 106 in the image is a bidirectional inter-frame prediction image, in which the reference images (i.e., images 104 and 108) are in two time directions relative to image 104.

[0083] Referring again to the forward path of process 200B, after spatial prediction stage 2042 and temporal prediction stage 2044, in mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra-frame prediction or inter-frame prediction) for the current iteration of process 200B. For example, the encoder can perform a rate distortion optimization technique, whereby the encoder selects a prediction mode to minimize the value of the cost function based on the bit rate of the candidate prediction modes and the distortion of the reconstructed reference image under the candidate prediction modes. Based on the selected prediction mode, the encoder can generate the corresponding prediction BPU 208 and prediction data 206.

[0084] In the reconstruction path of process 200B, if intra-frame prediction mode is selected in the forward path, the encoder can directly provide prediction reference 224 to spatial prediction stage 2042 for later use (e.g., for interpolation of the next BPU in the current image) after generating prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current image). If inter-frame prediction mode is selected in the forward path, the encoder can provide prediction reference 224 to loop filtering stage 232 after generating prediction reference 224 (e.g., the current image where all BPUs have been encoded and reconstructed), whereby the encoder can apply loop filtering to prediction reference 224 to reduce or eliminate distortions introduced by inter-frame prediction (e.g., block artifacts). The encoder can apply various loop filtering techniques in loop filtering stage 232, such as deblocking, sample adaptive offset, adaptive loop filtering, etc. The loop-filtered reference image can be stored in buffer 234 (or "decoded image buffer") for later use (e.g., as an inter-frame prediction reference image for future images of video sequence 202). The encoder may store one or more reference images in buffer 234 for use in the time prediction stage 2044. In some embodiments, the encoder may encode parameters of the loop filter (e.g., loop filter strength), as well as quantized transform coefficients 216, prediction data 206, and other information in the binary encoding stage 226.

[0085] Figure 3A The illustration shows a schematic diagram of an example decoding process 300A consistent with embodiments of the present disclosure. Process 300A may be corresponding to... Figure 2A The compression process 200A in the video stream is followed by the decompression process. In some embodiments, process 300A can be similar to the reconstruction path of process 200A. The decoder can decode the video bitstream 228 into video stream 304 according to process 300A. Video stream 304 can be very similar to video sequence 202. However, due to information loss during compression and decompression (e.g., Figure 2A-2B In the quantization stage 214), video stream 304 is typically not the same as video sequence 202. Figure 2A-2B Similar to processes 200A and 200B, the decoder can perform process 300A on each image encoded in the video bitstream 228 at the Basic Processing Unit (BPU) level. For example, the decoder can perform process 300A in an iterative manner, where the decoder can decode one BPU in one iteration of process 300A. In some embodiments, the decoder can perform process 300A in parallel on regions (e.g., regions 114-118) of each image encoded in the video bitstream 228.

[0086] exist Figure 3AIn this process, the decoder may provide a portion of the video bitstream 228 associated with a basic processing unit (referred to as an "encoded BPU") of the encoded image to the binary decoding stage 302. In the binary decoding stage 302, the decoder may decode this portion into prediction data 206 and quantized transform coefficients 216. The decoder may provide the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may provide the prediction data 206 to the prediction stage 204 to generate a prediction BPU 208. The decoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded image buffer in computer memory). The decoder may provide the prediction reference 224 to the prediction stage 204 to perform a prediction operation in the next iteration of process 300A.

[0087] The decoder can iteratively execute process 300A to decode each encoded BPU of the encoded image and generate a prediction reference 224 for decoding the next encoded BPU of the encoded image. After decoding all encoded BPUs of the encoded image, the decoder can output the image to video stream 304 for display and continue decoding the next encoded image in video bit stream 228.

[0088] In binary decoding stage 302, the decoder can perform the inverse operation of the binary encoding technique used by the encoder (e.g., entropy coding, variable-length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder can decode other information in binary decoding stage 302, such as the prediction mode, parameters of the prediction operation, transform type, quantization parameter process (e.g., quantization parameters), encoder control parameters (e.g., bit rate control parameters), etc. In some embodiments, if the video bitstream 228 is transmitted over the network in packets, the decoder can unpack it before feeding the video bitstream 228 to binary decoding stage 302.

[0089] Figure 3B A schematic diagram of another example decoding process 300B consistent with embodiments of this disclosure is shown. Process 300B can be modified from process 300A. For example, process 300B can be used by a decoder conforming to a hybrid video coding standard (e.g., H.26x series). Compared to process 300A, process 300B further divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and further includes a loop filtering stage 232 and a buffer 234.

[0090] In process 300B, for the encoded basic processing unit (referred to as the "current BPU") of the decoded encoded image (referred to as the "current image"), the prediction data 206 decoded by the decoder in the binary decoding stage 302 can contain various types of data, depending on the prediction mode used by the encoder to encode the current BPU. For example, if the encoder uses intra-frame prediction to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra-frame prediction, parameters for the intra-frame prediction operation, etc. Parameters for the intra-frame prediction operation may include, for example, the positions (e.g., coordinates) of one or more neighboring BPUs used as references, the sizes of neighboring BPUs, interpolation parameters, the orientation of neighboring BPUs relative to the original BPU, etc. As another example, if the encoder uses inter-frame prediction to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter-frame prediction, parameters for the inter-frame prediction operation, etc. The parameters of the inter-frame prediction operation may include, for example, the number of reference images associated with the current BPU, the weights associated with each reference image, the positions (e.g., coordinates) of one or more matching regions in each reference image, and one or more motion vectors associated with each matching region.

[0091] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra-frame prediction) in the spatial prediction phase 2042 or temporal prediction (e.g., inter-frame prediction) in the temporal prediction phase 2044. Figure 2B The details of performing this spatial or temporal prediction are described herein and will not be repeated below. After performing this spatial or temporal prediction, the decoder can generate a prediction BPU 208. The decoder can then add the prediction BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, as shown below. Figure 3A As described in [the text].

[0092] In process 300B, the decoder can provide prediction reference 224 to either spatial prediction stage 2042 or temporal prediction stage 2044 for performing prediction operations in the next iteration of process 300B. For example, if the current BPU is decoded using intra-frame prediction in spatial prediction stage 2042, the decoder can provide prediction reference 224 directly to spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current image) after generating prediction reference 224 (e.g., the decoded current BPU). If the current BPU is decoded using inter-frame prediction in temporal prediction stage 2044, the encoder can provide prediction reference 224 to loop filtering stage 232 to reduce or eliminate distortion (e.g., block artifacts) after generating prediction reference 224 (e.g., a reference image where all BPUs have been decoded). The decoder can... Figure 2BThe loop filtering described herein applies to prediction reference 224. The loop-filtered reference image can be stored in buffer 234 (e.g., a decoded image buffer in computer memory) for later use (e.g., as an inter-frame prediction reference image used as a future encoded image of video bitstream 228). The decoder can store one or more reference images in buffer 234 for use in the temporal prediction stage 2044. In some embodiments, the prediction data can further include parameters of the loop filtering (e.g., loop filter strength) when the prediction mode indicator of the prediction data 206 indicates that inter-frame prediction is used to encode the current BPU.

[0093] Figure 4 This is a block diagram of an example apparatus 400 for encoding or decoding video, consistent with embodiments of this disclosure. Figure 4 As shown, device 400 may include processor 402. When processor 402 executes the instructions described herein, device 400 may become a dedicated machine for video encoding or decoding. Processor 402 may be any type of circuit capable of manipulating or processing information. For example, processor 402 may include any number of central processing units (or “CPU”), graphics processing units (or “GPU”), neural processing units (“NPU”), microcontroller units (“MCU”), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general-purpose array logic (GALs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), system-on-a-chip (SoCs), application-specific integrated circuits (ASICs), and any combination thereof. In some embodiments, processor 402 may also be a group of processors grouped into individual logic components. For example, such as Figure 4 As shown, processor 402 may include multiple processors, including processor 402a, processor 402b and processor 402n.

[0094] Device 400 may also include a memory 404 configured to store data (e.g., a set of instructions, computer code, intermediate data, etc.). For example, as shown in the figure. Figure 4As shown, the stored data may include program instructions (e.g., program instructions for implementing stages in processes 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 can access the program instructions and data for processing (e.g., via bus 410) and execute program instructions to manipulate or control the data for processing. Memory 404 may include a high-speed random access memory device or a non-volatile memory device. In some embodiments, memory 404 may include any number of random access memories (RAM), read-only memories (ROM), optical discs, magnetic disks, hard disks, solid-state drives, flash drives, secure digital cards (SD cards), memory sticks, compact flash memory (CF cards), etc. Memory 404 may also be a group of memories grouped into single logical components. Figure 4 (Not shown in the image).

[0095] Bus 410 may be a communication device for transmitting data between components within device 400, such as an internal bus (e.g., CPU-memory bus), an external bus (e.g., a Universal Serial Bus port, a Peripheral Component Interconnect Fast Port), etc.

[0096] For ease of explanation and to avoid ambiguity, the processor 402 and other data processing circuitry are collectively referred to as "data processing circuitry" in this disclosure. The data processing circuitry can be implemented entirely as hardware, or as a combination of software, hardware, or firmware. Furthermore, the data processing circuitry can be a single, independent module, or it can be wholly or partially integrated into any other component of the device 400.

[0097] Device 400 may also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, intranet, local area network, mobile communication network, etc.). In some embodiments, network interface 406 may include any combination of any number of network interface controllers (NICs), radio frequency (RF) modules, repeaters, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication (“NFC”) adapters, cellular network chips, etc.

[0098] In some embodiments, the device 400 may optionally include a peripheral interface 408 to provide connectivity to one or more peripheral devices. Figure 4 As shown, peripheral devices may include, but are not limited to, cursor control devices (such as mice, touchpads, or touchscreens), keyboards, displays (such as cathode ray tube displays, liquid crystal displays, or light-emitting diode displays), video input devices (e.g., cameras or input interfaces that are coupled to video files for communication), etc.

[0099] It should be noted that the video codec (e.g., the codec for executing processes 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules in device 400. For example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more software modules of device 400, such as program instructions that can be loaded into memory 404. As another example, some or all stages of processes 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of device 400, such as dedicated data processing circuitry (e.g., FPGA, ASIC, NPU, etc.).

[0100] In the quantization and dequantization function blocks (e.g., Figure 2A or Figure 2B Quantization 214 and inverse quantization 218, Figure 3A or Figure 3B The inverse quantization (218) uses quantization parameters (QP) to determine the amount of quantization (and inverse quantization) applied to the prediction residual. The initial QP value for encoding an image or slice can be signaled at a higher level, for example, using the `init_qp_minus26` syntax element in the Image Parameter Set (PPS) and the `slice_qp_delta` syntax element in the slice header. Furthermore, the QP value can be adjusted at the local level for each CU using incremental QP values ​​sent at the granularity of quantization groups.

[0101] To improve the accuracy of motion vectors (MVs) in merged modes, Versatile Video Coding (VVC) draft 6 employs decoder-side motion vector correction (DMVR) based on bilateral matching (BM). In bilateral operation, corrected MVs are searched around the initial MVs in reference image lists L0 and L1. BM-based DMVR calculates the distortion between two candidate blocks in reference image lists L0 and L1. Figure 5 An example process 500 of decoder-side motion vector correction (DMVR) according to some embodiments of the present disclosure is illustrated. Figure 5The diagram illustrates the current image 502, the first reference image 504 in the first reference image list L0, and the second reference image 506 in the second reference image list L1. It also shows a first initial MV 508 pointing from the current block 510 in the current image 502 to the first initial reference block 512 in the first reference image 504, and a second initial MV 514 pointing from the current block 510 to the second initial reference block 516 in the second reference image 506. The process 500 can perform a BM-based DMVR to determine a first candidate reference block 518 in the first reference image 504, a second candidate reference block 520 in the second reference image 506, a first candidate MV 522 connecting the current block 510 and the first candidate reference block 518, and a second candidate MV 524 connecting the current block 510 and the second candidate reference block 516. (The last sentence appears to be a fragment and doesn't translate directly.) Figure 5 As shown, the first candidate MV 522 and the second candidate MV 524 are located close to the first initial MV 508 and the second initial MV 514, respectively. In some embodiments, process 500 may calculate the sum of absolute differences (SAD) between the first initial reference block 518 and the second initial reference block 520 based on each MV candidate (e.g., the first candidate MV 522 or the second candidate MV 524) around the initial MV (e.g., the first initial MV 508 or the second initial MV 514). The first candidate MV 522 and the second candidate MV 524 with the lowest SAD can become the corrected MV for generating a bidirectional prediction signal.

[0102] In some embodiments, as described in VVC Draft 6, DMVR is applied to a CU that satisfies all of the following conditions: (1) the merging mode is at the CU level with bidirectional prediction MV; (2) the block is predicted using bidirectional prediction MV with equal weights (e.g., by not applying bidirectional prediction with weighted average (BWA) to the block; (3) relative to the current image (e.g., current image 502), one reference image is in the past (e.g., first reference image 504) and another reference image is in the future (e.g., second reference image 506); (4) the two reference images are at the same distance (e.g., image order count (POC) difference) from the current image; and (5) the block has at least 128 luminance samples and the width and height of the block are both at least 8 luminance samples.

[0103] The revised MV derived from the DMVR process (e.g., process 500) can be used to generate inter-frame prediction samples and temporal motion vector predictions for future image coding. The initial MV (e.g., first initial MV 508 or second initial MV 514) can be used for deblocking and spatial MV prediction for future CU coding within the current image (e.g., current image 502).

[0104] like Figure 5As shown, the first MV offset 526 represents the corrected offset between the first initial MV 508 and the first candidate MV 522, and the second MV offset 528 represents the corrected offset between the second initial MV 514 and the second candidate MV 524. The first MV offset 526 and the second MV offset 528 may have the same size and opposite directions. In some embodiments, the search point may be around the initial MV (e.g., the first initial MV 508 and the second initial MV 514), and the MV offsets (e.g., the first MV offset 526 and the second MV offset 528) may follow the MV difference mirroring rule. For example, the point checked by DMVR may be represented by a pair of candidate MVs MV0 (e.g., the first candidate MV 522) and MV1 (e.g., the second candidate MV 524) based on equations (1) and (2):

[0105] Equation (1) is: MV0′=MV0+MV_offset

[0106] Equation (2) is: MV1′=MV1-MV_offset

[0107] In equations (1) and (2), MV0' (e.g., first initial MV 508) and MV1' (e.g., second candidate MV 514) represent a pair of initial MVs, and MV_offset represents the correction offset (e.g., first MV offset 526 or second MV offset 52) ​​between the initial MV (e.g., MV0' or MV1') and the correction MV (e.g., MV0 or MV1). Note that MV_offset is a vector with motion displacement (e.g., in the X and Y dimensions). In some embodiments, as described in VVC draft 5, the correction search range (e.g., the search range of DMVR) can be two integer luminance samples from the initial MVs (e.g., first initial MV 508 and second initial MV 514) in the horizontal and vertical directions.

[0108] Figure 6 An example DMVR search process 600 according to some embodiments of the present disclosure is illustrated. In some embodiments, process 600 may be performed by a codec (e.g., Figure 2A-2B encoder or Figures 3A-3B The codec is executed by the decoder in the codec. For example, the codec can be implemented as a device for encoding or transcoding video sequences (e.g., ...). Figure 4 The device 400 in the process comprises one or more software or hardware components. In some embodiments, process 600 may be an example of the DMVR search process described in VVC draft 4. Figure 6As shown, process 600 includes a stage 602 for integer sample offset search and a stage 604 for fractional sample correction. To reduce search complexity, in some embodiments, a fast search method with an early termination mechanism is applied in stage 602. For example, a two-iteration search scheme can be applied in stage 602 instead of using a 25-point full search to reduce SAD checkpoints.

[0109] like Figure 6 As shown, stage 604 can occur after stage 602. To save computational complexity, in some embodiments, the fractional sample correction for stage 604 can be derived using parametric error surface equations instead of performing an additional search involving SAD comparisons. Stage 604 can be conditionally invoked based on the output of stage 602.

[0110] Figure 7 An example mode 700 for DMVR integer luminance sample search according to some embodiments of the present disclosure is illustrated. DMVR integer luminance sample search can determine the point with the minimum SAD in the search samples. For example, DMVR integer luminance sample search can be implemented as follows: Figure 6 The process 600 includes a phase for integer sample offset search (e.g., phase 602) and a phase for fractional sample correction search (e.g., phase 604), wherein each phase can be performed in at least one iteration. In some embodiments, up to six SADs can be checked in the first iteration of the DMVR integer luminance sample search. Figure 7 For example, in the first iteration, the SAD of five points 702-710 (represented as black blocks) can be compared, with point 702 serving as the search center point. If the SAD of the center point (i.e., point 702) is the smallest, the integer sampling phase of the DMVR can be terminated. Otherwise, another point 712 (represented as a shadow block) determined by the SAD distribution of points 704-710 can be examined. In the second iteration of the DMVR integer luminance sample search, the point with the smallest SAD among points 704-712 can be selected as the new search center point. In some embodiments, the second iteration can be performed in the same manner as the first iteration. In some embodiments, the SAD calculated in the first iteration can be reused in the second iteration, so only further calculation of the SAD of additional points is required.

[0111] In some embodiments, as described in VVC Draft 6, the following can be removed: Figure 7 The two-iteration search described in the text. Then, in the integer sample offset search phase (e.g., Figure 6 In stage 602, the SAD of all 25 points can be calculated in one iteration. Figure 8An example mode 800 is illustrated for a stage of integer sample offset search in DMVR integer luminance sample search according to some embodiments of the present disclosure. For example, the stage of integer sample offset search may be... Figure 6 Phase 602. Figure 8 The initial MV 802 and 25 points are shown, and their SAD can be calculated together. In some embodiments, the SAD of the initial MV 802 can be reduced (e.g., reduced by a quarter) to adjust the initial MV 802. In some embodiments, this can be done in a stage for partial sample correction (e.g., in...). Figure 6 In stage 604, further corrections are made to the position with the minimum SAD. The fractional sample correction stage can be conditionally invoked based on the position with the minimum SAD. For example, as... Figure 8 As shown, if the position of minimum SAD is one of the nine points surrounding the initial MV 802 (as shown in box 804), the stage for fractional sample correction can be invoked to determine the corrected MV as the output of the DMVR integer luminance sample search. If the position of minimum SAD is not any of the nine points surrounding the initial MV 802, the position of minimum SAD can be directly used as the output of the DMVR integer luminance sample search.

[0112] Figure 9 This is a schematic diagram illustrating an example mode 900 for estimating the error surface of DMVR parameters according to some embodiments of the present disclosure. Figure 8 The initial MV 902 and 25 points are shown. The initial MV 902 is connected to the center point 904 with the minimum SAD. In sub-pixel offset estimation based on the parametric error surface, as... Figure 9 As shown, the sum of the absolute difference (SAD) cost of center point 904 and the SAD costs of the four adjacent points 906-912 around center point 904 can be used to fit the equation of the two-dimensional parabolic error surface. For example, the equation of the two-dimensional parabolic error surface can be based on equation (3).

[0113] E(x,y)=((A(xx mim ) 2 +B(yy min ) 2 +)>> mvShift)+E(0,0) Equation (3)

[0114] In equation (3), (x min y min E(x, y) corresponds to the fractional position with the lowest SAD cost, E(x, y) corresponds to the SAD cost of the center point 904 and the four adjacent points 906-912, mvShift can be set to 4 as in VVC (in VVC, the MV precision is 1 / 16 pixel), and A and B can be determined according to equations (4) and (5) respectively:

[0115]

[0116] By solving equations (3) to (5) using the SAD cost values ​​of the five search points (i.e., points 904-912), (x) can be determined according to equations (6) and (7). min y min ).

[0117]

[0118] In some embodiments, x min and y min The value can be automatically limited to between -8 and 8 (e.g., in 1 / 16 sample precision) because all SAD cost values ​​are positive and the minimum is E(0,0), which corresponds to a half-pixel offset in VVC with 1 / 16 pixel MV precision. The calculated score (x...) min y min This can be added to the integer distance correction MV to give the correction MV subpixel precision.

[0119] The Two-Way Optical Flow (BDOF) tool is included in VVC. As the name suggests, the BDOF mode is based on the concept of optical flow, assuming that the motion of objects is smooth. BDOF, formerly known as BIO, is also included in the Joint Video Exploration Model (JEM) software. Compared to BIO in JEM, BDOF in VVC is a simpler version, especially in terms of the number of multiplications and the size of the multiplier, requiring significantly less computation.

[0120] BDOF can be used to correct the bidirectional prediction signal of the CU at the 4×4 sub-block level. In some embodiments, BDOF is applied to the CU under the following conditions: (1) the height of the CU is not 4, and the size of the CU is not 4×8; (2) the CU is not encoded using affine mode or Advanced Time Motion Vector Prediction (ATMVP) merging mode; (3) the CU is encoded using a “true” bidirectional prediction mode, where one of the two reference images (e.g., Figure 5 The first reference image 504 in the image is displayed in the current image (e.g., ...). Figure 5 The current image 502 in the image is before another one (e.g., Figure 5 The second reference image 506 is placed after the current image in the display order. In some embodiments, BDOF can be applied to the luminance component.

[0121] In some embodiments, when BDOF is used to correct the bidirectional prediction signal of the CU at the 4×4 sub-block level, for each 4×4 sub-block, motion correction (v x v yThis can be calculated by minimizing the difference between predicted samples in two reference image lists L0 and L1. x v y This can then be used to adjust the bidirectional prediction sample values ​​in the 4×4 sub-block.

[0122] In some embodiments, the following steps are applied during the BDOF process: First, the horizontal and vertical gradients of the two predicted signals at k=0,1 can be determined based on calculating the difference between two adjacent samples. and As shown in equations (8) and (9):

[0123]

[0124] In equations (8) and (9), I (k) (i,j) is the sample value of the predicted signal in list k at coordinate (i,j) when k = 0, 1. shift1 is calculated based on the luminance bit depth (“bitDepth”), as shown in equation (10):

[0125] shift1 = max(2, 14-bitDepth) Equation (10)

[0126] Then, the autocorrelation and cross-correlation of gradients S1, S2, S3, S5 and S6 can be determined according to equations (11) to (15):

[0127]

[0128] For equations (11) to (15), ψ can be determined based on equations (16) to (18). x (i,j),ψ y The values ​​of (i,j) and θ(i,j)

[0129]

[0130] θ(i,j)=(I (1) (i,j)>>n b )-(I (0) (i,j)>>n b Equation (18)

[0131] In equations (11) to (18), Ω is the 6×6 window surrounding the 4×4 sub-block, and n a and n b The values ​​are set by equations (19) and (20) respectively:

[0132] n a =min(5,bitDepth-7) Equation (19)

[0133] n b =min(8,bitDepth-4) Equation (20)

[0134] Then, based on the cross-correlation and autocorrelation terms, equations (21) and (22) are used to derive the motion correction (v). x ,v y ):

[0135]

[0136] In equations (21) and (22), th′ BIO =2 13-BD

[0137] Represents the floor function, and

[0138] Based on motion correction and gradient, the following adjustment b(x,y) for each sample in the 4×4 sub-block can be determined using equation (23):

[0139]

[0140] Finally, the BDOF samples of CU can be determined by adjusting the bidirectional prediction samples according to equation (24):

[0141] pred BDoF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+o offset >> Shift Equation (24)

[0142] In some embodiments, the values ​​in equations (8) to (23) can be selected such that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit width of the intermediate parameters in the BDOF process can be kept within 32 bits.

[0143] To derive the gradient values, some predicted samples I outside the current CU boundary can be generated in a list k (k = 0, 1). (k) (i,j). Figure 10 This is a schematic diagram illustrating an example of an extended coding unit (CU) region 1000 used in BDOF according to some embodiments of this disclosure. Figure 10As shown, the 4×4 block 1002 (surrounded by the solid black line) used in BDOF is surrounded by an extended row or column around the boundary of block 1002 (represented by the dashed black line), forming a bounding region 1004. To control the computational complexity of generating prediction samples outside the boundary, prediction samples within the extended region 1006 (represented by the white box) can be generated by directly obtaining reference samples at nearby integer positions (using floor operations on the coordinates) without interpolation, and prediction samples can be generated within CU 1008 (represented by the gray box) using ordinary 8-tap motion-compensated interpolation filtering. These extended sample values ​​can only be used for gradient calculation. For the remaining steps of the BDOF process, if any samples and gradient values ​​outside the boundary of CU 1008 are needed, they can be populated (or repeated) from their nearest neighbors.

[0144] At the JVET conference, a coding tool called Optical Flow Prediction Correction (PROF) was adopted. PROF improves the accuracy of affine motion compensation prediction by using sub-block-based affine motion compensation predictions with optical flow correction. Affine motion model parameters can be used to derive the motion vector at each sample location in the CU. However, due to the high complexity and memory access bandwidth required to generate per-sample affine motion compensation predictions, affine prediction in VVC uses a sub-block-based affine motion compensation method, where one CU is divided into 4×4 sub-blocks, each assigned an MV derived from the control point MV of the affine CU. Sub-block-based affine motion compensation is a trade-off between coding efficiency, complexity, and memory access bandwidth. Because it is based on sub-block predictions rather than motion compensation predictions based on theoretical samples, it loses some prediction accuracy.

[0145] To achieve finer granularity of affine motion compensation, in some embodiments, PROF can be applied after affine motion compensation based on conventional sub-blocks. Sample-based corrections can be obtained based on optical flow equations, such as equation (25):

[0146] ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y Equation (25) (i,j)

[0147] In equation (25), g x (i,j) and g y (i,j) is the spatial gradient of the sample position (i,j), and Δv is the motion offset from the sub-block-based motion vector to the sample-based motion vector derived from the affine model parameters.

[0148] Figure 11This is a schematic diagram illustrating examples of sub-block-based translational motion and sample-based affine motion according to some embodiments of this disclosure. Figure 11 As shown, V(i,j) is the theoretical motion vector of the sample position (i,j) derived using an affine model. SB It is based on the motion vector of the sub-block, and ΔV(i,j) (represented by the dashed arrow) is the sum of V(i,j) and V SB The differences between them.

[0149] Then, the prediction correction ΔI(i,j) can be added to the sub-block prediction I(i,j). The final prediction I' can be generated based on equation (26):

[0150] I'(i,j)=I(i,j)+ΔI(i,j) Equation (26)

[0151] Consistent with the disclosed embodiments, both DMVR and BDOF can have control flags at two levels within the syntax structure. The first control flag can be sent in the sequence parameter set (SPS) at the sequence level, and the second control flag can be sent in the slice header at the slice level. Figure 12 Table 1 is shown, illustrating example syntax structures for the Sequence Parameter Set (SPS) of control flags for implementing DMVR and BDOF according to some embodiments of this disclosure. Figure 12 As shown in Table 1, `sps_bdof_enabled_flag` and `sps_dmvr_enabled_flag` are the control flags for BDOF and DMVR, respectively, and are sent at the sequence level in the SPS. When either `sps_bdof_enabled_flag` or `sps_dmvr_enabled_flag` is false, BDOF or DMVR can be disabled throughout the entire video sequence referencing this SPS. When both `sps_bdof_enabled_flag` and `sps_dmvr_enabled_flag` are true, BDOF or DMVR can be enabled for the current video sequence. In this case, another flag, `sps_bdof_dmvr_slice_present_flag`, can be further signaled to indicate whether slice-level control of BDOF and DMVR is enabled.

[0152] Figure 13Table 2 illustrates example syntax structures for slice headers implementing control flags for DMVR and BDOF according to some embodiments of this disclosure. As shown in Table 2, when sps_bdof_dmvr_slice_present_flag set in Table 1 is true, slice_disable_bdof_dmvr_flag can indicate in the slice header whether BDOF and DMVR are disabled for the current slice using signals.

[0153] Figure 12-13 A two-level control mechanism of DMVR and BDOF is illustrated. By using this mechanism, the encoder (e.g., implements...) Figure 2A-2B The encoders of process 200A or 200B can use the slice-level flag slice_disable_bdof_dmvr_flag to enable or disable the DMVR and BDOF of individual slices. This slice-level adaptation has two benefits: (1) disabling at least one of the DMVR or BDOF when it is useless to the current slice can improve encoding performance; (2) disabling the DMVR and BDOF can reduce the encoding and decoding complexity of the current slice, as both DMVR and BDOF have relatively high computational complexity.

[0154] In the disclosed embodiments, the control flags for PROF can also be used at the sequence level and slice level. In some embodiments, three separate flags can be signaled in the SPS to indicate whether DMVR, BDOF, and PROF are enabled, respectively. If any of them are enabled, a corresponding lower-level control enable flag can be signaled to indicate whether the enabled tool is controlled at a lower level. The lower level can be the slice level or the image level. If slice-level or image-level control is enabled, a slice-level or image-level disable flag can be signaled in each slice header or image header to indicate that the enabled tool for the current slice or image is disabled.

[0155] Consistent with the disclosed embodiments, Figure 14A Table 3A is shown, illustrating example syntax structures for the sequence parameter set (SPS) of slice-level control flags for DMVR, BDOF, and PROF according to some embodiments of this disclosure. Figure 14BTable 3B illustrates an example syntax structure of the SPS for implementing image-level control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As shown in Tables 3A-3B, highlighted in italics, sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag are flags that signal in the SPS, indicating whether BDOF, DMVR, and PROF are enabled for the video sequence, respectively. If BDOF, DMVR, or PROF is enabled, as shown in Table 3A, sps_bdof_slice_present_flag, sps_dmvr_slice_present_flag, or sps_affine_prof_slice_present_flag can further signal whether slice-level control for BDOF, DMVR, and PROF is enabled, respectively. If BDOF, DMVR, or PROF is enabled, as shown in Table 3B, sps_bdof_picture_present_flag, sps_dmvr_picture_present_flag, or sps_affine_prof_picture_present_flag can further emit signals to indicate whether image level control for BDOF, DMVR, and PROF is enabled.

[0156] Figure 15A Table 4A is illustrated, showing an example syntax structure of the slice header for implementing control flags for DMVR, BDOF, and PROF according to some embodiments of this disclosure. As highlighted in italics in Table 4A, if any of sps_bdof_slice_present_flag, sps_dmvr_slice_present_flag, or sps_affine_prof_slice_present_flag in Table 3A is set to true, then slice_disable_bdof_flag, slice_disable_dmvr_flag, or slice_disable_affine_prof_flag can be signaled to indicate whether BDOF, DMVR, or PROF is disabled for the current slice. Figure 15BTable 4B illustrates example syntax structures for image headers implementing control flags for DMVR, BDOF, and PROF according to some embodiments of this disclosure. As highlighted in italics in Table 4B, if any of sps_bdof_picture_present_flag, sps_dmvr_picture_present_flag, or sps_affine_prof_picture_present_flag in Table 3B is set to true, then ph_disable_bdof_flag, ph_disable_dmvr_flag, or ph_disable_affine_prof_flag can be signaled to indicate whether BDOD, DMVR, or PROF is disabled for the current image.

[0157] In some embodiments, DMVR, BDOF, and PROF can have three separate sequence-level enable flags but share the same slice-level control enable flag. For example, a slice-level disable flag can be sent to DMVR, BDOF, and PROF. In another example, three slice-level disable flags can be signaled separately for DMVR, BDOF, and PROF. As yet another example, two slice-level disable flags can be sent to DMVR, BDOF, and PROF. It should be noted that the control of DMVR, BDOF, and PROF can be implemented in various syntaxes at the sequence level and at levels below the sequence level (here referred to as "lower levels," such as slice level or image level), and is not limited to the examples described herein.

[0158] Figure 16 Table 5 illustrates an example syntax structure for a Sequence Parameter Set (SPS) of separate sequence-level control flags for DMVR, BDOF, and PROF, according to some embodiments of this disclosure. As shown in Table 5, highlighted in italics, three separate flags, sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag, are signaled in the SPS to indicate whether DMVR, BDOF, and PROF are enabled. If at least one of DMVR, BDOF, or PROF is enabled, the slice control enable flag sps_bdof_dmvr_affine_prof_slice_present_flag can be signaled to indicate whether at least one of DMVR, BDOF, or PROF enabled at the sequence level is controlled at a lower level.

[0159] Figure 17Table 6 illustrates an example syntax structure of a slice header for implementing joint control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As highlighted in italics in Table 6, if slice-level control is enabled as described in Table 5 (e.g., `sps_bdof_dmvr_affine_prof_slice_present_flag` is true), then in each slice header, the slice-level disable flag `slice_disable_bdof_dmvr_affine_prof_flag` can be signaled to indicate whether at least one of the DMVR, BDOF, or PROF enabled at the sequence level is disabled for the current slice. In the syntax structure shown in Table 6, if multiple DMVR, BDOF, and PROF are enabled at the sequence level, then, with slice-level control enabled, the multiple enablements can be jointly controlled at the slice level.

[0160] Figure 18Table 7 is illustrated, showing an example syntax structure of the slice header implementing separate control flags for DMVR, BDOF, and PROF according to some embodiments of this disclosure. As highlighted in italics in Table 7, if the slice-level controls described in Table 5 are enabled (e.g., sps_bdof_dmvr_affine_prof_slice_present_flag is true), then for one of the DMVR, BDOF, and PROF enabled at the sequence level, a slice-level disable flag can be signaled to indicate whether one of the aforementioned enablements is disabled for the current slice. For example, if sps_bdof_enabled_flag and sps_bdof_dmvr_affine_prof_slice_present_flag in Table 5 are set to true, then slice_disable_bdof_flag in Table 7 can be signaled to indicate whether BDOF is disabled for the current slice. As another example, if `sps_dmvr_enabled_flag` and `sps_bdof_dmvr_affine_prof_slice_present_flag` in Table 5 are set to true, then `slice_disable_dmvr_flag` in Table 7 can be signaled to indicate whether DMVR is disabled for the current slice. In yet another example, if `sps_affine_prof_enabled_flag` and `sps_bdof_dmvr_affine_prof_slice_present_flag` in Table 5 are set to true, then `slice_disable_affine_prof_flag` in Table 7 can be signaled to indicate whether PROF is disabled for the current slice. In Table 7, if slice-level control is enabled, each of DMVR, BDOF, and PROF can be individually controlled at the slice level.

[0161] Given that both BDOF and PROF use optical flow to correct the inter-frame predictor, in some embodiments, BDOF and PROF can share the same slice-level control flags, while DMVR can use separate slice-level control flags. Figure 19Table 8 illustrates an example syntax structure for a slice header implementing mixed control flags for DMVR, BDOF, and PROF according to some embodiments of this disclosure. In the syntax structure of Table 8, BDOF and PROF can share the same slice-level disable flags, while DMVR can use a different slice-level disable flag. As highlighted in italics in Table 8, if slice-level control is enabled in Table 5 (e.g., sps_bdof_dmvr_affine_prof_slice_present_flag is true), two slice-level disable flags can be signaled to indicate whether DMVR, BDOF, and PROF are disabled for the current slice. For example, if at least one of `sps_bdof_enabled_flag` and `sps_affine_prof_enabled_flag` in Table 5 is set to true, and `sps_bdof_dmvr_affine_prof_slice_present_flag` in Table 5 is set to true, then `slice_disable_bdof_affine_prof_flag` in Table 8 can be signaled to indicate whether at least one of BDOF or PROF enabled at the sequence level is disabled for the current slice. In another example, if both `sps_dmvr_enabled_flag` and `sps_bdof_dmvr_affine_prof_slice_present_flag` in Table 5 are set to true, then `slice_disable_dmvr_flag` in Table 8 can be signaled to indicate whether DMVR is disabled for the current slice. In Table 8, if slice-level control is enabled, BDOF and PROF are jointly controlled at the slice level, while DMVR is controlled separately from BDOF and PROF. Figure 20Table 9 illustrates an example syntax structure for a sequence parameter set (SPS) of mixed sequence-level control flags for DMVR, BDOF, and PROF, according to some embodiments of the present disclosure. In the syntax structure of Table 9, three slice-level disable flags can be issued for DMVR, BDOF, and PROF, respectively. As shown in Table 9, highlighted in italics, three separate flags, sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag, can be signaled in the SPS to indicate whether DMVR, BDOF, or PROF is enabled, respectively. If sps_dmvr_enabled_flag is true, the slice-level control enable flag sps_dmvr_slice_present_flag can be signaled to indicate whether DMVR is controlled at the slice level. If at least one of sps_bdof_enabled_flag or sps_affine_prof_enabled_flag is true, the slice-level control enable flag sps_bdof_affine_prof_slice_present_flag can be signaled to indicate whether at least one of BDOF or PROF is controlled at the slice level.

[0162] Figure 21 Table 10 illustrates another example syntax structure of the slice header for implementing mixed control flags for DMVR, BDOF, and PROF according to some embodiments of this disclosure. As highlighted in italics in Table 10, if `sps_dmvr_slice_present_flag` in Table 9 is set to true, the slice-level disable flag `slice_disable_dmvr_flag` can be signaled to indicate whether DMVR is disabled for the current slice. If `sps_bdof_affine_prof_slice_present_flag` in Table 9 is set to true, the slice-level disable flag `slice_disable_bdof_affine_prof_flag` can be sent to indicate whether at least one of the BDOF or PROF enabled at the sequence level (as described in Table 8) can be disabled for the current slice. In the syntax structure of Table 9, BDOF and PROF can be jointly controlled at the slice level, and DMVR can be controlled individually at the slice level if slice-level control is enabled.

[0163] Figure 22Table 11 is shown, illustrating another example syntax structure of the slice header for implementing separate control flags for DMVR, BDOF, and PROF according to some embodiments of this disclosure. As highlighted in italics in Table 11, if `sps_dmvr_enabled_flag` in Table 9 is set to true, the slice-level disable flag `slice_disable_dmvr_flag` can be signaled to indicate whether DMVR is disabled for the current slice. If `sps_bdof_enabled_flag` and `sps_bdof_affine_prof_slice_present_flag` in Table 9 are set to true, the slice-level disable flag `slice_disable_bdof_flag` can be signaled to indicate whether BDOF is disabled for the current slice. If `sps_affine_prof_enabled_flag` and `sps_bdof_affine_prof_slice_present_flag` in Table 9 are set to true, the slice-level disable flag `slice_disable_affine_prof_flag` can be signaled to indicate whether PROF is disabled for the current slice. In the syntax structure of Table 11, if slice-level control is enabled, each of DMVR, BDOF, and PROF is controlled separately at the slice level.

[0164] Figure 23-26 A flowchart illustrating example processes 2300-2600 for controlling video encoding / decoding modes according to some embodiments of the present disclosure is shown. In some embodiments, processes 2300-2600 may be controlled by a codec (e.g., Figure 2A-2B encoder or Figures 3A-3B The codec is executed by the decoder in the video sequence. For example, the codec can be implemented as one or more software or hardware components of a means (e.g., means 400) for controlling the encoding or decoding modes of a video sequence.

[0165] Figure 23 A flowchart illustrating an example process 2300 for controlling a video decoding mode according to some embodiments of the present disclosure is shown. In step 2302, the codec (e.g., Figures 3A-3B The decoder in the video stream can receive bitstreams of video data (e.g., Figures 3A-3B The video bitstream 228 in process 300A or 300B.

[0166] In step 2304, the codec can enable or disable the video sequence based on a first flag in the bitstream (e.g., Figures 3A-3BThe encoding mode of the video stream 304 in process 300A or 300B. For example, the encoding mode may be at least one of bidirectional optical flow (BDOF) mode, optical flow prediction correction (PROF) mode, or decoder-side motion vector correction (DMVR) mode. In some embodiments, the codec may detect a first flag in the sequence parameter set (SPS) of the video sequence. For example, the first flag may be such as Figure 12 , 14A The flags described in -14B, 16, or 20: sps_bdof_enabled_flag, sps_dmvr_enabled_flag, or sps_affine_prof_enabled_flag.

[0167] In step 2306, the codec may determine whether to enable or disable encoding mode control at a level below the sequence level based on a second flag in the bitstream. Levels below the sequence level may include slice level or image level. In some embodiments, the codec may detect the second flag in the SPS of the video sequence in response to enabling an encoding mode for the video sequence. For example, the second flag may be as follows: Figure 12 , 14A The flags described in -14B, 16, or 20 are: sps_bdof_dmvr_slice_present_flag, sps_bdof_slice_present_flag, sps_dmvr_slice_present_flag, sps_affine_prof_slice_present_flag, sps_bdof_picture_present_flag, sps_dmvr_picture_present_flag, sps_affine_prof_picture_present_flag, sps_bdof_affine_prof_slice_present_flag, or sps_bdof_dmvr_affine_prof_slice_present_flag.

[0168] In some embodiments, after step 2306, in response to control enabling an encoding mode at a level below the sequence level, the codec may enable or disable an encoding mode for a target low-level region based on a third flag in the bitstream. The target low-level region may be a target slice or a target image. If the low level is the slice level, in some embodiments, the codec may detect the third flag in the slice header of the target slice. If the low level is the image level, in some embodiments, the codec may detect the third flag in the image header of the target image. For example, the third flag may be as follows: Figure 13 , 15A The flags described in -15B, 17-19, or 21-22 are slice_disable_bdof_dmvr_flag, slice_disable_bdof_flag, slice_disable_dmvr_flag, slice_disable_affine_prof_flag, ph_disable_bdof_flag, ph_disable_dmvr_flag, ph_disable_affine_prof_flag, slice_disable_bdof_dmvr_affine_prof_flag, or slice_disable_bdof_affine_prof_flag.

[0169] Figure 24 A flowchart of another example process 2400 for controlling a video decoding mode according to some embodiments of the present disclosure is shown. In step 2402, the codec (e.g., Figures 3A-3B The decoder in the video stream can receive bitstreams of video data (e.g., Figures 3A-3B The video bitstream 228 in process 300A or 300B.

[0170] In step 2404, the codec can enable or disable the video sequence based on a first flag in the bitstream (e.g., Figures 3A-3B The first encoding mode of the video stream (304) in process 300A or 300B. In step 2406, the codec can enable or disable a second encoding mode of the video sequence based on a second flag in the bitstream. The first and second encoding modes can be two different encoding modes, selectable from bidirectional optical flow (BDOF) mode, optical flow prediction correction (PROF) mode, and decoder-side motion vector correction (DMVR) mode. For example, the first encoding mode and the second encoding mode can be bidirectional optical flow (BDOF) mode and optical flow prediction correction (PROF) mode, respectively.

[0171] In some embodiments, the codec can detect first and second flags in the sequence parameter set (SPS) of the video sequence. For example, the first and second flags can be obtained from... Figure 12 , 14A Choose from the flags described in -14B, 16, or 20: sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag. As another example, if the first encoding mode and the second encoding mode are BDOF mode and PROF mode respectively, then the first flag and the second flag can be respectively... Figure 12 , 14A The flags sps_bdof_enabled_flag and the tag sps_affine_prof_enabled_flag described in -14B, 16 or 20.

[0172] In step 2408, the codec may determine, based on a third flag in the bitstream, whether to enable control over at least one of the first or second coding modes at a level below the sequence level. Levels below the sequence level may include the slice level or the picture level. In some embodiments, the codec may detect the third flag in the SPS of the video sequence in response to enabling at least one of the first or second coding modes for the video sequence. For example, the third flag can be the flag sps_bdof_dmvr_slice_present_flag, as described in Figures 12, 14A-14B, 16, or 20.

[0173] In some embodiments, after step 2408, in response to a first flag indicating that the video sequence enables a first encoding mode (e.g., ... Figure 20The third flag (as described in the document) indicates the control of enabling at least one of the first or second encoding modes (e.g., PROF) at a level below the sequence level. Figure 20 The codec can be based on the fourth flag in the bitstream (as described in sps_bdof_affine_prof_slice_present_flag). Figure 22 The `slice_disable_bdof_flag` described in the document enables or disables a first encoding mode (e.g., BDOF) for a target low-level region. The target low-level region can be a target slice or a target image. If the target low-level region is a target slice, in some embodiments, the codec can detect a fourth flag in the slice header of the target slice. If the target low-level region is a target image, in some embodiments, the codec can detect a fourth flag in the image header of the target image. For example, the fourth flag can be... Figure 13 , 15A The flags described in -15B, 17-19, or 21-22 are slice_disable_bdof_dmvr_flag, slice_disable_bdof_flag, slice_disable_dmvr_flag, slice_disable_affine_prof_flag, ph_disable_bdof_flag, ph_disable_dmvr_flag, ph_disable_affine_prof_flag, slice_disable_bdof_dmvr_affine_prof_flag, or slice_disable_bdof_affine_prof_flag.

[0174] In some embodiments, after step 2408, in response to control enabling at least one of a first encoding mode or a second encoding mode at a level below the sequence level, the codec may base its codec on a fourth flag in the bitstream (e.g., Figure 21 The slice_disable_bdof_affine_prof_flag shown enables or disables the first encoding mode (e.g., BDOF) and the second encoding mode (e.g., PROF) for the target lower-level region. For example, a third flag (e.g., such as...) Figure 20 The sps_bdof_affine_prof_slice_present_flag described in the document can indicate that at least one of the first encoding mode or the second encoding mode is enabled at a lower level (e.g., the slice level).

[0175] In some embodiments, after enabling or disabling the first and second coding modes of a lower-level target region (e.g., a target slice or target image) based on a fourth flag in the bitstream, the codec can further enable or disable a third coding mode of the video sequence based on the second flag in the bitstream, and determine whether to enable control of the third coding mode at a level below the sequence level based on a fifth flag in the bitstream. For example, the first, second, and third coding modes can be BDOF mode, PROF mode, and DMVR mode, respectively. In this example, the fourth flag can be as follows: Figure 21 The second flag can be as described in the slice_disable_bdof_affine_prof_flag. Figure 20 The fifth flag described in the document, sps_dmvr_enabled_flag, can be as follows: Figure 20 The flag described in [the document] is sps_dmvr_slice_present_flag.

[0176] In some embodiments, in response to enabling a third coding mode at a lower level (e.g., slice level or image level), the codec may further enable or disable a third coding mode for a target lower-level region (e.g., target slice or target image) based on a sixth flag in the bitstream. For example, when the first, second, and third coding modes can be BDOF mode, PROF mode, and DMVR mode, respectively, the sixth flag can be as follows: Figure 21 The slice_disable_dmvr_flag mentioned above.

[0177] Figure 25 A flowchart illustrating an example process 2500 for controlling a video encoding mode according to some embodiments of the present disclosure is shown. In step 2502, the codec (e.g., Figure 2A-2B The encoder in the video can receive video sequences (e.g., Figure 2A-2B The video sequence 202 in process 200A or 200B, the first flag and the second flag. For example, the first flag can be as follows: Figure 12 , 14A The flags described in -14B, 16, or 20 are sps_bdof_enabled_flag, sps_dmvr_enabled_flag, or sps_affine_prof_enabled_flag. As another example, the second flag could be as shown in... Figure 12 , 14AThe flags described in -14B, 16, or 20 are: sps_bdof_dmvr_slice_present_flag, sps_bdof_slice_present_flag, sps_dmvr_slice_present_flag, sps_affine_prof_slice_present_flag, sps_bdof_picture_present_flag, sps_dmvr_picture_present_flag, sps_affine_prof_picture_present_flag, sps_bdof_affine_prof_slice_present_flag, or sps_bdof_dmvr_affine_prof_slice_present_flag.

[0178] In step 2504, the codec can enable or disable the video bitstream based on a first flag in the bitstream (e.g., Figure 2A-2B The encoding mode of the video bitstream 228 in process 200A or 200B. For example, the encoding mode may be at least one of bidirectional optical flow (BDOF) mode, optical flow prediction correction (PROF) mode, or decoder-side motion vector correction (DMVR) mode.

[0179] In step 2506, the codec can enable or disable control over the encoding mode at a level below the sequence level based on the second flag. Levels below the sequence level can include slice level or image level.

[0180] Figure 26 A flowchart of another example process 2600 for controlling a video encoding mode according to some embodiments of the present disclosure is shown. In step 2602, the codec (e.g., Figure 2A-2B The encoder in the video can receive video sequences (e.g., Figure 2A-2B The process 200A or 200B includes video sequence 202), a first flag, a second flag, and a third flag. For example, the first and second flags can be derived from, for example... Figure 12 , 14A Choose from the flags sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag as described in -14B, 16, or 20. As another example, the third flag could be as follows: Figure 12 , 14AThe flags described in -14B, 16, or 20 are: sps_bdof_dmvr_slice_present_flag, sps_bdof_slice_present_flag, sps_dmvr_slice_present_flag, sps_affine_prof_slice_present_flag, sps_bdof_picture_present_flag, sps_dmvr_picture_present_flag, sps_affine_prof_picture_present_flag, sps_bdof_affine_prof_slice_present_flag, or sps_bdof_dmvr_affine_prof_slice_present_flag.

[0181] In step 2604, the codec can base its code on the first flag for the video bitstream (e.g., Figure 2A-2B In process 200A or 200B, the video bitstream 228) enables or disables a first coding mode. In step 2606, the codec can enable or disable a second coding mode for the video bitstream based on a second flag. The first coding mode and the second coding mode can be two different coding modes, selectable from a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode, and a decoder-side motion vector correction (DMVR) mode. For example, the first coding mode and the second coding mode can be a bidirectional optical flow (BDOF) mode and an optical flow prediction correction (PROF) mode, respectively.

[0182] In step 2608, the codec may enable or disable control over at least one of the first or second encoding modes at a level below the sequence level based on a third flag. The level below the sequence level may include the slice level or the image level.

[0183] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by a device (e.g., the disclosed encoder and decoder) to perform the methods described above. Common forms of non-transitory media include, for example, floppy disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a perforated pattern, RAM, PROMs and EPROMs, FLASH-EPROMs or any other flash memory, NVRAM, caches, registers, any other memory chips or cassette memories, and their network versions. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memory.

[0184] The embodiments may be further described using the following terms:

[0185] 1. A computer-implemented method, comprising:

[0186] Receive video data bitstream;

[0187] Enabling or disabling the encoding mode for the video sequence based on a first flag in the bitstream; and

[0188] Based on the second flag in the bitstream, it is determined whether to enable or disable control over the encoding mode at a level below the sequence level.

[0189] 2. The computer implementation method according to Clause 1, wherein the level below the sequence level includes the slice level or the image level.

[0190] 3. The computer-implemented method according to any one of clauses 1-2 further includes:

[0191] In response to enabling control of the encoding mode at a level below the sequence level, the encoding mode for a target low-level region is enabled or disabled based on a third flag in the bitstream.

[0192] 4. The computer-implemented method according to Clause 3 further includes:

[0193] Detect the third flag in the slice header of the target slice, wherein the target slice is a low-level region of the target; or

[0194] The third marker is detected in the image header of the target image, wherein the target image is the target low-level region.

[0195] 5. A computer-implemented method according to any one of clauses 1-4, wherein the encoding mode is at least one of the following:

[0196] Bidirectional optical flow (BDOF) mode;

[0197] Optical flow prediction correction (PROF) mode; or

[0198] Decoder-side Motion Vector Correction (DMVR) mode.

[0199] 6. The computer-implemented method according to any one of clauses 1-5 further includes:

[0200] Detect the first flag in the sequence parameter set (SPS) of the video sequence.

[0201] 7. The computer-implemented method according to any one of clauses 1-6 further includes:

[0202] The response enables the encoding mode for the video sequence and detects the second flag in the SPS of the video sequence.

[0203] 8. A computer-implemented method, comprising:

[0204] Receive video data bitstream;

[0205] A first encoding mode for a video sequence is enabled or disabled based on a first flag in the bitstream;

[0206] A second encoding mode for the video sequence is enabled or disabled based on a second flag in the bitstream; and

[0207] Based on a third flag in the bitstream, it is determined whether to enable control over at least one of the first encoding mode or the second encoding mode at a level below the sequence level.

[0208] 9. The computer implementation method according to Clause 8, wherein the level below the sequence level includes the slice level or the image level.

[0209] 10. The computer-implemented method according to any one of clauses 8-9 further includes:

[0210] In response to enabling at least one of the first encoding mode or the second encoding mode at a level below the sequence level, the first encoding mode and the second encoding mode are enabled or disabled for a target low-level region based on a fourth flag in the bitstream.

[0211] 11. The computer-implemented method according to Clause 10 further includes:

[0212] Detect the fourth flag in the slice header of the target slice, where the target slice is a low-level region of the target; or

[0213] The fourth marker in the image header of the target image is detected, wherein the target image is the target low-level region.

[0214] 12. The computer-implemented method according to any one of clauses 10-11 further includes:

[0215] Based on the second flag in the bitstream, enable or disable the third encoding mode for the video sequence; and

[0216] Based on the fifth flag in the bitstream, it is determined whether to enable control of the third encoding mode at the level below the sequence level.

[0217] 13. The computer-implemented method according to Clause 12 further includes:

[0218] In response to enabling the third coding mode at a level below the sequence level, the third coding mode for the target low-level region is enabled or disabled based on a sixth flag in the bitstream.

[0219] 14. The computer implementation method according to any one of Clauses 12 and 13, wherein the first encoding mode, the second encoding mode and the third encoding mode are respectively a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode and a decoder-side motion vector correction (DMVR) mode.

[0220] 15. The computer-implemented method according to any one of clauses 8-14 further includes:

[0221] In response to a first flag indicating the activation of a first coding mode for a video sequence and a third flag indicating the activation of control over at least one of the first coding mode or the second coding mode at a level below the sequence level, the first coding mode for a target lower-level region is enabled or disabled based on a fourth flag in the bitstream.

[0222] 16. The computer implementation method according to any one of clauses 8-15, wherein the first encoding mode and the second encoding mode are selected from two different encoding modes:

[0223] Bidirectional optical flow (BDOF) mode;

[0224] Optical flow prediction correction (PROF) mode; and

[0225] Decoder-side Motion Vector Correction (DMVR) mode.

[0226] 17. A computer-implemented method according to any one of Clauses 8-16, wherein the first encoding mode and the second encoding mode are a bidirectional optical flow (BDOF) mode and an optical flow prediction correction (PROF) mode, respectively.

[0227] 18. The computer-implemented method according to any one of clauses 8-17 further includes:

[0228] Detect the first and second flags in the sequence parameter set (SPS) of the video sequence.

[0229] 19. The computer implementation method according to any one of clauses 8-18 further includes:

[0230] In response to enabling at least one of the first encoding mode or the second encoding mode for the video sequence, a third flag in the SPS of the video sequence is detected.

[0231] 20. A computer-implemented method, comprising:

[0232] Receive the video sequence, the first flag, and the second flag;

[0233] Enabling or disabling the encoding mode for the video bitstream based on the first flag; and

[0234] Based on the second flag, control over the encoding mode is enabled or disabled at a level below the sequence level.

[0235] 21. The computer implementation method according to Clause 20, wherein the level below the sequence level includes the slice level or the image level.

[0236] 22. The computer-implemented method according to any one of clauses 20-21 further includes:

[0237] Receive third flag; and

[0238] In response to enabling or disabling control of the encoding mode at a level below the sequence level, the encoding mode for a target low-level region is enabled or disabled based on the third flag.

[0239] 23. The computer-implemented method according to Clause 22 further includes:

[0240] The third flag is stored in the slice header of the target slice, wherein the target slice is the target low-level region; or

[0241] The third flag is stored in the image header of the target image, wherein the target image is the target low-level region.

[0242] 24. The computer-implemented method according to any one of clauses 20-23, wherein the encoding mode is at least one of the following:

[0243] Bidirectional optical flow (BDOF) mode;

[0244] Optical flow prediction correction (PROF) mode; or

[0245] Decoder-side motion vector correction (DMVR) mode.

[0246] 25. The computer-implemented method according to any one of clauses 20-24 further includes:

[0247] The first flag is stored in the sequence parameter set (SPS) of the video bitstream.

[0248] 26. The computer-implemented method according to any one of clauses 20-25 further includes:

[0249] In response to enabling or disabling the encoding mode of the video bitstream, the second flag is stored in the SPS of the video bitstream.

[0250] 27. A computer-implemented method, comprising:

[0251] Receive a video sequence, a first flag, a second flag, and a third flag; enable or disable a first encoding mode for the video bitstream based on the first flag;

[0252] Based on the second flag, a second encoding mode for the video bitstream is enabled or disabled; and

[0253] Based on the third flag, control over at least one of the first encoding mode or the second encoding mode is enabled or disabled at a level below the sequence level.

[0254] 28. The computer implementation method according to Clause 27, wherein the level below the sequence level includes the slice level or the image level.

[0255] 29. The computer-implemented method according to any one of clauses 27-28 further includes:

[0256] Receive the fourth flag; and

[0257] In response to enabling control of the first encoding mode of the video bitstream based on the first flag and enabling control of at least one of the first encoding mode or the second encoding mode at a level below the sequence level based on the third flag, the first encoding mode for a target low-level region is enabled or disabled according to the fourth flag.

[0258] 30. The computer implementation method according to any one of clauses 27-28 further includes:

[0259] Accept the fourth sign; and

[0260] In response to enabling or disabling control of at least one of the first encoding mode or the second encoding mode at a level below the sequence level, the first encoding mode and the second encoding mode for a target low-level region are enabled or disabled based on the fourth flag.

[0261] 31. The computer implementation method according to any one of clauses 29-30 further includes:

[0262] The fourth flag is stored in the slice header of the target slice, wherein the target slice is the target low-level region; or

[0263] The fourth flag is stored in the image header of the target image, wherein the target image is a low-level target region.

[0264] 32. The computer-implemented method according to Clause 31 further includes:

[0265] Accept the fifth sign;

[0266] The third encoding mode for the video bitstream is enabled or disabled based on the second flag; and

[0267] Based on the fifth flag, control over the third encoding mode is enabled or disabled at the level below the sequence level.

[0268] 33. The computer-implemented method according to any one of clauses 31 and 32 further includes:

[0269] Accept the sixth sign; and

[0270] In response to enabling or disabling control of the third encoding mode at a level below the sequence level, the third encoding mode for the target low-level region is enabled or disabled based on the sixth flag.

[0271] 34. The computer implementation method according to any one of Clauses 27-33, wherein the first encoding mode, the second encoding mode and the third encoding mode are respectively a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode and a decoder-side motion vector correction (DMVR) mode.

[0272] 35. The computer implementation method according to any one of clauses 27-33, wherein the first encoding mode and the second encoding mode are selected from two different encoding modes:

[0273] Bidirectional optical flow (BDOF) mode;

[0274] Optical flow prediction correction (PROF) mode; and

[0275] Decoder-side motion vector refinement (DMVR) mode.

[0276] 36. The computer-implemented method according to any one of clauses 27-35, wherein the first encoding mode and the second encoding mode are respectively a bidirectional optical flow (BDOF) mode and an optical flow prediction correction (PROF) mode.

[0277] 37. The computer implementation method according to any one of clauses 27-36 further includes:

[0278] The first and second flags are stored in the sequence parameter set (SPS) of the video bitstream.

[0279] 38. The computer implementation method according to any one of clauses 27-37 further includes:

[0280] In response to enabling at least one of the first encoding mode or the second encoding mode for the video bitstream, the third flag is stored in the SPS of the video sequence.

[0281] 39. A non-transitory computer-readable medium storing an instruction set executable by at least one processor of a device to cause the device to perform a method comprising:

[0282] Receive video data bitstream;

[0283] Enabling or disabling the encoding mode for the video sequence based on a first flag in the bitstream; and

[0284] Based on the second flag in the bitstream, it is determined whether to enable or disable control over the encoding mode at a level below the sequence level.

[0285] 40. The non-transitory computer-readable medium as described in Clause 39, wherein the level below the sequence level includes the slice level or the image level.

[0286] 41. A non-transitory computer-readable medium according to any one of clauses 39-40, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0287] In response to enabling control of the encoding mode at a level below the sequence level, an encoding mode for a target low-level region is enabled or disabled based on a third flag in the bitstream.

[0288] 42. A non-transitory computer-readable medium as described in Clause 41, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0289] Detect the third flag in the slice header of the target slice, wherein the target slice is a low-level region of the target; or

[0290] The third marker is detected in the image header of the target image, wherein the target image is a low-level region of the target.

[0291] 43. A non-transitory computer-readable medium according to any one of clauses 39-42, wherein the encoding mode is at least one of the following:

[0292] Bidirectional optical flow (BDOF) mode;

[0293] Optical flow prediction correction (PROF) mode; or

[0294] Decoder-side motion vector correction (DMVR) mode.

[0295] 44. A non-transitory computer-readable medium according to any one of clauses 39-43, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0296] The first flag is detected in the sequence parameter set (SPS) of the video sequence.

[0297] 45. A non-transitory computer-readable medium according to any one of clauses 39-44, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0298] In response to enabling the encoding mode for the video sequence, the second flag in the SPS of the video sequence is detected.

[0299] 46. ​​A non-transitory computer-readable medium storing an instruction set executable by at least one processor of a device to cause the device to perform a method, the method comprising:

[0300] Receive video data bitstream;

[0301] A first encoding mode for the video sequence is enabled or disabled based on a first flag in the bitstream;

[0302] A second encoding mode for the video sequence is enabled or disabled based on a second flag in the bitstream; and

[0303] Based on a third flag in the bitstream, it is determined whether to enable control over at least one of the first encoding mode or the second encoding mode at a level below the sequence level.

[0304] 47. The non-transitory computer-readable medium as described in Clause 46, wherein the level below the sequence level includes the slice level or the image level.

[0305] 48. A non-transitory computer-readable medium according to any one of clauses 46-47, wherein the set of instructions executable by the at least one processor of the device causes the device to further perform:

[0306] In response to enabling control of at least one of the first encoding mode or the second encoding mode at a level below the sequence level, the first encoding mode and the second encoding mode are enabled or disabled for a target low-level region based on a fourth flag in the bitstream.

[0307] 49. A non-transitory computer-readable medium as described in Clause 48, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0308] Detect the fourth flag in the slice header of the target slice, wherein the target slice is a low-level region of the target; or

[0309] The fourth marker is detected in the image header of the target image, wherein the target image is the target low-level region.

[0310] 50. A non-transitory computer-readable medium according to any one of clauses 48-49, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0311] Based on the second flag in the bitstream, enable or disable the third encoding mode for the video sequence; and

[0312] Based on the fifth flag in the bitstream, it is determined whether to enable control of the third encoding mode at the level below the sequence level.

[0313] 51. A non-transitory computer-readable medium as described in Clause 50, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0314] In response to enabling the third coding mode at a level below the sequence level, the third coding mode for the target low-level region is enabled or disabled based on a sixth flag in the bitstream.

[0315] 52. A non-transitory computer-readable medium according to any one of clauses 50 and 51, wherein the first encoding mode, the second encoding mode and the third encoding mode are respectively a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode and a decoder-side motion vector correction (DMVR) mode.

[0316] 53. A non-transitory computer-readable medium according to any one of clauses 46-52, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0317] In response to a first flag indicating the activation of the first coding mode for the video sequence and a third flag indicating the activation of control over at least one of the first coding mode or the second coding mode at a level below the sequence level, the first coding mode for a target lower-level region is enabled or disabled based on a fourth flag in the bitstream.

[0318] 54. A non-transitory computer-readable medium according to any one of clauses 46-53, wherein the first encoding mode and the second encoding mode are selected from two different encoding modes:

[0319] Bidirectional optical flow (BDOF) mode;

[0320] Optical flow prediction correction (PROF) mode; and

[0321] Decoder-side motion vector correction (DMVR) mode.

[0322] 55. A non-transitory computer-readable medium according to any one of clauses 46-54, wherein the first encoding mode and the second encoding mode are a bidirectional optical flow (BDOF) mode and an optical flow prediction correction (PROF) mode, respectively.

[0323] 56. A non-transitory computer-readable medium according to any one of clauses 46-55, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0324] Detect the first and second flags in the sequence parameter set (SPS) of the video sequence.

[0325] 57. A non-transitory computer-readable medium according to any one of clauses 46-56, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0326] In response to enabling at least one of the first encoding mode or the second encoding mode for the video sequence, the third flag in the SPS of the video sequence is detected.

[0327] 58. A non-transitory computer-readable medium storing a set of instructions executable by at least one processor of a device to cause the device to perform a method, the method comprising:

[0328] Receive the video sequence, the first flag, and the second flag;

[0329] Enabling or disabling the encoding mode for the video bitstream based on the first flag; and

[0330] Based on the second flag, control over the encoding mode is enabled or disabled at a level below the sequence level.

[0331] 59. The non-transitory computer-readable medium as described in Clause 58, wherein the level below the sequence level includes the slice level or the image level.

[0332] 60. A non-transitory computer-readable medium according to any one of clauses 58-59, wherein a set of instructions executable by at least one processor of the device causes the device to further perform:

[0333] Receive third flag; and

[0334] In response to enabling or disabling control of the encoding mode at a level below the sequence level, the encoding mode for a target low-level region is enabled or disabled based on the third flag.

[0335] 61. A non-transitory computer-readable medium as described in Clause 60, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0336] The third flag is stored in the slice header of the target slice, wherein the target slice is the target low-level region; or

[0337] The third flag is stored in the image header of the target image, wherein the target image is the target low-level region.

[0338] 62. A non-transitory computer-readable medium according to any one of clauses 58-61, wherein the encoding mode is at least one of the following:

[0339] Bidirectional optical flow (BDOF) mode;

[0340] Optical flow prediction correction (PROF) mode; or

[0341] Decoder-side motion vector correction (DMVR) mode.

[0342] 63. A non-transitory computer-readable medium according to any one of clauses 58-62, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0343] The first flag is stored in the sequence parameter set (SPS) of the video bitstream.

[0344] 64. A non-transitory computer-readable medium according to any one of clauses 58-63, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0345] In response to enabling or disabling the encoding mode for the video bitstream, the second flag is stored in the SPS of the video bitstream.

[0346] 65. A non-transitory computer-readable medium storing a set of instructions executable by at least one processor of a device to cause the device to perform a method, the method comprising:

[0347] Receive a video sequence, a first flag, a second flag, and a third flag; enable or disable a first encoding mode for the video bitstream based on the first flag;

[0348] Based on the second flag, a second encoding mode for the video bitstream is enabled or disabled; and

[0349] Based on the third flag, control over at least one of the first encoding mode or the second encoding mode is enabled or disabled at a level below the sequence level.

[0350] 66. The non-transitory computer-readable medium as described in Clause 65, wherein the level below the sequence level includes the slice level or the image level.

[0351] 67. A non-transitory computer-readable medium according to any one of clauses 65-66, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0352] Receive the fourth flag; and

[0353] In response to control enabling the first encoding mode for the video bitstream based on the first flag and enabling control of at least one of the first encoding mode or the second encoding mode at a level below the sequence level based on the third flag, the first encoding mode for a target low-level region is enabled or disabled based on the fourth flag.

[0354] 68. A non-transitory computer-readable medium according to any one of clauses 65-66, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0355] Receive the fourth flag; and

[0356] In response to enabling or disabling control of at least one of the first encoding mode or the second encoding mode at a level below the sequence level, the first encoding mode and the second encoding mode for a target low-level region are enabled or disabled based on the fourth flag.

[0357] 69. A non-transitory computer-readable medium according to any one of clauses 67-68, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0358] The fourth flag is stored in the slice header of the target slice, wherein the target slice is the target low-level region; or

[0359] The fourth flag is stored in the image header of the target image, wherein the target image is the target low-level region.

[0360] 70. A non-transitory computer-readable medium as described in Clause 69, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0361] Receive the fifth flag;

[0362] The third encoding mode for the video bitstream is enabled or disabled based on the second flag; and

[0363] Based on the fifth flag, control over the third encoding mode is enabled or disabled at the level below the sequence level.

[0364] 71. A non-transitory computer-readable medium according to any one of clauses 69 and 70, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0365] Receive the sixth flag; and

[0366] In response to enabling or disabling control of the third encoding mode at a level below the sequence level, the third encoding mode for the target low-level region is enabled or disabled based on the sixth flag.

[0367] 72. A non-transitory computer-readable medium according to any one of clauses 65-71, wherein the first encoding mode, the second encoding mode and the third encoding mode are respectively a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode and a decoder-side motion vector correction (DMVR) mode.

[0368] 73. A non-transitory computer-readable medium according to any one of clauses 65-72, wherein the first encoding mode and the second encoding mode are selected from two different encoding modes:

[0369] Bidirectional optical flow (BDOF) mode;

[0370] Optical flow prediction correction (PROF) mode; and

[0371] Decoder-side motion vector refinement (DMVR) mode.

[0372] 74. A non-transitory computer-readable medium according to any one of clauses 65-73, wherein the first encoding mode and the second encoding mode are a bidirectional optical flow (BDOF) mode and an optical flow prediction correction (PROF) mode, respectively.

[0373] 75. A non-transitory computer-readable medium according to any one of clauses 65-74, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0374] The first and second flags are stored in the sequence parameter set (SPS) of the video bitstream.

[0375] 76. A non-transitory computer-readable medium according to any one of clauses 65-75, wherein a set of instructions executable by the at least one processor of the device causes the device to further perform:

[0376] In response to enabling at least one of the first encoding mode or the second encoding mode for the video bitstream, the third flag is stored in the SPS of the video sequence.

[0377] 77. An apparatus comprising:

[0378] Memory, configured to store instruction sets; and

[0379] One or more processors, which are communicatively coupled to the memory and configured to execute the instruction set to enable the device to:

[0380] Receive video data bitstream;

[0381] Enabling or disabling the encoding mode for the video sequence based on a first flag in the bitstream; and

[0382] Based on the second flag in the bitstream, it is determined whether to enable or disable control over the encoding mode at a level below the sequence level.

[0383] 78. The apparatus according to Clause 77, wherein the level below the sequence level includes the slice level or the image level.

[0384] 79. The apparatus according to any one of clauses 77-78, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0385] In response to enabling control of the encoding mode at a level below the sequence level, an encoding mode for a target low-level region is enabled or disabled based on a third flag in the bitstream.

[0386] 80. The apparatus according to clause 79, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0387] Detect the third flag in the slice header of the target slice, wherein the target slice is a low-level region of the target; or

[0388] The third marker is detected in the image header of the target image, wherein the target image is a low-level region of the target.

[0389] 81. The apparatus according to any one of clauses 77-80, wherein the encoding mode is at least one of the following:

[0390] Bidirectional optical flow (BDOF) mode;

[0391] Optical flow prediction correction (PROF) mode; or

[0392] Decoder-side motion vector correction (DMVR) mode.

[0393] 82. The apparatus according to any one of clauses 77-81, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0394] The first flag is detected in the sequence parameter set (SPS) of the video sequence.

[0395] 83. The apparatus according to any one of clauses 77-82, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0396] In response to enabling the encoding mode for the video sequence, the second flag in the SPS of the video sequence is detected.

[0397] 84. An apparatus comprising:

[0398] Memory, configured to store instruction sets; and

[0399] One or more processors, which are communicatively coupled to the memory and configured to execute the instruction set to enable the device to:

[0400] Receive video data bitstream;

[0401] A first encoding mode for the video sequence is enabled or disabled based on a first flag in the bitstream;

[0402] A second encoding mode for the video sequence is enabled or disabled based on a second flag in the bitstream; and

[0403] Based on a third flag in the bitstream, it is determined whether to enable control over at least one of the first encoding mode or the second encoding mode at a level below the sequence level.

[0404] 85. The apparatus according to Clause 84, wherein the level below the sequence level includes a slice level or an image level.

[0405] 86. The apparatus according to any one of clauses 84-85, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0406] In response to enabling control of at least one of the first encoding mode or the second encoding mode at a level below the sequence level, the first encoding mode and the second encoding mode are enabled or disabled for a target low-level region based on a fourth flag in the bitstream.

[0407] 87. The apparatus according to clause 86, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0408] Detect the fourth flag in the slice header of the target slice, wherein the target slice is a low-level region of the target; or

[0409] The fourth marker is detected in the image header of the target image, wherein the target image is the target low-level region.

[0410] 88. The apparatus according to any one of clauses 86-87, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0411] Based on the second flag in the bitstream, enable or disable the third encoding mode for the video sequence; and

[0412] Based on the fifth flag in the bitstream, it is determined whether to enable control of the third encoding mode at the level below the sequence level.

[0413] 89. The apparatus according to clause 88, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0414] In response to enabling the third coding mode at a level below the sequence level, the third coding mode for the target low-level region is enabled or disabled based on a sixth flag in the bitstream.

[0415] 90. The apparatus according to any one of clauses 88 and 89, wherein the first encoding mode, the second encoding mode and the third encoding mode are respectively a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode and a decoder-side motion vector correction (DMVR) mode.

[0416] 91. The apparatus according to any one of clauses 84-90, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0417] In response to a first flag indicating the activation of the first coding mode for the video sequence and a third flag indicating the activation of control over at least one of the first coding mode or the second coding mode at a level below the sequence level, the first coding mode for a target lower-level region is enabled or disabled based on a fourth flag in the bitstream.

[0418] 92. The apparatus according to any one of clauses 84-91, wherein the first encoding mode and the second encoding mode are selected from two different encoding modes:

[0419] Bidirectional optical flow (BDOF) mode;

[0420] Optical flow prediction correction (PROF) mode; and

[0421] Decoder-side motion vector correction (DMVR) mode.

[0422] 93. The apparatus according to any one of clauses 84-92, wherein the first encoding mode and the second encoding mode are a bidirectional optical flow (BDOF) mode and an optical flow prediction correction (PROF) mode, respectively.

[0423] 94. The apparatus according to any one of clauses 84-93, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0424] Detect the first and second flags in the sequence parameter set (SPS) of the video sequence.

[0425] 95. The apparatus according to any one of clauses 84-94, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0426] In response to enabling at least one of a first encoding mode or a second encoding mode for the video sequence, the third flag in the SPS of the video sequence is detected.

[0427] 96. An apparatus comprising:

[0428] Memory, configured to store instruction sets; and

[0429] One or more processors, communicatively coupled to the aforementioned memory and configured to execute the instruction set to enable the device to:

[0430] Receive the video sequence, the first flag, and the second flag;

[0431] Enabling or disabling the encoding mode for the video bitstream based on the first flag; and

[0432] Based on the second flag, control over the encoding mode is enabled or disabled at a level below the sequence level.

[0433] 97. The apparatus according to Clause 96, wherein the level below the sequence level includes a slice level or an image level.

[0434] 98. The apparatus according to any one of clauses 96-97, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0435] Receive the third flag; and

[0436] In response to enabling or disabling control of the encoding mode at a level below the sequence level, the encoding mode for a target low-level region is enabled or disabled based on the third flag.

[0437] 99. The apparatus according to clause 98, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0438] The third flag is stored in the slice header of the target slice, wherein the target slice is the target low-level region; or

[0439] The third flag is stored in the image header of the target image, wherein the target image is the target low-level region.

[0440] 100. The apparatus according to any one of clauses 96-99, wherein the encoding mode is at least one of the following:

[0441] Bidirectional optical flow (BDOF) mode;

[0442] Optical flow prediction correction (PROF) mode; or

[0443] Decoder-side motion vector correction (DMVR) mode.

[0444] 101. The apparatus according to any one of clauses 96-100, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0445] The first flag is stored in the sequence parameter set (SPS) of the video bitstream.

[0446] 102. The apparatus according to any one of clauses 96-101, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0447] In response to enabling or disabling the encoding mode for the video bitstream, the second flag is stored in the SPS of the video bitstream.

[0448] 103. An apparatus comprising:

[0449] Memory, configured to store instruction sets; and

[0450] One or more processors, communicatively coupled to the aforementioned memory and configured to execute the instruction set to enable the device to:

[0451] Receive video sequence, first flag, second flag, and third flag;

[0452] The first encoding mode for the video bitstream is enabled or disabled based on the first flag.

[0453] Based on the second flag, a second encoding mode for the video bitstream is enabled or disabled; and

[0454] Based on the third flag, control over at least one of the first encoding mode or the second encoding mode is enabled or disabled at a level below the sequence level.

[0455] 104. The apparatus according to clause 102, wherein the level below the sequence level includes a slice level or an image level.

[0456] 105. The apparatus according to any one of clauses 103-104, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0457] Receive the fourth flag; and

[0458] In response to control enabling the first encoding mode for the video bitstream based on the first flag and enabling control of at least one of the first encoding mode or the second encoding mode at a level below the sequence level based on the third flag, the first encoding mode for a target low-level region is enabled or disabled based on the fourth flag.

[0459] 106. The apparatus according to any one of clauses 103-104, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0460] Receive the fourth flag; and

[0461] In response to enabling or disabling control of at least one of the first encoding mode or the second encoding mode at a level below the sequence level, the first encoding mode and the second encoding mode for a target low-level region are enabled or disabled based on the fourth flag.

[0462] 107. The apparatus according to any one of clauses 105-106, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0463] The fourth flag is stored in the slice header of the target slice, wherein the target slice is the target low-level region; or

[0464] The fourth flag is stored in the image header of the target image, wherein the target image is the target low-level region.

[0465] 108. The apparatus according to clause 107, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0466] Receive the fifth flag;

[0467] The third encoding mode for the video bitstream is enabled or disabled based on the second flag; and

[0468] Based on the fifth flag, control over the third encoding mode is enabled or disabled at the level below the sequence level.

[0469] 109. The apparatus according to any one of clauses 107 and 108, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0470] Receive the sixth flag; and

[0471] In response to enabling or disabling control of the third encoding mode at a level below the sequence level, the third encoding mode for the target low-level region is enabled or disabled based on the sixth flag.

[0472] 110. The apparatus according to any one of clauses 103-109, wherein the first encoding mode, the second encoding mode and the third encoding mode are respectively a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode and a decoder-side motion vector correction (DMVR) mode.

[0473] 111. The apparatus according to any one of clauses 103-110, wherein the first encoding mode and the second encoding mode are selected from two different encoding modes:

[0474] Bidirectional optical flow (BDOF) mode;

[0475] Optical flow prediction correction (PROF) mode; and

[0476] Decoder-side motion vector refinement (DMVR) mode.

[0477] 112. The apparatus according to any one of clauses 103-110, wherein the first encoding mode and the second encoding mode are a bidirectional optical flow (BDOF) mode and an optical flow prediction correction (PROF) mode, respectively.

[0478] 113. The apparatus according to any one of clauses 103-112, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0479] The first and second flags are stored in the sequence parameter set (SPS) of the video bitstream.

[0480] 114. The apparatus according to any one of clauses 103-113, wherein the one or more processors are further configured to execute the instruction set to cause the apparatus to:

[0481] In response to enabling at least one of the first encoding mode or the second encoding mode for the video bitstream, the third flag is stored in the SPS of the video sequence.

[0482] It should be noted that the relational terms such as "first" and "second" in this document are used only to distinguish one entity or operation from another, and do not require or imply any actual relationship or order between these entities or operations. Furthermore, words such as "contains," "has," "includes," and "includes," as well as other similar forms, are intended to convey the same meaning and are open-ended, as one or more items following any of these words are not intended to be an exhaustive list of such items or items, or to be limited to the listed items or items.

[0483] As used herein, unless otherwise expressly stated, the term "or" covers all possible combinations unless impractical. For example, if a component is declared to include A or B, then unless otherwise expressly stated or impractical, the component may include A, or B, or A and B. As a second example, if a component is declared to include A, B, or C, then unless otherwise expressly stated or impractical, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.

[0484] It is understood that the above embodiments can be implemented by hardware, software (program code), or a combination of hardware and software. If implemented by software, it can be stored in the above-described computer-readable medium. When executed by a processor, the software can perform the disclosed methods. The computing units and other functional units described in this invention can be implemented by hardware, by software, or by a combination of hardware and software. Those skilled in the art will also understand that multiple of the above modules / units can be combined into one module / unit, and each of the above modules / units can be further divided into multiple sub-modules / sub-units.

[0485] In the foregoing specification, embodiments have been described with reference to numerous specific details, which may vary depending on the implementation. Certain modifications and changes may be made to the described embodiments. Other embodiments will be apparent to those skilled in the art in consideration of the specification and practice of the invention disclosed herein. The specification and examples are to be considered as examples only, and the true scope and spirit of the invention are indicated by the appended claims. The sequence of steps shown in the figures is also intended for illustrative purposes only and is not intended to limit to any particular order of steps. Therefore, those skilled in the art will understand that these steps may be performed in a different order in carrying out the same method.

[0486] Example embodiments have been disclosed in the accompanying drawings and description. However, many variations and modifications can be made to these embodiments. Therefore, although specific terms are used, they are for general and descriptive purposes only and not for limiting purposes.

Claims

1. A computer-implemented decoding method, comprising: Receive the bitstream associated with the video sequence; Decode the first flag and the second flag in the Sequence Parameter Set (SPS) of the bitstream, wherein the first flag indicates that at least one of a plurality of encoding modes is enabled or disabled at the sequence level; and the second flag indicates whether the encoding mode enabled at the sequence level is controlled at a lower level below the sequence level. Based on the value of the second flag, determine whether a third flag exists in the bit stream; wherein, the third flag indicates whether the encoding mode enabled at the sequence level is enabled or disabled at a lower level below the sequence level; The various coding modes include optical flow prediction correction modes; When the third flag is present in the bit stream, the bit stream is decoded based on the value of the third flag.

2. The computer-implemented decoding method according to claim 1, wherein the lower level below the sequence level is a stripe level or an image level.

3. The computer-implemented decoding method according to claim 1, further comprising: In response to the second flag having a first value, it is determined that the third flag exists in the strip header or image header of the bit stream.

4. The computer-implemented decoding method according to claim 3, wherein the first value is 1.

5. A computer-implemented encoding method, comprising: A first flag and a second flag in the Sequence Parameter Set (SPS) of the bitstream associated with the video sequence are encoded, wherein the first flag indicates that at least one of a plurality of coding modes is enabled or disabled at the sequence level; and the second flag indicates whether the coding mode enabled at the sequence level is controlled at a lower level below the sequence level. Based on the value of the second flag, determine whether to send the third flag by signal in the bit stream; The third flag indicates whether the encoding mode enabled at the sequence level is enabled or disabled at a lower level below the sequence level; The various coding modes include optical flow prediction correction modes; When it is determined that the third flag in the bit stream is to be sent by signal, the bit stream is encoded based on the value of the third flag.

6. The computer-implemented encoding method according to claim 5, wherein, The lower level below the sequence level is either the strip level or the image level.

7. The computer-implemented encoding method according to claim 5, further comprising: In response to the second flag having a first value, encoding is performed in the strip header or image header of the bit stream based on the third flag.

8. The computer-implemented encoding method according to claim 7, wherein the first value is 1.

9. A non-transitory computer-readable storage medium storing a bitstream associated with a video sequence, the bitstream being generated by a method performed by a video processing apparatus, the method comprising: A first flag and a second flag in the Sequence Parameter Set (SPS) of the bitstream associated with the video sequence are encoded, wherein the first flag indicates that at least one of a plurality of coding modes is enabled or disabled at the sequence level; and the second flag indicates whether the coding mode enabled at the sequence level is controlled at a lower level below the sequence level. Based on the value of the second flag, determine whether to send the third flag by signal in the bit stream; The third flag indicates whether the encoding mode enabled at the sequence level is enabled or disabled at a lower level below the sequence level; The various coding modes include optical flow prediction correction modes; When it is determined that the third flag in the bit stream is to be sent by signal, the bit stream is encoded based on the value of the third flag.

10. The non-transitory computer-readable storage medium of claim 9, wherein the lower level below the sequence level is a stripe level or an image level.

11. The non-transitory computer-readable storage medium of claim 9, wherein the method further comprises: In response to the second flag having a first value, encoding is performed in the strip header or image header of the bit stream based on the third flag.

12. The non-transitory computer-readable storage medium according to claim 11, wherein, The first value is 1.