Computer-implemented encoding and decoding method, decoding device and computer readable storage medium

By receiving flags in the video data bitstream and dynamically managing the encoding mode of the video sequence, the problem of inflexible encoding mode adjustment in the prior art is solved, and more efficient and flexible video encoding is achieved.

CN120128702APending Publication Date: 2025-06-10HFI INNOVATION INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510428488.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2019-09-12
Filing Date
2020-08-20
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing video encoding technology lacks flexibility in controlling the encoding mode, and it is difficult to dynamically adjust the encoding mode according to the needs of different video sequences, resulting in poor encoding efficiency and quality.

Method used

Flexible management of encoding mode is achieved by receiving flags in the bitstream of video data, dynamically enabling or disabling the encoding mode of the video sequence, and enabling or disabling control of the encoding mode below the sequence level (such as the slice level or the image level).

Benefits of technology

It improves the flexibility and efficiency of video encoding, and can dynamically adjust the encoding mode according to the needs of specific video sequences, thereby improving encoding quality and reducing resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128702A_ABST
    Figure CN120128702A_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and apparatus for controlling an encoding mode for video data. The method and apparatus include receiving a bitstream of video data; enabling or disabling a coding mode for a video sequence based on a first flag in the bitstream; and determining whether to enable or disable control of the encoding mode at a level lower than a sequence level based on a second flag in the bitstream.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - Reference to Related Applications

[0001] This disclosure claims priority to U.S. Provisional Application No. 62 / 899,169, filed on September 12, 2019, the entire content of which is incorporated herein by reference. Background Art

[0002] Video is a set of static images (or "frames") that capture visual information. To reduce storage memory and transmission bandwidth, video can be compressed before storage or transmission and then decompressed before display. The compression process is generally referred to as encoding, and the decompression process is generally referred to as decoding. There are various video coding formats using standardized video coding techniques, and the most common ones are based on prediction, transformation, quantization, entropy coding, and loop filtering. Video coding standards, such as the High Efficiency Video Coding (HEVC / H.265) standard, the Versatile Video Coding (VVC / H.266) standard, and the AVS standard, specify specific video coding formats and are developed by standardization organizations. As more and more advanced video coding techniques are adopted in video standards, the coding efficiency of new video coding standards is also getting higher and higher. Summary of the Invention

[0003] Embodiments of the present invention provide a method and an apparatus for controlling an encoding mode for video data. In one example embodiment, a method includes: receiving a bitstream of video data; enabling or disabling an encoding mode for a video sequence based on a first flag in the bitstream; and determining, based on a second flag in the bitstream, whether to enable or disable control of the encoding mode at a level lower than the sequence level.

[0004] In another example embodiment, a method includes: receiving a bitstream of video data; enabling or disabling a first encoding mode for a video sequence based on a first flag in the bitstream; enabling or disabling a second encoding mode for the video sequence based on a second flag in the bitstream; and determining, based on a third flag in the bitstream, whether to enable control of at least one of the first encoding mode or the second encoding mode at a level lower than the sequence level.

[0005] In another example embodiment, a method includes: receiving a video sequence, a first flag, and a second flag; enabling or disabling an encoding mode for a video bitstream based on the first flag; and enabling or disabling control of the encoding mode at a level lower than the sequence level based on the second flag.

[0006] In another example embodiment, a method includes: receiving a video sequence, a first flag, a second flag, and a third flag; enabling or disabling a first coding mode for a video bitstream based on the first flag; enabling or disabling a second coding mode for the video bitstream based on the second flag; and enabling or disabling control of at least one of the first coding mode or the second coding mode at a level below the sequence level based on the third flag.

[0007] In another example embodiment, a non-transitory computer-readable medium stores a set of instructions executable by at least one processor of a device to cause the device to perform a method that includes: receiving a video data bitstream; enabling or disabling a coding mode for a video sequence based on a first flag in the bitstream; and determining whether to enable or disable control of the coding mode at a level below the sequence level based on a second flag in the bitstream.

[0008] In another example embodiment, a non-transitory computer-readable medium stores a set of instructions executable by at least one processor of a device to cause the device to perform a method that includes: receiving a video data bitstream; enabling or disabling a first coding mode for a video sequence based on a first flag in the bitstream; enabling or disabling a second coding mode for the video sequence based on a second flag in the bitstream; and determining whether to enable control of at least one of the first coding mode or the second coding mode at a level below the sequence level based on a third flag in the bitstream.

[0009] In another embodiment, a device includes a memory configured to store a set of instructions and one or more processors communicatively coupled to the memory, the one or more processors configured to execute the set of instructions to cause the device to: receive a bitstream of video data; enable or disable a coding mode for a video sequence based on a first flag in the bitstream; and determine whether to enable or disable control of the coding mode at a level below the sequence level based on a second flag in the bitstream.

[0010] In another embodiment, a device includes a memory configured to store a set of instructions and one or more processors communicatively coupled to the memory, the one or more processors configured to execute the set of instructions to cause the device to: receive a bitstream of video data; enable or disable a first coding mode for a video sequence based on a first flag in the bitstream; enable or disable a second coding mode for the video sequence based on a second flag in the bitstream; and determine whether to enable control of at least one of the first coding mode or the second coding mode at a level below the sequence level based on a third flag in the bitstream. Description of the Drawings

[0011] Embodiments and various aspects of the present disclosure are illustrated in the following detailed description and the accompanying drawings. The various features shown in the drawings are not drawn to scale.

[0012] Figure 1 is a schematic diagram showing the structure of an example video sequence according to some embodiments of the present disclosure.

[0013] Figure 2A is a schematic diagram showing an example encoding process of a hybrid video coding system consistent with embodiments of the present disclosure.

[0014] Figure 2B is a schematic diagram showing another example encoding process of a hybrid video coding system consistent with embodiments of the present disclosure.

[0015] Figure 3A is a schematic diagram showing an example decoding process of a hybrid video coding system consistent with embodiments of the present disclosure.

[0016] Figure 3B is a schematic diagram showing another example decoding process of a hybrid video coding system consistent with embodiments of the present disclosure.

[0017] Figure 4 is a block diagram showing an example apparatus for encoding or decoding video according to some embodiments of the present disclosure.

[0018] Figure 5 is a schematic diagram showing an example process of decoder-side motion vector refinement (DMVR) according to some embodiments of the present disclosure.

[0019] Figure 6 is a schematic diagram showing an example DMVR search process according to some embodiments of the present disclosure.

[0020] Figure 7 is a schematic diagram illustrating an example pattern for DMVR integer luminance sample search according to some embodiments of the present disclosure.

[0021] Figure 8 is a schematic diagram showing another example pattern of the stage of integer sample offset search in DMVR integer luminance sample search according to some embodiments of the present disclosure.

[0022] Figure 9 is a schematic diagram illustrating an example pattern for estimating the DMVR parameter error surface according to some embodiments of the present disclosure.

[0023] Figure 10 is a schematic diagram of an example of an extended coding unit (CU) region used in bidirectional optical flow (BDOF) according to some embodiments of the present disclosure.

[0024] Figure 11is a schematic diagram of examples of sub-block-based affine motion and sample-based affine motion according to some embodiments of the present disclosure.

[0025] Figure 12 Table 1 is shown, which shows an example syntax structure of a sequence parameter set (SPS) implementing control flags for DMVR and BDOF according to some embodiments of the present disclosure.

[0026] Figure 13 Table 2 is shown, which shows an example syntax structure of a slice header implementing control flags for DMVR and BDOF according to some embodiments of the present disclosure.

[0027] Figure 14A Table 3A is shown, which shows an example syntax structure of an SPS that implements slice level control flags for DMVR, BDOF, and optical flow prediction correction (PROF) according to some embodiments of the present disclosure.

[0028] Figure 14B Table 3B is shown, which shows an example syntax structure of an SPS implementing image level control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.

[0029] Figure 15A Table 4A is shown, which shows an example syntax structure of a silce header implementing control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.

[0030] Figure 15B Table 4B is shown, which shows an example syntax structure of an image header implementing control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.

[0031] Figure 16 Table 5 is shown, which shows an example syntax structure of an SPS implementing separate sequence level control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.

[0032] Figure 17 Table 6 is shown, which shows an example syntax structure of a slice head implementing a joint control flag for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.

[0033] Figure 18 Table 7 is shown, which shows an example syntax structure of a slice head that implements separate control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.

[0034] Figure 19 Table 8 is shown, which shows an example syntax structure of a slice head implementing hybrid control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.

[0035] Figure 20 Table 9 is shown, which shows an example syntax structure of an SPS implementing hybrid sequence level control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.

[0036] Figure 21 Table 10 is shown, which shows another example syntax structure of a slice head implementing hybrid control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.

[0037] Figure 22 Table 11 is shown, which shows another example syntax structure of a slice head implementing separate control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure.

[0038] Figure 23 A flowchart is shown of an example process for controlling a video decoding mode according to some embodiments of the present disclosure.

[0039] Figure 24 A flowchart is shown of another example process for controlling a video decoding mode according to some embodiments of the present disclosure.

[0040] Figure 25 A flowchart is shown of an example process for controlling a video encoding mode according to some embodiments of the present disclosure.

[0041] Figure 26 A flowchart is shown of another example process for controlling a video encoding mode according to some embodiments of the present disclosure. DETAILED DESCRIPTION

[0042] Reference can now be made in detail to example embodiments, examples of which are shown in the accompanying drawings. The following description refers to the accompanying drawings, in which the same numbers in different drawings represent the same or similar elements unless otherwise specified. The embodiments set forth in the following description of example embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with aspects related to the present invention as described in the appended claims. Specific aspects of the present disclosure are described in more detail below. In the event of a conflict with terms and / or definitions incorporated by reference, the terms and definitions provided herein shall prevail.

[0043] The Joint Video Experts Group (JVET) of the ITU-T Video Coding Experts Group (ITU-T VCEG) and the ISO / IEC Moving Picture Experts Group (ISO / IEC MPEG) is currently developing the Versatile Video Coding (VVC / H.266) standard. The VVC standard aims to double the compression efficiency of its predecessor, the High Efficiency Video Coding (HEVC / H.265) standard. In other words, VVC aims to achieve the same subjective quality as HEVC / H.265 using half the bandwidth.

[0044] In order to achieve the same subjective quality as HEVC / H.265 using half the bandwidth, JVET has been developing technologies beyond HEVC using the Joint Exploration Model (JEM) reference software. As coding technologies are incorporated into JEM, JEM achieves higher coding performance than HEVC.

[0045] The VVC standard was developed recently and continues to include more coding techniques that provide better compression performance. VVC is based on the same hybrid video coding system used in modern video compression standards such as HEVC, H.264 / AVC, MPEG2, H.263, etc.

[0046] Video is a set of static images (or "frames") arranged in time sequence to store visual information. Video acquisition devices (e.g., cameras) can be used to acquire and store these images in time sequence, and video playback devices (e.g., televisions, computers, smartphones, tablets, video players, or any end-user terminal with display capabilities) can be used to display such images in time sequence. In addition, in some applications, video acquisition devices can transmit the acquired video to video playback devices (e.g., computers with monitors) in real time, such as for monitoring, conferencing, or live broadcasting.

[0047] In order to reduce the storage space and transmission bandwidth required for such applications, the video can be compressed before storage and transmission, and decompressed before display. Compression and decompression can be implemented by software executed by a processor (e.g., a processor of a general-purpose computer) or dedicated hardware. The module for compression is generally referred to as an "encoder", and the module for decompression is generally referred to as a "decoder". Encoders and decoders can be collectively referred to as "codecs". Encoders and decoders can be implemented as any of a variety of suitable hardware, software, or combinations thereof. For example, the hardware implementation of encoders and decoders can include circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, or any combination thereof. The software implementation of encoders and decoders can include program code, computer executable instructions, firmware, or any suitable computer-implemented algorithm or process fixed in a computer-readable medium. Video compression and decompression can be implemented by various algorithms or standards, such as MPEG-1, MPEG-2, MPEG-4, H.26x series, etc. In some applications, a codec can decompress a video from a first coding standard and recompress the decompressed video using a second coding standard, in which case the codec can be referred to as a "transcoder".

[0048] The video encoding process can identify and retain useful information that can be used to reconstruct the image, while ignoring unimportant information for reconstruction. If the ignored, unimportant information cannot be fully reconstructed, the encoding process can be called "lossy". Otherwise, it can be called "lossless". Most encoding processes are lossy, which is a trade-off made to reduce the required storage space and transmission bandwidth.

[0049] Useful information about the image being encoded (referred to as the "current image") includes changes relative to a reference image (e.g., a previously encoded and reconstructed image). Such changes can include changes in pixel position, brightness, or color, with changes in position being of greatest interest. Changes in the position of a group of pixels representing an object can reflect the motion of the object between the reference image and the current image.

[0050] A picture that is encoded without reference to another picture (i.e., it is its own reference picture) is called an "I-picture." A picture that is encoded using a previous picture as a reference picture is called a "P-picture." A picture that is encoded using both a previous picture and a future picture as reference pictures (i.e., the reference is "bidirectional") is called a "B-picture."

[0051] Figure 1The structure of an example video sequence 100 according to some embodiments of the present disclosure is illustrated. The video sequence 100 may be a live video or a video that has been captured and archived. The video 100 may be a real video, a computer-generated video (e.g., a computer game video), or a combination thereof (e.g., a real video with an augmented reality effect). The video sequence 100 may be obtained from a video capture device (e.g., a camera), a video archive containing previously captured videos (e.g., a video file stored in a storage device), or a video providing interface (e.g., a video broadcast transceiver) that receives video input from a video content provider.

[0052] like Figure 1 As shown, video sequence 100 may include a series of images arranged in time along a time axis, including images 102, 104, 106, and 108. Images 102-106 are consecutive, and there are more images between images 106 and 108. Figure 1 , image 102 is an I image, and its reference image is image 102 itself. Image 104 is a P image, and its reference image is image 102, as indicated by the arrow. Image 106 is a B image, and its reference images are images 104 and 108, as indicated by the arrow. In some embodiments, the reference image of an image (e.g., image 104) may not be immediately before or after the image. For example, the reference image of image 104 may be an image before image 102. It should be noted that the reference images of images 102-106 are only examples, and the present disclosure does not limit the embodiments of the reference images to those in Figure 1 The example shown in .

[0053] Typically, video codecs do not encode or decode the entire image at once due to the computational complexity of the encoding and decoding tasks. Instead, they may divide the image into basic segments and encode or decode the image segment by segment. Such basic segments are referred to as basic processing units ("BPUs") in this disclosure. For example, Figure 1Structure 110 in shows an example structure of an image (e.g., any of images 102-108) of video sequence 100. In structure 110, the image is divided into 4×4 basic processing units, whose boundaries are shown as dashed lines. In some embodiments, the basic processing unit may be referred to as a "macroblock" in some video coding standards (e.g., the MPEG family, H.261, H.263, or H.264 / AVC), or as a "coding tree unit" ("CTU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing units in an image may have different sizes, such as 128×128, 64×64, 32×32, 16×16, 4×8, 16×32, or pixels of arbitrary shapes and sizes. The size and shape of the basic processing unit for an image may be selected based on a balance between coding efficiency and the level of detail to be retained in the basic processing unit.

[0054] A basic processing unit may be a logical unit that may include a set of different types of video data stored in a computer memory (e.g., in a video frame buffer). For example, a basic processing unit for a color image may include a luma component (Y) representing colorless brightness information, one or more chroma components (e.g., Cb and Cr) representing color information, and associated syntax elements, where luma and chroma components may have basic processing units of the same size. In some video coding standards (e.g., H.265 / HEVC or H.266 / VVC), luma and chroma components may be referred to as "coding tree blocks" ("CTBs"). Any operation performed on a basic processing unit may be repeated for each of its luma and chroma components.

[0055] Video encoding has several stages of operation, examples of which are given in Figures 2A - 2B and Figures 3A - 3BAs shown in . For each stage, the size of the basic processing unit may still be too large to be processed, so it can be further divided into segments, referred to as "basic processing subunits" in this disclosure. In some embodiments, the basic processing subunit may be called a "block" in some video coding standards (e.g., MPEG series, H.261, H.263, or H.264 / AVC), or a "coding unit" ("CU") in some other video coding standards (e.g., H.265 / HEVC or H.266 / VVC). The basic processing subunit may have the same size as the basic processing unit or a smaller size than the basic processing subunit. Similar to the basic processing unit, the basic processing subunit is also a logical unit, which may include a set of different types of video data (e.g., Y, Cb, Cr, and related syntax elements) stored in a computer memory (e.g., in a video frame buffer). Any operation performed on the basic processing subunit can be repeated for each of its luminance and chrominance components. It should be noted that this division can be performed to a further level according to processing needs. It should also be noted that different stages can use different schemes to divide the basic processing units.

[0056] For example, in the mode decision phase (an example of which is in Figure 2B ), the encoder can decide what prediction mode to use for the basic processing unit (e.g., intra-image prediction or inter-image prediction), and the basic processing unit may be too large to make this decision. The encoder can split the basic processing unit into multiple basic processing sub-units (e.g., CUs in H.265 / HEVC or H.266 / VVC) and determine the prediction type for each individual basic processing sub-unit.

[0057] For example, in the prediction stage (the example is Figures 2A - 2B As shown in Figure 1, the encoder can perform prediction operations at the basic processing sub-unit (e.g., CU) level. However, in some cases, the basic processing sub-unit may still be too large to process. The encoder can further split the basic processing sub-unit into smaller fragments (e.g., called "prediction blocks" or "PBs" in H.265 / HEVC or H.266 / VVC), and prediction operations can be performed at this level.

[0058] For another example, in the transformation phase (the example in Figures 2A - 2B). The encoder can perform transform operations on the residual basic processing sub-unit (e.g., CU). However, in some cases, the basic processing sub-unit may still be too large to process. The encoder can further divide the basic processing sub-unit into smaller segments (e.g., called "transform blocks" or "TBs" in H.265 / HEVC or H.266 / VVC), and the transform operation can be performed at this level. It should be noted that the division scheme of the same basic processing sub-unit can be different in the prediction stage and the transform stage. For example, in H.265 / HEVC or H.266 / VVC, the prediction blocks and transform blocks of the same CU can have different sizes and numbers.

[0059] exist Figure 1 In the structure 110, the basic processing unit 112 is further divided into 3×3 basic processing sub-units, whose boundaries are indicated by dotted lines. Different basic processing units of the same image can be divided into basic processing sub-units of different schemes.

[0060] In some embodiments, in order to provide parallel processing and fault tolerance for video encoding and decoding, the image can be divided into multiple regions for processing, so that for one region of the image, the encoding or decoding process does not have to depend on information from any other region of the image. In other words, each region of the image can be processed independently. By doing so, the codec can process different regions of the image in parallel, thereby improving coding efficiency. In addition, when the data of one region is damaged during processing or lost in network transmission, the codec can correctly encode or decode other regions of the same image without relying on the damaged or lost data, thereby providing fault tolerance. In some video coding standards, the image can be divided into different types of regions. For example, H.265 / HEVC and H.266 / VVC provide two types of regions: "slices" and "tiles". It should also be noted that different images of the video sequence 100 can have different partitioning schemes for dividing the image into multiple regions.

[0061] For example, in Figure 1 In FIG. 1 , structure 110 is divided into three regions 114, 116, and 118, whose boundaries are shown as solid lines within structure 110. Region 114 includes four basic processing units. Each of regions 116 and 118 includes six basic processing units. It should be noted that Figure 1 The basic processing units, basic processing sub-units, and regions of the structure 110 in FIG. 1 are merely examples, and the present disclosure does not limit embodiments thereof.

[0062] Figure 2A Schematic diagram of an example encoding process 200A consistent with an embodiment of the present disclosure is illustrated. For example, the encoding process 200A may be performed by an encoder.Figure 2A As shown, the encoder may encode the video sequence 202 into a video bitstream 228 according to process 200A. Figure 1 In the video sequence 100, the video sequence 202 may include a set of images (referred to as "original images") arranged in time sequence. Figure 1 In the structure 110 in FIG. 1 , each original image of the video sequence 202 can be divided into basic processing units, basic processing sub-units, or regions by the encoder for processing. In some embodiments, the encoder can perform process 200A at the basic processing unit level for each original image of the video sequence 202. For example, the encoder can perform process 200A in an iterative manner, wherein the encoder can encode the basic processing unit in one iteration of process 200A. In some embodiments, the encoder can perform process 200A in parallel for regions (e.g., regions 114-118) of each original image of the video sequence 202.

[0063] exist Figure 2A 2, an encoder may provide a basic processing unit (referred to as an "original BPU") of an original image of a video sequence 202 to a prediction stage 204 to generate prediction data 206 and a prediction BPU 208. The encoder may subtract the prediction BPU 208 from the original BPU to generate a residual BPU 210. The encoder may provide the residual BPU 210 to a transform stage 212 and a quantization stage 214 to generate quantized transform coefficients 216. The encoder may provide the prediction data 206 and the quantized transform coefficients 216 to a binary encoding stage 226 to generate a video bitstream 228. The components 202, 204, 206, 208, 210, 212, 214, 216, 226, and 228 may be referred to as a "forward path." During process 200A, after the quantization stage 214, the encoder may provide the quantized transform coefficients 216 to an inverse quantization stage 218 and an inverse transform stage 220 to generate a reconstructed residual BPU 222. The encoder may add the reconstruction residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224, which is used in the prediction stage 204 of the next iteration of the process 200A. The components 218, 220, 222, and 224 of the process 200A may be referred to as a "reconstruction path." The reconstruction path may be used to ensure that both the encoder and the decoder use the same reference data for prediction.

[0064] The encoder can iteratively perform process 200A to encode each original BPU of the original image (in the forward path) and generate the next original BPU prediction reference 224 (in the reconstruction path) for encoding the original image. After encoding all the original BPUs of the original image, the encoder can continue to encode the next image in the video sequence 202.

[0065] Referring to process 200A, an encoder may receive a video sequence 202 generated by a video acquisition device (eg, a camera). The term "receiving" as used herein may refer to any action of receiving, inputting, acquiring, retrieving, obtaining, reading, accessing, or in any manner inputting data.

[0066] In the prediction phase 204, in the current iteration, the encoder may receive the original BPU and the prediction reference 224 and perform a prediction operation to generate the prediction data 206 and the prediction BPU 208. The prediction reference 224 may be generated from the reconstruction path of the previous iteration of the process 200A. The purpose of the prediction phase 204 is to reduce information redundancy by extracting the prediction data 206, which may be used to reconstruct the original BPU into the prediction BPU 208 from the prediction data 206 and the prediction reference 224.

[0067] Ideally, the predicted BPU 208 can be the same as the original BPU. However, due to non-ideal prediction and reconstruction operations, the predicted BPU 208 is usually slightly different from the original BPU. In order to record this difference, after generating the predicted BPU 208, the encoder can subtract it from the original BPU to generate a residual BPU 210. For example, the encoder can subtract the value of the corresponding pixel of the predicted BPU 208 (e.g., grayscale value or RGB value) from the pixel value of the original BPU 208. Each pixel of the residual BPU 210 can have a residual value, which is the result of the subtraction of the corresponding pixel of the original BPU and the predicted BPU 208. Compared with the original BPU, the predicted data 206 and the residual BPU 210 can have fewer bits, but they can be used to reconstruct the original BPU without significantly reducing the quality. Therefore, the original BPU is compressed.

[0068] To further compress the residual BPU 210, in the transform stage 212, the encoder can reduce the spatial redundancy of the residual BPU 210 by decomposing it into a set of two-dimensional "basic patterns", each of which is associated with a "transform coefficient". The basic patterns can have the same size (e.g., the size of the residual BPU 210). Each basic pattern can represent a component of the frequency of change (e.g., the frequency of brightness change) of the residual BPU 210. Any basic pattern cannot be reproduced from any combination (e.g., linear combination) of any other basic patterns. In other words, the decomposition can decompose the changes of the residual BPU 210 into the frequency domain. This decomposition is similar to the discrete Fourier transform of a function, where the basic patterns are similar to the basic functions of the discrete Fourier transform (e.g., trigonometric functions), and the transform coefficients are similar to the coefficients associated with the basic functions.

[0069] Different transform algorithms may use different basic modes. Various transform algorithms may be used in the transform stage 212, such as discrete cosine transform, discrete sine transform, etc. The transform in the transform stage 212 is reversible. That is, the encoder may restore the residual BPU 210 by an inverse operation of the transform (referred to as an "inverse transform"). For example, in order to restore the pixels of the residual BPU 210, the inverse transform may be to multiply the values ​​of the corresponding pixels of the basic mode by their respective correlation coefficients and add the products to produce a weighted sum. For video coding standards, both the encoder and the decoder may use the same transform algorithm (and therefore the same basic mode). Therefore, the encoder may only record the transform coefficients, and the decoder may reconstruct the residual BPU 210 from these coefficients without receiving the basic mode from the encoder. The transform coefficients may have fewer bits than the residual BPU 210, but they may be used to reconstruct the residual BPU 210 without significantly reducing the quality. Therefore, the residual BPU 210 is further compressed.

[0070] The encoder may further compress the transform coefficients in the quantization stage 214. During the transform process, different basic modes may represent different frequencies of change (e.g., frequency of brightness changes). Since the human eye is generally better at recognizing low-frequency changes, the encoder may ignore information about high-frequency changes without causing a significant decrease in decoding quality. For example, in the quantization stage 214, the encoder may generate quantized transform coefficients 216 by dividing each transform coefficient by an integer value (called a "quantization parameter") and rounding the quotient to its nearest integer. After such an operation, some transform coefficients of high-frequency basic modes may be converted to zero, while transform coefficients of low-frequency basic modes may be converted to smaller integers. The encoder may ignore zero-valued quantized transform coefficients 216, further compressing the transform coefficients through this operation. The quantization process is also reversible, where the quantized transform coefficients 216 can be reconstructed into transform coefficients in the inverse operation of quantization (called "inverse quantization").

[0071] Because the encoder ignores the remainder of this division in the rounding operation, the quantization stage 214 may be lossy. Generally, the quantization stage 214 causes the greatest information loss in the process 200A. The greater the information loss, the fewer bits may be required for the quantized transform coefficients 216. To achieve different degrees of information loss, the encoder may use different quantization parameter values ​​or any other parameters of the quantization process.

[0072] In the binary encoding stage 226, the encoder may encode the prediction data 206 and the quantized transform coefficients 216 using a binary encoding technique, such as entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless or lossy compression algorithm. In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the encoder may encode other information in the binary encoding stage 226, such as the prediction mode used in the prediction stage 204, the parameters of the prediction operation, the transform type in the transform stage 212, the parameters of the quantization process (e.g., quantization parameters), encoder control parameters (e.g., bit rate control parameters), etc. The encoder may use the output data of the binary encoding stage 226 to generate a video bitstream 228. In some embodiments, the video bitstream 228 may be further packaged for network transmission.

[0073] Referring to the reconstruction path of process 200A, at the inverse quantization stage 218, the encoder may perform inverse quantization on the quantized transform coefficients 216 to generate reconstructed transform coefficients. At the inverse transform stage 220, the encoder may generate a reconstructed residual BPU 222 based on the reconstructed transform coefficients. The encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224 to be used in the next iteration of process 200A.

[0074] It should be noted that other variations of process 200A may be employed in encoding video sequence 202. In some embodiments, the stages of process 200A may be performed by the encoder in a different order. In some embodiments, one or more stages of process 200A may be combined into a single stage. In some embodiments, a single stage of process 200A may be divided into multiple stages. For example, transform stage 212 and quantization stage 214 may be combined into a single stage. In some embodiments, process 200A may include additional stages. In some embodiments, process 200A may be omitted. Figure 2A one or more stages in a process.

[0075] Figure 2B A schematic diagram of another example encoding process 200B consistent with an embodiment of the present disclosure is illustrated. Process 200B can be modified from process 200A. For example, process 200B can be used by an encoder that conforms to a hybrid video coding standard (e.g., the H.26x series). Compared to process 200A, the forward path of process 200B additionally includes a mode decision stage 230 and divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044. The reconstruction path of process 200B additionally includes a loop filter stage 232 and a buffer 234.

[0076] In general, prediction techniques can be divided into two types: spatial prediction and temporal prediction. Spatial prediction (e.g., intra-image prediction or "intra-frame prediction") can use pixels from one or more encoded adjacent BPUs in the same image to predict the current BPU. That is, the prediction reference 224 in spatial prediction may include adjacent BPUs. Spatial prediction can reduce the spatial redundancy inherent in the image. Temporal prediction (e.g., inter-image prediction or "inter-frame prediction") can use regions from one or more encoded images to predict the current BPU. That is, the prediction reference 224 in temporal prediction may include encoded images. Temporal prediction can reduce the inherent temporal redundancy of the image.

[0077] Referring to process 200B, in the forward path, the encoder performs prediction operations in the spatial prediction stage 2042 and the temporal prediction stage 2044. For example, in the spatial prediction stage 2042, the encoder may perform intra-frame prediction. For the original BPU of the image being encoded, the prediction reference 224 may include one or more neighboring BPUs that have been encoded (in the forward path) and reconstructed (in the reconstruction path) in the same image. The encoder may generate the predicted BPU 208 by interpolating the neighboring BPUs. Interpolation techniques may include, for example, linear extrapolation or interpolation, polynomial extrapolation or interpolation, etc. In some embodiments, the encoder may perform interpolation at the pixel level, for example, by interpolating the values ​​of the corresponding pixels for each pixel of the predicted BPU 208. The neighboring BPUs used for interpolation may be located in various directions relative to the original BPU, such as in the vertical direction (e.g., at the top of the original BPU), the horizontal direction (e.g., on the left side of the original BPU), the diagonal direction (e.g., the lower left, lower right, upper left, or upper right of the original BPU), or any direction defined in the video coding standard used. For intra prediction, the prediction data 206 may include, for example, the position (eg, coordinates) of the used neighboring BPU, the size of the used neighboring BPU, interpolation parameters, the direction of the used neighboring BPU relative to the original BPU, and the like.

[0078] In another example, in the temporal prediction stage 2044, the encoder may perform inter-frame prediction. For the original BPU of the current image, the prediction reference 224 may include one or more images (referred to as "reference images") that have been encoded (in the forward path) and reconstructed (in the reconstruction path). In some embodiments, the reference images may be encoded and reconstructed BPU by BPU. For example, the encoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a reconstructed BPU. When all reconstructed BPUs of the same image are generated, the encoder may generate the reconstructed image as a reference image. The encoder may perform a "motion estimation" operation to search for a matching area in a range of the reference image (referred to as a "search window"). The position of the search window in the reference image may be determined based on the position of the original BPU in the current image. For example, the search window may be centered at a position in the reference image having the same coordinates as the original BPU in the current image and may extend a predetermined distance. When the encoder identifies (e.g., by using a pixel recursive algorithm, a block matching algorithm, etc.) an area in the search window that is similar to the original BPU, the encoder may determine such an area as a matching area. The matching region may have a different size than the original BPU (e.g., smaller than, equal to, larger than, or having a different shape than the original BPU). Because the reference image and the current image are temporally separated on the time axis (e.g., as Figure 1 ), so the matching area can be considered to "move" to the location of the original BPU over time. The encoder can record the direction and distance of this movement as a "motion vector". Figure 1 The encoder can search for matching areas and determine the motion vector associated with each reference image. In some embodiments, the encoder can assign weights to the pixel values ​​of the matching areas of each matching reference image.

[0079] Motion estimation may be used to identify various types of motion, such as translation, rotation, scaling, etc. For inter-frame prediction, the prediction data 206 may include, for example, the location (e.g., coordinates) of the matching region, the motion vector associated with the matching region, the number of reference images, the weights associated with the reference images, etc.

[0080] To generate the predicted BPU 208, the encoder may perform a "motion compensation" operation. Motion compensation may be used to reconstruct the predicted BPU 208 based on the prediction data 206 (e.g., motion vectors) and the prediction reference 224. For example, the encoder may move a matching area of ​​a reference image according to the motion vector, where the encoder may predict the original BPU of the current image. Figure 1In some embodiments, if the encoder has assigned weights to the pixel values ​​of the matching regions of the matching reference images, the encoder can add the weighted sum of the pixel values ​​of the moving matching regions.

[0081] In some embodiments, inter-frame prediction can be unidirectional or bidirectional. Unidirectional inter-frame prediction can use one or more reference images in the same temporal direction relative to the current image. For example, Figure 1 The picture 104 in is a unidirectional inter-frame prediction picture, where the reference picture (i.e., picture 102) precedes the picture 104. Bidirectional inter-frame prediction can use one or more reference pictures in two temporal directions relative to the current picture. For example, Figure 1 Picture 106 in is a bi-directional inter-predicted picture, where the reference pictures (ie, pictures 104 and 108) are relative to picture 104 in both temporal directions.

[0082] Still referring to the forward path of process 200B, after spatial prediction stage 2042 and temporal prediction stage 2044, at mode decision stage 230, the encoder can select a prediction mode (e.g., one of intra prediction or inter prediction) for the current iteration of process 200B. For example, the encoder can perform a rate-distortion optimization technique, in which the encoder can select a prediction mode to minimize the value of a cost function based on the bit rate of a candidate prediction mode and the distortion of a reconstructed reference image under the candidate prediction mode. Based on the selected prediction mode, the encoder can generate a corresponding prediction BPU 208 and prediction data 206.

[0083] In the reconstruction path of process 200B, if intra-prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current BPU that has been encoded and reconstructed in the current image), the encoder can directly provide the prediction reference 224 to the spatial prediction stage 2042 for later use (e.g., for interpolation of the next BPU of the current image). If inter-prediction mode is selected in the forward path, after generating the prediction reference 224 (e.g., the current image in which all BPUs have been encoded and reconstructed), the encoder can provide the prediction reference 224 to the loop filtering stage 232, where the encoder can apply loop filtering to the prediction reference 224 to reduce or eliminate distortion (e.g., block artifacts) introduced by inter-prediction. The encoder can apply various loop filtering techniques in the loop filtering stage 232, such as deblocking, sample adaptive offset, adaptive loop filtering, etc. The loop-filtered reference image can be stored in a buffer 234 (or "decoded image buffer") for later use (e.g., used as an inter-prediction reference image for future images of the video sequence 202). The encoder may store one or more reference pictures in a buffer 234 for use in a temporal prediction stage 2044. In some embodiments, the encoder may encode loop filter parameters (e.g., loop filter strength) in a binary encoding stage 226, as well as quantized transform coefficients 216, prediction data 206, and other information.

[0084] Figure 3A A schematic diagram of an example decoding process 300A consistent with an embodiment of the present disclosure is illustrated. Process 300A may correspond to Figure 2A 300A. In some embodiments, process 300A may be similar to the reconstruction path of process 200A. The decoder may decode video bitstream 228 into video stream 304 according to process 300A. Video stream 304 may be very similar to video sequence 202. However, due to information loss during compression and decompression (e.g., Figures 2A - 2B Typically, the video stream 304 is not identical to the video sequence 202. Figures 2A - 2B Similar to processes 200A and 200B in the above, the decoder may perform process 300A at a basic processing unit (BPU) level for each picture encoded in the video bitstream 228. For example, the decoder may perform process 300A in an iterative manner, where the decoder may decode one basic processing unit in one iteration of process 300A. In some embodiments, the decoder may perform process 300A in parallel for regions (e.g., regions 114-118) of each picture encoded in the video bitstream 228.

[0085] exist Figure 3A, the decoder may provide a portion of the video bitstream 228 associated with the basic processing unit of the encoded picture (referred to as the "encoded BPU") to the binary decoding stage 302. In the binary decoding stage 302, the decoder may decode the portion into prediction data 206 and quantized transform coefficients 216. The decoder may provide the quantized transform coefficients 216 to the inverse quantization stage 218 and the inverse transform stage 220 to generate a reconstructed residual BPU 222. The decoder may provide the prediction data 206 to the prediction stage 204 to generate the prediction BPU 208. The decoder may add the reconstructed residual BPU 222 to the prediction BPU 208 to generate a prediction reference 224. In some embodiments, the prediction reference 224 may be stored in a buffer (e.g., a decoded picture buffer in a computer memory). The decoder may provide the prediction reference 224 to the prediction stage 204 to perform a prediction operation in the next iteration of the process 300A.

[0086] The decoder can iteratively perform process 300A to decode each coded BPU of the coded picture and generate prediction reference 224 for decoding the next coded BPU of the coded picture. After decoding all coded BPUs of the coded picture, the decoder can output the picture to the video stream 304 for display and continue to decode the next coded picture in the video bitstream 228.

[0087] In the binary decoding stage 302, the decoder may perform the inverse operation of the binary coding technique used by the encoder (e.g., entropy coding, variable length coding, arithmetic coding, Huffman coding, context-adaptive binary arithmetic coding, or any other lossless compression algorithm). In some embodiments, in addition to the prediction data 206 and the quantized transform coefficients 216, the decoder may decode other information in the binary decoding stage 302, such as prediction mode, parameters of the prediction operation, transform type, quantization parameter process (e.g., quantization parameter), encoder control parameters (e.g., bit rate control parameter), etc. In some embodiments, if the video bitstream 228 is transmitted in the form of packets over the network, the decoder may depacketize the video bitstream 228 before feeding it to the binary decoding stage 302.

[0088] Figure 3B A schematic diagram of another example decoding process 300B consistent with an embodiment of the present disclosure is shown. Process 300B can be modified from process 300A. For example, process 300B can be used by a decoder that conforms to a hybrid video coding standard (e.g., H.26x series). Compared to process 300A, process 300B additionally divides the prediction stage 204 into a spatial prediction stage 2042 and a temporal prediction stage 2044, and additionally includes a loop filtering stage 232 and a buffer 234.

[0089] In process 300B, for an encoded basic processing unit (referred to as a "current BPU") of a decoded encoded image (referred to as a "current image"), the prediction data 206 decoded by the decoder in the binary decoding stage 302 may contain various types of data, depending on what prediction mode the encoder used to encode the current BPU. For example, if the encoder uses intra-frame prediction to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating intra-frame prediction, parameters of the intra-frame prediction operation, etc. The parameters of the intra-frame prediction operation may include, for example, the position (e.g., coordinates) of one or more neighboring BPUs used as a reference, the size of the neighboring BPU, interpolation parameters, the direction of the neighboring BPU relative to the original BPU, etc. For another example, if the encoder uses inter-frame prediction to encode the current BPU, the prediction data 206 may include a prediction mode indicator (e.g., a flag value) indicating inter-frame prediction, parameters of the inter-frame prediction operation, etc. The parameters of the inter-frame prediction operation may include, for example, the number of reference images associated with the current BPU, the weights associated with the reference images respectively, the positions (e.g., coordinates) of one or more matching regions in each reference image, one or more motion vectors associated with each matching region, etc.

[0090] Based on the prediction mode indicator, the decoder can decide whether to perform spatial prediction (e.g., intra prediction) in the spatial prediction stage 2042 or temporal prediction (e.g., inter prediction) in the temporal prediction stage 2044. Figure 2B The details of performing this spatial prediction or temporal prediction are described in , and will not be repeated below. After performing such spatial prediction or temporal prediction, the decoder can generate a prediction BPU 208. The decoder can add the prediction BPU 208 and the reconstructed residual BPU 222 to generate a prediction reference 224, such as Figure 3A As described in.

[0091] In process 300B, the decoder may provide the prediction reference 224 to the spatial prediction stage 2042 or the temporal prediction stage 2044 for performing a prediction operation in the next iteration of process 300B. For example, if the current BPU is decoded using intra-frame prediction in the spatial prediction stage 2042, then after generating the prediction reference 224 (e.g., the decoded current BPU), the decoder may provide the prediction reference 224 directly to the spatial prediction stage 2042 for later use (e.g., for interpolating the next BPU of the current image). If the current BPU is decoded using inter-frame prediction in the temporal prediction stage 2044, then after generating the prediction reference 224 (e.g., a reference image in which all BPUs have been decoded), the encoder may provide the prediction reference 224 to the loop filtering stage 232 to reduce or eliminate distortion (e.g., blocking artifacts). The decoder may Figure 2BIn-loop filtering is applied to the prediction reference 224 in the manner described in . The loop-filtered reference picture may be stored in a buffer 234 (e.g., a decoded picture buffer in a computer memory) for later use (e.g., as an inter-prediction reference picture for a future encoded picture of the video bitstream 228). The decoder may store one or more reference pictures in the buffer 234 for use in the temporal prediction stage 2044. In some embodiments, when the prediction mode indicator of the prediction data 206 indicates that inter-prediction was used to encode the current BPU, the prediction data may further include parameters for loop filtering (e.g., loop filter strength).

[0092] Figure 4 4 is a block diagram of an example apparatus 400 for encoding or decoding a video consistent with an embodiment of the present disclosure. Figure 4 As shown, the device 400 may include a processor 402. When the processor 402 executes the instructions described herein, the device 400 may become a special-purpose machine for video encoding or decoding. The processor 402 may be any type of circuit capable of manipulating or processing information. For example, the processor 402 may include any number of central processing units (or "CPUs"), graphics processing units (or "GPUs"), neural processing units ("NPUs"), microcontroller units ("MCUs"), optical processors, programmable logic controllers, microcontrollers, microprocessors, digital signal processors, intellectual property (IP) cores, programmable logic arrays (PLAs), programmable array logic (PALs), general array logic (GALs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), systems on chips (SoCs), application-specific integrated circuits (ASICs), and the like. In some embodiments, the processor 402 may also be a group of processors grouped into a single logical component. For example, as Figure 4 As shown, processor 402 may include multiple processors, including processor 402a, processor 402b, and processor 402n.

[0093] The device 400 may also include a memory 404 configured to store data (eg, a set of instructions, computer code, intermediate data, etc.). For example, as shown in FIG. Figure 4As shown, the stored data may include program instructions (e.g., program instructions for implementing stages in process 200A, 200B, 300A, or 300B) and data for processing (e.g., video sequence 202, video bitstream 228, or video stream 304). Processor 402 may access program instructions and data for processing (e.g., via bus 410) and execute program instructions to operate or manipulate the data for processing. Memory 404 may include a high-speed random access storage device or a non-volatile storage device. In some embodiments, memory 404 may include any number of random access memories (RAM), read-only memories (ROM), optical disks, magnetic disks, hard drives, solid-state drives, flash drives, secure digital (SD) cards, memory sticks, compact flash (CF) cards, etc. Memory 404 may also be a group of memories grouped into a single logical component ( Figure 4 not shown).

[0094] The bus 410 may be a communication device that transmits data between components inside the apparatus 400 , such as an internal bus (eg, a CPU-memory bus), an external bus (eg, a universal serial bus port, a peripheral component interconnect express port), and the like.

[0095] For ease of explanation and without ambiguity, the processor 402 and other data processing circuits are collectively referred to as "data processing circuits" in this disclosure. The data processing circuits may be fully implemented as hardware, or a combination of software, hardware, or firmware. In addition, the data processing circuit may be a single independent module, or may be fully or partially combined into any other component of the device 400.

[0096] The device 400 may also include a network interface 406 to provide wired or wireless communication with a network (e.g., the Internet, an intranet, a local area network, a mobile communication network, etc.). In some embodiments, the network interface 406 may include any number of network interface controllers (NICs), radio frequency (RF) modules, repeaters, transceivers, modems, routers, gateways, wired network adapters, wireless network adapters, Bluetooth adapters, infrared adapters, near field communication (“NFC”) adapters, cellular network chips, etc.

[0097] In some embodiments, the apparatus 400 may optionally further include a peripheral interface 408 to provide a connection to one or more peripheral devices. Figure 4 As shown, peripheral devices may include, but are not limited to, a cursor control device (such as a mouse, touchpad, or touch screen), a keyboard, a display (such as a cathode ray tube display, a liquid crystal display, or a light emitting diode display), a video input device (for example, a camera or an input interface coupled to a video archive), etc.

[0098] It should be noted that the video codec (e.g., a codec that performs the process 200A, 200B, 300A, or 300B) can be implemented as any combination of any software or hardware modules in the device 400. For example, some or all stages of the process 200A, 200B, 300A, or 300B can be implemented as one or more software modules of the device 400, such as program instructions that can be loaded into the memory 404. As another example, some or all stages of the process 200A, 200B, 300A, or 300B can be implemented as one or more hardware modules of the device 400, such as dedicated data processing circuits (e.g., FPGA, ASIC, NPU, etc.).

[0099] In the quantization and dequantization blocks (e.g. Figure 2A or Figure 2B quantization 214 and inverse quantization 218, Figure 3A or Figure 3B The quantization parameter (QP) is used to determine the amount of quantization (and inverse quantization) applied to the prediction residual. The initial QP value for encoding of a picture or slice can be signaled at a high level, for example, using the init_qp_minus26 syntax element in the picture parameter set (PPS) and using the slice_qp_delta syntax element in the slice header. In addition, the QP value can be adjusted for each CU at a local level using a delta QP value sent at the granularity of the quantization group.

[0100] To improve the accuracy of the motion vector (MV) in merge mode, decoding-side motion vector modification (DMVR) based on bilateral matching (BM) is adopted in Versatile Video Coding (VVC) draft 6. In the bilateral operation, the modified MV is searched around the initial MV in reference picture list L0 and reference picture list L1. BM-based DMVR calculates the distortion between two candidate blocks in reference picture list L0 and list L1. Figure 5 An example process 500 of decoding-side motion vector revision (DMVR) according to some embodiments of the present disclosure is illustrated. Figure 5A current image 502, a first reference image 504 in a first reference image list L0, and a second reference image 506 in a second reference image list L1 are shown, a first initial MV 508 pointing from a current block 510 in the current image 502 to a first initial reference block 512 in the first reference image 504, and a second initial MV 514 pointing from the current block 510 to a second initial reference block 516 in the second reference image 506. The process 500 may perform BM-based DMVR to determine a first candidate reference block 518 in the first reference image 504, a second candidate reference block 520 in the second reference image 506, a first candidate MV 522 connecting the current block 510 and the first candidate reference block 518, and a second candidate MV 524 connecting the current block 510 and the second candidate reference block 516. Figure 5 As shown, the first candidate MV 522 and the second candidate MV 524 are close to the first initial MV 508 and the second initial MV 514, respectively. In some embodiments, the process 500 may calculate the sum of absolute differences (SAD) between the first initial reference block 518 and the second initial reference block 520 based on each MV candidate (e.g., the first candidate MV 522 or the second candidate MV 524) around the initial MV (e.g., the first initial MV 508 or the second initial MV 514). One of the first candidate MV 522 and the second candidate MV 524 having the lowest SAD may become the modified MV for generating the bidirectional prediction signal.

[0101] In some embodiments, as described in VVC draft 6, DMVR is applied to a CU that satisfies all of the following conditions: (1) the merge mode is at the CU level with a bidirectional prediction MV; (2) the block is predicted using a bidirectional prediction MV with equal weights (e.g., by not applying bidirectional prediction with weighted averaging (BWA) to the block; (3) relative to a current image (e.g., current image 502), one reference image is in the past (e.g., first reference image 504) and the other reference image is in the future (e.g., second reference image 506); (4) the distances (e.g., picture order count (POC) differences) of the two reference images to the current image are the same; (5) the block has at least 128 luma samples, and the width and height of the block are both at least 8 luma samples.

[0102] The modified MV obtained by the DMVR process (e.g., process 500) can be used to generate inter-frame prediction samples and temporal motion vector prediction for future image encoding. The initial MV (e.g., first initial MV 508 or second initial MV 514) can be used for deblocking process and spatial MV prediction for future CU encoding within the current image (e.g., current image 502).

[0103] like Figure 5As shown, the first MV offset 526 represents the corrected offset between the first initial MV 508 and the first candidate MV 522, and the second MV offset 528 represents the corrected offset between the second initial MV 514 and the second candidate MV 524. The first MV offset 526 and the second MV offset 528 may have the same magnitude and opposite directions. In some embodiments, the search point may be around the initial MVs (e.g., the first initial MV 508 and the second initial MV 514), and the MV offsets (e.g., the first MV offset 526 and the second MV offset 528) may follow the MV difference mirror rule. For example, a point checked by DMVR may be represented by a pair of candidate MVs MV0 (e.g., the first candidate MV 522) and MV1 (e.g., the second candidate MV 524) based on equations (1) and (2): MV0′=MV0+MV_offset Equation (1) MV1′=MV1-MV_offset Equation (2)

[0104] In equations (1) and (2), MV0' (e.g., the first initial MV 508) and MV1' (e.g., the second candidate MV 514) represent a pair of initial MVs, and MV_offset represents a modified offset (e.g., the first MV offset 526 or the second MV offset 52) ​​between the initial MV (e.g., MV0' or MV1') and the modified MV (e.g., MV0 or MV1). Note that MV_offset is a vector with motion displacement (e.g., in X and Y dimensions). In some embodiments, as described in VVC draft 5, the modified search range (e.g., the search range of DMVR) can be two integer luma samples from the initial MVs (e.g., the first initial MV 508 and the second initial MV 514) in the horizontal direction and the vertical direction.

[0105] Figure 6 An example DMVR search process 600 is illustrated according to some embodiments of the present disclosure. In some embodiments, process 600 may be performed by a codec (e.g., Figures 2A - 2B The encoder in Figures 3A - 3B For example, a codec may be implemented as a device for encoding or transcoding a video sequence (e.g., Figure 4 In some embodiments, process 600 may be an example of a search process for a DMVR as described in VVC draft 4. Figure 6As shown, process 600 includes stage 602 for integer sample offset search and stage 604 for fractional sample correction. To reduce search complexity, in some embodiments, a fast search method with an early termination mechanism is applied in stage 602. For example, a 2-iteration search scheme can be applied in stage 602 instead of using a 25-point full search to reduce SAD checkpoints.

[0106] like Figure 6 As shown, stage 604 may follow stage 602. To save computational complexity, in some embodiments, the fractional sample corrections of stage 604 may be derived using the parameter error surface equation rather than performing an additional search involving SAD comparisons. Stage 604 may be conditionally called based on the output of stage 602.

[0107] Figure 7 An example mode 700 for DMVR integer luma sample search according to some embodiments of the present disclosure is illustrated. The DMVR integer luma sample search may determine the point in the search sample where the SAD is the smallest. For example, the DMVR integer luma sample search may be implemented as Figure 6 The process 600 in FIG. 6 includes a stage for integer sample offset search (e.g., stage 602) and a stage for fractional sample correction search (e.g., stage 604), wherein each stage may be performed in at least one iteration. In some embodiments, in the first iteration of the DMVR integer luma sample search, up to 6 SADs may be checked. Figure 7 For example, in the first iteration, the SADs of five points 702-710 (represented as black blocks) can be compared, where point 702 can be used as the center point of the search. If the SAD of the center point (i.e., point 702) is the smallest, the integer sampling phase of the DMVR can be terminated. Otherwise, another point 712 (represented as a shaded block) determined by the SAD distribution of points 704-710 can be checked. In the second iteration of the DMVR integer brightness sample search, the point with the smallest SAD among points 704-712 can be selected as the new search center point. In some embodiments, the second iteration can be performed in the same manner as the first iteration. In some embodiments, the SAD calculated in the first iteration can be reused in the second iteration, so only the SAD of the additional points needs to be further calculated.

[0108] In some embodiments, as described in VVC Draft 6, the Figure 7 Then, in the integer sample offset search phase (e.g., Figure 6 In stage 602), the SAD of all 25 points can be calculated in one iteration. Figure 8An example pattern 800 is illustrated for the phase of integer sample offset search in a DMVR integer luma sample search according to some embodiments of the present disclosure. For example, the phase of integer sample offset search may be Figure 6 in stage 602. Figure 8 The initial MV 802 and 25 points are shown, and their SADs can be calculated together. In some embodiments, the SAD of the initial MV 802 can be reduced (e.g., reduced by one-fourth) to adjust the initial MV 802. In some embodiments, the SAD of the initial MV 802 can be reduced (e.g., reduced by one-fourth) in a stage for partial sample correction (e.g., in Figure 6 The fractional sample correction stage may be conditionally called based on the position where the SAD is the smallest. For example, Figure 8 As shown, if the position where the SAD is the smallest is one of the nine points around the initial MV 802 (as shown in block 804), the stage for fractional sample correction can be called to determine the corrected MV as the output of the DMVR integer brightness sample search. If the position where the SAD is the smallest is not any of the nine points around the initial MV 802, the position where the SAD is the smallest can be directly used as the output of the DMVR integer brightness sample search.

[0109] Figure 9 is a schematic diagram illustrating an example mode 900 for estimating a DMVR parameter error surface according to some embodiments of the present disclosure. Figure 8 The initial MV 902 and 25 points are shown. The initial MV 902 is connected to the center point 904 with the minimum SAD. In the sub-pixel offset estimation based on the parameter error surface, as Figure 9 As shown, the sum of the value of absolute difference (SAD) cost of the center point 904 and the SAD costs of four neighboring points 906-912 around the center point 904 can be used to fit a two-dimensional parabolic error surface equation. For example, the two-dimensional parabolic error surface equation can be based on equation (3). E(x,y)=((A(xx min ) 2 +B(yy min ) 2 +)>> mvShift)+E(0,0) Equation (3)

[0110] In equation (3), (x min ,y min ) corresponds to the fractional position with the lowest SAD cost, E(x, y) corresponds to the SAD cost of the center point 904 and the four neighboring points 906-912, mvShift can be set to 4 as in VVC (in VVC, the MV accuracy is 1 / 16 pixel), and A and B can be determined according to equations (4) and (5), respectively:

[0111] By solving equations (3) to (5) using the SAD cost values ​​of the five search points (i.e., points 904-912), (x min ,y min ).

[0112] In some embodiments, x min and min The value of can be automatically clamped between –8 and 8 (e.g., in 1 / 16 sample precision) because all SAD cost values ​​are positive and the minimum is E(0,0), which corresponds to a half-pixel offset with 1 / 16 pixel MV precision in VVC. The calculated fraction (x min ,y min ) can be added to the integer distance correction MV to make the correction MV have sub-pixel accuracy.

[0113] The Bidirectional Optical Flow (BDOF) tool is included in VVC. As the name suggests, the BDOF mode is based on the concept of optical flow and assumes that the motion of objects is smooth. BDOF, formerly known as BIO, is also included in the Joint Video Exploration Model (JEM) software. Compared to BIO in JEM, BDOF in VVC is a simpler version, especially in terms of the number of multiplications and the size of the multipliers, requiring much less computation.

[0114] BDOF can be used to modify the bidirectional prediction signal of a CU at the 4×4 sub-block level. In some embodiments, BDOF is applied to a CU that meets the following conditions: (1) the height of the CU is not 4 and the size of the CU is not 4×8; (2) the CU is not encoded using the affine mode or the advanced temporal motion vector prediction (ATMVP) merge mode; (3) the CU is encoded using the "true" bidirectional prediction mode, in which one of the two reference pictures (e.g., Figure 5 The first reference image 504 in the display order is in the current image (eg, Figure 5 502), and another one (e.g., Figure 5 The second reference image 506 in the display order follows the current image. In some embodiments, BDOF can be applied to the luminance component.

[0115] In some embodiments, when BDOF is used to modify the bidirectional prediction signal of a CU at the 4×4 sub-block level, for each 4×4 sub-block, the motion correction (v x , v y ) can be calculated by minimizing the difference between the prediction samples in the two reference picture lists L0 and L1. (v x , vy ) can then be used to adjust the bidirectionally predicted sample values ​​in the 4×4 sub-block.

[0116] In some embodiments, the following steps are applied in the BDOF process: 1. The horizontal and vertical gradients of the two prediction signals at k=0,1 can be determined based on calculating the difference between two adjacent samples. and As shown in equations (8) and (9):

[0117] In equations (8) and (9), I (k) (i,j) is the sample value of the prediction signal at coordinate (i,j) in list k when k=0,1. shift1 is calculated based on the luma bit depth (“bitDepth”), as shown in equation (10): shift1=max(2,14-bitDepth) Equation (10)

[0118] Then, the gradient S 1 ,S 2 ,S 3 ,S 5 and S 6 The autocorrelation and cross-correlation of can be determined according to equations (11) to (15):

[0119] For equations (11) to (15), ψ can be determined according to equations (16) to (18): x (i,j),ψ y (i,j), the value of θ(i,j) θ(i,j)=(I (1) (i,j)>>n b )-(I (0) (i,j)>>n b ) Equation (18)

[0120] In equations (11) to (18), Ω is a 6×6 window around the 4×4 sub-block, n a and n b The values ​​of are set by equations (19) and (20) respectively: n a =min(5,bitDepth-7) Equation (19) n b =min(8,bitDepth-4) Equation (20)

[0121] The motion correction (v) is then derived based on the cross-correlation and autocorrelation terms using equations (21) and (22). x ,v y ):

[0122] In equations (21) and (22), th′ BIO =2 13-B represents the floor function, and

[0123] Based on the motion correction and the gradient, the following adjustment b(x,y) for each sample in the 4×4 sub-block can be determined using equation (23):

[0124] Finally, the BDOF samples of the CU can be determined by adjusting the bidirectional prediction samples according to equation (24): pred BDOF (x,y)=(I (0) (x,y)+I (1) (x,y)+b(x,y)+o offset )>>shift Equation (24)

[0125] In some embodiments, the values ​​in equations (8) to (23) can be selected so that the multiplier in the BDOF process does not exceed 15 bits, and the maximum bit width of the intermediate parameters in the BDOF process can be kept within 32 bits.

[0126] In order to derive the gradient value, some prediction samples I outside the current CU boundary in list k (k = 0, 1) can be generated (k) (i,j). Figure 10 1 is a schematic diagram of an example of an extended coding unit (CU) region 1000 used in BDOF according to some embodiments of the present disclosure. Figure 10As shown, the 4×4 block 1002 used in BDOF (surrounded by a solid black line) is surrounded by an extended row or column around the boundary of the block 1002 (represented by a dashed black line), forming a surrounding area 1004. In order to control the computational complexity of generating prediction samples outside the boundary, the prediction samples within the extended area 1006 (represented by a white box) can be generated by directly obtaining reference samples at nearby integer positions (using a floor operation on the coordinates) without interpolation, and the prediction samples can be generated within the CU 1008 (represented by a gray box) using a normal 8-tap motion compensated interpolation filter. These extended sample values ​​can only be used for gradient calculations. For the remaining steps of the BDOF process, if any samples and gradient values ​​outside the boundary of the CU 1008 are needed, they can be filled (or repeated) from their nearest neighbors.

[0127] At the JVET conference, a coding tool called Optical Flow Prediction Correction (PROF) was adopted. PROF improves the accuracy of affine motion compensation prediction by correcting sub-block based affine motion compensation prediction using optical flow. The affine motion model parameters can be used to derive the motion vector for each sample position in the CU. However, due to the complexity and high memory access bandwidth of generating sample-by-sample affine motion compensation prediction, affine prediction in VVC uses a sub-block based affine motion compensation method, where one CU is divided into 4×4 sub-blocks, each assigned an MV derived from the control point MV of the affine CU. Sub-block based affine motion compensation is a compromise between coding efficiency, complexity, and memory access bandwidth. Due to its sub-block based prediction rather than theoretical sample based motion compensation prediction, it loses some prediction accuracy.

[0128] To achieve finer affine motion compensation granularity, in some embodiments, PROF can be applied after conventional sub-block based affine motion compensation. The sample-based correction can be obtained based on the optical flow equation, such as equation (25): ΔI(i,j)=g x (i,j)*Δv x (i,j)+g y (i,j)*Δv y (i,j) Equation (25)

[0129] In equation (25), g x (i,j) and g y (i, j) is the spatial gradient at sample location (i, j), and Δv is the motion offset from the sub-block based motion vector to the sample based motion vector derived from the affine model parameters.

[0130] Figure 11 FIG. 1 is a schematic diagram of an example of sub-block-based translation motion and sample-based affine motion according to some embodiments of the present disclosure.Figure 11 As shown, V(i, j) is the theoretical motion vector of the sample position (i, j) derived using the affine model, V SB is a sub-block based motion vector, and ΔV(i,j) (indicated by a dotted arrow) is the sum of V(i,j) and V SB The difference between.

[0131] Then, the prediction correction ΔI(i,j) can be added to the sub-block prediction I(i,j). The final prediction I′ can be generated based on equation (26): I′(i,j)=I(i,j)+ΔI(i,j) Equation (26)

[0132] Consistent with the disclosed embodiments, both DMVR and BDOF may have control flags at two levels in the syntax structure. The first control flag may be sent in a sequence parameter set (SPS) at the sequence level, and the second control flag may be sent in a slice header at the slice level. Figure 12 Table 1 is shown, which shows an example syntax structure of a sequence parameter set (SPS) of control flags for implementing DMVR and BDOF according to some embodiments of the present disclosure. Figure 12 As shown in Table 1, sps_bdof_enabled_flag and sps_dmvr_enabled_flag are control flags for BDOF and DMVR, respectively, at the sequence level sent in the SPS. When sps_bdof_enabled_flag or sps_dmvr_enabled_flag is false, BDOF or DMVR can be disabled in the entire video sequence that references this SPS. When sps_bdof_enabled_flag and sps_dmvr_enabled_flag are true, BDOF or DMVR can be enabled for the current video sequence. In this case, another flag sps_bdof_dmvr_slice_present_flag can be further signaled to indicate whether slice level control of BDOF and DMVR is enabled.

[0133] Figure 13 Table 2 is shown, which shows an example syntax structure of a slice header implementing control flags for DMVR and BDOF according to some embodiments of the present disclosure. As shown in Table 2, when sps_bdof_dmvr_slice_present_flag set in Table 1 is true, slice_disable_bdof_dmvr_flag can signal in the slice header whether the current slice disables BDOF and DMVR.

[0134] Figures 12 - 13 A two-level control mechanism for DMVR and BDOF is shown. By using this mechanism, an encoder (e.g., Figures 2A - 2B The encoder of process 200A or 200B in the example above may use the slice-level flag slice_disable_bdof_dmvr_flag to turn on or off DMVR and BDOF for each slice. This slice-level adaptation may have two benefits: (1) when at least one of DMVR or BDOF is not useful for the current slice, turning it off may improve encoding performance; (2) the computational complexity of DMVR and BDOF is relatively high, and turning them off may reduce the encoding and decoding complexity of the current slice.

[0135] In the disclosed embodiments, the control flags for PROF may also be used at the sequence level and slice level. In some embodiments, three separate flags may be signaled in the SPS to indicate whether DMVR, BDOF, and PROF are enabled, respectively. If any of them are enabled, the corresponding lower level control enable flag may be signaled to indicate whether the enabled tool is controlled at a lower level. The lower level may be a slice level or an image level. If slice level or image level control is enabled, a slice level or image level disable flag may be signaled in each slice header or image header to indicate that the enabled tool for the current slice or image is disabled.

[0136] Consistent with the disclosed embodiments, Figure 14A Table 3A is shown, which shows an example syntax structure of a sequence parameter set (SPS) implementing slice level control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. Figure 14BTable 3B is shown, which shows an example syntax structure of an SPS for implementing picture level control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As shown in Tables 3A-3B, with emphasis in italics, sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag are flags that are signaled in the SPS to indicate whether BDOF, DMVR, and PROF are enabled for the video sequence, respectively. If BDOF, DMVR, or PROF is enabled, as shown in Table 3A, sps_bdof_slice_present_flag, sps_dmvr_slice_present_flag, or sps_affine_prof_slice_present_flag may further signal to indicate whether slice level control of BDOF, DMVR, and PROF is enabled, respectively. If BDOF, DMVR, or PROF is enabled, as shown in Table 3B, sps_bdof_picture_present_flag, sps_dmvr_picture_present_flag, or sps_affine_prof_picture_present_flag may further signal, respectively, to indicate whether picture level control of BDOF, DMVR, and PROF is enabled.

[0137] Figure 15A Table 4A is illustrated, which shows an example syntax structure of a slice header implementing control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As shown in Table 4A highlighted in italics, if any of sps_bdof_slice_present_flag, sps_dmvr_slice_present_flag, or sps_affine_prof_slice_present_flag in Table 3A is set to true, then slice_disable_bdof_flag, slice_disable_dmvr_flag, or slice_disable_affine_prof_flag may be signaled to indicate whether BDOF, DMVR, or PROF is disabled for the current slice, respectively. Figure 15BTable 4B is shown, which shows an example syntax structure of a picture header implementing control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As shown in Table 4B highlighted in italics, if any one of sps_bdof_picture_present_flag, sps_dmvr_picture_present_flag, or sps_affine_prof_picture_present_flag in Table 3B is set to true, then ph_disable_bdof_flag, ph_disable_dmvr_flag, or ph_disable_affine_prof_flag may be signaled to indicate whether BDOD, DMVR, or PROF is disabled for the current picture, respectively.

[0138] In some embodiments, DMVR, BDOF, and PROF may have three separate sequence-level enable flags but share the same slice-level control enable flag. For example, a slice-level disable flag may be sent for DMVR, BDOF, and PROF. In another example, three slice-level disable flags may be signaled for DMVR, BDOF, and PROF, respectively. As another example, two slice-level disable flags may be sent for DMVR, BDOF, and PROF. It should be noted that the control of DMVR, BDOF, and PROF may implement a variety of grammatical structures at the sequence level and at levels below the sequence level (referred to herein as "lower levels", such as slice levels or image levels), which are not limited to the examples described herein.

[0139] Figure 16 Table 5 is illustrated, which shows an example syntax structure of a sequence parameter set (SPS) that implements separate sequence level control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As shown in Table 5, highlighted in italics, three separate flags sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag are signaled in the SPS to indicate whether DMVR, BDOF, and PROF are enabled, respectively. If at least one of DMVR, BDOF, or PROF is enabled, a slice control enable flag sps_bdof_dmvr_affine_prof_slice_present_flag may be signaled to indicate whether at least one of DMVR, BDOF, or PROF enabled at the sequence level is controlled in a lower level.

[0140] Figure 17Table 6 is shown, which shows an example syntax structure of a slice header implementing a joint control flag for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As shown in Table 6 with italic highlights, if the slice level control is enabled as described in Table 5 (e.g., sps_bdof_dmvr_affine_prof_slice_present_flag is true), then in each slice header, a slice level disable flag slice_disable_bdof_dmvr_affine_prof_flag may be signaled to indicate whether at least one of DMVR, BDOF, or PROF enabled at the sequence level is disabled for the current slice. In the syntax structure shown in Table 6, if multiple DMVR, BDOF, and PROF are enabled at the sequence level, the multiple enables may be jointly controlled at the slice level if the slice level control is enabled.

[0141] Figure 18Table 7 is illustrated, which shows an example syntax structure of a slice header that implements separate control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As highlighted in italics in Table 7, if the slice level control described in Table 5 is enabled (e.g., sps_bdof_dmvr_affine_prof_slice_present_flag is true), then for one of DMVR, BDOF, and PROF enabled at the sequence level, a slice level disable flag may be signaled to indicate whether one of the aforementioned enables is disabled for the current slice. For example, if sps_bdof_enabled_flag and sps_bdof_dmvr_affine_prof_slice_present_flag in Table 5 are set to true, then slice_disable_bdof_flag in Table 7 may be signaled to indicate whether BDOF is disabled for the current slice. As another example, if sps_dmvr_enabled_flag and sps_bdof_dmvr_affine_prof_slice_present_flag in Table 5 are set to true, slice_disable_dmvr_flag in Table 7 may be signaled to indicate whether DMVR is disabled for the current slice. In yet another example, if sps_affine_prof_enabled_flag and sps_bdof_dmvr_affine_prof_slice_present_flag in Table 5 are set to true, slice_disable_affine_prof_flag in Table 7 may be signaled to indicate whether PROF is disabled for the current slice. In Table 7, if slice-level control is enabled, each of DMVR, BDOF, and PROF may be individually controlled at the slice level.

[0142] Considering the fact that both BDOF and PROF use optical flow to modify the inter-frame predictor, in some embodiments, BDOF and PROF may share the same slice level control flags, and DMVR may use separate slice level control flags. Figure 19Table 8 is shown, which shows an example syntax structure of a slice header implementing a hybrid control flag for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. In the syntax structure of Table 8, BDOF and PROF can share the same slice level disable flag, and DMVR can use another slice level disable flag. As shown in Table 8 highlighted in italics, if slice level control is enabled in Table 5 (e.g., sps_bdof_dmvr_affine_prof_slice_present_flag is true), two slice level disable flags can be signaled to indicate whether DMVR, BDOF, and PROF are disabled for the current slice. For example, if at least one of sps_bdof_enabled_flag and sps_affine_prof_enabled_flag in Table 5 is set to true, and sps_bdof_dmvr_affine_prof_slice_present_flag in Table 5 is set to true, then slice_disable_bdof_affine_prof_flag in Table 8 may be signaled to indicate whether at least one of BDOF or PROF enabled at the sequence level is disabled for the current slice. In another example, if sps_dmvr_enabled_flag and sps_bdof_dmvr_affine_prof_slice_present_flag in Table 5 are set to true, then slice_disable_dmvr_flag in Table 8 may be signaled to indicate whether DMVR is disabled for the current slice. In Table 8, if slice-level control is enabled, BDOF and PROF are jointly controlled at the slice level, and DMVR is controlled separately from BDOF and PROF. Figure 20Table 9 is shown, which shows an example syntax structure of a sequence parameter set (SPS) that implements hybrid sequence level control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. In the syntax structure of Table 9, three slice level disable flags can be issued for DMVR, BDOF, and PROF, respectively. As shown in Table 9, highlighted in italics, three separate flags sps_bdof_enabled_flag, sps_dmvr_enabled_flag, and sps_affine_prof_enabled_flag can be signaled in the SPS to indicate whether DMVR, BDOF, or PROF are enabled, respectively. If sps_dmvr_enabled_flag is true, the slice level control enable flag sps_dmvr_slice_present_flag can be signaled to indicate whether DMVR is controlled at the slice level. If at least one of sps_bdof_enabled_flag or sps_affine_prof_enabled_flag is true, a slice level control enable flag sps_bdof_affine_prof_slice_present_flag may be signaled to indicate whether at least one of BDOF or PROF is controlled in a slice level.

[0143] Figure 21 Table 10 is shown, which shows another example syntax structure of a slice header implementing a hybrid control flag for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As shown in the emphasis shown in italics in Table 10, if sps_dmvr_slice_present_flag in Table 9 is set to true, the slice level disable flag slice_disable_dmvr_flag can be signaled to indicate whether DMVR is disabled for the current slice. If sps_bdof_affine_prof_slice_present_flag in Table 9 is set to true, the slice level disable flag slice_disable_bdof_affine_prof_flag can be sent to indicate whether at least one of BDOF or PROF enabled in the sequence level (as described in Table 8) can be disabled for the current slice. In the syntax structure of Table 9, BDOF and PROF can be jointly controlled at the slice level, and if slice level control is enabled, DMVR can be controlled separately at the slice level.

[0144] Figure 22Table 11 is shown, which shows another example syntax structure of a slice header that implements separate control flags for DMVR, BDOF, and PROF according to some embodiments of the present disclosure. As shown in Table 11 with italic highlights, if sps_dmvr_enabled_flag in Table 9 is set to true, the slice level disable flag slice_disable_dmvr_flag can be signaled to indicate whether DMVR is disabled for the current slice. If sps_bdof_enabled_flag and sps_bdof_affine_prof_slice_present_flag in Table 9 are set to true, the slice level disable flag slice_disable_bdof_flag can be signaled to indicate whether BDOF is disabled for the current slice. If sps_affine_prof_enabled_flag and sps_bdof_affine_prof_slice_present_flag in Table 9 are set to true, the slice level disable flag slice_disable_affine_prof_flag may be signaled to indicate whether PROF is disabled for the current slice. In the syntax structure of Table 11, if slice level control is enabled, each of DMVR, BDOF, and PROF is controlled separately in the slice level.

[0145] Figures 23 - 26 Flowcharts showing example processes 2300-2600 for controlling video codec modes according to some embodiments of the present disclosure. In some embodiments, processes 2300-2600 may be performed by a codec (e.g., Figures 2A - 2B The encoder in Figures 3A - 3B For example, the codec may be implemented as one or more software or hardware components of an apparatus (eg, apparatus 400) for controlling a codec mode for encoding or decoding a video sequence.

[0146] Figure 23 FIG. 2 is a flowchart showing an example process 2300 for controlling a video decoding mode according to some embodiments of the present disclosure. In step 2302, a codec (e.g., Figures 3A - 3B A decoder in a processor may receive a bit stream of video data (e.g., Figures 3A - 3B 228 in process 300A or 300B).

[0147] At step 2304, the codec may enable or disable the video sequence based on a first flag in the bitstream (e.g., Figures 3A - 3BThe encoding mode of the video stream 304 in process 300A or 300B in the embodiment of the present invention can be selected from the encoding mode of the video stream 304 in process 300A or 300B in the embodiment of the present invention. For example, the encoding mode can be at least one of a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode, or a decoder-side motion vector correction (DMVR) mode. In some embodiments, the codec can detect a first flag in a sequence parameter set (SPS) of the video sequence. For example, the first flag can be such as Figure 12 , 14A - the flag sps_bdof_enabled_flag, the flag sps_dmvr_enabled_flag or the flag sps_affine_prof_enabled_flag described in 14B, 16 or 20.

[0148] At step 2306, the codec may determine whether to enable or disable control of the coding mode at a level below the sequence level based on a second flag in the bitstream. Levels below the sequence level may include slice level or picture level. In some embodiments, the codec may detect a second flag in the SPS of the video sequence in response to enabling the coding mode for the video sequence. For example, the second flag may be such as Figure 12 , 14A -14B, 16 or 20 describes the flag sps_bdof_dmvr_slice_present_flag, the flag sps_bdof_slice_present_flag, the flag sps_dmvr_slice_present_flag, the flag sps_affine_prof_slice_present_flag, the flag sps_bdof_picture_present_flag, the flag sps_dmvr_picture_present_flag, the flag sps_affine_prof_picture_present_flag, the flag sps_bdof_affine_prof_slice_present_flag, or the flag sps_bdof_dmvr_affine_prof_slice_present_flag.

[0149] In some embodiments, after step 2306, in response to the control of enabling the coding mode at a level lower than the sequence level, the codec may enable or disable the coding mode for the target low-level region based on a third flag in the bitstream. The target low-level region may be a target slice or a target picture. If the low level is a slice level, in some embodiments, the codec may detect the third flag in a slice header of the target slice. If the low level is a picture level, in some embodiments, the codec may detect the third flag in a picture header of the target picture. For example, the third flag may be Figure 13 , 15A - The flag slice_disable_bdof_dmvr_flag, the flag slice_disable_bdof_flag, the flag slice_disable_dmvr_flag, the flag slice_disable_affine_prof_flag, the flag ph_disable_bdof_flag, the flag ph_disable_dmvr_flag, the flag ph_disable_affine_prof_flag, the flag slice_disable_bdof_dmvr_affine_prof_flag, or the flag slice_disable_bdof_affine_prof_flag described in -15B, 17-19, or 21-22.

[0150] Figure 24 FIG. 24 is a flowchart of another example process 2400 for controlling a video decoding mode according to some embodiments of the present disclosure. In step 2402, a codec (e.g., Figures 3A - 3B A decoder in a processor may receive a bit stream of video data (e.g., Figures 3A - 3B 228 in process 300A or 300B).

[0151] At step 2404, the codec may enable or disable a video sequence based on a first flag in the bitstream (e.g., Figures 3A - 3B In step 2406, the codec may enable or disable a second coding mode for the video sequence based on a second flag in the bitstream. The first and second coding modes may be two different coding modes, which may be selected from a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode, and a decoder-side motion vector correction (DMVR) mode. For example, the first coding mode and the second coding mode may be a bidirectional optical flow (BDOF) mode and an optical flow prediction correction (PROF) mode, respectively.

[0152] In some embodiments, the codec may detect the first and second flags in a sequence parameter set (SPS) of the video sequence. For example, the first and second flags may be from Figure 12 , 14A -14B, 16 or 20. As another example, if the first encoding mode and the second encoding mode are BDOF mode and PROF mode, respectively, the first flag and the second flag may be respectively Figure 12 , 14A - The flag sps_bdof_enabled_flag and the flag sps_affine_prof_enabled_flag described in 14B, 16 or 20.

[0153] In step 2408, the codec may determine whether to enable control of at least one of the first coding mode or the second coding mode at a level below the sequence level based on a third flag in the bitstream. The level below the sequence level may include a slice level or a picture level. In some embodiments, the codec may detect a third flag in the SPS of the video sequence in response to enabling at least one of the first coding mode or the second coding mode for the video sequence. For example, the third flag may be a flag sps_bdof_dmvr_slice_present_flag, a flag sps_bdof_slice_present_flag, a flag sps_dmvr_slice_present_flag, a flag sps_affine__slice_present_flag, a flag sps_bdof_picture_present_flag, a flag sps_dmvr_picture_present_flag, a flag sps_affine_prof_picture_present_flag, a flag sps_affine_prof_slice_present_flag, or a flag sps_bdof_dmvr_affine_prof_slice_present_flag as described in FIGS. 12 , 14A-14B, 16 , or 20 .

[0154] In some embodiments, after step 2408, in response to a first flag (e.g., Figure 20) and a third flag (e.g., sps_bdof_enabled_flag) indicating control of enabling at least one of the first encoding mode or the second encoding mode (e.g., PROF) at a level lower than the sequence level. Figure 20 sps_bdof_affine_prof_slice_present_flag) described in , the codec can be based on a fourth flag in the bitstream (e.g., Figure 22 The target lower level region may be a target slice or a target picture. If the target lower level is a target slice, in some embodiments, the codec may detect a fourth flag in a slice header of the target slice. If the target lower level is a target picture, in some embodiments, the codec may detect a fourth flag in a picture header of the target picture. For example, the fourth flag may be Figure 13 , 15A - The flag slice_disable_bdof_dmvr_flag, the flag slice_disable_bdof_flag, the flag slice_disable_dmvr_flag, the flag slice_disable_affine_prof_flag, the flag ph_disable_bdof_flag, the flag ph_disable_dmvr_flag, the flag ph_disable_affine_prof_flag, the flag slice_disable_bdof_dmvr_affine_prof_flag, or the flag slice_disable_bdof_affine_prof_flag described in -15B, 17-19, or 21-22.

[0155] In some embodiments, after step 2408, in response to enabling control of at least one of the first encoding mode or the second encoding mode at a level lower than the sequence level, the codec may, based on a fourth flag (e.g., Figure 21 slice_disable_bdof_affine_prof_flag) enables or disables a first coding mode (e.g., BDOF) and a second coding mode (e.g., PROF) for a target lower level region. Figure 20 sps_bdof_affine_prof_slice_present_flag described in may indicate that at least one of the first encoding mode or the second encoding mode is enabled at a lower level (eg, a slice level).

[0156] In some embodiments, after enabling or disabling the first encoding mode and the second encoding mode of the target lower-level area (e.g., the target slice or the target image) based on the fourth flag in the bitstream, the codec can further enable or disable the third encoding mode of the video sequence according to the second flag in the bitstream, and determine whether to turn on the control of the third encoding mode at a level lower than the sequence level based on the fifth flag in the bitstream. For example, the first encoding mode, the second encoding mode, and the third encoding mode can be BDOF mode, PROF mode, and DMVR mode, respectively. In this example, the fourth flag can be as follows: Figure 21 The flag slice_disable_bdof_affine_prof_flag described in , the second flag can be as follows Figure 20 The fifth flag may be as follows: Figure 20 The flag described in sps_dmvr_slice_present_flag.

[0157] In some embodiments, in response to enabling the third coding mode at a lower level (e.g., slice level or picture level), the codec may further enable or disable the third coding mode for a target lower level region (e.g., target slice or target picture) based on a sixth flag in the bitstream. For example, when the first, second, and third coding modes may be BDOF mode, PROF mode, and DMVR mode, respectively, the sixth flag may be as follows: Figure 21 The slice_disable_dmvr_flag described.

[0158] Figure 25 FIG. 25 shows a flowchart of an example process 2500 for controlling a video encoding mode according to some embodiments of the present disclosure. In step 2502, a codec (e.g., Figures 2A - 2B The encoder in ) can receive a video sequence (for example, Figures 2A - 2B The video sequence 202 in the process 200A or 200B in FIG. 200A and FIG. 200B ), the first mark and the second mark. For example, the first mark can be as follows: Figure 12 , 14A -14B, 16 or 20. As another example, the second flag may be as in Figure 12 , 14A- the flag sps_bdof_dmvr_slice_present_flag, the flag sps_bdof_slice_present_flag, the flag sps_dmvr_slice_present_flag, the flag sps_affine_prof_slice_present_flag, the flag sps_bdof_picture_present_flag, the flag sps_dmvr_picture_present_flag, the flag sps_affine_prof_picture_present_flag, the flag sps_bdof_affine_prof_slice_present_flag, or the flag sps_bdof_dmvr_affine_prof_slice_present_flag as described in 14B, 16 or 20.

[0159] At step 2504, the codec may enable or disable the video bitstream based on a first flag in the bitstream (e.g., Figures 2A - 2B The encoding mode may be a coding mode of the video bitstream 228 in process 200A or 200B in the embodiment of the present invention. For example, the coding mode may be at least one of a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode, or a decoder-side motion vector correction (DMVR) mode.

[0160] At step 2506, the codec may enable or disable control of the encoding mode at a level lower than the sequence level based on the second flag. The level lower than the sequence level may include a slice level or a picture level.

[0161] Figure 26 FIG. 2 is a flowchart of another example process 2600 for controlling a video encoding mode according to some embodiments of the present disclosure. In step 2602, a codec (e.g., Figures 2A - 2B The encoder in ) can receive a video sequence (for example, Figures 2A - 2B In the process 200A or 200B of the video sequence 202), the first mark, the second mark, and the third mark. For example, the first mark and the second mark can be obtained from Figure 12 , 14A -14B, 16 or 20. As another example, the third flag may be selected from the flag sps_bdof_enabled_flag, the flag sps_dmvr_enabled_flag and the flag sps_affine_prof_enabled_flag described in FIG. Figure 12 , 14A- the flag sps_bdof_dmvr_slice_present_flag, the flag sps_bdof_slice_present_flag, the flag sps_dmvr_slice_present_flag, the flag sps_affine_prof_slice_present_flag, the flag sps_bdof_picture_present_flag, the flag sps_dmvr_picture_present_flag, the flag sps_affine_prof_picture_present_flag, the flag sps_bdof_affine_prof_slice_present_flag, or the flag sps_bdof_dmvr_affine_prof_slice_present_flag described in 14B, 16 or 20.

[0162] At step 2604, the codec may provide a video bitstream (eg, Figures 2A - 2B The first coding mode is enabled or disabled based on the video bitstream 228 in process 200A or 200B in step 2606. At step 2607, the codec may enable or disable the second coding mode for the video bitstream based on the second flag. The first coding mode and the second coding mode may be two different coding modes, which may be selected from a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode, and a decoder-side motion vector correction (DMVR) mode. For example, the first coding mode and the second coding mode may be a bidirectional optical flow (BDOF) mode and an optical flow prediction correction (PROF) mode, respectively.

[0163] At step 2608, the codec may enable or disable control of at least one of the first encoding mode or the second encoding mode at a level below the sequence level based on the third flag. The level below the sequence level may include a slice level or a picture level.

[0164] In some embodiments, a non-transitory computer-readable storage medium including instructions is also provided, and the instructions can be executed by a device (e.g., the disclosed encoder and decoder) to perform the above method. Common forms of non-transitory media include, for example, floppy disks, disks, hard disks, solid-state drives, tapes or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with hole patterns, RAM, PROM and EPROM, FLASH-EPROM or any other flash memory, NVRAM, cache, registers, any other memory chip or cartridge memory, and network versions thereof. The device may include one or more processors (CPUs), input / output interfaces, network interfaces, and / or memories.

[0165] The embodiments may be further described using the following terms: 1. A computer-implemented method comprising: receiving a bit stream of video data; enabling or disabling a coding mode for a video sequence based on a first flag in the bitstream; and Based on a second flag in the bitstream, it is determined whether to enable or disable control of the encoding mode at a level below the sequence level. 2. A computer-implemented method according to clause 1, wherein the level below the sequence level includes a slice level or a picture level. 3. A computer-implemented method according to any of clauses 1-2, further comprising: In response to enabling control of the encoding mode at the level lower than the sequence level, the encoding mode for a target low-level region is enabled or disabled based on a third flag in the bitstream. 4. The computer-implemented method of clause 3, further comprising: detecting the third flag in a slice header of a target slice, wherein the target slice is the target low-level region; or The third mark in the image header of a target image is detected, wherein the target image is the target low-level area. 5. A computer-implemented method according to any of clauses 1 to 4, wherein the encoding mode is at least one of: Bidirectional optical flow (BDOF) mode; Optical Flow Prediction Correction (PROF) mode; or Decoder-side motion vector correction (DMVR) mode. 6. A computer-implemented method according to any one of clauses 1 to 5, further comprising: A first marker in a sequence parameter set (SPS) of the video sequence is detected. 7. A computer-implemented method according to any one of clauses 1 to 6, further comprising: In response to enabling the encoding mode for the video sequence, the second flag is detected in the SPS of the video sequence. 8. A computer-implemented method comprising: receiving a bit stream of video data; enabling or disabling a first coding mode for a video sequence based on a first flag in the bitstream; enabling or disabling a second coding mode for the video sequence based on a second flag in the bitstream; and Based on a third flag in the bitstream, it is determined whether to enable control of at least one of the first encoding mode or the second encoding mode at a level below a sequence level. 9. A computer-implemented method according to clause 8, wherein the level below the sequence level comprises a slice level or a picture level. 10. A computer-implemented method according to any of clauses 8-9, further comprising: In response to enabling at least one of the first encoding mode or the second encoding mode at the level lower than the sequence level, enabling or disabling the first encoding mode and the second encoding mode for a target low-level area based on a fourth flag in the bitstream. 11. The computer-implemented method of clause 10, further comprising: detecting a fourth flag in a slice header of a target slice, wherein the target slice is the target low-level region; or A fourth marker in an image header of a target image is detected, wherein the target image is the target low-level region. 12. A computer-implemented method according to any of clauses 10-11, further comprising: enabling or disabling a third coding mode for the video sequence based on the second flag in the bitstream; and Based on a fifth flag in the bitstream, it is determined whether control of the third encoding mode is enabled at the level lower than the sequence level. 13. The computer-implemented method of clause 12, further comprising: In response to enabling the third coding mode at the level lower than the sequence level, enabling or disabling the third coding mode for the target low-level region based on a sixth flag in the bitstream. 14. A computer-implemented method according to any one of clauses 12 and 13, wherein the first encoding mode, the second encoding mode and the third encoding mode are respectively a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode and a decoder-side motion vector correction (DMVR) mode. 15. A computer-implemented method according to any of clauses 8 to 14, further comprising: In response to enabling a first flag indicating a first encoding mode for a video sequence and a third flag indicating enabling control of at least one of the first encoding mode or the second encoding mode at the level lower than the sequence level, the first encoding mode for a target lower level area is enabled or disabled based on a fourth flag in the bitstream. 16. A computer-implemented method according to any of clauses 8 to 15, wherein the first encoding mode and the second encoding mode are two different encoding modes selected from the following: Bidirectional optical flow (BDOF) mode; Optical Flow Prediction Correction (PROF) mode; and Decoder-side motion vector correction (DMVR) mode. 17. A computer-implemented method according to any of clauses 8-16, wherein the first encoding mode and the second encoding mode are a Bidirectional Optical Flow (BDOF) mode and a Prediction Correction of Optical Flow (PROF) mode, respectively. 18. A computer-implemented method according to any of clauses 8 to 17, further comprising: A first flag and a second flag in a sequence parameter set (SPS) of a video sequence are detected. 19. A computer-implemented method according to any one of clauses 8 to 18, further comprising: In response to enabling at least one of the first encoding mode or the second encoding mode for the video sequence, detecting a third flag in an SPS of the video sequence. 20. A computer-implemented method comprising: receiving a video sequence, a first marker, and a second marker; enabling or disabling a coding mode for a video bitstream based on the first flag; and Based on the second flag, control of the encoding mode is enabled or disabled at a level below the sequence level. 21. A computer-implemented method according to clause 20, wherein the level below the sequence level comprises a slice level or a picture level. 22. A computer-implemented method according to any of clauses 20-21, further comprising: receiving a third flag; and In response to enabling or disabling control of the encoding mode at the level lower than the sequence level, the encoding mode for a target low-level region is enabled or disabled based on the third flag. 23. The computer-implemented method of clause 22, further comprising: storing the third flag in a slice header of a target slice, wherein the target slice is the target low-level region; or The third flag is stored in an image header of a target image, wherein the target image is the target low-level area. 24. A computer-implemented method according to any of clauses 20-23, wherein the encoding mode is at least one of: Bidirectional optical flow (BDOF) mode; Optical Flow Prediction Correction (PROF) mode; or Decoder side motion vector correction (DMVR) mode. 25. A computer-implemented method according to any of clauses 20-24, further comprising: The first flag is stored in a sequence parameter set (SPS) of the video bitstream. 26. A computer-implemented method according to any of clauses 20-25, further comprising: In response to enabling or disabling the encoding mode of the video bitstream, the second flag is stored in the SPS of the video bitstream. 27. A computer-implemented method comprising: receiving a video sequence, a first flag, a second flag, and a third flag; enabling or disabling a first coding mode for a video bitstream based on the first flag; enabling or disabling a second encoding mode for the video bitstream based on the second flag; and Based on the third flag, control of at least one of the first encoding mode or the second encoding mode is enabled or disabled at a level below a sequence level. 28. A computer-implemented method according to clause 27, wherein the level below the sequence level comprises a slice level or a picture level. 29. A computer-implemented method according to any of clauses 27-28, further comprising: receiving a fourth flag; and In response to enabling control of the first encoding mode of the video bitstream based on the first flag and enabling control of at least one of the first encoding mode or the second encoding mode at the level below the sequence level based on the third flag, enabling or disabling the first encoding mode for a target low-level area according to the fourth flag. 30. A computer-implemented method according to any of clauses 27-28, further comprising: Accept the fourth sign; and In response to enabling or disabling control of at least one of the first encoding mode or the second encoding mode at the level lower than the sequence level, enabling or disabling the first encoding mode and the second encoding mode for a target low-level area based on the fourth flag. 31. A computer-implemented method according to any of clauses 29-30, further comprising: storing the fourth flag in a slice header of a target slice, wherein the target slice is the target low-level region; or The fourth flag is stored in an image header of a target image, wherein the target image is a target low-level area. 32. The computer-implemented method of clause 31, further comprising: Accept the fifth sign; enabling or disabling a third encoding mode for the video bitstream based on the second flag; and Based on the fifth flag, control of the third encoding mode is enabled or disabled at the level lower than the sequence level. 33. The computer-implemented method of any one of clauses 31 and 32, further comprising: Accept the Sixth Mark; and In response to enabling or disabling control of the third encoding mode at a level lower than the sequence level, the third encoding mode for the target low-level region is enabled or disabled based on the sixth flag. 34. A computer-implemented method according to any of clauses 27-33, wherein the first encoding mode, the second encoding mode and the third encoding mode are respectively a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode and a decoder-side motion vector correction (DMVR) mode. 35. A computer-implemented method according to any of clauses 27-33, wherein the first encoding mode and the second encoding mode are two different encoding modes selected from the group consisting of: Bidirectional optical flow (BDOF) mode; Optical Flow Prediction Correction (PROF) mode; and Decoder side motion vector refinement (DMVR) mode. 36. A computer-implemented method according to any of clauses 27-35, wherein the first encoding mode and the second encoding mode are a Bidirectional Optical Flow (BDOF) mode and a Prediction Correction of Optical Flow (PROF) mode, respectively. 37. A computer-implemented method according to any of clauses 27-36, further comprising: The first flag and the second flag are stored in a sequence parameter set (SPS) of the video bitstream. 38. A computer-implemented method according to any of clauses 27-37, further comprising: In response to enabling at least one of the first encoding mode or the second encoding mode for the video bitstream, storing the third flag in an SPS of the video sequence. 39. A non-transitory computer readable medium storing a set of instructions executable by at least one processor of a device to cause the device to perform a method comprising: receiving a bit stream of video data; enabling or disabling a coding mode for a video sequence based on a first flag in the bitstream; and Based on a second flag in the bitstream, it is determined whether to enable or disable control of the encoding mode at a level below the sequence level. 40. The non-transitory computer-readable medium of clause 39, wherein the level below the sequence level comprises a slice level or a picture level. 41. The non-transitory computer-readable medium of any of clauses 39-40, wherein the set of instructions executable by the at least one processor of the apparatus causes the apparatus to further perform: In response to enabling control of the coding mode at the level lower than the sequence level, a coding mode for a target low-level region is enabled or disabled based on a third flag in the bitstream. 42. The non-transitory computer-readable medium of clause 41, wherein the set of instructions executable by the at least one processor of the device causes the device to further perform: detecting the third flag in a slice header of a target slice, wherein the target slice is the target low-level region; or The third mark in the image header of a target image is detected, wherein the target image is a target low-level area. 43. The non-transitory computer-readable medium of any of clauses 39-42, wherein the encoding mode is at least one of: Bidirectional optical flow (BDOF) mode; Optical Flow Prediction Correction (PROF) mode; or Decoder side motion vector correction (DMVR) mode. 44. The non-transitory computer-readable medium of any of clauses 39-43, wherein the set of instructions executable by the at least one processor of the apparatus causes the apparatus to further perform: The first flag is detected in a sequence parameter set (SPS) of the video sequence. 45. The non-transitory computer-readable medium of any of clauses 39-44, wherein the set of instructions executable by the at least one processor of the apparatus causes the apparatus to further perform: In response to enabling the coding mode for the video sequence, detecting the second flag in the SPS of the video sequence. 46. ​​A non-transitory computer readable medium storing a set of instructions executable by at least one processor of a device to cause the device to perform a method comprising: receiving a bit stream of video data; enabling or disabling a first coding mode for a video sequence based on a first flag in the bitstream; enabling or disabling a second coding mode for the video sequence based on a second flag in the bitstream; and Based on a third flag in the bitstream, it is determined whether to enable control of at least one of the first encoding mode or the second encoding mode at a level below a sequence level. 47. The non-transitory computer-readable medium of clause 46, wherein the level below the sequence level comprises a slice level or a picture level. 48. The non-transitory computer-readable medium of any of clauses 46-47, wherein the set of instructions executable by the at least one processor of the apparatus causes the apparatus to further perform: In response to enabling control of at least one of the first encoding mode or the second encoding mode at the level lower than the sequence level, the first encoding mode and the second encoding mode for a target low-level area are enabled or disabled based on a fourth flag in the bitstream. 49. The non-transitory computer-readable medium of clause 48, wherein the set of instructions executable by the at least one processor of the device causes the device to further perform: detecting the fourth flag in a slice header of a target slice, wherein the target slice is the target low-level region; or The fourth mark in an image header of a target image is detected, wherein the target image is the target low-level area. 50. The non-transitory computer-readable medium of any of clauses 48-49, wherein the set of instructions executable by the at least one processor of the apparatus causes the apparatus to further perform: enabling or disabling a third coding mode for a video sequence based on the second flag in the bitstream; and Based on a fifth flag in the bitstream, it is determined whether control of the third encoding mode is enabled at the level lower than the sequence level. 51. The non-transitory computer-readable medium of clause 50, wherein the set of instructions executable by the at least one processor of the device causes the device to further perform: In response to enabling the third coding mode at the level lower than the sequence level, the third coding mode for the target low-level region is enabled or disabled based on a sixth flag in the bitstream. 52. A non-transitory computer-readable medium according to any one of clauses 50 and 51, wherein the first encoding mode, the second encoding mode and the third encoding mode are respectively a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode and a decoder-side motion vector correction (DMVR) mode. 53. The non-transitory computer-readable medium of any of clauses 46-52, wherein the set of instructions executable by the at least one processor of the apparatus causes the apparatus to further perform: In response to the first flag indicating enabling of the first encoding mode for the video sequence and the third flag indicating enabling of control of at least one of the first encoding mode or the second encoding mode at the level lower than the sequence level, the first encoding mode for a target lower level area is enabled or disabled based on a fourth flag in the bitstream. 54. The non-transitory computer-readable medium of any of clauses 46-53, wherein the first encoding mode and the second encoding mode are two different encoding modes selected from the group consisting of: Bidirectional optical flow (BDOF) mode; Optical Flow Prediction Correction (PROF) mode; and Decoder side motion vector correction (DMVR) mode. 55. The non-transitory computer-readable medium of any of clauses 46-54, wherein the first encoding mode and the second encoding mode are a Bidirectional Optical Flow (BDOF) mode and a Prediction Correction of Optical Flow (PROF) mode, respectively. 56. The non-transitory computer-readable medium of any of clauses 46-55, wherein the set of instructions executable by the at least one processor of the apparatus causes the apparatus to further perform: A first flag and a second flag in a sequence parameter set (SPS) of the video sequence are detected. 57. The non-transitory computer-readable medium of any of clauses 46-56, wherein the set of instructions executable by the at least one processor of the apparatus causes the apparatus to further perform: In response to enabling at least one of the first encoding mode or the second encoding mode for the video sequence, detecting the third flag in the SPS of the video sequence. 58. A non-transitory computer readable medium storing a set of instructions executable by at least one processor of a device to cause the device to perform a method comprising: receiving a video sequence, a first marker, and a second marker; enabling or disabling a coding mode for a video bitstream based on the first flag; and Based on the second flag, control of the encoding mode is enabled or disabled at a level below the sequence level. 59. The non-transitory computer-readable medium of clause 58, wherein the level below the sequence level comprises a slice level or a picture level. 60. The non-transitory computer-readable medium of any of clauses 58-59, wherein the set of instructions executable by at least one processor of the device causes the device to further perform: receiving a third flag; and In response to enabling or disabling control of the encoding mode at the level lower than the sequence level, an encoding mode for a target low-level region is enabled or disabled based on the third flag. 61. The non-transitory computer-readable medium of clause 60, wherein the set of instructions executable by the at least one processor of the device causes the device to further perform: storing the third flag in a slice header of a target slice, wherein the target slice is the target low-level region; or The third flag is stored in an image header of a target image, wherein the target image is the target low-level area. 62. The non-transitory computer-readable medium of any of clauses 58-61, wherein the encoding mode is at least one of: Bidirectional optical flow (BDOF) mode; Optical Flow Prediction Correction (PROF) mode; or Decoder side motion vector correction (DMVR) mode. 63. The non-transitory computer-readable medium of any of clauses 58-62, wherein the set of instructions executable by the at least one processor of the apparatus causes the apparatus to further perform: The first flag is stored in a sequence parameter set (SPS) of the video bitstream. 64. The non-transitory computer-readable medium of any of clauses 58-63, wherein the set of instructions executable by the at least one processor of the apparatus causes the apparatus to further perform: Responsive to enabling or disabling a coding mode for the video bitstream, storing the second flag in an SPS of the video bitstream. 65. A non-transitory computer readable medium storing a set of instructions executable by at least one processor of a device to cause the device to perform a method comprising: receiving a video sequence, a first flag, a second flag, and a third flag; enabling or disabling a first coding mode for a video bitstream based on the first flag; enabling or disabling a second encoding mode for the video bitstream based on the second flag; and Based on the third flag, control of at least one of the first encoding mode or the second encoding mode is enabled or disabled at a level below a sequence level. 66. The non-transitory computer-readable medium of clause 65, wherein the level below the sequence level comprises a slice level or a picture level. 67. The non-transitory computer-readable medium of any of clauses 65-66, wherein the set of instructions executable by the at least one processor of the apparatus causes the apparatus to further perform: receiving a fourth flag; and In response to enabling control of the first encoding mode for the video bitstream based on the first flag and enabling control of at least one of the first encoding mode or the second encoding mode at the level below the sequence level based on the third flag, enabling or disabling the first encoding mode for a target low-level area based on the fourth flag. 68. The non-transitory computer-readable medium of any of clauses 65-66, wherein the set of instructions executable by the at least one processor of the apparatus causes the apparatus to further perform: receiving a fourth flag; and In response to enabling or disabling control of at least one of the first encoding mode or the second encoding mode at the level lower than the sequence level, enabling or disabling the first encoding mode and the second encoding mode for a target low-level area based on the fourth flag. 69. The non-transitory computer-readable medium of any of clauses 67-68, wherein the set of instructions executable by the at least one processor of the apparatus causes the apparatus to further perform: storing the fourth flag in a slice header of a target slice, wherein the target slice is the target low-level region; or The fourth flag is stored in an image header of a target image, wherein the target image is the target low-level area. 70. The non-transitory computer-readable medium of clause 69, wherein the set of instructions executable by the at least one processor of the device causes the device to further perform: Receive the fifth sign; enabling or disabling a third encoding mode for the video bitstream based on the second flag; and Based on the fifth flag, control of the third encoding mode is enabled or disabled at the level lower than the sequence level. 71. The non-transitory computer-readable medium of any of clauses 69 and 70, wherein the set of instructions executable by the at least one processor of the apparatus causes the apparatus to further perform: receiving a sixth flag; and In response to enabling or disabling control of the third encoding mode at a level lower than the sequence level, the third encoding mode for the target low-level region is enabled or disabled based on the sixth flag. 72. A non-transitory computer-readable medium according to any one of clauses 65-71, wherein the first encoding mode, the second encoding mode and the third encoding mode are respectively a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode and a decoder-side motion vector correction (DMVR) mode. 73. The non-transitory computer-readable medium of any of clauses 65-72, wherein the first encoding mode and the second encoding mode are two different encoding modes selected from: Bidirectional optical flow (BDOF) mode; Optical Flow Prediction Correction (PROF) mode; and Decoder side motion vector refinement (DMVR) mode. 74. The non-transitory computer-readable medium of any of clauses 65-73, wherein the first encoding mode and the second encoding mode are a bidirectional optical flow (BDOF) mode and a prediction correction of optical flow (PROF) mode, respectively. 75. The non-transitory computer-readable medium of any of clauses 65-74, wherein the set of instructions executable by the at least one processor of the apparatus causes the apparatus to further perform: The first flag and the second flag are stored in a sequence parameter set (SPS) of the video bitstream. 76. The non-transitory computer-readable medium of any of clauses 65-75, wherein the set of instructions executable by the at least one processor of the apparatus causes the apparatus to further perform: In response to enabling at least one of the first encoding mode or the second encoding mode for the video bitstream, storing the third flag in an SPS of the video sequence. 77. A device comprising: a memory configured to store a set of instructions; and one or more processors communicatively coupled to the memory and configured to execute the set of instructions to cause the apparatus to: receiving a bit stream of video data; enabling or disabling a coding mode for a video sequence based on a first flag in the bitstream; and Based on a second flag in the bitstream, it is determined whether to enable or disable control of the encoding mode at a level below the sequence level. 78. An apparatus as described in clause 77, wherein the level below the sequence level includes a slice level or a picture level. 79. An apparatus as described in any of clauses 77-78, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: In response to enabling control of the coding mode at the level lower than the sequence level, a coding mode for a target low-level region is enabled or disabled based on a third flag in the bitstream. 80. The apparatus of clause 79, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: detecting the third flag in a slice header of a target slice, wherein the target slice is the target low-level region; or The third mark in the image header of a target image is detected, wherein the target image is a target low-level area. 81. Apparatus according to any of clauses 77 to 80, wherein the encoding mode is at least one of: Bidirectional optical flow (BDOF) mode; Optical Flow Prediction Correction (PROF) mode; or Decoder side motion vector correction (DMVR) mode. 82. An apparatus as described in any of clauses 77-81, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: The first flag is detected in a sequence parameter set (SPS) of the video sequence. 83. An apparatus as described in any of clauses 77-82, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: In response to enabling the coding mode for the video sequence, detecting the second flag in the SPS of the video sequence. 84. A device comprising: a memory configured to store a set of instructions; and one or more processors communicatively coupled to the memory and configured to execute the set of instructions to cause the apparatus to: receiving a bit stream of video data; enabling or disabling a first coding mode for a video sequence based on a first flag in the bitstream; enabling or disabling a second coding mode for the video sequence based on a second flag in the bitstream; and Based on a third flag in the bitstream, it is determined whether to enable control of at least one of the first encoding mode or the second encoding mode at a level below a sequence level. 85. An apparatus as described in clause 84, wherein the level below the sequence level comprises a slice level or a picture level. 86. An apparatus as described in any of clauses 84-85, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: In response to enabling control of at least one of the first encoding mode or the second encoding mode at the level lower than the sequence level, the first encoding mode and the second encoding mode for a target low-level area are enabled or disabled based on a fourth flag in the bitstream. 87. An apparatus as described in clause 86, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: detecting the fourth flag in a slice header of a target slice, wherein the target slice is the target low-level region; or The fourth mark in an image header of a target image is detected, wherein the target image is the target low-level area. 88. An apparatus as described in any of clauses 86-87, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: enabling or disabling a third coding mode for a video sequence based on the second flag in the bitstream; and Based on a fifth flag in the bitstream, it is determined whether control of the third encoding mode is enabled at the level lower than the sequence level. 89. The apparatus of clause 88, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: In response to enabling the third coding mode at the level lower than the sequence level, enabling or disabling the third coding mode for the target low-level region based on a sixth flag in a bitstream. 90. An apparatus according to any one of clauses 88 and 89, wherein the first coding mode, the second coding mode and the third coding mode are respectively a bidirectional optical flow (BDOF) mode, an optical flow prediction correction (PROF) mode and a decoder-side motion vector correction (DMVR) mode. 91. An apparatus as described in any of clauses 84-90, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: In response to the first flag indicating enabling of the first encoding mode for the video sequence and the third flag indicating enabling of control of at least one of the first encoding mode or the second encoding mode at the level lower than the sequence level, the first encoding mode for a target lower level area is enabled or disabled based on a fourth flag in the bitstream. 92. Apparatus according to any of clauses 84 to 91, wherein the first encoding mode and the second encoding mode are two different encoding modes selected from: Bidirectional optical flow (BDOF) mode; Optical Flow Prediction Correction (PROF) mode; and Decoder side motion vector correction (DMVR) mode. 93. Apparatus according to any of clauses 84-92, wherein the first encoding mode and the second encoding mode are respectively a Bidirectional Optical Flow (BDOF) mode and a Prediction Correction of Optical Flow (PROF) mode. 94. An apparatus as described in any of clauses 84-93, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: A first flag and a second flag in a sequence parameter set (SPS) of the video sequence are detected. 95. An apparatus as described in any of clauses 84-94, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: In response to enabling at least one of the first encoding mode or the second encoding mode for the video sequence, the third flag in the SPS of the video sequence is detected. 96. A device comprising: a memory configured to store a set of instructions; and one or more processors, the one or more processors being communicatively coupled to the memory and configured to execute the set of instructions so that the apparatus: receiving a video sequence, a first marker, and a second marker; enabling or disabling a coding mode for a video bitstream based on the first flag; and Based on the second flag, control of the encoding mode is enabled or disabled at a level below the sequence level. 97. An apparatus as described in clause 96, wherein the level below the sequence level includes a slice level or a picture level. 98. An apparatus as described in any of clauses 96-97, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: receiving the third flag; and In response to enabling or disabling control of the encoding mode at the level lower than the sequence level, an encoding mode for a target low-level region is enabled or disabled based on the third flag. 99. The apparatus of clause 98, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: storing the third flag in a slice header of a target slice, wherein the target slice is the target low-level region; or The third flag is stored in an image header of a target image, wherein the target image is the target low-level area. 100. Apparatus according to any of clauses 96 to 99, wherein the encoding mode is at least one of: Bidirectional optical flow (BDOF) mode; Optical Flow Prediction Correction (PROF) mode; or Decoder side motion vector correction (DMVR) mode. 101. An apparatus as described in any of clauses 96-100, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: The first flag is stored in a sequence parameter set (SPS) of the video bitstream. 102. An apparatus as described in any of clauses 96-101, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: Responsive to enabling or disabling a coding mode for the video bitstream, storing the second flag in an SPS of the video bitstream. 103. A device comprising: a memory configured to store a set of instructions; and one or more processors, the one or more processors being communicatively coupled to the memory and configured to execute the set of instructions so that the apparatus: receiving a video sequence, a first marker, a second marker, and a third marker; enabling or disabling a first encoding mode for a video bitstream based on the first flag; enabling or disabling a second encoding mode for the video bitstream based on the second flag; and Based on the third flag, control of at least one of the first encoding mode or the second encoding mode is enabled or disabled at a level below a sequence level. 104. An apparatus as described in clause 102, wherein the level below the sequence level comprises a slice level or a picture level. 105. An apparatus as described in any of clauses 103-104, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: receiving a fourth flag; and In response to enabling control of the first encoding mode for the video bitstream based on the first flag and enabling control of at least one of the first encoding mode or the second encoding mode at the level below the sequence level based on the third flag, enabling or disabling the first encoding mode for a target low-level area based on the fourth flag. 106. An apparatus as described in any of clauses 103-104, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: receiving a fourth flag; and In response to enabling or disabling control of at least one of the first encoding mode or the second encoding mode at the level lower than the sequence level, enabling or disabling the first encoding mode and the second encoding mode for a target low-level area based on the fourth flag. 107. An apparatus as described in any of clauses 105-106, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: storing the fourth flag in a slice header of a target slice, wherein the target slice is the target low-level region; or The fourth flag is stored in an image header of a target image, wherein the target image is the target low-level area. 108. The apparatus of clause 107, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: Receive the fifth sign; enabling or disabling a third encoding mode for the video bitstream based on the second flag; and Based on the fifth flag, control of the third encoding mode is enabled or disabled at the level lower than the sequence level. 109. An apparatus as described in any of clauses 107 and 108, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: receiving a sixth flag; and In response to enabling or disabling control of the third encoding mode at a level lower than the sequence level, the third encoding mode for the target low-level region is enabled or disabled based on the sixth flag. 110. An apparatus according to any of clauses 103-109, wherein the first coding mode, the second coding mode and the third coding mode are respectively a Bidirectional Optical Flow (BDOF) mode, an Optical Flow Prediction Correction (PROF) mode and a Decoder-Side Motion Vector Correction (DMVR) mode. 111. An apparatus as described in any of clauses 103-110, wherein the first encoding mode and the second encoding mode are two different encoding modes selected from the following: Bidirectional optical flow (BDOF) mode; Optical Flow Prediction Correction (PROF) mode; and Decoder side motion vector refinement (DMVR) mode. 112. An apparatus according to any of clauses 103-110, wherein the first encoding mode and the second encoding mode are a Bidirectional Optical Flow (BDOF) mode and a Prediction Correction of Optical Flow (PROF) mode, respectively. 113. An apparatus as described in any of clauses 103-112, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: The first flag and the second flag are stored in a sequence parameter set (SPS) of the video bitstream. 114. An apparatus as described in any of clauses 103-113, wherein the one or more processors are further configured to execute the set of instructions to cause the apparatus to: In response to enabling at least one of the first encoding mode or the second encoding mode for the video bitstream, storing the third flag in an SPS of the video sequence.

[0166] It should be noted that the relational terms such as "first" and "second" in this article are only used to distinguish one entity or operation from another entity or operation, and do not require or imply any actual relationship or order between these entities or operations. In addition, the words "including", "having", "including" and "comprising" and other similar forms are intended to mean the same and are open-ended, in that the one or more items following any of these words are not intended to be an exhaustive list of such items or items, or to be limited to the listed items or items.

[0167] As used herein, unless expressly stated otherwise, the term "or" encompasses all possible combinations unless not feasible. For example, if it is stated that a component may include A or B, then unless expressly stated otherwise or not feasible, the component may include A, or B, or A and B. As a second example, if it is stated that a component may include A, B, or C, then unless expressly stated otherwise or not feasible, the component may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.

[0168] It can be understood that the above embodiments can be implemented by hardware, software (program code) or a combination of hardware and software. If implemented by software, it can be stored in the above computer-readable medium. The software can execute the disclosed method when executed by a processor. The computing unit and other functional units described in the present invention can be implemented by hardware, software, or a combination of hardware and software. It can also be understood by those of ordinary skill in the art that multiple of the above modules / units can be combined into one module / unit, and each of the above modules / units can be further divided into multiple sub-modules / sub-units.

[0169] In the foregoing description, the embodiments have been described with reference to many specific details, which may vary from implementation to implementation. Certain modifications and changes may be made to the described embodiments. In view of the description and practice of the invention disclosed herein, other embodiments will be apparent to those skilled in the art. The description and examples are to be regarded as examples only, and the true scope and spirit of the invention are indicated by the appended claims. The order of steps shown in the figures is also intended to be used for illustrative purposes only and is not intended to be limited to any particular order of steps. Therefore, it will be appreciated by those skilled in the art that these steps may be performed in different orders when implementing the same method.

[0170] In the drawings and the specification, exemplary embodiments have been disclosed. However, many variations and modifications may be made to these embodiments. Therefore, although specific terms are used, they are used only for a general and descriptive purpose and not for a limiting purpose.

Claims

1. A computer-implemented decoding method, comprising: receiving a bitstream associated with a video sequence; decoding a first flag in a sequence parameter set (SPS) of the bitstream; respectively determining whether a second flag exists in the bitstream based on a value of the first flag; wherein the second flag indicates whether an encoding mode is enabled or disabled at a lower level below the sequence level; the encoding mode includes a bi-directional optical flow (BDOF) mode; dividing an encoding block into sub-blocks, and when the second flag exists in the bitstream, decoding based on a value of the second flag.

2. The computer-implemented decoding method according to claim 1, wherein, the second flag includes: a control flag associated with the BDOF encoding mode; wherein the method further includes: when the control flag associated with the BDOF encoding mode exists in the bitstream, enabling or disabling the BDOF mode at the lower level below the sequence level based on the control flag associated with the BDOF encoding mode.

3. The computer-implemented decoding method according to claim 1, further comprising: decoding a third flag in the SPS of the bitstream; enabling or disabling the encoding mode at the lower level below the sequence level respectively based on a value of the third flag.

4. The computer-implemented decoding method according to claim 1, wherein the lower level below the sequence level is a slice level or a picture level.

5. The computer-implemented decoding method according to claim 1, the method further comprising: in response to the first flag having a first value, determining that the second flag exists in a slice header or a picture header of the bitstream.

6. The computer-implemented decoding method according to claim 5, wherein the first value is 1.

7. A decoding apparatus, comprising: a memory configured to store an instruction set; and one or more processors configured to execute the instruction set to cause the apparatus to perform: receiving a bitstream associated with a video sequence; decoding a first flag in a sequence parameter set (SPS) of the bitstream; respectively determining whether a second flag exists in the bitstream based on a value of the first flag; wherein the second flag indicates whether an encoding mode is enabled or disabled at a lower level below the sequence level; the encoding mode includes a bi-directional optical flow (BDOF) mode; dividing an encoding block into sub-blocks, and when the second flag exists in the bitstream, decoding based on a value of the second flag.

8. The decoding apparatus according to claim 7, wherein the second flag includes: a control flag associated with the BDOF encoding mode; wherein one or more of the processors are further configured to execute the instruction set to cause the apparatus to perform: when the control flag associated with the BDOF encoding mode exists in the bitstream, enabling or disabling the BDOF mode at the lower level below the sequence level based on the control flag associated with the BDOF encoding mode.

9. The decoding apparatus according to claim 7, wherein one or more of the processors are further configured to execute the instruction set to cause the apparatus to perform: Decode a third flag in the SPS of the bitstream; Enable or disable an encoding mode at a lower level below the sequence level based on the value of the third flag.

10. The decoding apparatus according to claim 7, wherein the lower level below the sequence level is a slice level or a picture level.

11. The decoding apparatus according to claim 7, wherein one or more of the processors are further configured to execute the instruction set to cause the apparatus to perform: In response to the first flag having a first value, determine that the second flag is present in the slice header or picture header of the bitstream.

12. The decoding apparatus according to claim 11, wherein the first value is 1.

13. A non-transitory computer-readable storage medium storing a bitstream associated with a video sequence, the non-transitory computer-readable storage medium being part of a computing device configured to execute a set of instructions to cause the computing device to decode the bitstream according to operations including the following: Receive a bitstream associated with a video sequence; Decode a first flag in a sequence parameter set (SPS) of the bitstream; Based on the value of the first flag, determine whether a second flag is present in the bitstream; wherein the second flag indicates whether an encoding mode is enabled or disabled at a lower level below the sequence level; The encoding mode includes a bidirectional optical flow (BDOF) mode; Divide an encoded block into sub-blocks, and when the second flag is present in the bitstream, decode based on the value in the second flag.

14. The non-transitory computer-readable storage medium according to claim 13, wherein, the operations further include: Decode a third flag in the SPS of the bitstream; Enable or disable an encoding mode at a lower level below the sequence level based on the value of the third flag.

15. The non-transitory computer-readable storage medium according to claim 13, wherein the lower level below the sequence level is a slice level or a picture level.

16. The non-transitory computer-readable storage medium according to claim 13, wherein the SPS of the bitstream further includes: A picture header, wherein the second flag is included in the picture header.

17. The non-transitory computer-readable storage medium according to claim 13, wherein the SPS of the bitstream further includes: A slice header, wherein the flag is included in the slice header.

18. A computer-implemented encoding method, including: Encode a first flag in a sequence parameter set (SPS) of a bitstream associated with a video sequence; Based on the value of the first flag respectively, determine whether to signal a second flag in the bitstream; The second flag indicates whether an encoding mode is enabled or disabled at a lower level below the sequence level; The encoding mode includes a bidirectional optical flow (BDOF) mode; Divide an encoded block into sub-blocks, and when it is determined to signal the second flag in the bitstream, encode based on the value of the second flag.

19. The computer-implemented encoding method according to claim 18, wherein the second flag includes: A control flag associated with the BDOF encoding mode; Wherein, the method further comprises: When it is determined to signal a control flag associated with the BDOF coding mode, enabling or disabling the BDOF mode at a lower level below the sequence level.

20. The computer-implemented coding method according to claim 18, further comprises: Encoding a third flag in the SPS of the bitstream; The third flag respectively indicates whether the coding mode is enabled or disabled for the video sequence.

21. The computer-implemented coding method according to claim 18, wherein, The lower level below the sequence level is the slice level or the picture level.

22. The computer-implemented coding method according to claim 18, the method further comprises: In response to the first flag having a first value, encoding in the slice header or picture header of the bitstream based on the second flag.

23. The computer-implemented coding method according to claim 22, wherein the first value is 1.

Citation Information

Patent Citations

  • Methods of palette coding with inter-prediction in video coding

    CN107409227A

  • Method and System for Lossless Coding Mode in Video Coding

    US20130077696A1