Enhancements to block-adaptive weighted predictions

Enhancements to block adaptive weighted prediction (BAWP) address local illumination issues in video coding and decoding, enhancing compression and transmission efficiency by modeling local illumination compensation (LIC) using scaling factor lookup tables and linear equations.

JP2026518092APending Publication Date: 2026-06-04TENCENT AMERICA LLC

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2023-09-12
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing video coding and decoding technologies struggle to effectively handle local illumination variations, leading to inefficiencies in compression and transmission of video data.

Method used

Enhancements to block adaptive weighted prediction (BAWP) are introduced to model local illumination compensation (LIC) through the use of scaling factor lookup tables and linear equations for predicting current blocks based on reference blocks, incorporating memory and processor-based decoding devices and methods.

Benefits of technology

Improves video coding efficiency by compensating for local illumination variations, reducing redundancy and optimizing data processing and transmission bandwidth.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026518092000001_ABST
    Figure 2026518092000001_ABST
Patent Text Reader

Abstract

This disclosure generally relates to video coding / decoding, and in particular to enhancing block-adaptive weighted prediction. The method involves receiving a coded video bitstream, which includes the current block of the current frame and a first syntax element indicating the prediction mode of the current block, wherein a plurality of scaling factor lookup tables are stored and include different ranges of scaling factors. In multiple scaling factor lookup tables, the step size or precision of the scaling factor in each lookup table is the same, and A step of determining a prediction mode based on the value of a first syntax element, wherein the prediction mode is used to predict the current block based on the reference block of the reference frame. The steps include determining a scaling factor from one of several scaling factor lookup tables, The process includes the step of reconstructing the current block based on a reference block and a determined scaling factor.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority based on U.S. Provisional Application No. 63 / 468,487, filed on May 23, 2023, which is hereby incorporated by reference in its entirety. This application also claims the benefit of priority based on U.S. Non - Provisional Patent Application No. 18 / 461,706, filed on September 6, 2023, which is hereby incorporated by reference in its entirety.

[0002] This disclosure describes a set of advanced video / streaming coding / decoding techniques. More specifically, the disclosed techniques include enhancements to block adaptive weighted prediction (BAWP) for compensating local illumination variations.

Background Art

[0003] Uncompressed digital video can include a series of pictures and may include specific bit - rate requirements for memory, data processing, and transmission bandwidth in streaming applications. One purpose of video coding and decoding can be the reduction of the redundancy of an uncompressed input video signal by various compression techniques.

Summary of the Invention

Means for Solving the Problems

[0004] This disclosure describes various embodiments of a method, an apparatus, and a computer - readable storage medium for enhancing block adaptive weighted prediction (BAWP) to model local illumination compensation (LIC).

[0005] In one aspect, an embodiment of the present disclosure provides a method for decoding the current block of the current frame in a coded video bitstream. The method comprises a decoding device having a memory for storing instructions and a processor communicating with the memory, the method comprising receiving a coded video bitstream including the current block of the current frame and a first syntax element indicating a predictive mode of the current block, The decoding device's memory further stores multiple scaling factor lookup tables, where the multiple scaling factor lookup tables contain scaling factors of different ranges, and the step size or precision of the scaling factors for each lookup table within the multiple scaling factor lookup tables is the same, and A decoding device determines a prediction mode based on the value of a first syntax element, wherein the prediction mode is used to predict the current block based on the reference block of the reference frame. The decoding device determines the scaling factor from one of several scaling factor lookup tables, The process includes the steps of: reconstructing the current block based on a reference block and a scaling factor determined according to a linear equation, using a decoding device.

[0006] In another aspect, embodiments of the present disclosure provide a device for processing the current block of the current frame in a coded video bitstream. The device includes a memory for storing instructions and a processor for communicating with the memory. When the processor executes an instruction, the processor is configured to cause the device to perform the above-described method for video decoding and / or encoding. In another aspect, embodiments of the present disclosure provide a non-temporary computer-readable medium for storing instructions that cause a computer to perform the above-described method for video decoding and / or encoding when executed by the computer for video decoding and / or encoding.

[0007] The above and other aspects and embodiments thereof will be described in further detail in the drawings, specification and claims.

[0008] Further features, properties, and various advantages of the disclosed subject matter will become clearer from the detailed description and accompanying drawings below. [Brief explanation of the drawing]

[0009] [Figure 1] This is a schematic diagram showing a simplified block diagram of a communication system 100 according to an exemplary embodiment. [Figure 2] This is a schematic diagram showing a simplified block diagram of a communication system 200 according to an exemplary embodiment. [Figure 3] This is a schematic diagram showing a simplified block diagram of a video decoder according to an exemplary embodiment. [Figure 4] This is a schematic diagram showing a simplified block diagram of a video encoder according to an exemplary embodiment. [Figure 5] This is a block diagram of a video encoder according to another exemplary embodiment. [Figure 6] This is a block diagram of a video decoder according to another exemplary embodiment. [Figure 7]This figure shows a coding block division method according to an exemplary embodiment of the present disclosure. [Figure 8] This figure shows another method of coding block partitioning according to the exemplary embodiments of the present disclosure. [Figure 9] This figure shows another method of coding block partitioning according to the exemplary embodiments of the present disclosure. [Figure 10] This is a diagram showing the combined motion compensation. [Figure 11] This figure shows an example of an interpolated reference frame for motion compensation. [Figure 12] This diagram shows an example of a template for a current block and a referenced block. [Figure 13] This diagram shows an illustrative logical flow of the method described herein. [Figure 14] This is a schematic diagram of a computer system according to an exemplary embodiment of the present disclosure. [Modes for carrying out the invention]

[0010] Next, the present invention will be described in detail below with reference to the accompanying drawings, which form part of the present invention and illustrate specific examples of embodiments. However, it should be noted that the present invention may be embodied in various different forms, and therefore the subject matter covered or claimed is intended to be construed as not being limited to any of the embodiments described below. It should also be noted that the present invention may be embodied as a method, device, component, or system. Thus, embodiments of the present invention may take the form of, for example, hardware, software, firmware, or any combination thereof.

[0011] Throughout this specification and the claims, terms may have nuances implied or suggested in context beyond their expressly stated meanings. The phrases “in one embodiment” or “in several embodiments” used in this disclosure do not necessarily refer to the same embodiment, and the phrases “in another embodiment” or “in other embodiments” used in this disclosure do not necessarily refer to different embodiments. Similarly, the phrases “in one embodiment” or “in several embodiments” used in this specification do not necessarily refer to the same embodiment, and the phrases “in another embodiment” or “in other embodiments” used in this specification do not necessarily refer to different embodiments. For example, the claimed subject matter is intended to include all or some combinations of exemplary embodiments / implementations.

[0012] In general, terms can be understood at least partially from their usage in context. For example, terms such as “and,” “or,” or “and / or” as used herein may have various meanings that may depend at least partially on the context in which such terms are used. Typically, when “or” is used to relate a list such as A, B, or C, it is intended to mean A, B, and C, used here in an inclusive sense, as well as A, B, or C, used here in an exclusive sense. In addition, the terms “one or more” or “at least one” as used herein may be used at least partially on context to describe any feature, structure, or characteristic in a singular sense, or to describe a combination of features, structures, or characteristics in a plural sense. Similarly, terms such as “a,” “an,” or “the” may also be understood, at least partially on context, to convey either a singular usage or a plural usage. In addition, the terms "based on" or "determined by" are sometimes understood not to necessarily convey an exclusive set of factors, but rather, depending at least partially on the context, may allow for the existence of further factors that are not necessarily explicitly described.

[0013] As shown in FIG. 1, the terminal device may be implemented as a server, a personal computer, and a smartphone, but the applicability of the basic principle of the present disclosure may not be limited thereto. Embodiments of the present disclosure may be implemented in a desktop computer, a laptop computer, a tablet computer, a media player, a wearable computer, a dedicated video conferencing device, and the like. The network (150) represents any number or type of network that transmits coded video data between terminal devices, including, for example, a wired (wired) and / or wireless communication network. The communication network (150) can exchange data over circuit switching, packet switching, and / or other types of channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0014] FIG. 2 shows the arrangement of a video encoder and a video decoder in a video streaming environment as an example of an application for the disclosed subject matter. The disclosed subject matter may be equally applicable to other video applications, including, for example, video conferencing, digital television broadcasting, gaming, virtual reality, storage of compressed video on digital media including CDs, DVDs, memory sticks, and the like.

[0015] As shown in Figure 2, a video streaming system may include a video capture subsystem (213) which may include a video source (201), such as a digital camera, to create a stream (202) of uncompressed video pictures or images. In one example, the stream (202) of video pictures includes samples recorded by the digital camera of the video source (201). The stream (202) of video pictures, drawn as a thick line to emphasize the high data volume compared to the encoded video data (204) (or encoded video bitstream), may be processed by an electronic device (220) which includes a video encoder (203) coupled to the video source (201). The video encoder (203) may include hardware, software, or a combination thereof, enabling or implementing embodiments of the disclosed subject, as will be described in more detail below. The encoded video data (204) (or encoded video bitstream 204), shown as a thin line to emphasize its smaller data size compared to the uncompressed video picture stream (202), may be stored in the streaming server 205 for future use or directly in a downstream video device (not shown). One or more streaming client subsystems, such as client subsystems (206) and (208) in Figure 2, can access the streaming server (205) to retrieve copies (207) and (209) of the encoded video data (204). The client subsystem (206) may include, for example, a video decoder (210) in an electronic device (230). The video decoder (210) decodes the input copy (207) of the encoded video data and generates an output stream of a video picture (211) that is uncompressed and can be rendered on a display (212) (e.g., a display screen) or other rendering device (not shown).

[0016] Figure 3 shows a block diagram of a video decoder (310) of an electronic device (330) according to any embodiment of the present disclosure below. The electronic device (330) may include a receiver (331) (e.g., a receiving circuit). The video decoder (310) can be used instead of the video decoder (210) in the example of Figure 2. As shown in Figure 3, the receiver (331) may receive one or more coded video sequences from a channel (301). A buffer memory (315) may be placed between the receiver (331) and an entropy decoder / parser (320) (hereinafter "parser (320)") to counteract network jitter and / or handle playback timing. The parser (320) may reconstruct symbols (321) from the coded video sequences. Categories of these symbols include information used to manage the operation of the video decoder (310) and potential information for controlling rendering devices such as a display (312) (e.g., a display screen). The parser (320) may analyze / entropy decode the coded video sequence. The parser (320) may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder. Subgroups may include picture groups (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), predictive units (PU), etc. The parser (320) may also extract information from the coded video sequence such as transform factors (e.g., Fourier transform factors), quantizer parameter values, and motion vectors. The reconstruction of the symbol (321) may include several different processing or function units. The units involved and how they are involved may be controlled by the parser (320) by the subgroup control information analyzed from the coded video sequence. The first unit may include a scaler / inverse transform unit (351).The scaler / inverse unit (351) can receive control information from the parser (320) including the quantized transformation factors, information indicating which type of inverse transformation should be used, block size, quantization factors / parameters, quantization scaling matrix, and state as symbol (321). The scaler / inverse unit (351) can output a block containing sample values ​​that can be input to the aggregator (355).

[0017] In some cases, the output samples of the scaler / inverse transform (351) may relate to intracoded blocks, i.e., blocks that do not use prediction information from previously reconstructed pictures but can use prediction information from previously reconstructed portions of the current picture. Such prediction information can be provided by the intrapicture prediction unit (352). In some cases, the intrapicture prediction unit (352) can generate blocks of the same size and shape as the block being reconstructed, using surrounding block information that has already been restored and stored in the current picture buffer (358). The current picture buffer (358) buffers, for example, partially reconstructed current pictures and / or fully reconstructed current pictures. In some implementations, the aggregator (355) may, sample by sample, add the prediction information generated by the intraprediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351).

[0018] In other cases, the output samples of the scaler / inverse unit (351) may relate to an intercoded, potentially motion-compensated block. In such cases, the motion-compensated prediction unit (353) can access the reference picture memory (357) based on the motion vector to fetch samples to be used for inter-picture prediction. After motion-compensating the fetched reference samples according to the symbols (321) associated with the block, these samples can be added to the output of the scaler / inverse unit (351) by the aggregator (355) to generate output sample information (the output of unit 351 may be referred to as residual samples or residual signal).

[0019] The output samples of the aggregator (355) can be subjected to various loop filtering techniques within a loop filter unit (356) which includes several types of loop filters. The output of the loop filter unit (356) can be a sample stream that can be output to the rendering device (312) as well as stored in a reference picture memory (357) for use in future interpicture prediction.

[0020] Figure 4 shows a block diagram of a video encoder (403) according to an exemplary embodiment of the present disclosure. The video encoder (403) may be included in an electronic device (420). The electronic device (420) may further include a transmitter (440) (e.g., a transmitting circuit). The video encoder (403) can be used instead of the video encoder (403) in the example of Figure 4. The video encoder (403) may receive video samples from a video source (401). According to some exemplary embodiments, the video encoder (403) can encode and compress pictures of a source video sequence into a coded video sequence (443) in real time or under other temporal constraints required by the application. Implementing an appropriate coding speed constitutes one function of the controller (450). In some embodiments, the controller (450) can be functionally coupled with and controlled by other functional units, as described below. The parameters set by the controller (450) may include rate control-related parameters (picture skip, quantizer, lambda value for rate distortion optimization technique, etc.), picture size, picture group (GOP) layout, maximum motion vector search range, etc.

[0021] In some exemplary embodiments, the video encoder (403) may be configured to operate in a coding loop. The coding loop may include a source coder (430) and a (local) decoder (433) built into the video encoder (403). (Since any compression between symbols and the coded video bitstream in entropy coding may be reversible in the video compression techniques considered in the disclosed subject) the decoder (433) reconstructs the symbols to create sample data in a manner similar to that created by a (remote) decoder, even if the built-in decoder 433 processes the video stream coded by the source coder (430) without entropy coding. At this point, it can be said that any decoder technique other than syntax analysis / entropy decoding, which may only exist within the decoder, may also necessarily need to exist in substantially the same functional form in the corresponding encoder. For this reason, the disclosed subject may focus on decoder operation, which is similar to the decoding portion of the encoder. Thus, a description of encoder techniques can be omitted, as it is the inverse of a comprehensive description of decoder techniques. A more detailed description of the encoder is provided below, but only in specific areas or embodiments.

[0022] In operation in some exemplary implementations, the source coder (430) may perform motion-compensated predictive coding, which predictively codes an input picture by referencing one or more previously coded pictures from a video sequence designated as a “reference picture”.

[0023] The local video decoder (433) may decode the coded video data of a picture that may be designated as a reference picture. The local video decoder (433) may replicate the decoding process that may be performed by the video decoder on the reference picture and store the reconstructed reference picture in the reference picture cache (434). In this way, the video encoder (403) can locally store a copy of the reconstructed reference picture that has content in common with the reconstructed reference picture obtained by the far-end (remote) video decoder (without transmission error).

[0024] The predictor (435) can perform predictive searches on the coding engine (432). That is, for a new picture to be coded, the predictor (435) can search the reference picture memory (434) for specific metadata such as sample data (as candidate reference pixel blocks) or reference picture motion vectors, block shapes, etc., which can serve as appropriate predictive references for the new picture.

[0025] The controller (450) can manage the encoding operation of the source coder (430), including, for example, setting parameters and subgroup parameters used to encode video data.

[0026] The outputs of all the aforementioned functional units can be entropy-coded by an entropy coder (445). A transmitter (440) may buffer the encoded video sequence created by the entropy coder (445) and prepare it for transmission over a communication channel (460), which may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (440) may merge the encoded video data from the video coder (403) with other data being transmitted, such as encoded audio data and / or auxiliary data streams (sources not shown).

[0027] The controller (450) can manage the operation of the video encoder (403). During coding, the controller (450) may assign a specific coding picture type to each coded picture, which may affect the coding technique that can be applied to each picture. For example, a picture may often be assigned as one of the following picture types: intra-picture (I-picture), predictive picture (P-picture), bidirectional predictive picture (B-picture), or multi-predictive picture. A source picture may generally be spatially subdivided into multiple sample coding blocks, as will be described in more detail below.

[0028] Figure 5 shows a diagram of a video encoder (503) according to another exemplary embodiment of the present disclosure. The video encoder (503) is configured to receive a processing block (e.g., a prediction block) of sample values ​​in the current video picture within a sequence of video pictures, and to code the processing block into a coded picture which is part of a coded video sequence. The exemplary video encoder (503) may be used instead of the video encoder (403) in the example of Figure 4.

[0029] For example, the video encoder (503) receives a matrix of sample values ​​for a processing block. The video encoder (503) then determines whether the processing block is best coded using, for example, rate distortion optimization (RDO) in intra-mode, inter-mode, or bi-predictive mode.

[0030] In the example shown in Figure 5, the video encoder (503) includes an interencoder (530), an intraencoder (522), a residual calculator (523), a switch (526), ​​a residual encoder (524), a general-purpose controller (521), and an entropy encoder (525), all coupled together as shown in the exemplary configuration of Figure 5.

[0031] The interencoder (530) is configured to receive a sample of the current block (e.g., a processing block), compare the block with one or more reference blocks in the reference picture (e.g., blocks in the previous and subsequent pictures in display order), generate interprediction information (e.g., descriptions of redundant information by interencoding techniques, motion vectors, merge mode information), and compute interprediction results (e.g., predicted blocks) based on the interprediction information using any appropriate technique.

[0032] The intra encoder (522) is also configured to receive a sample of the current block (e.g., a processing block), compare the block with a block already coded in the same picture, generate a quantization factor after the transformation, and optionally generate intra prediction information (e.g., intra prediction direction information by one or more intra encoding techniques).

[0033] The general-purpose controller (521) may be configured, for example, to determine a prediction mode for a block, to determine general-purpose control data to provide a control signal to a switch (526) based on the prediction mode, and to control other components of the video encoder (503) based on the general-purpose control data.

[0034] A residual calculator (523) may be configured to calculate the difference (residual data) between the received block and the predicted result of a block selected from an intra-encoder (522) or interencoder (530). A residual encoder (524) may be configured to encode the residual data to generate a transformation factor. The transformation factor is then subjected to a quantization process to obtain a quantized transformation factor. In various exemplary embodiments, the video encoder (503) also includes a residual decoder (528). The residual decoder (528) is configured to perform an inverse transform to generate decoded residual data. An entropy encoder (525) may be configured to format the bitstream to include the encoded block and perform entropy coding.

[0035] Figure 6 shows a diagram of an exemplary video decoder (610) according to another embodiment of the present disclosure. The video decoder (610) is configured to receive a coded picture, which is part of a coded video sequence, and to decode the coded picture to produce a reconstructed picture. In one example, the video decoder (610) may be used instead of the video decoder (410) in the example of Figure 4.

[0036] In the example shown in Figure 6, the video decoder (610) includes an entropy decoder (671), an interdecoder (680), a residual decoder (673), a reconstruction module (674), and an intradecoder (672), all coupled together as shown in the exemplary configuration of Figure 6.

[0037] The entropy decoder (671) can be configured to reconstruct specific symbols representing the syntax elements that make up the coded picture from the coded picture. The interdecoder (680) can be configured to receive interprediction information and generate interprediction results based on the interprediction information. The intradecoder (672) can be configured to receive intraprediction information and generate prediction results based on the intraprediction information. The residual decoder (673) can be configured to perform inverse quantization to extract the inversely quantized transformation factor and process the inversely quantized transformation factor to convert the residual from the frequency domain to the spatial domain. The reconstruction module (674) in the spatial domain, the residual decoder (673) The residuals output by and the prediction results (which may be output by the inter-prediction module or intra-prediction module) can be combined to form a reconstructed block that forms part of the reconstructed picture as part of the reconstructed video.

[0038] It should be noted that the video encoders (203), (403), and (503), as well as the video decoders (210), (310), and (610), can be implemented using any suitable technique. In some embodiments, the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610) can be implemented using one or more integrated circuits. In other embodiments, the video encoders (203), (403), and (503), as well as the video decoders (210), (310), and (610), can be implemented using one or more processors that execute software instructions.

[0039] Focusing on block partitioning used for coding and decoding, a typical partition can start from a base block and follow a predetermined set of rules, a specific pattern, a partition tree, or some partition structure or scheme. The partitioning may be hierarchical and recursive. A final set of partitions or coding blocks may be obtained after separating or partitioning the base block according to one of the exemplary partitioning procedures described below, or other procedures, or a combination thereof. Each of these partitions may be at one of various partitioning levels within the partitioning hierarchy and may be of various shapes. Each partition may be called a coding block (CB). In the various exemplary partitioning implementations further described below, each resulting CB may be any CB of an acceptable size and partitioning level. Such partitions are called coding blocks because several basic coding / decoding decisions can be made for them, forming units for which coding / decoding parameters can be optimized, determined, and signaled in the encoded video bitstream. The highest or deepest level in the final partition represents the depth of the coding block partitioning structure in the tree. Coding blocks may be luma coding blocks or chroma coding blocks. The CB tree structure for each color is sometimes called a coding block tree (CBT). The coding blocks for all color channels are sometimes collectively called a coding unit (CU). The hierarchical structure for all color channels is sometimes collectively called a coding tree unit (CTU). The division patterns or structures of the various color channels within a CTU may or may not be the same.

[0040] In some implementations, the partitioning tree scheme or structure used for luma channels and chroma channels do not necessarily have to be the same. In other words, luma channels and chroma channels may have separate coding tree structures or patterns. Furthermore, whether the same coding partitioning tree structure is used for luma and chroma channels or different coding partitioning tree structures, and the actual coding partitioning tree structure used, may depend on whether the slice being coded is a P slice, a B slice, or an I slice. For example, in an I slice, chroma channels and luma channels may have separate coding partitioning tree structures or coding partitioning tree structure modes, while in a P slice or B slice, luma channels and chroma channels may share the same coding partitioning tree scheme. When separate coding partitioning tree structures or modes are applied, a luma channel may be partitioned into CBs by one coding partitioning tree structure, and a chroma channel may be partitioned into chroma CBs by another coding partitioning tree structure.

[0041] Figure 7 shows an exemplary predefined 10-way partition structure / pattern that allows recursive partitioning to form a partition tree. The root block can start from a predefined level (e.g., from a base block at a 128x128 or 64x64 level). The exemplary partition structure in Figure 7 includes various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. In some exemplary implementations, none of the rectangular partitions in Figure 7 can be further subdivided. A coding tree depth may be further defined to indicate the partition depth from the root node or root block. For example, the coding tree depth relative to the root node or root block may be set to 0, and after the root block is partitioned one more time according to Figure 7, the coding tree depth increases by 1. In some implementations, only 710 square partitions may allow recursive partitioning to the next level of the partition tree following the pattern in Figure 7.

[0042] In some other exemplary implementations for coding block partitioning, a quadtree structure may be used. Such a quadtree partition may be applied hierarchically and recursively to any square partition. Whether the base block or intermediate block or partition is further quadtree-partitioned can be adapted to various local characteristics of the base block or intermediate block / partition.

[0043] In several other examples, a ternary pattern may be used to partition a base block or any intermediate block, as shown in Figure 8. The ternary pattern may be implemented vertically, as shown in 802 of Figure 8, or horizontally, as shown in 804 of Figure 13. The exemplary partition ratio in Figure 8 is shown as 1:2:1, but other ratios may be predefined. In some implementations, two or more different ratios may be predefined. In some implementations, the width and height of the partitions in the exemplary ternary tree are always powers of 2 to avoid further transformations.

[0044] The partitioning schemes described above can be combined in any way at different partitioning levels. For example, the quadtree and binary partitioning schemes described above may be combined to partition a base block into a quadtree-binary (QTBT) structure. In such a scheme, the base block or intermediate block / partition may be either a quadtree partition or a binary partition, if specified, according to a set of predefined conditions. A particular example is shown in Figure 9, where the base block is a quadtree that is initially quadtree partitioned, as indicated by 902, 904, 906, and 908. Each of the resulting partitions is then either quadtree partitioned into four further partitions (such as 908), or binary partitioned into two further partitions at the next level (e.g., horizontal or vertical, such as 902 or 906, both of which are symmetric), or not partitioned at all (such as 904). Binary or quadtree partitioning may be recursively possible for square partitions, as shown by the overall exemplary partition pattern in 910 and the corresponding tree structure / representation in 920, where solid lines represent quadtree partitioning and dashed lines represent binary partitioning. A flag may be used for each binary node (non-leaf binary partition) to indicate whether the binary is horizontal or vertical. For example, as shown in 920, which matches the partition structure in 910, a flag "0" may represent horizontal binary and a flag "1" may represent vertical binary. In the case of quadtree partitioning, there is no need to indicate the partition type, as quadtree partitioning always divides a block or partition both horizontally and vertically to produce four subblocks / partitions of equal size. In some implementations, a flag "1" may represent horizontal binary and a flag "0" may represent vertical binary.

[0045] In some exemplary embodiments of QTBT, the quadtree and binary rule set may be represented by the following predefined parameters and their associated corresponding functions. -CTU size: The size of the root node of the quadtree (the size of the base block). -MinQTSize: Minimum allowable quadtree leaf node size -MaxBTSize: Maximum allowable binary tree root node size -MaxBTDepth: Maximum allowable binary tree depth -MinBTSize: Minimum allowable binary tree leaf node size

[0046] In some exemplary embodiments of the QTBT partitioning structure, the CTU size may be set as a 128x128 chromasample with two corresponding 64x64 blocks of chromasample (when exemplary chroma subsampling is considered and used), MinQTSize may be set as 16x16, MaxBTSize may be set as 64x64, MinBTSize may be set as 4x4 (for both width and height), and MaxBTDepth may be set as 4. A quadtree partition may be applied to the CTU first to generate a quadtree leaf node. A quadtree leaf node can have a size from its minimum allowable size (i.e., MinQTSize) of 16x16 to 128x128 (i.e., CTU size). If a node is 128x128, it will not be partitioned by the binary tree first because its size exceeds MaxBTSize (i.e., 64x64). Otherwise, nodes that do not exceed MaxBTSize may be partitioned by the binary tree. In the example in Figure 9, the base block is 128x128. The base block can only be quadtree-partitioned according to a predefined set of rules. The base block has a partitioning depth of 0. Each of the four resulting partitions is 64x64, not exceeding MaxBTSize, and may be further quadtree-partitioned or binary-partitioned at level 1. The process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), further partitioning is not considered. When the width of a binary tree node is equal to MinBTSize (i.e., 4), further horizontal partitioning is not considered. Similarly, when the height of a binary tree node is equal to MinBTSize, further vertical partitioning is not considered.

[0047] In some exemplary implementations, the above QTBT scheme may be configured to support the flexibility for lumens and chromians to have the same QTBT structure or separate QTBT structures. For example, in the case of P-slice and B-slice, the lumens CTB and chromens CTB within one CTU may share the same QTBT structure. However, in the case of I-slice, the lumens CTB may be divided into CBs by a QTBT structure, and the chromens CTB may be divided into chromens CBs by another QTBT structure. This means that CUs may be used to refer to different color channels within an I-slice, for example, an I-slice may consist of a coding block for the lumens component or a coding block for two chromens components, and a CU in a P-slice or B-slice may consist of a coding block for all three color components.

[0048] The various CB partitioning schemes and further partitioning of the CB into PB described above may be combined in any way. The following specific implementations are provided as non-limiting examples.

[0049] Interpretation can be performed, for example, in single-reference mode or composite-reference mode. In some implementations, a skip flag may be initially included in the bitstream of the current block (or at a higher level) to indicate whether the current block is being intercoded and not to be skipped. If the current block is being intercoded, another flag may be further included in the bitstream as a signal indicating whether single-reference mode or composite-reference mode is being used to predict the current block. In single-reference mode, one reference block may be used to generate the predicted block for the current block. In composite-reference mode, two or more reference blocks may be used, for example, by a weighted average to generate the predicted block. One or more reference blocks may be identified using one or more reference frame indices, and further using one or more corresponding motion vectors indicating the position relative to the frame, e.g., the shift in horizontal and vertical pixels, between one or more reference blocks and the current block. For example, the current block's interpretation block may be generated from a single reference block identified by a single motion vector in the reference frame, as a prediction block in single reference mode. However, in composite reference mode, the prediction block may be generated by a weighted average of two reference blocks in two reference frames, indicated by two reference frame indices and two corresponding motion vectors. Motion vectors can be coded in various ways and included in the bitstream.

[0050] In some exemplary implementations, one or more reference picture lists, including the identification of short-term and long-term reference frames for interpretation, may be formed based on information in a reference picture set (RPS). For example, a single picture reference list may be formed for unidirectional interpretation, denoted as L0 reference (or reference list 0), and two picture reference lists may be formed for bidirectional interpretation, denoted as L0 (or reference list 0) and L1 (or reference list 1) for each of the two prediction directions. The reference frames included in the L0 and L1 lists may be ordered in various predetermined ways. The lengths of the L0 and L1 lists may be signaled in the video bitstream. Unidirectional interpretation can be either single-reference mode or decoded-reference mode, provided that the multiple references for generating prediction blocks by weighted averaging in composite prediction mode are on the same side of the frame in which the block to be predicted is located. Bidirectional interpretation can only be composite mode, in that bidirectional interpretation includes at least two reference blocks.

[0051] In some implementations, a merge mode (MM) may be implemented for interpretation. Generally, in merge mode, one or more motion vectors in a single reference prediction or a composite reference prediction of the current PB may be derived from other motion vectors rather than being computed and signaled independently. For example, in an encoding system, the current motion vector of the current PB can be represented by the difference between the current motion vector and one or more other already encoded motion vectors (called reference motion vectors). Such a difference of motion vectors, rather than the entire current motion vector, may be encoded and included in the bitstream and linked to the reference motion vectors. Correspondingly, in a decoding system, the motion vector corresponding to the current PB may be derived based on the decoded motion vector difference and the decoded reference motion vector linked to it. As a specific form of general merge mode (MM) interpretation, such interpretation based on motion vector differences is sometimes called merge mode with motion vector differences (MMVD). Thus, general MM, or MMVD in particular, may be implemented to improve coding efficiency by leveraging correlations between motion vectors associated with different PBs. For example, neighboring PBs may have similar motion vectors, and therefore their MVDs may be small and can be coded efficiently. In another example, motion vectors can be correlated temporally (between frames) for blocks that are similarly positioned / placed in space.

[0052] In some exemplary implementations of MMVD, a list of reference motion vectors (RMVs) or MV predictor candidates can be formed for a block to be predicted. The list of RMV candidates can contain a predetermined number (e.g., two) of MV predictor candidate blocks whose motion vectors could be used to predict the current motion vector. RMV candidate blocks can include blocks selected from neighboring blocks and / or time blocks within the same frame (e.g., blocks that are identically located in the current or subsequent frames). These options represent blocks that are spatially or temporally located relative to the current block and are likely to have a similar or identical motion vector to the current block. The size of the list of MV predictor candidates may be predetermined. For example, the list may contain two or more candidates. In order to be on the list of RMV candidates, candidate blocks may need to have, and must exist, the same reference frame (or multiple frames) as the current block (e.g., boundary checking must be performed if the current block is near the edge of a frame), must have been encoded during the encoding process, and / or must have been decoded during the decoding process. In some implementations, the list of merge candidates may be filled first with spatially neighboring blocks (traversed in a specific predefined order) if available and satisfying the above conditions, and then with time blocks if space is still available in the list. Neighboring RMV candidate blocks can be selected, for example, from the blocks to the left and above the current box. The list of RMV predictor candidates may be dynamically formed as a dynamic reference list (DRL) at various levels (sequence, picture, frame, slice, superblock, etc.). The DRL may be signaled with a bitstream.

[0053] In some implementations, the actual MV predictor candidate currently being used as the reference motion vector for predicting the motion vector of the block may be signaled. If the RMV candidate list contains two candidates, a one-bit flag called the merge candidate flag may be used to indicate the selection of the reference merge candidate. For the current block being predicted in composite mode, each of the multiple motion vectors predicted using the MV predictor may be associated with a reference motion vector from the merge candidate list. The encoder can determine which RMV candidate is used for the closer prediction of the MV of the current coding block, and the selected one can be signaled as an index and placed in the DRL.

[0054] In some exemplary implementations of MMVD, an RMV candidate is selected and used as the base motion vector predictor for the motion vector to be predicted, after which the motion vector difference (MVD or deltaMV representing the difference between the motion vector to be predicted and the reference candidate motion vector) can be calculated in the encoding system. Such an MVD may contain information representing the magnitude and direction of the MV difference, both of which can be signaled in the bitstream in various ways.

[0055] In some exemplary implementations of MMVD, a distance index can be used to specify the magnitude of the motion vector difference, indicating one of a set of predefined offsets that represent a predefined motion vector difference from a starting point (reference motion vector). The MV offset corresponding to the signaled index can then be added to either the horizontal or vertical component of the starting (reference) motion vector. Exemplary predefined relationships between distance indices and predefined offsets are specified in Table 1.

[0056] [Table 1]

[0057] In some exemplary implementations of MMVD, the direction index may be further signaled and used to represent the direction of the MVD relative to the reference motion vector. In some implementations, the direction may be restricted to either the horizontal or vertical direction. Exemplary 2-bit direction indices are shown in Table 2. In the examples in Table 2, the interpretation of the MVD may vary depending on the information of the start / reference MV. For example, if the start / reference MV corresponds to a single prediction block, or if both reference frame lists correspond to two prediction blocks pointing to the same side of the current picture (i.e., the POCs of both reference pictures are both greater than the POC of the current picture, or both are less than the POC of the current picture), the sign in Table 2 may specify the sign (direction) of the MV offset applied to the start / reference MV. If the start / referenced MV corresponds to a biprediction block having two reference pictures on different sides of the current picture (i.e., the POC of one reference picture is greater than the POC of the current picture, and the POC of the other reference picture is less than the POC of the current picture), and the difference between the reference POC in picture reference list 0 and the current frame is greater than the difference between the reference POC in picture reference list 1 and the current frame, then the sign in Table 2 can specify the sign of the MV offset applied to the reference MV corresponding to the reference picture in picture reference list 0, and the sign of the offset of the MV corresponding to the reference picture in picture reference list 1 can have the opposite value (opposite sign of the offset). Otherwise, if the difference between the reference POC in picture reference list 1 and the current frame is greater than the difference between the reference POC in picture reference list 0 and the current frame, then the sign in Table 2 can specify that the sign of the MV offset applied to the reference MV associated with picture reference list 1 and the sign of the offset to the reference MV associated with picture reference list 0 have opposite values.

[0058] [Table 2]

[0059] In some exemplary implementations, the MVD may be scaled according to the difference in POCs in each direction. If the difference in POCs in both lists is the same, scaling is not necessary. If, instead, the difference in POCs in reference list 0 is greater than the difference in reference list 1, the MVD of reference list 1 is scaled. If the difference in POCs in reference list 1 is greater than that of list 0, the MVD of list 0 may be scaled similarly. If the initial MV is single predicted, the MVD is added to the available or reference MVs.

[0060] In some exemplary implementations of MVD coding and signaling for bidirectional composite prediction, in addition to coding and signaling two MVDs separately, or instead, symmetric MVD coding may be implemented such that only one MVD requires signaling and the other MVD can be derived from the signaled MVD. In such implementations, motion information containing both reference picture indices in List 0 and List 1 is not signaled. Specifically, at the slice level, a flag called "mvd_l1_0_flag" may be included in the bitstream to indicate whether reference List 1 is not signaled in the bitstream. If this flag is 1, indicating that reference List 1 is equal to 0 (and therefore not signaled), then a bidirectional prediction flag called "BiDirPredFlag" may be set to 0, which means there is no bidirectional prediction. If mvd_l1_0_flag is 0, then BiDirPredFlag may be set to 1 if the nearest reference picture in List 0 and the nearest reference picture in List 1 form a forward-reverse or reverse-reverse-reference picture pair, and both reference pictures in List 0 and List 1 are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. A BiDirPredFlag of 1 may indicate that a symmetric mode flag is additionally signaled in the bitstream. The decoder may extract the symmetric mode flag from the bitstream if BiDirPredFlag is 1. The symmetric mode flag may be signaled, for example, at the CU level (if necessary) to indicate whether a symmetric MVD coding mode is being used for the corresponding CU.If the symmetric mode flag is 1, it indicates the use of the symmetric MVD coding mode, where only the reference picture indices in both List 0 and List 1 (referred to as "mvp_l0_flag" and "mvp_l1_flag") are signaled by the MVD associated with List 0 (referred to as "MVD0"), and the other motion vector difference "MVD1" should be derived rather than signaled. For example, MVD1 might be derived as -MVD0. Thus, in the exemplary symmetric MVD mode, only one MVD is signaled.

[0061] In several other exemplary implementations for MV prediction, harmonic schemes may be used to implement the general merge-mode MMVD for both single-reference-mode and compound-reference-mode MV prediction, and several other types of MV prediction. Various syntactic elements may be used to signal how the MV of the current block is predicted. For example, in single-reference mode, the following MV prediction modes may be signaled: NEARMV - Uses one of the motion vector predictors (MVPs) in a list indicated by a direct DRL (Dynamic Reference List) index, without any MVDs. Use one of the motion vector predictors (MVPs) in the list notified by the NEWMV-DRL index as a reference and apply a delta to the MVP (e.g., use MVD). GLOBALMV - Uses motion vectors based on global motion parameters at the frame level.

[0062] Similarly, in the case of a composite reference interpretation mode that uses two reference frames corresponding to the two MVs to be predicted, the following MV prediction modes may be signaled: NEAR_NEARMV - For each of the two motion vectors to be predicted, one of the motion vector predictors (MVPs) in the list signaled by the DRL index without MVD is used. NEAR_NEWMV - To predict the first of the two motion vectors, one of the motion vector predictors (MVPs) in the list signaled by the DRL index without MVD is used as the reference MV, and to predict the second of the two motion vectors, one of the motion vector predictors (MVPs) in the list signaled by the DRL index is used as the reference MV in conjunction with an additionally signaled delta MV (MVD). NEW_NEARMV - Uses one of the motion vector predictors (MVPs) in a list signaled by a DRL index without MVD as the reference MV to predict the second motion vector of two motion vectors, and uses one of the motion vector predictors (MVPs) in a list signaled by a DRL index as the reference MV in conjunction with an additionally signaled delta MV (MVD) to predict the first motion vector of two motion vectors. NEW_NEWMV - Uses one of the motion vector predictors (MVPs) in a list signaled by a DRL index as the reference MV, and uses it in conjunction with an additionally signaled delta MV to predict for each of the two MVs. GLOBAL_GLOBALMV - Uses MVs from each reference based on frame-level global motion parameters.

[0063] Therefore, the term "NEAR" above refers to MV prediction using a reference MV with no MVD as a general merge mode, while the term "NEW" refers to MV prediction that uses a reference MV and involves offsetting it with a signaled or derived MVD, as in the MMVD mode. In the case of composite interpretation, both the reference base motion vector and motion vector delta above may generally be different or independent between the two references or two MVDs, for example, the two MVDs may be correlated, and such correlation can be used to reduce the amount of information required to signal the two motion vector deltas. To take advantage of such correlation, co-signaling of the two MVDs may be implemented and represented in a bitstream, as will be described in more detail below.

[0064] In some implementations of MVD, the default pixel resolution of MVD may be used. For example, a motion vector accuracy (or precision) of 1 / 8 of a pixel may be acceptable. The MVD described above can be constructed and signaled in various ways with various MV prediction modes. In some implementations, various syntax elements can be used to signal the above motion vector difference in reference frame list 0 or list 1.

[0065] For example, a syntax element called "mv_joint" can specify which component of the associated motion vector difference is non-zero. For instance, an mv_joint with a value of 0 can indicate that there are no non-zero MVDs along either the horizontal or vertical direction. 1 can indicate that there are non-zero MVDs only along the horizontal direction. 2 can indicate that there are non-zero MVDs only along the vertical direction. And / or 3 can indicate that there are non-zero MVDs along both the horizontal and vertical directions.

[0066] If the “mv_joint” syntax element for MVD signals that there are no non-zero MVD components, no further MVD information can be signaled. However, if the “mv_joint” syntax signals that there are one or two non-zero components, additional syntax elements can further signal each of the non-zero MVD components, as described below.

[0067] For example, a syntax element called "mv_sign" may be used to further specify whether the corresponding motion vector difference component is positive or negative.

[0068] In another example, a syntax element called "mv_class" can be used to specify the class of motion vector differences between a predefined set of classes for corresponding non-zero MVD components. These predefined classes for motion vector differences can be used, for example, to separate a continuous size space of motion vector differences into non-overlapping class ranges. Thus, the signaled MVD class indicates the size range of the corresponding MVD components. In the exemplary implementation shown in Table 3 below, higher classes correspond to motion vector differences with larger size ranges. The symbol (n,m) is used to represent the range of motion vector differences greater than n pixels and less than or equal to m pixels.

[0069] [Table 3]

[0070] In some other implementations, a syntax element called “mv_bit” may be used to specify the integer part of the offset between the non-zero motion vector difference component and the magnitude of the start of the correspondingly signaled MV class size range. In some other implementations, a syntax element called “mv_fr” may be used to specify the first two fractional bits of the motion vector difference of the corresponding non-zero MVD component, and a syntax element called “mv_hp” may be used to specify a third fractional bit (high resolution bit) of the motion vector difference of the corresponding non-zero MVD component. Two “mv_fr” bits essentially provide a 1 / 4 pixel MVD resolution, while “mv_hp” bits can further provide a 1 / 8 pixel resolution. In some other implementations, two or more “mv_hp” bits may be used to provide MVD pixel resolutions finer than 1 / 8 pixels. In some exemplary implementations, additional flags may be signaled at one or more of various levels to indicate whether MVD resolutions of 1 / 8 pixels or higher are supported. If an MVD resolution is not applicable to a particular coding unit, the above syntax elements for the corresponding unsupported MVD resolution may not be signaled.

[0071] However, in some other exemplary implementations, the resolution of motion vector differences across different MVD size classes may be differentiated or adapted. Specifically, higher resolution MVDs for larger MVD sizes in higher MVD classes may not result in a statistically significant improvement in compression efficiency or coding gain. Thus, MVDs may be coded with a resolution that decreases or does not increase (integer pixel resolution or fractional pixel resolution) for larger MVD ranges corresponding to higher MVD size classes. The term “resolution” may be further referred to as “pixel resolution.” Adaptive MVD resolution may be implemented in various ways, as described by the exemplary embodiments below, to achieve better overall compression efficiency. In particular, the reduction in the number of signaling bits by aiming for lower precision MVDs may exceed the additional bits required to code inter-predictive residuals as a result of such lower precision MVDs, due to the statistical observation that non-adaptively treating the MVD resolution of large or high-class MVDs at the same level as low-size or low-class MVDs does not significantly increase the inter-predictive residual coding efficiency of Bock with large or high-class MVDs. In other words, using a higher MVD resolution for large-scale or high-class MVD may not result in more coding gains than using a lower MVD resolution.

[0072] In some implementations, the precision of the MVD in NEW_NEARMV and NEAR_NEWMV modes depends on the relevant class and the size of the MVD. Firstly, fractional MVDs may only be allowed if the size of the MVD is one pixel or less. Secondly, when the value of the relevant MV class is MV_CLASS_1 or greater, only one MVD value may be allowed, and the MVD values ​​for each MV class are derived as 4, 8, 16, 32, and 64 for MV classes 1 (MV_CLASS_1), 2 (MV_CLASS_2), 3 (MV_CLASS_3), 4 (MV_CLASS_4), or 5 (MV_CLASS_5). In some examples, the single allowed value may be the upper limit of the respective ranges for these MV classes in Table 3. The allowed MVD values ​​for each MV class can be shown in Table 4.

[0073] [Table 4]

[0074] In some exemplary implementations, each MVD class may be associated with a single allowed resolution. In some other implementations, one or more MVD classes may each be associated with two or more optional MVD pixel resolutions. For example, adaptively allowed MVD pixel resolutions may include, but are not limited to, 1 / 64pel (pixel), 1 / 32pel, 1 / 16pel, 1 / 8pel, 1-4pel, 1 / 2pel, 1pel, 2pel, 4pel… (in descending order of resolution).

[0075] In some other exemplary implementations, for MV classes above a threshold MV class, only a single MVD value may be permitted. For example, such a threshold MV class may be MV_CLASS 2. Therefore, for MV_CLASS_2 and above, only having a single MVD value and not having fractional pixel resolution may be permitted.

[0076] Looking at the various composite interprediction modes in which each MV is predicted by a reference motion vector and can be coded by an MVD, the two MVDs can be signaled separately or jointly in the bitstream, as described above. Thus, in some exemplary implementations, in addition to the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes described above, another interprediction mode called JOINT_NEWMV can be introduced for a mode in which an MVD is introduced for joint signaling of reference lists 0 and 1. Specifically, when the interprediction mode is indicated as NEW_NEWMV, the MVDs of reference lists 0 and 1 are signaled separately, but when the interprediction mode is indicated as JOINT_NEWMV mode, the MVDs of reference lists 0 and 1 are signaled jointly. In particular, in the case of joint MVD, only one MVD called joint_delta_mv may need to be signaled and transmitted within the bitstream, and the MVDs in reference lists 0 and 1 can be derived from joint_delta_mv. The derived MVD can then be combined with the reference motion vector in reference list 0 or 1 to generate two motion vectors for locating the reference block for composite interpretation.

[0077] In some implementations of composite interpretation, the JOINT_NEWMV mode may be signaled along with the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes. In such implementations, the syntax may be included in the bitstream to indicate any one of these alternative composite interpretation modes at any of the various signaling levels (e.g., sequence level, picture level, frame level, slice level, tile level, superblock level, etc.). Alternatively, the JOINT_NEWMV mode may be implemented as a submode of the NEW_NEWMV mode. In other words, under the NEW_NEWMV mode, two MVDs of two reference blocks are either jointly signaled (hence the JOINT_NEWMV submode) or not signaled (another submode of the NEW_NEWMV mode). In such an implementation, a first syntax element may be included in the bitstream to indicate one of the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes, and if the first syntax element indicates that the NEW_NEWMV mode has been selected for the coding block, a second syntax element may be further included in the bitstream, which can be extracted by the decoder to indicate whether the MVD of the coding block is signaled separately or jointly.

[0078] In some exemplary embodiments, when the JOINT_NEWMV mode is signaled and the POC distances between two reference frames and the current frame are different, the MVD may be scaled for reference list 0 or reference list 1 based on the POC distance. Specifically, the distance between reference frame list 0 and the current frame may be denoted as td0, and the distance between reference frame list 1 and the current frame may be denoted as td1. If td0 is greater than or equal to td1, joint_mvd may be used directly for reference list 0, and the MVD for reference list 1 may be derived from joint_mvd based on equation (1).

number

[0079] Instead, if td1 is greater than or equal to td0, joint_mvd is used directly in reference list 1, and the MVD of reference list 0 is derived from joint_mvd based on equation (2).

number

[0080] In some exemplary implementations, another intercoded mode called AMVDMV may be added to a single reference case. When the AMVDMV mode is selected, it indicates that AMVD (Adaptive Motion Vector Difference) is applied to the signal MVD. For example, one flag named amvd_flag may be added under the JOINT_NEWMV mode to indicate whether AMVD is applied to the Joint MVD coding mode. When Adaptive MVD Resolution is applied to the Joint MVD coding mode, also called Joint AMVD coding mode, the MVDs of two reference frames are jointly signaled, and the accuracy of the MVD is implicitly determined by the magnitude of the MVD. Otherwise, the MVDs of two (or more) reference frames are jointly signaled, and conventional MVD coding is applied. Alternatively, one new interpredictive mode called JOINT_AMVDNEWMV is added to indicate that AMVD is applied to the Joint MVD coding mode.

[0081] Turning to the composite intermodes, as shown in Figure 10, these modes are two different reference frames F i-1 and F i+1 By combining two hypotheses about motion vectors MV0 and MV1 from the current frame F i This generates predictions for the blocks within. Therefore, two motion information components (e.g., motion vectors) can be signaled in the bitstream for each block.

[0082] Alternatively, as shown in Figure 11, two reference frames, F i-1 and F i+1 The information is combined and interpolated, and the current frame F i A frame may be projected and interpolated at the same time. Multiple TIP modes may be supported. In one TIP mode, the interpolated frame may be used as an additional reference frame. Current frame F iThe coding block directly references the interpolated frame, utilizing information from two different references with only the overhead cost of a single interprediction mode. In another TIP mode, the interpolated frame can now be directly assigned as the output of the decoding process of frame Fi, skipping any other conventional coding steps. This mode can offer significant coding and complexity advantages, especially for low-bitrate applications.

[0083] In some implementations, Adaptive Motion Vector Resolution (AMVR) may be supported. In certain exemplary implementations, a total of 7 MV accuracies (8, 4, 2, 1, 1 / 2, 1 / 4, 1 / 8) may be supported. For each prediction block, the AMVR encoder can search all supported accuracy values ​​and signal the best accuracy to the decoder.

[0084] To reduce encoder execution time, two MV precision sets are supported. Each precision set may include, for example, four predefined precisions. The precision set may be adaptively selected at the frame level based on the maximum precision value of the frame. The maximum precision may be signaled in the frame header. Table 5 summarizes exemplary supported precision values ​​based on exemplary frame-level maximum precision.

[0085] [Table 5]

[0086] In some exemplary implementations of AMVR, there may be a frame-level flag indicating whether the frame's MV contains a precision greater than one pixel. AMVR can only be enabled if the value of the cur_frame_force_integer_mv flag is 0. In AMVR, if the block precision is lower than the maximum precision, the motion model and interpolation filter may not be signaled. If the block precision is lower than the maximum precision, the motion mode is inferred to translational motion, and the interpolation filter is inferred to a REGULAR interpolation filter. Similarly, if the block precision is either 4 pixels or 8 pixels, the intra-mode may not be signaled and may be inferred to be 0.

[0087] In some implementations, when a block is coded as a joint MVD coding mode, JOINT_NEWMV or JOINT_AMVDNEWMV, a new syntax called mvd_scaling_factor_idx can be signaled to the bitstream to explicitly indicate the scaling factor of the MVD between reference frame 0 and reference frame 1. As shown in Tables 6 and 7, two predefined lookup tables can be used to store separately the supported / allowed scaling factors for JOINT_NEWMV or JOINT_AMVDNEWMV. The relevant entry index of the selected scaling factor in the lookup table can be signaled in the bitstream. In JOINT_AMVDNEWMV mode, the same scaling factor applies to both the vertical and horizontal components of the MVD in reference frame lists 0 and / or 1. In JOINT_NEWMV mode, the scaling factor for one component of the MVD (either the vertical or horizontal component) may be restricted to 1, while the scaling factor for the other component of the MVD may be any other value, e.g., 2 or 1 / 2. In one example, the MVD (mvd_ref0 or mvd_ref1) of reference framelist 0 or 1 is calculated using the following formula:

number

[0088] [Table 6]

[0089] [Table 7]

[0090] In some exemplary embodiments, biprediction with CU-level weights (BCW) can generate a biprediction signal by averaging two prediction signals obtained from two different reference pictures and / or by using two different motion vectors. In some other implementations, the biprediction mode may be extended beyond a simple average to allow a weighted average of the two prediction signals. For example, P bi-pred =((8-w)*P0+w*P1+4)≫3. In weighted average biprediction, five weights, w∈{-2,3,4,5,10} may be allowed. If w is equal to 4, equal weight factors are used to perform the weighted average of the two predicted samples. For each biprediction CU, the weight w can be determined in one of two ways: (1) for non-merged CUs, the weight index is signaled after the motion vector difference; and / or (2) for merged CUs, the weight index is inferred from neighboring blocks based on the merge candidate index. BCW can only be applied to CUs with 256 or more luma samples (i.e., CU width × CU height is 256 or more). For low-latency pictures, all five weights are used. For non-low-latency pictures, only three weights (w∈{3,4,5}) are used.

[0091] In another embodiment, the COMPOUND_AVERAGE mode can be extended to allow a weighted average of the prediction signals. This extended mode is called Composite Weighted Prediction (CWP). Specifically, if the two reference frames are from different directions, five weight factors w∈{8,10,6,12,4} are supported, and if they are from the same direction, five different weight factors w∈{8,12,4,20,-4} are supported. As a result, the above equation can be modified to P(x,y)=(W×P0(x,y)+(16-W)×P1(x,y)+8)≫4.

[0092] In some implementations, the index of the selected weight factor may be signaled when all of the following conditions are met: (i) the composite type is COMPOUND_AVERAGE; (ii) the inter-prediction mode is NEAR_NEARMV, JOINT_NEWMV, or JOINT_AMVDNEWMV; and (iii) the MVD scaling factor (1,1) is used in joint MVD mode. In some implementations, the index of the selected weight factor may be signaled when only some of the above conditions are met. In some implementations, in skip mode, the index of the selected weight factor is inferred from the neighboring block based on the motion vector predictor's DRL index.

[0093] In some embodiments, block-adaptive weighted prediction (BAWP) may include block-level weighted prediction to model local illumination variations between a current block and its predicted block as a function of local illumination variations between a current block template (or the sample causing the current block) and a reference block template. The template for the current block (1210) (or referred to as the current template, 1212) and the template for the reference block (1220) (or referred to as the reference template, 1222) are shown in Figure 12. The reference block may be indicated or determined by a motion vector (MV, 1230). The current block may be in the current picture (or current frame), and the reference block may be in the reference picture (or reference frame). In some implementations, the function may be a linear function. The parameters of the function may be expressed by a scaling factor α and an offset β, which form a linear equation for compensating for illumination changes, i.e., α*p[x]+β, where p[x] is the reference sample pointed to by the MV at position x on the reference picture. The scaling factor α may be the value (factor) of a multiplicative factor with respect to the reference block, and the offset β may be the value (parameter) of a subtractive (or additive) offset with respect to the reference block. In some implementations, α and β may be derived based on the current block template and the reference block template, and no signaling overhead is required for them. In some implementations, the BAWP flag may be signaled for single inter prediction mode to indicate the use of BAWP. BAWP is applied to blocks of 8x8 size or larger and may be coded in single inter prediction mode. In some implementations, BAWP may be applied only to the luma component.

[0094] In some implementations, when a block is coded as BAWP mode, another flag may be signaled to indicate whether explicit or implicit signaling of scaling factors is used. When implicit signaling of BAWP scaling factors is used, the BAWP scaling factor and offset value may be derived from a linear equation based on the template of the current block and the template of the reference block pointed to by the motion vector of the current block. When explicit signaling of BAWP scaling factors is used, another flag is signaled to indicate which scaling factor is used for the current block, and the offset is derived as 0. Supported scaling factors may be stored in one or more predefined lookup tables, and the index of the scaling factor in the lookup table is signaled to the bitstream and parsed on the decoder side.

[0095] In some implementations, there are several challenges or problems related to the BAWP method, particularly the lookup table that stores the scaling factor. For example, if explicit signaling for BAWP is used, and only one fixed lookup table is used, the coding / decoding efficiency may not be optimal. In another example, signaling with flags indicating explicit and implicit signaling does not consider the correlation between the current block and its neighboring blocks, which may not be optimal in terms of coding / decoding efficiency. This disclosure describes various embodiments for enhancing BAWP, addressing at least one of the challenges or problems described above, improving coding / decoding efficiency, and advancing video codec technology.

[0096] The various embodiments and / or embodiments described herein may be performed separately or combined in any order. Furthermore, each of the methods (or embodiments), encoders, and decoders may be implemented by processing circuits (e.g., one or more processors or one or more integrated circuits). One or more processors execute a program stored in a non-temporary computer-readable medium. In this disclosure, the term "block" may be interpreted as a prediction block, coding block, or coding unit (CU).

[0097] Figure 13 shows a flowchart 1300 of an exemplary method following the underlying principles of the above implementation for augmenting block adaptive weighted prediction (BAWP) to compensate for local illumination variations. The exemplary decoding method flow begins at 1301 and may include some or all of the following steps. S1310, a step of receiving a coded video bitstream by a decoding device including a memory for storing instructions and a processor for communicating with the memory, wherein the coded video bitstream includes a current block of the current frame and a first syntax element indicating a prediction mode of the current block, and the memory of the decoding device further stores a plurality of scaling factor lookup tables, the plurality of scaling factor lookup tables include scaling factors of different ranges, and the step size of the scaling factor or the precision of the scaling factor is the same for each lookup table in the plurality of scaling factor lookup tables. S1320, a step in which a decoding device determines a prediction mode based on the value of a first syntax element, wherein the prediction mode is used to predict the current block based on the reference block of the reference frame. S1330, a step in which a decoding device determines a scaling factor from one of several scaling factor lookup tables. and / or, S1340, a step in which the decoding device reconstructs the current block based on the reference block and a scaling factor determined according to a linear equation. An exemplary method terminates at S1399.

[0098] In any part or combination of the above implementation forms, multiple scaling factor lookup tables contain the same number of scaling factors.

[0099] In any part or combination of the above implementation forms, the step of determining a scaling factor from one of several scaling factor lookup tables includes the step of determining a scaling factor either explicitly based on a second syntax element or implicitly based on a local illumination value.

[0100] In some implementations, local illumination variation may refer to intensity variations in blocking (e.g., luminous and / or chromatic channels) due to local illumination (e.g., illumination from light posts or passing vehicles). Local illumination value may refer to at least one of the following: local illumination variation between the current block template and the reference block template, local illumination variation between the sample causing the current block and the reference block template, and / or LIC parameters.

[0101] In any part or combination of the above implementation forms, the step of determining a scaling factor from one of several scaling factor lookup tables includes the step of selecting a scaling factor lookup table from several scaling factor lookup tables based on the derived scaling factor, and / or the step of determining a scaling factor from a scaling factor lookup table.

[0102] In some implementations, the range of the scaling factor lookup table may refer to the value between the minimum scaling factor in the scaling factor lookup table and the maximum scaling factor in the scaling factor lookup table. For example, if we consider {+1,-1,+3,-3,+6,-6} as an unrestricted exemplary scaling factor lookup table, the range may be 12, which is the value between -6 (the minimum scaling factor in the scaling factor lookup table) and +6 (the maximum scaling factor in the scaling factor lookup table).

[0103] In some implementations, the range of the scaling factor lookup table may refer to a value between the minimum absolute scaling factor in the scaling factor lookup table and the maximum absolute scaling factor in the scaling factor lookup table. If we consider {+1,-1,+3,-3,+6,-6} as an unrestricted exemplary scaling factor lookup table, the range may also be 5, which is a value between 1 (the minimum absolute scaling factor in the scaling factor lookup table) and 6 (the maximum absolute scaling factor in the scaling factor lookup table).

[0104] Another exemplary decoding method is: The steps include receiving a coded video bitstream, A step of determining a prediction mode for predicting the current block based on the reference block of the reference frame, based on the coded video bitstream, A step of selecting a scaling factor lookup table from multiple scaling factor lookup tables based on derived scaling factors, wherein the multiple scaling factor lookup tables include different ranges. The steps include determining the scaling factor from a selected scaling factor lookup table based on the coded video bitstream, The steps include: reconstructing the current block based on the reference block and the scaling factor determined according to the linear equation, It may include part or all of it.

[0105] In any part or combination of the above implementation forms, multiple scaling factor lookup tables contain the same number of scaling factors.

[0106] In any part or combination of the above implementation forms, for each lookup table within multiple scaling factor lookup tables, the step size of the scaling factor in each lookup table is the same, or the precision of the scaling factor in each lookup table is the same.

[0107] In any part or combination of the above implementation forms, the step size of the scaling factor for different lookup tables or the precision of the scaling factor for different lookup tables will differ among multiple scaling factor lookup tables.

[0108] In any part or combination of the above implementation forms, the step size of the scaling factor is the same for one lookup table and different for another lookup table among multiple scaling factor lookup tables.

[0109] In any part or combination of the above embodiments, the multiple scaling factor lookup tables include a first lookup table and a second lookup table. The step of selecting a scaling factor lookup table from multiple scaling factor lookup tables based on the derived scaling factors is: The steps include determining whether the derived scaling factor is greater than a predefined threshold, In response to the determination that the derived scaling factor is smaller than a predefined threshold, the steps include selecting a first lookup table and / or, In response to the determination that the derived scaling factor is not smaller than a predefined threshold, the steps include selecting a second lookup table, Includes.

[0110] In any part or combination of the above embodiments, the step of determining a prediction mode for predicting the current block based on a reference block of a reference frame, based on a coded video bitstream, includes the step of extracting a flag from the coded video bitstream, wherein the flag indicates a prediction mode for predicting the current block based on a reference block of a reference frame.

[0111] In any part or combination of the above embodiments, the method may further include the step of deriving a derived scaling factor based on the current template of the current block and the reference template of the reference block.

[0112] In any part or combination of the above embodiments, the method may further include the step of determining a reference block of a reference frame based on a motion vector.

[0113] In any part or combination of the above implementation forms, the step of reconstructing the current block based on a reference block and a scaling factor determined according to a linear equation includes the step of calculating the pixel value of the current block as a*p+b, where a is the determined scaling factor, b is the determined offset, and p is the reference pixel value at the reference point determined by the motion vector.

[0114] In this disclosure, the orientation of a reference frame may be determined by whether the reference frame is before the current frame in the display order or after the current frame in the display order.

[0115] In various embodiments, one block is coded as BAWP (or LIC, or Composite Weighted Prediction (CWP)) mode, and will be referred to as BAWP below for simplicity of explanation. A flag called bawp_type may be signaled to indicate whether explicit or implicit signaling of BAWP scaling factors is used. If explicit signaling of BAWP scaling factors is used for the current block, the scaling factor or its corresponding index in the lookup table(s) may be signaled to the bitstream and analyzed on the decoder side. The offset value β, as the only parameter to be derived, may, as an unrestricted example, be derived from a linear equation between the templates of the reference block and the current block by least squares fitting. In some implementations, if implicit signaling of BAWP scaling factors is used for the current block, the two parameters to be derived, the scaling factor α and the offset value β, may, as an unrestricted example, be derived from a linear equation between the templates of the reference block and the current block by least squares fitting.

[0116] In various embodiments, the context for signaling bawp_type may depend on at least one of the coded information of the current block and neighboring blocks, and / or the number of neighboring blocks using BAWP / LIC / CWP mode. In a non-restrictive example, a first context is used when none of the neighboring blocks are using BAWP / LIC / CWP mode, a second context is used when one of the neighboring blocks is using BAWP / LIC / CWP mode, and / or a third context is used when two or more neighboring blocks are using BAWP / LIC / CWP mode.

[0117] In some implementations, the context for signaling `bawp_type` may depend on the reference frame index of the current block and / or neighboring blocks. In a non-restrictive example, the first context is used when the reference frame index of the current block and / or neighboring blocks are the same, and / or the second context is used when the reference frame index of the current block and / or neighboring blocks are different.

[0118] In some implementations, the context for signaling `bawp_type` may depend on whether the current block is coded in single-reference or compound predictive mode. As a non-restrictive example, the first context is used when the current block is coded in single-reference mode, and / or the second context is used when the current block is coded in compound predictive mode.

[0119] In some implementations, when BAWP is applied to a composite prediction block, the context for signaling bawp_type may depend on syntax values ​​related to the application of adaptive composite prediction weights.

[0120] In various embodiments, one flag in the bitstream may indicate that the current block is using explicit signaling, while another flag called bawp_sign may be signaled in the bitstream to indicate whether the scaling factor of the current block is greater than or less than one threshold (TH), which may be predefined.

[0121] In some implementations, the value of TH may be signaled in high-level syntax (HLS), such as in sequence parameter sets (SPS), picture parameter sets (PPS), picture headers, slice headers, tile headers, or CTU headers, in non-restrictive examples. In some implementations, TH is set as a fixed value for all video sequences, such as 0 or 1. In some embodiments, TH is set to a derived scale α (or a quantized value of derived scale α), which is calculated based on local illumination variations between the template of the current block (or the sample that gives rise to the current block) and the template of the reference block. In some implementations, the context for signaling bawp_sign may depend on the bawp_sign value of neighboring blocks. In some implementations, after signaling bawp_sign, another flag may be signaled in the bitstream to indicate the magnitude of the scaling factor or adjustment of the scaling factor.

[0122] In various embodiments, multiple scaling factor lookup tables may be included, and the selection between different lookup tables may be implicitly determined based on the derived scaling factor values ​​from the current block template and the reference block template. In some implementations, there may be predefined scaling factor offsets and predefined factors, so the actual scaling factor used to reconstruct the current block is calculated by adding the predefined scaling factor offsets and then dividing by the predefined factor: ((scaling factor determined from multiple scaling factor lookup tables) + (predefined scaling factor offset)) / (predefined factor). In a non-restrictive example, if the predefined scaling factor offset is +1 and the predefined factor is 16, then the actual scaling factor = ((scaling factor determined from multiple scaling factor lookup tables) + 1) / 16.

[0123] In some implementations, the number of scaling factors is the same across multiple lookup tables. In non-restrictive examples, the number of scaling factors in a lookup table may be 2, 3, 4, 5, 6, 7, 8, or 10.

[0124] In some implementations, the step size and / or precision (or absolute value of the scaling factor) of the scaling factor within each lookup table are the same, but differ between different lookup tables. In one unrestricted example, one lookup table is {+2,-2,+3,-3,+4,-4} and another lookup table is {+2,-2,+4,-4,+6,-6}. In another unrestricted example, one lookup table is {+1,-1,+2,-2,+3,-3} and another lookup table is {+2,-2,+4,-4,+6,-6}.

[0125] In some implementations, the step size and / or precision (or absolute value of the scaling factor) of the scaling factor may be the same in one lookup table and different in other lookup tables. In one non-restrictive example, two lookup tables are used. The first lookup table has a fixed step size and contains the values ​​{+2,-2,+3,-3,+4,-4}. The second lookup table has different step sizes and consists of the values ​​{+2,-2,+4,-4,+8,-8}.

[0126] In some implementations, two scaling factor lookup tables are supported, and the choice between these two lookup tables may depend on the proximity of the derived scaling factor to a predefined value TH. In a non-restrictive example, if the derived scaling factor is close to the predefined value TH (e.g., the derived scaling factor is smaller than the predefined value TH), the lookup table with the smaller step size is selected. Otherwise, the other lookup table is selected. For example, TH is set to 1.

[0127] In various embodiments, multiple scaling factor lookup tables are supported, and the selection between different lookup tables may be implicitly determined based on the picture order count (POC) distance between the reference frame and the current frame. In some implementations, two lookup tables are supported. When the POC distance between the reference frame and the current frame falls within a predefined value TH (e.g., less than TH), the lookup table with the smaller step size is selected. Otherwise, the other lookup table is selected. For example, the TH value is set to 4.

[0128] In various embodiments, only one predefined lookup table is supported, and the step size of the absolute value of the scaling factor in the lookup table increases as the index of the scaling factor increases. In some implementations, the absolute value of the scaling factor in the lookup table can only be a power of 2, such as 1, 2, 4, 8, 16, etc. In one non-restrictive example, the scaling factor in the lookup table is {+1,-1,+3,-3,+6,-6,+8,-8}. In another non-restrictive example, the scaling factor in the lookup table is {+1,-1,+3,-3,+6,-6}.

[0129] Various embodiments of this disclosure may include methods for encoding a current block into a video bitstream, which is performed by an encoder and includes inverse processing as any part or all of the processing described for a decoder.

[0130] The operations described above can be combined or arranged in any quantity or order as needed. Two or more steps and / or operations may be performed in parallel. Embodiments and implementations of this disclosure may be used individually or combined in any order. Furthermore, each of the methods (or embodiments), encoders and decoders may be implemented by processing circuits (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored on a non-temporary computer-readable medium. Embodiments of this disclosure may be applied to luma blocks or chroma blocks. The term “block” may be interpreted as a prediction block, coding block, or coding unit (i.e., CU). The term “block” in this disclosure may also be used to refer to a transformation block. In the following sections, when we refer to block size, it may refer to the width or height of the block, or the maximum width and height, or the minimum width and height, or the size of the area (width * height), or the aspect ratio of the block (width:height, or height:width).

[0131] The techniques described above can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, Figure 14 shows a computer system (1800) suitable for carrying out a particular embodiment of the disclosed subject matter.

[0132] Computer software can be coded using any suitable machine code or computer language that can undergo mechanisms such as assembly, compilation, and linking to create code that contains instructions that can be executed directly by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or that can be executed via interpretation, microcode execution, etc.

[0133] Instructions can be executed on various types of computers or computer components, including, for example, personal computers, tablet computers, servers, smartphones, game consoles, and Internet of Things devices.

[0134] The components shown in Figure 14 for the computer system (1800) are illustrative in nature and are not intended to imply any limitation on the scope of use or functionality of computer software implementing embodiments of the present disclosure. Furthermore, the configuration of the components should not be construed as having any dependency or requirement relating to any one or combination of components shown in the exemplary embodiments of the computer system (1800).

[0135] The computer system (1800) may include certain human interface input devices. The input human interface devices may include one or more of the following: keyboard (1801), mouse (1802), trackpad (1803), touch screen (1810), data glove (not shown), joystick (1805), microphone (1806), scanner (1807), and camera (1808) (only one of each is shown).

[0136] The computer system (1800) may also include several human interface output devices. Such human interface output devices can stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., touch screens (1810), tactile feedback via data gloves (not shown) or joysticks (1805); however, tactile feedback devices that are not used as input devices may also exist), sound output devices (e.g., speakers (1809), headphones (not shown)), visual output devices (e.g., screens (1810), including CRT screens, LCD screens, plasma screens, and OLED screens; each may or may not have touch screen input functionality and tactile feedback functionality. Some of the above screens may have the ability to output two-dimensional visual output or output three or more dimensions through means such as stereographic output. They may also include virtual reality glasses (not shown), hologram displays, smoke tanks (not shown)), and printers (not shown).

[0137] A computer system (1800) may also include optical media (1821), including CD / DVD ROM / RW (1820) using CD / DVD or similar media; thumb drives (1822); removable hard drives and solid-state drives (1823); legacy magnetic media such as tapes and floppy disks (not shown); and specialized ROM / ASIC / PLD-based devices such as security dongles (not shown), as well as related media that can be directly operated by humans.

[0138] Those skilled in the art will also understand that the term "computer-readable medium" as used in relation to the protected subject matter disclosed herein does not include transmission media, carrier waves, or other transient signals.

[0139] A computer system (1800) may also include an interface (1854) to one or more communication networks (1855). These networks may be, for example, wireless, wired, or optical networks. Networks may further be local, wide-area, metropolitan, automotive, and industrial, real-time, or latency-tolerant. Examples of networks include local area networks such as Ethernet, cellular networks including wireless LAN, GSM, 3G, 4G, 5G, and LTE, wired television or wireless wide-area digital networks including cable television, satellite television, and terrestrial television, and automotive and industrial networks including CAN bus.

[0140] The aforementioned human interface devices, memory devices that can be directly operated by humans, and network interfaces can be attached to the core (1840) of the computer system (1800).

[0141] The core (1840) may include one or more central processing units (CPUs) (1841), graphics processing units (GPUs) (1842), specialized programmable processing units in the form of field-programmable gate areas (FPGAs) (1843), hardware accelerators for specific tasks (1844), graphics adapters (1850), and the like. These devices may be connected via a system bus (1848) along with read-only memory (ROM) (1845), random access memory (1846), and internal mass storage such as built-in hard drives and SSDs (1847) that are not accessible to the user. In some computer systems, the system bus (1848) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs and GPUs, etc. Peripheral devices can be attached directly to the core's system bus (1848) or via a peripheral bus (1849). In one example, a screen (1810) may be connected to a graphics adapter (1850). Peripheral bus architectures include PCI, USB, and others.

[0142] A computer-readable medium may contain computer code for performing various computer implementation operations. The medium and computer code may be specifically designed and constructed for the purposes of this disclosure, or they may be of a type that is well known and available to those skilled in the computer software technology.

[0143] While this disclosure has described several exemplary embodiments, there are many modifications, substitutions, and alternative equivalents that fall within the scope of this disclosure. Those skilled in the art will therefore understand that numerous systems and methods not expressly shown or described herein can be devised to embody the principles of this disclosure and thus fall within the spirit and scope of this disclosure. [Explanation of symbols]

[0144] 100 Communication System, 150 Communication Network, 200 Communication System, 201 Video Source, 202 Stream, 203 Video Encoder, 204 Video Bitstream, Video Data, 205 Streaming Server, 206 Client Subsystem, 207 Copy, 210 Video Decoder, 211 Video Picture, 212 Display, 213 Video Capture Subsystem, 220 Electronic Device, 230 Electronic Device, 301 Channel, 310 Video Decoder, 312 Rendering Device, Display, 315 Buffer Memory, 320 Parser, 321 Symbol, 330 Electronic Device, 331 Receiver, 351 Inverse Unit, 352 Intra-Picture Prediction Unit, 353 Motion Compensation Prediction Unit, 355 Aggregator, 356 Loop Filter Unit, 357 Reference Picture Memory, 358 Picture Buffer, 401 Video Source, 403 Video Encoder, Video Coder, 410 Video decoder, 420 Electronic device, 430 Source coder, 432 Coding engine, 433 Decoder, 433 Embedded decoder, Local video decoder, 434 Reference picture cache, Reference picture memory, 435 Predictor, 440 Transmitter, 443 Video sequence, 445 Entropy coder, 450 Controller, 460 Communication channel, 503 Video encoder, 521 General-purpose controller, 522 Intra encoder, 523 Residual calculator, 524 Residual encoder, 525 Entropy encoder, 526 Switch, 528 Residual decoder, 530 Interencoder, 610 Video decoder, 671 Entropy decoder, 672 Intra decoder, 673 Residual decoder, 674 Reconstruction module, 680 Interdecoder, 1210 Current block, 1212 Current template, 1220 Reference block, 1222 Reference template, 1230 Motion vector, 1300 flowchart, 1800 computer system, 1801 keyboard, 1802 mouse, 1803 trackpad, 1805 joystick, 1806 microphone, 1807 scanner, 1808 camera, 1809 audio output device speaker, 1810 touch screen, 1821Optical media, 1823 Solid-state drive, 1840 Core, 1843 Field-programmable gate area (FPGA), 1844 Hardware accelerator, 1845 Read-only memory (ROM), 1846 Random-access memory, 1847 Internal mass storage, 1848 System bus, 1849 Peripheral bus, 1850 Graphics adapter, 1854 Interface, 1855 Communication network

Claims

1. A method for decoding the current block of the current frame in a coded video bitstream, A decoding device comprising a memory for storing instructions and a processor for communicating with the memory receives a coded video bitstream, wherein the coded video bitstream comprises the current block of the current frame and a first syntax element indicating a prediction mode of the current block, the memory of the decoding device further stores a plurality of scaling factor lookup tables, the plurality of scaling factor lookup tables include different ranges of scaling factors, and for each lookup table in the plurality of scaling factor lookup tables, the step size of the scaling factor or the precision of the scaling factor is the same; The steps include: determining the prediction mode based on the value of the first syntax element using the decoding device, wherein the prediction mode is used to predict the current block based on the reference block of the reference frame; The decoding device determines a scaling factor from one of the multiple scaling factor lookup tables, The decoding device reconstructs the current block based on the reference block and the scaling factor determined according to the linear equation, Methods that include...

2. The aforementioned multiple scaling factor lookup tables contain the same number of scaling factors, The method according to claim 1.

3. The step of determining the scaling factor from one of the plurality of scaling factor lookup tables is, The method according to claim 1, comprising the step of determining the scaling factor either explicitly based on a second syntax element or implicitly based on a local illumination value.

4. The step of determining the scaling factor from one of the plurality of scaling factor lookup tables is, The steps include selecting a scaling factor lookup table from multiple scaling factor lookup tables based on the derived scaling factors, The steps include determining the scaling factor from the scaling factor lookup table, The method according to claim 1, including the method described in claim 1.

5. The plurality of scaling factor lookup tables include a first lookup table and a second lookup table, The step of selecting the scaling factor lookup table from the plurality of scaling factor lookup tables based on the derived scaling factor is, The steps include determining whether the derived scaling factor is greater than a predefined threshold, In response to the determination that the derived scaling factor is smaller than the predefined threshold, the step of selecting the first lookup table, In response to the determination that the derived scaling factor is greater than or equal to the predefined threshold, the steps include selecting the second lookup table, The method according to claim 4, including the method described in claim 4.

6. Among the plurality of scaling factor lookup tables, the step size of the scaling factor differs for different lookup tables, or the precision of the scaling factor differs for different lookup tables. The method according to claim 1.

7. The step of determining the prediction mode for predicting the current block based on the reference block of the reference frame, based on the coded video bitstream, A step of extracting a flag from the coded video bitstream, wherein the flag indicates the prediction mode for predicting the current block based on the reference block of the reference frame, The method according to claim 1, including the method according to claim 1.

8. The device performs the step of deriving the derived scaling factor based on the current template of the current block and the reference template of the reference block. The method according to claim 1, further comprising:

9. The device performs the step of determining the reference block of the reference frame based on the motion vector. The method according to claim 1, further comprising:

10. The step of reconstructing the current block based on the reference block and the scaling factor determined according to the linear equation is: The step includes calculating the pixel value of the current block as a*p+b, where a is the determined scaling factor, b is the determined offset, and p is the reference pixel value at the reference point determined by the motion vector. The method according to claim 1.

11. A device for decoding the current block of the current frame in a coded video bitstream, Memory for storing instructions, A processor that communicates with the memory, wherein when the processor executes the instruction, the processor is configured to cause the device to carry out the method according to any one of claims 1 to 10, A device equipped with the following features.

12. A non-temporary computer-readable storage medium for storing instructions, wherein when an instruction is executed by a processor, the instruction is configured to cause the processor to perform the method according to any one of claims 1 to 10. Non-temporary computer-readable storage medium.