Enhancement for block adaptive weighted prediction

Through block adaptive weighted prediction technology, multiple scaling factors lookup tables and linear equations are used to enhance the local lighting compensation model, solving the problem of low encoding efficiency of video encoding under local lighting changes, and achieving more efficient video encoding and decoding.

CN120359507APending Publication Date: 2025-07-22TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380085913.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-09-06
Filing Date
2023-09-12
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When existing video encoding technology deals with local lighting changes, it is difficult to effectively compress and decode, resulting in low encoding efficiency.

Method used

Block adaptive weighted prediction (BAWP) technology is used to enhance the local lighting compensation (LIC) model by using multiple scaling factors to predict and reconstruct video blocks.

Benefits of technology

Improve the efficiency and quality of video encoding, especially when dealing with local lighting changes, reduce data redundancy and improve coding efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359507A_ABST
    Figure CN120359507A_ABST
Patent Text Reader

Abstract

The present disclosure generally relates to video encoding / decoding, in particular for enhanced block adaptive weighted prediction. In one embodiment, a method includes receiving an encoded video bitstream, the encoded video bitstream including a current block of a current frame and a first syntax element indicating a prediction mode for the current block, storing a plurality of scaling factor lookup tables, the plurality of scaling factor lookup tables including different ranges of scaling factors, in the plurality of scaling factor lookup tables, the step lengths of the scaling factors in each lookup table are the same, or the precision of the scaling factors is the same; determining, based on the value of the first syntax element, a prediction mode for predicting the current block based on a reference block of the reference frame; determining a scaling factor from one of a plurality of scaling factor lookup tables; and reconstructing the current block based on the reference block and the determined scaling factor.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Incorporation by reference

[0002] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 468,487, filed on May 23, 2023, which is hereby incorporated by reference in its entirety. This application also claims the benefit of priority to U.S. Non - Provisional Patent Application No. 18 / 461,706, filed on September 6, 2023, which is hereby incorporated by reference in its entirety. Technical Field

[0003] The present disclosure describes a set of advanced video / stream encoding / decoding techniques. More specifically, the disclosed techniques relate to enhancements to block - adaptive weighted prediction (BAWP) for compensating local illumination variations. Background Art

[0004] Uncompressed digital video can include a sequence of pictures and may have specific bit - rate requirements for storage, data - processing, and transmission bandwidth in streaming applications. One purpose of video encoding and decoding can be to reduce redundancy in the uncompressed input video signal through various compression techniques. Summary of the Invention

[0005] The present disclosure describes various embodiments of methods, apparatuses, and computer - readable storage media for enhancing block - adaptive weighted prediction (BAWP) to model local illumination compensation (LIC).

[0006] According to one aspect, embodiments of the present disclosure provide a method for decoding a current block of a current frame in an encoded video bitstream. The method includes: a decoding device receives the encoded video bitstream, the encoded video bitstream includes the current block of the current frame and a first syntax element that indicates a prediction mode for the current block, the decoding device includes a memory storing instructions and a processor communicatively coupled to the memory, the memory of the decoding device further stores a plurality of scaling - factor lookup tables, wherein the plurality of scaling - factor lookup tables include scaling factors in different ranges, and for each lookup table in the plurality of scaling - factor lookup tables, the step size of the scaling factors is the same, or the precision of the scaling factors is the same; the decoding device determines the prediction mode based on the value of the first syntax element, wherein the prediction mode is used to predict the current block based on a reference block of a reference frame; the decoding device determines a scaling factor from one of the plurality of scaling - factor lookup tables; the decoding device reconstructs the current block based on the reference block and the determined scaling factor according to a linear equation.

[0007] According to another aspect, embodiments of the present disclosure provide an apparatus for processing a current block of a current frame in an encoded video bitstream. The apparatus includes a memory storing instructions; and a processor communicatively coupled to the memory. When the processor executes the instructions, the processor is configured to cause the apparatus to perform the above-described method for video decoding and / or encoding. In another aspect, embodiments of the present disclosure provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding and / or encoding, cause the computer to perform the above-described method for video decoding and / or encoding.

[0008] The above and other aspects and their implementations are described in more detail in the drawings, the description, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Other features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0010] Figure 1 A simplified block diagram schematic of a communication system (100) according to an example embodiment is shown;

[0011] Figure 2 A simplified block diagram schematic of a communication system (200) according to an example embodiment is shown;

[0012] Figure 3 A simplified block diagram schematic of a video decoder according to an example embodiment is shown;

[0013] Figure 4 A simplified block diagram schematic of a video encoder according to an example embodiment is shown;

[0014] Figure 5 A block diagram of a video encoder according to another example embodiment is shown;

[0015] Figure 6 A block diagram of a video decoder according to another example embodiment is shown;

[0016] Figure 7 An encoding block partitioning scheme according to an example embodiment of the present disclosure is shown;

[0017] Figure 8 Another encoding block partitioning scheme according to an example embodiment of the present disclosure is shown;

[0018] Figure 9 Yet another encoding block partitioning scheme according to an example embodiment of the present disclosure is shown;

[0019] Figure 10 Compound motion compensation is illustrated;

[0020] Figure 11Illustrates an example interpolated reference frame for motion compensation;

[0021] Figure 12 Shows an example template of a current block and a reference block;

[0022] Figure 13 Shows an example logic flow of the method in the present disclosure; and

[0023] Figure 14 Shows a schematic diagram of a computer system according to an example embodiment of the present disclosure. Detailed Description

[0024] The present invention will now be described in detail below with reference to the accompanying drawings, which form a part of the present invention and illustrate specific examples of embodiments by way of illustration. However, note that the present invention can be embodied in various different forms, and thus the subject matter covered or claimed is intended to be construed as not limited to any of the embodiments to be set forth below. Also note that the present invention can be embodied as a method, a device, a component, or a system. Therefore, the embodiments of the present invention can take, for example, the form of hardware, software, firmware, or any combination thereof.

[0025] Throughout the specification and claims, terms may have nuanced meanings that are implied or implicit in the context beyond the explicitly stated meaning. As used herein, the phrase "in one embodiment" or "in some embodiments" does not necessarily refer to the same embodiment, and the phrase "in another embodiment" or "in other embodiments" does not necessarily refer to different embodiments. Similarly, as used herein, the phrase "in one implementation" or "in some implementations" does not necessarily refer to the same implementation, and the phrase "in another implementation" or "in other implementations" does not necessarily refer to different implementations. For example, the claimed subject matter is intended to include combinations of all or part of the exemplary embodiments / implementations.

[0026] Generally, terms may be understood, at least in part, based on their use in context. For example, terms such as (e.g., "and", "or", or "and / or") as used herein may include a variety of meanings that may depend, at least in part, on the context in which such terms are used. Generally, "or" when used in connection with a list (e.g., A, B, or C) is intended to mean A, B, and C (used herein in an inclusive sense) as well as A, B, or C (used herein in an exclusive sense). Additionally, depending at least in part on the context, the terms "one or more" or "at least one" as used herein may be used to describe any feature, structure, or property in the singular sense or may be used to describe a combination of features, structures, or properties in the plural sense. Similarly, depending at least in part on the context, terms such as "a", "an", or "the" may also be understood to convey a singular usage or to convey a plural usage. Further, the terms "based on" or "determined by" may be understood to not necessarily convey a set of exclusive factors but may allow for additional factors that are not necessarily explicitly described, again, depending at least in part on the context.

[0027] As Figure 1 shown, the terminal device may be implemented as a server, a personal computer, and a smart phone, but the applicability of the basic principles disclosed in this application is not limited thereto. The embodiments disclosed in this application may be implemented in a desktop computer, a laptop computer, a tablet computer, a media player, a wearable computer, a dedicated video conferencing device, etc. The network (150) represents any number or type of network that transfers encoded video data between terminal devices, such as including wired (wired) and / or wireless communication networks. The communication network (150) may exchange data in circuit-switched, packet-switched, and / or other types of channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0028] As an example of an application of the subject matter disclosed in this application, Figure 2 illustrates the placement of a video encoder and a video decoder in a streaming environment. The subject matter disclosed in this application is equally applicable to other video applications, including, for example, video conferencing, digital television broadcasting, gaming, virtual reality, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0029] As Figure 2As shown, a video streaming system may include a video capture subsystem (213), which may include a video source (201) such as a digital camera. The video source creates an uncompressed video picture stream or picture stream (202). In an embodiment, the video picture stream (202) includes samples recorded by the digital camera of the video source (201). Compared with the encoded video data (204) (or encoded video bitstream), the video picture stream (202) is depicted as a thick line to emphasize the high data volume of the video picture stream. The video picture stream (202) may be processed by an electronic device (220), which includes a video encoder (203) coupled to the video source (201). The video encoder (203) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. Compared with the uncompressed video picture stream (202), the encoded video data (204) (or encoded video bitstream (204)) is depicted as a thin line to emphasize the lower data volume, which may be stored on a streaming server (205) for future use or directly stored to a downstream video device (not shown). One or more streaming client subsystems, such as Figure 2 the client subsystem (206) and the client subsystem (208) in, may access the streaming server (205) to retrieve copies (207) and (209) of the encoded video data (204). The client subsystem (206) may include, for example, a video decoder (210) in an electronic device (230). The video decoder (210) decodes an input copy (207) of the encoded video data and creates an output video picture (211) stream, which is uncompressed and may be rendered on a display (212) (e.g., a display screen) or other display device (not depicted).

[0030] Figure 3 A block diagram of a video decoder (310) of an electronic device (330) according to an embodiment disclosed in any of the following applications is shown. The electronic device (330) may include a receiver (331) (e.g., a receiving circuit). The video decoder (310) may be used to replace Figure 2 the video decoder (210) in the embodiment. As shown, in Figure 3In this case, a receiver (331) may receive one or more encoded video sequences from a channel (301). To counteract network jitter and / or handle playback timing, a buffer memory (315) may be provided between the receiver (331) and an entropy decoder / parser (320) (hereinafter referred to as "parser (320)"). The parser (320) may reconstruct symbols (321) based on the encoded video sequences. The categories of these symbols include information for managing the operations of the video decoder (310), and information that may be used to control a display device such as a display (e.g., a display screen). The parser (320) may perform parsing / entropy decoding on the encoded video sequences. The parser (320) may extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the encoded video sequences. Subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), etc. The parser (320) may also extract information from the encoded video sequence information, such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc. The reconstruction of the symbols (321) may involve multiple different processing units or functional units. Which units are involved and the manner of involvement may be controlled by subgroup control information parsed by the parser (320) from the encoded video sequences. The first unit may include a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) may receive, from the parser (320), quantized transform coefficients as symbols (321) and control information, including information indicating the type of inverse transform to be used, block size, quantization factor / quantization parameter, quantization scaling matrix, etc. The scaler / inverse transform unit (351) may output a block containing sample values, which may be input into an aggregator (355).

[0031] In some cases, the output samples of the scaler / inverse transform (351) may belong to an intra-coded block, i.e., a block that does not use prediction information from a previously reconstructed picture but may use prediction information from a previously reconstructed portion of the current picture. Such predictive information may be provided by an intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) may generate a block of the same size and shape as the block being reconstructed by using surrounding block information that has been reconstructed and stored in the current picture buffer (358). For example, the current picture buffer (358) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In certain embodiments, the aggregator (355) may add, on a per-sample basis, the prediction information generated by the intra-picture prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351).

[0032] In other cases, the output samples of the scaler / inverse transform unit (351) may belong to an inter-coded and potentially motion-compensated block. In this case, the motion compensation prediction unit (353) may access the reference picture memory (357) based on the motion vector to extract samples for inter-picture prediction. After motion-compensating the extracted reference samples according to the symbol (321) associated with the block, these samples may be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) (the output of unit 351 may be referred to as residual samples or a residual signal) to generate output sample information.

[0033] The output samples of the aggregator (355) may be employed by various loop filtering techniques in a loop filter unit (356) including several types of loop filters. The output of the loop filter unit (356) may be a sample stream, which may be output to the display device (312) and stored in the reference picture memory (357) for subsequent inter-picture prediction.

[0034] Figure 4 A block diagram of a video encoder (403) according to an embodiment disclosed in the present application is shown. The video encoder (403) may be included in an electronic device (420). The electronic device (420) may also include a transmitter (440) (e.g., a transmission circuit). The video encoder (403) may replace Figure 4 the video encoder (403) in the embodiment. The video encoder (403) may receive video samples from a video source (401). According to some example embodiments, the video encoder (403) may encode and compress pictures of a source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by an application. Implementing an appropriate encoding speed constitutes a function of the controller (450). In some embodiments, the controller (450) controls and is functionally coupled to other functional units as described below. The parameters set by the controller (450) may include parameters related to rate control (picture skip, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc.

[0035] In some example embodiments, a video encoder (403) may operate in an encoding loop. The encoding loop may include a source encoder (430) and a (local) decoder (433) embedded in the video encoder (403). The decoder (433) reconstructs symbols to create sample data in a manner similar to how a (remote) decoder creates sample data, even though the embedded decoder 433 processes the video stream encoded by the source encoder 430 without performing entropy encoding (because in the video compression techniques contemplated in this application, any compression between symbols and the encoded video bitstream may be lossless during entropy encoding). At this point, it can be observed that any decoder technique other than parsing / entropy decoding that may only exist in the decoder must also exist in the corresponding encoder in substantially the same functional form. For this reason, this application may sometimes focus on decoder operations, which are related to the decoding part of the encoder. Thus, the description of encoder techniques can be simplified because encoder techniques are reciprocal to the decoder techniques described comprehensively. Only in certain areas or aspects is a more detailed description of the encoder provided below.

[0036] During operation, in some example embodiments, the source encoder (430) may perform motion-compensated predictive coding, predicting and encoding an input picture with reference to one or more previously encoded pictures designated as "reference pictures" in a video sequence.

[0037] The local video decoder (433) decodes the encoded video data of pictures that may be designated as reference pictures. The local video decoder (433) replicates the decoding process that may be performed by a video decoder on a reference picture and may cause the reconstructed reference picture to be stored in the reference picture cache (434). In this way, the video encoder (403) may locally store a copy of the reconstructed reference picture that has the same content (absent transmission errors) as the reconstructed reference picture that will be obtained by a distal (remote) video decoder.

[0038] The predictor (435) may perform a prediction search for the encoding engine (432). That is, for a new picture to be encoded, the predictor (435) may search the reference picture memory (434) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., that may serve as an appropriate prediction reference for the new picture.

[0039] The controller (450) may manage the encoding operations of the source encoder (430), including, for example, setting parameters and subgroup parameters for encoding video data.

[0040] The outputs of all the above functional units can be entropy encoded in an entropy encoder (445). The transmitter (440) can buffer the encoded video sequence created by the entropy encoder (445) to prepare for transmission over a communication channel (460), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (440) can merge the encoded video data from the video encoder (403) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0041] The controller (450) can manage the operation of the video encoder (403). During encoding, the controller (450) can assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, a picture can typically be assigned to one of the following picture types: an intra picture (I picture), a predictive picture (P picture), a bi-predictive picture (B picture), a multi-predictive picture. The source picture can typically be spatially subdivided into a plurality of sample coding blocks, as described in further detail below.

[0042] Figure 5 is a diagram of a video encoder (503) according to another exemplary embodiment disclosed in the present application. The video encoder (503) is configured to receive sample values in a processing block (e.g., a prediction block) within a current video picture in a sequence of video pictures, and encode the processing block into an encoded picture that is part of an encoded video sequence. An exemplary video encoder (503) can be used to replace Figure 4 the video encoder (403) in the embodiment.

[0043] For example, the video encoder (503) receives a matrix of sample values for a processing block. Then, the video encoder (503) uses, for example, rate-distortion (RD) optimization to determine whether to use an intra mode, an inter mode, or a bi-predictive mode to optimally encode the processing block.

[0044] In Figure 5 the example, the video encoder (503) includes an inter encoder (530), an intra encoder (522), a residual calculator (523), a switch (526), a residual encoder (524), a general controller (521), and an entropy encoder (525) coupled together as shown in Figure 5 .

[0045] The inter-frame encoder (530) is configured to receive samples of a current block (e.g., a processing block), compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a later picture in display order), generate inter-frame prediction information (e.g., a description of redundant information according to inter-frame coding techniques, a motion vector, merge mode information), and calculate an inter-frame prediction result (e.g., a predicted block) based on the inter-frame prediction information using any suitable technique.

[0046] The intra-frame encoder (522) is configured to receive samples of a current block (e.g., a processing block), compare the block with encoded blocks in the same picture, generate quantization coefficients after transformation, and in some cases also generate intra-frame prediction information (e.g., intra-frame prediction direction information according to one or more intra-frame coding techniques).

[0047] The general controller (521) is configured to determine general control data, and based on the general control data, control other components of the video encoder (503) to determine a prediction mode of a block, and based on the prediction mode, provide a control signal to the switch (526).

[0048] The residual calculator (523) may be used to calculate the difference (residual data) between a received block and a prediction result of a block selected from the intra-frame encoder (522) or the inter-frame encoder (530). The residual encoder (524) may be used to encode the residual data to generate transform coefficients. The transform coefficients are then subjected to quantization processing to obtain quantized transform coefficients. In various embodiments, the video encoder (503) further includes a residual decoder (528). The residual decoder (528) is configured to perform an inverse transformation and generate decoded residual data. The entropy encoder (525) is configured to format the bitstream to produce an encoded block and perform entropy coding.

[0049] Figure 6 FIG. shows an example video decoder (610) according to another embodiment disclosed in the present application. The video decoder (610) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In an embodiment, the video decoder (610) may replace Figure 4 the video decoder (410) in the embodiment.

[0050] In Figure 6 the embodiment, the video decoder (610) includes an entropy decoder (671), an inter-frame decoder (680), a residual decoder (673), a reconstruction module (674), and an intra-frame decoder (672) coupled together as shown in the example arrangement of Figure 6 FIG.

[0051] An entropy decoder (671) can be used to reconstruct certain symbols from an encoded picture, where these symbols represent syntax elements that make up the encoded picture. The inter-frame decoder (680) can be used to receive inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information. The intra-frame decoder (672) can be used to receive intra-frame prediction information and generate a prediction result based on the intra-frame prediction information. The residual decoder (673) can be used to perform inverse quantization to extract dequantized transform coefficients and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The reconstruction module (674) can be used to combine, in the spatial domain, the residual output by the residual decoder (673) and the prediction result (optionally, output as an inter-frame prediction module or an intra-frame prediction module) to form a reconstructed block, which can be part of a reconstructed picture, and the reconstructed picture can be part of a reconstructed video.

[0052] It should be noted that any suitable technique can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610). In some example embodiments, one or more integrated circuits can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610). In another embodiment, one or more processors executing software instructions can be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610).

[0053] Turning to block partitioning for encoding and decoding, the general partitioning can start from a basic block and can follow a predefined set of rules, a specific pattern, a partitioning tree, or any partitioning structure or scheme. The partitioning can be hierarchical and recursive. After partitioning or dividing the basic block according to any one or a combination of the example partitioning procedures or other procedures described below, a final set of partitions or a final set of encoded blocks can be obtained. Each of these partitions can be at one of various partitioning levels in the partitioning hierarchy and can have various shapes. Each partition can be referred to as a coding block (CB). For the various example partitioning implementations described further below, each resulting CB can have any allowed size and partitioning level. Such partitions are called coding blocks because they can form units for which some basic encoding / decoding decisions can be made and for which encoding / decoding parameters can be optimized, determined, and signaled in the encoded video bitstream. The highest or deepest level in the final partitioning represents the depth of the tree-like coding block partitioning structure. The coding blocks can be luminance coding blocks or chrominance coding blocks. The CB tree structure for each color can be referred to as a coding block tree (CBT). The coding blocks for all color channels can be collectively referred to as coding units (CUs). The hierarchical structure for all color channels can be collectively referred to as coding tree units (CTUs). The partitioning pattern or partitioning structure for the various color channels in a CTU may or may not be the same.

[0054] In some embodiments, the partitioning tree scheme or partitioning tree structure for the luminance channel and the chrominance channel may not need to be the same. In other words, the luminance channel and the chrominance channel can have separate coding tree structures or coding tree patterns. Additionally, whether the luminance channel and the chrominance channel use the same or different coding partitioning tree structures, and what the actual coding partitioning tree structure to be used is, can depend on whether the slice being encoded is a P-slice, a B-slice, or an I-slice. For example, for an I-slice, the chrominance channel and the luminance channel can have separate coding partitioning tree structures or coding partitioning tree structure patterns, while for a P-slice or a B-slice, the luminance channel and the chrominance channel can share the same coding partitioning tree scheme. When applying separate coding partitioning tree structures or patterns, the luminance channel can be partitioned into CBs by one coding partitioning tree structure, and the chrominance channel can be partitioned into chrominance CBs by another coding partitioning tree structure.

[0055] Figure 7 An example predefined 10-way partitioning structure / partitioning pattern that allows recursive partitioning to form a partitioning tree is shown. The root block can start from a predefined level (e.g., starting from a basic block at the 128×128 level or the 64×64 level). Figure 7 The example partitioning structure includes various 2:1 / 1:2 rectangular partitions and 4:1 / 1:4 rectangular partitions. In some example embodiments, Figure 7The rectangular partitions are not allowed to be further subdivided. The coding tree depth can be further defined to indicate the depth of the split starting from the root node or root block. For example, the coding tree depth of the root node or root block can be set to 0, and after further splitting the root block once according to Figure 7 the coding tree depth increases by 1. In some embodiments, only the full square partitions in 710 can be recursively partitioned into the next level of the partition tree according to Figure 7 the pattern.

[0056] In some other exemplary embodiments of the coded block partitioning, a quadtree structure can be used. Such quadtree splitting can be applied hierarchically and recursively to any square partition. Whether to further perform quadtree splitting on the base block or intermediate block or intermediate partition can be adapted according to various local characteristics of the base block or intermediate block / intermediate partition.

[0057] In still some other examples, a ternary partitioning scheme can be used to partition the base block or any intermediate block, as shown in Figure 8 . The ternary pattern can be implemented vertically, as shown in 802, or horizontally, as shown in 804. Although Figure 8 the example split ratio shown is 1:2:1, other ratios can be predefined. In some embodiments, two or more different ratios can be predefined. In some embodiments, the partition width and height of the example ternary tree are always powers of 2 to avoid performing additional transformations.

[0058] The above partitioning schemes can be combined in any way at different partitioning levels. As an example, the above quadtree partitioning scheme and binary partitioning scheme can be combined to partition the base block into a quadtree - binary tree (QTBT) structure. In such a scheme, quadtree splitting or binary splitting can be performed on the base block or intermediate block / intermediate partition, depending on a set of predefined conditions (if specified). Figure 9Specific examples are illustrated in which a basic block quadtree is first divided into four partitions as shown at 902, 904, 906, and 908. Thereafter, each of the resulting partitions is either quadtree partitioned into four additional partitions (such as 908), or binarized into two additional partitions at the next level (either horizontally or vertically, such as 902 or 906, both being symmetric), or is not partitioned (such as 904). For square partitions, recursive binarization or quadtree partitioning can be allowed, as shown by the overall example partition pattern 910 and the corresponding tree structure / representation 920, where solid lines represent quadtree partitioning and dashed lines represent binarization. For each binarization node (non-leaf binary partition), a flag can be used to indicate whether the binarization is horizontal or vertical. For example, as shown at 920 and consistent with the partition structure of 910, the flag "0" can represent a horizontal binarization and the flag "1" can represent a vertical binarization. For quadtree partitioned partitions, no indication of the partition type is needed since quadtree partitioning always divides the block or partition both horizontally and vertically to produce 4 sub-blocks / partitions of equal size. In some embodiments, the flag "1" can represent a horizontal binarization and the flag "0" can represent a vertical binarization.

[0059] In some example embodiments of the QTBT, the quadtree and binarization rule sets can be represented by the following predefined parameters and their associated corresponding functions:

[0060] - CTU size: the size of the root node of the quadtree (the size of the basic block)

[0061] - MinQTSize: the minimum allowed quadtree leaf node size

[0062] - MaxBTSize: the maximum allowed binary tree root node size

[0063] - MaxBTDepth: the maximum allowed binary tree depth

[0064] - MinBTSize: the minimum allowed binary tree leaf node size

[0065] In some example embodiments of the QTBT partitioning structure, the CTU size can be set to 128×128 luminance samples, and two corresponding 64×64 chrominance sample blocks (when considering and using example chrominance subsampling). MinQTSize can be set to 16×16, MaxBTSize can be set to 64×64, MinBTSize (for both width and height) can be set to 4×4, and MaxBTDepth can be set to 4. The quadtree partitioning can be applied to the CTU first to generate quadtree leaf nodes. The size of the quadtree leaf nodes can range from its minimum allowed size of 16×16 (i.e., MinQTSize) to 128×128 (i.e., the CTU size). If the node is 128×128, then the node will not be split by the binary tree first because the size exceeds MaxBTSize (i.e., 64×64). Otherwise, nodes not exceeding MaxBTSize can be partitioned by the binary tree. In Figure 9 the example of Figure 9 , the basic block is 128×128. The basic block can only be split by the quadtree according to a predefined set of rules. The partitioning depth of the basic block is 0. Each of the four resulting partitions is 64×64, which does not exceed MaxBTSize, and can be further split by the quadtree or the binary tree at level 1. This process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), further splitting can be disregarded. When the width of a binary tree node is equal to MinBTSize (i.e., 4), further horizontal splitting can be disregarded. Similarly, when the height of a binary tree node is equal to MinBTSize, further vertical splitting is not considered.

[0066] In some example embodiments, the above QTBT scheme can be used to support the flexibility of having the same QTBT structure for luminance and chrominance or separate QTBT structures. For example, for P slices and B slices, the luminance CTB and chrominance CTB in a CTU can share the same QTBT structure. However, for I slices, the luminance CTB can be partitioned into CUs by one QTBT structure, and the chrominance CTB can be partitioned into chrominance CUs by another QTBT structure. This means that a CU can be used to refer to different color channels in an I slice. For example, an I slice can consist of coded blocks of the luminance component or coded blocks of two chrominance components, and a CU in a P slice or B slice can consist of coded blocks of all three color components.

[0067] The above various CU partitioning schemes and further partitioning of CUs to PBs can be combined in any way. The following specific embodiments are provided as non-limiting examples.

[0068] Inter-frame prediction can be implemented, for example, in a single-reference mode or a composite-reference mode. In some embodiments, a skip flag for a current block (or at a higher level) may first be included in the bitstream to indicate whether the current block is inter-frame coded and not skipped. If the current block is inter-frame coded, another flag may be further included in the bitstream as a signal to indicate whether the single-reference mode or the composite-reference mode is used for prediction of the current block. For the single-reference mode, one reference block may be used to generate a prediction block for the current block. For the composite-reference mode, two or more reference blocks may be used, for example, to generate a prediction block by weighted averaging. One or more reference frame indices may be used, and additionally one or more corresponding motion vectors may be used to identify one or more reference blocks, and these one or more corresponding motion vectors indicate the position offset (e.g., in horizontal pixels and vertical pixels) between the one or more reference blocks and the current block with respect to the frame. For example, an inter-frame prediction block of the current block may be generated based on a single reference block identified by a motion vector in a reference frame as a prediction block in the single-reference mode, while for the composite-reference mode, the prediction block may be generated by weighted averaging of two reference blocks in two reference frames, and the two reference blocks are indicated by two reference frame indices and two corresponding motion vectors. The motion vectors may be encoded in various ways and included in the bitstream.

[0069] In some example embodiments, one or more reference picture lists may be formed based on information in a reference picture set (RPS), and the one or more reference picture lists contain identifications of short-term reference frames and long-term reference frames for inter-frame prediction. For example, for uni-directional inter-frame prediction, a single picture reference list, denoted as L0 reference (or reference list 0), may be formed, and for bi-directional inter-frame prediction, two picture reference lists, denoted as L0 (or reference list 0) and L1 (or reference list 1), may be formed for each of the two prediction directions. The reference frames included in the L0 list and the L1 list may be sorted in various predetermined ways. The lengths of the L0 list and the L1 list may be signaled in the video bitstream. Uni-directional inter-frame prediction may be performed in the single-reference mode, or when in the composite prediction mode and multiple references used for weighted averaging to generate a prediction block are on the same side of the frame where the block to be predicted is located, uni-directional inter-frame prediction may be performed in the composite-reference mode. Bi-directional inter-frame prediction can only be in the composite mode because bi-directional inter-frame prediction involves at least two reference blocks.

[0070] In some embodiments, a Merge Mode (MM) for inter - frame prediction can be implemented. Generally, for the Merge Mode, the motion vector in the single - reference prediction mode of the current PB or one or more motion vectors in the composite - reference prediction mode can be derived from one or more other motion vectors, rather than being calculated and signaled independently. For example, in an encoding system, one or more current motion vectors of the current PB can be represented by one or more differences between one or more current motion vectors and one or more already - encoded motion vectors (referred to as reference motion vectors). Such one or more motion - vector differences (instead of the one or more current motion vectors as a whole) can be encoded, included in the bitstream, and linked to one or more reference motion vectors. Correspondingly, in a decoding system, one or more motion vectors corresponding to the current PB can be derived based on one or more decoded motion - vector differences and one or more decoded reference motion vectors linked thereto. As a specific form of the conventional Merge Mode (MM) inter - frame prediction, this inter - frame prediction based on one or more motion - vector differences can be referred to as Merge Mode with Motion - Vector Difference (MMVD). Thus, the conventional MM, or particularly MMVD, can be implemented to utilize the correlation between motion vectors associated with different PBs and improve the encoding and decoding efficiency. For example, adjacent PBs may have similar motion vectors, so their MVDs can be small and can be effectively encoded. For another example, for blocks that are similar in position / location in space, the motion vectors may be temporally (between frames) correlated.

[0071] In some example embodiments of MMVD, a reference motion vector (RMV) list or a list of MV prediction value candidates for motion vector prediction can be formed for the block being predicted. The RMV candidate list can contain a predetermined number (e.g., two) of MV prediction value candidate blocks, whose motion vectors can be used to predict the current motion vector. The RMV candidate blocks can include blocks selected from adjacent blocks and / or temporal blocks (e.g., blocks at the same position in the previous or next frame of the current frame) in the same frame. These options represent blocks located at spatial or temporal positions relative to the current block, which may have similar or the same motion vectors as the current block. The size of the MV prediction value candidate list can be predetermined. For example, the list can contain two or more candidates. To appear in the RMV candidate list, it may be required that the candidate block has the same one or more reference frames as the current block, must exist (e.g., when the current block is near the edge of the frame, boundary checking needs to be performed), and must have been encoded during the encoding process and / or decoded during the decoding process. In some embodiments, the merge candidate list can first be filled with spatially adjacent blocks (scanned in a specific predefined order) (if available and meeting the above conditions), and then, if there is still space in the list, temporal blocks can be filled next. For example, adjacent RMV candidate blocks can be selected from the left and top blocks of the current block. The RMV prediction value candidate list can be formed dynamically at various levels (sequence, picture, frame, slice, superblock, etc.) as a dynamic reference list (DRL). The DRL can be signaled in the bitstream.

[0072] In some embodiments, the actual MV prediction value candidates can be signaled and used as reference motion vectors for predicting the motion vector of the current block. In the case where the RMV candidate list contains two candidates, a one-bit flag can be used to indicate the selected reference merge candidate, and this one-bit flag is called the merge candidate flag. For each of the multiple motion vectors predicted using the MV prediction value for the current block predicted in the composite mode, it can be associated with a reference motion vector from the merge candidate list. The encoder can determine which RMV candidate more closely predicts the MV of the current encoded block and signal the selection as an index of the DRL.

[0073] In some example embodiments of MMVD, after selecting an RMV candidate and using it as the basic motion vector prediction value of the motion vector to be predicted, a motion vector difference (MVD or delta MV, which represents the difference between the motion vector to be predicted and the reference candidate motion vector) can be calculated in the encoding system. This MVD can include information representing the magnitude and direction of the MV difference, and both can be signaled in the bitstream in various ways.

[0074] In some example embodiments of MMVD, a distance index may be used to specify the magnitude information of the motion vector difference and to indicate one of a set of predefined offsets that represent predefined motion vector differences from a starting point (reference motion vector). The MV offset according to the signaled index may then be added to the horizontal or vertical component of the starting (reference) motion vector. An example of the predefined relationship between the distance index and the predefined offsets is specified in Table 1.

[0075] Table 1 - Example relationship between distance index and predefined MV offsets

[0076]

[0077] In some example embodiments of MMVD, a direction index may further be signaled and used to represent the direction of the MVD relative to the reference motion vector. In some embodiments, the direction may be restricted to either the horizontal or vertical direction. An example of a 2-bit direction index is shown in Table 2. In the example of Table 2, the interpretation of the MVD may vary according to the information of the starting MV / reference MV. For example, when the starting MV / reference MV corresponds to a uni-directionally predicted block, or to a bi-directionally predicted block and both reference frame lists point to the same side of the current picture (i.e., the POCs of both reference pictures are greater than the POC of the current picture, or both are less than the POC of the current picture), the signs in Table 2 may specify the sign (direction) of the MV offset added to the starting MV / reference MV. When the starting MV / reference MV corresponds to a bi-directionally predicted block and the two reference pictures are on different sides of the current picture (i.e., the POC of one reference picture is greater than the POC of the current picture while the POC of the other reference picture is less than the POC of the current picture), and the difference between the reference POC in picture reference list 0 and the POC of the current frame is greater than the difference between the reference POC in picture reference list 1 and the POC of the current frame, the signs in Table 2 may specify the sign of the MV offset added to the reference MV corresponding to the reference picture in picture reference list 0, while the sign of the offset of the MV corresponding to the reference picture in picture reference list 1 may have an opposite value (opposite sign of the offset). Otherwise, if the difference between the reference POC in picture reference list 1 and the POC of the current frame is greater than the difference between the reference POC in picture reference list 0 and the POC of the current frame, then the signs in Table 2 may specify the sign of the MV offset added to the reference MV associated with picture reference list 1, and the sign of the offset added to the reference MV associated with picture reference list 0 has an opposite value.

[0078] Table 2 - Example embodiments of signs of MV offsets specified by the direction index

[0079] Direction Index 00 01 10 11 x - axis (Horizontal) + - Not applicable Not applicable y - axis (Vertical) Not applicable Not applicable + -

[0080] In some example embodiments, the MVD may be scaled according to the POC difference in each direction. If the POC differences in the two lists are the same, no scaling is required. Otherwise, if the POC difference in reference list 0 is greater than the POC difference in reference list 1, the MVD of reference list 1 is scaled. If the POC difference in reference list 1 is greater than the POC difference in list 0, the MVD of list 0 may be scaled in the same manner. If the starting MV is unidirectionally predicted, the MVD is added to the available MV or the reference MV.

[0081] In some example embodiments for MVD coding and signaling for bidirectional composite prediction, in addition to or as an alternative to separately coding and signaling the two MVDs, symmetric MVD coding may be implemented such that only one MVD needs to be signaled and the other MVD may be derived from the signaled MVD. In such embodiments, the motion information including the reference picture index of list 0 and the reference picture index of list 1 is not signaled simultaneously. Specifically, at the slice level, a flag (referred to as "mvd_l1_zero_flag") may be included in the bitstream to indicate whether reference list -1 is not signaled in the bitstream. If this flag is 1, indicating that reference list -1 is equal to zero (and thus not signaled), then the bidirectional prediction flag (referred to as "BiDirPredFlag") may be set to 0, meaning there is no bidirectional prediction. Otherwise, if mvd_l1_zero_flag is zero, if the nearest reference picture in list -0 and the nearest reference picture in list -1 form a forward reference picture and a backward reference picture pair, or a backward reference picture and a forward reference picture pair, then BiDirPredFlag may be set to 1 and both the list -0 reference picture and the list -1 reference picture are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. When BiDirPredFlag is 1, it may indicate that a symmetric mode flag is additionally signaled in the bitstream. When BiDirPredFlag is 1, the decoder may extract this symmetric mode flag from the bitstream. For example, the symmetric mode flag may be signaled at the CU level (if needed) and may indicate whether the symmetric MVD coding mode is being used on the corresponding CU. When the symmetric mode flag is 1, it indicates that the symmetric MVD coding mode is used and only the reference picture indices of list -0 and list -1 (referred to as "mvp_l0_flag" and "mvp_l1_flag") and the MVD associated with list -0 (referred to as "MVD0") are signaled, and the other motion vector difference "MVD1" will be derived rather than signaled. For example, MVD1 may be derived as -MVD0. In this way, in the example symmetric MVD mode, only one MVD is signaled.

[0082] In some other example embodiments of MV prediction, for both single-reference mode MV prediction and composite-reference mode MV prediction, a coordinated scheme can be used to implement the conventional merge mode MMVD and some other types of MV prediction. The way of MV prediction for the current block can be signaled using various syntax elements. For example, for the single-reference mode, the following MV prediction modes can be signaled: NEARMV - directly use a motion vector prediction value (MVP) in the list indicated by the DRL (Dynamic Reference List) index without any MVD; NEWMV - use a motion vector prediction value (MVP) in the list signaled by the DRL index as a reference and apply an increment to this MVP (e.g., use MVD); GLOBALMV - use a motion vector based on frame-level global motion parameters.

[0083] Similarly, for the composite reference inter prediction mode using two reference frames corresponding to the two MVs to be predicted, the following MV prediction modes can be signaled: NEAR_NEARMV - for each of the two MVs to be predicted, use a motion vector prediction value (MVP) in the list signaled by the DRL index without using MVD. NEAR_NEWMV - to predict the first of the two motion vectors, use a motion vector prediction value (MVP) in the list signaled by the DRL index as a reference MV without using MVD; to predict the second of the two motion vectors, use one of the motion vector prediction values (MVPs) in the list signaled by the DRL index as a reference MV in combination with an additionally signaled incremental MV (MVD). NEW_NEARMV - to predict the second of the two motion vectors, use a motion vector prediction value (MVP) in the list signaled by the DRL index as a reference MV without MVD; to predict the first of the two motion vectors, use a motion vector prediction value (MVP) in the list signaled by the DRL index as a reference MV in combination with an additionally signaled incremental MV (MVD). NEW_NEWMV – use a motion vector prediction value (MVP) in the list signaled by the DRL index as a reference MV and use it in combination with an additionally signaled incremental MV to predict each of the two MVs. GLOBAL_GLOBALMV - use the MV in each reference based on frame-level global motion parameters.

[0084] Accordingly, the term "NEAR" above refers to: using a reference MV without any MVD as the MV prediction for the regular merge mode, while the term "NEW" refers to: an MV prediction that involves using a reference MV and offsetting it with a signaled or derived MVD, such as in the MMVD mode. For composite inter prediction, between two references or two MVDs, both the reference base motion vector and the motion vector delta may generally be different or independent, even though the two MVDs may be related, for example, and this correlation can be exploited to reduce the amount of information needed to signal two motion vector deltas. To exploit this correlation, joint signaling of the two MVDs can be implemented and indicated in the bitstream, as described in further detail below.

[0085] In some example embodiments of the MVD, the MVD may be allowed to have a predefined pixel resolution. For example, 1 / 8 pixel motion vector precision (or accuracy) may be allowed. The MVDs in the various MV prediction modes described above can be constructed and signaled in various ways. In some embodiments, one or more of the above-mentioned motion vector differences in reference frame list 0 or list 1 can be signaled using various syntax elements.

[0086] For example, a syntax element called "mv_joint" can specify which components of the associated motion vector difference are non-zero. For example, mv_joint has the following values: 0, which can indicate that there is no non-zero MVD along the horizontal or vertical direction; 1, which can indicate that there is only a non-zero MVD along the horizontal direction; 2, which can indicate that there is only a non-zero MVD along the vertical direction; and / or 3, which can indicate that there are non-zero MVDs along both the horizontal and vertical directions.

[0087] When the "mv_joint" syntax element of the MVD signals the absence of non-zero MVD components, no further MVD information is signaled. However, if the "mv_joint" syntax element signals the presence of one or two non-zero components, additional syntax elements can be signaled for each non-zero MVD component, as described below.

[0088] For example, a syntax element called "mv_sign" can be used to additionally specify whether the corresponding motion vector difference component is positive or negative.

[0089] For another example, a syntax element called "mv_class" can be used to specify a motion vector difference level for a corresponding non-zero MVD component in a predefined set of levels. For example, predefined levels of motion vector differences can be used to divide the continuous magnitude space of the motion vector difference into non-overlapping ranges of levels. Thus, the signaled MVD level indicates the magnitude range of the corresponding MVD component. In the example embodiment shown in Table 3 below, higher levels correspond to motion vector differences with larger magnitude ranges. The notation (n,m] is used to denote a motion vector difference range greater than n pixels and less than or equal to m pixels.

[0090] Table 3: Magnitude levels of motion vector differences

[0091] MV Level Magnitude of MVD MV_CLASS_0 (0,2] MV_CLASS_1 (2,4] MV_CLASS_2 (4,8] MV_CLASS_3 (8,16] MV_CLASS_4 (16,32] MV_CLASS_5 (32,64] MV_CLASS_6 (64,128] MV_CLASS_7 (128,256] MV_CLASS_8 (256,512] MV_CLASS_9 (512,1024] MV_CLASS_10 (1024,2048]

[0092] In some other examples, a syntax element called "mv_bit" can be further used to specify the integer part of the offset between a non-zero motion vector difference component and the starting magnitude of the corresponding signaled MV level magnitude range. In some other examples, a syntax element called "mv_fr" can also be used to specify the first 2 fractional bits of the motion vector difference of the corresponding non-zero MVD component, while a syntax element called "mv_hp" can be used to specify the third fractional bit (high-resolution bit) of the motion vector difference of the corresponding non-zero MVD component. The two-bit "mv_fr" substantially provides 1 / 4 pixel MVD resolution, and the "mv_hp" bit can further provide 1 / 8 pixel resolution. In some other embodiments, more than one "mv_hp" bit can be used to provide a finer MVD pixel resolution than 1 / 8 pixel. In some example embodiments, additional flags can be signaled at one or more of the various levels to indicate whether 1 / 8 pixel or higher MVD resolution is supported. If the MVD resolution is not applied to a particular coding unit, the above syntax elements for the corresponding unsupported MVD resolution may not be signaled.

[0093] However, in some other exemplary embodiments, the resolution of the motion vector differences in each MVD magnitude level may vary or be adaptive. Specifically, for large MVD magnitudes in higher MVD levels, high-resolution MVD may not bring a statistically significant improvement in terms of compression efficiency or coding gain. Thus, for a larger range of MVD magnitudes (which corresponds to higher MVD levels), the MVD may be encoded with a decreasing resolution or a non-increasing resolution (integer pixel resolution or fractional pixel resolution). The term "resolution" may be further referred to as "pixel resolution". The adaptive MVD resolution can be implemented in various ways, as described in the following exemplary embodiments, to achieve better overall compression efficiency. Specifically, due to statistical observations, processing the MVD resolution of large magnitude MVDs or high-level MVDs in a non-adaptive manner, at a level similar to the MVD resolution of low magnitude MVDs or low-level MVDs, does not significantly improve the inter-frame prediction residual coding efficiency of blocks with large magnitude MVDs or high-level MVDs. Therefore, the number of signaling bits reduced by lowering the precision of the MVD may be greater than the number of bits required to encode the inter-frame prediction residual due to the lower MVD precision. In other words, using a higher MVD resolution for large magnitude MVDs or high-level MVDs may not result in much coding gain compared to using a lower MVD resolution.

[0094] In some embodiments, for the NEW_NEARMV and NEAR_NEWMV modes, the precision of the MVD depends on the associated MVD level and magnitude. First, fractional MVD may be allowed only when the MVD magnitude is equal to or less than one pixel. Second, when the value of the associated MV level is equal to or greater than MV_CLASS_1, only one MVD value may be allowed. For MV level 1 (MV_CLASS_1), MV level 2 (MV_CLASS_2), MV level 3 (MV_CLASS_3), MV level 4 (MV_CLASS_4), or MV level 5 (MV_CLASS_5), the MVD values in each MV level are derived as 4, 8, 16, 32, 64. In some examples, the single allowed value may be the higher-end value of each range of these MV levels in Table 3. The allowed MVD values in each MV level can be illustrated in Table 4.

[0095] Table 4: Adaptive MVD in Each MV Magnitude Level

[0096]

[0097]

[0098] In some example embodiments, each MVD level may be associated with a single allowed resolution. In some other embodiments, one or more MVD levels may each be associated with two or more optional MVD pixel resolutions. For example, the adaptive allowed MVD pixel resolutions may include, but are not limited to, 1 / 64 pixel, 1 / 32 pixel, 1 / 16 pixel, 1 / 8 pixel, 1 / 4 pixel, 1 / 2 pixel, 1 pixel, 2 pixels, 4 pixels... (in descending order of resolution).

[0099] In some other example embodiments, for MV levels equal to or higher than a threshold MV level, only a single MVD value may be allowed. For example, this threshold MV level may be MV_CLASS_2. Thus, MV_CLASS_2 and above levels may only allow a single MVD value and no fractional pixel resolution.

[0100] Turning to the various composite inter prediction modes, in the composite inter prediction mode, each MV may be predicted by a reference motion vector and encoded by an MVD, and the two MVDs may be signaled separately or jointly in the bitstream, as described above. Thus, in some example embodiments, in addition to the above NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes, for the mode of signaling the MVDs of reference list 0 and reference list 1 jointly, another inter prediction mode called JOINT_NEWMV may be introduced. Specifically, if the inter prediction mode is indicated as NEW_NEWMV, the MVDs of reference list 0 and reference list 1 are signaled separately, while when the inter prediction mode is indicated as the JOINT_NEWMV mode, the MVDs of reference list 0 and reference list 1 are signaled jointly. Particularly for the joint MVD, it may only be necessary to signal and transmit one MVD (called joint_delta_mv) in the bitstream, and the MVDs of reference list 0 and reference list 1 may be derived from joint_delta_mv. The derived MVDs may then be combined with the reference motion vectors in reference list 0 or reference list 1 to generate two motion vectors for locating the reference blocks for composite inter prediction.

[0101] In some embodiments of composite inter prediction, the JOINT_NEWMV mode may be signaled together with the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes. In such embodiments, a syntax may be included in the bitstream for indicating any one of these optional composite inter prediction modes at any one of various signaling levels (e.g., sequence level, picture level, frame level, slice level, tile level, superblock level, etc.). Alternatively, the JOINT_NEWMV mode may be implemented as a submode of the NEW_NEWMV mode. In other words, in the NEW_NEWMV mode, the two MVDs of two reference blocks may be jointly signaled (thus the JOINT_NEWMV submode), or the two MVDs of the two reference blocks may not be jointly signaled (another submode of the NEW_NEWMV mode). In such an embodiment, a first syntax element may be included in the bitstream for indicating any one of the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes, and when the first syntax element indicates that the NEW_NEWMV mode is selected for an encoded block, a second syntax element may further be included in the bitstream, which can be extracted by the decoder for indicating whether the MVD of the encoded block is signaled separately or jointly.

[0102] In some example embodiments, when the JOINT_NEWMV mode is signaled and the POC distances between the two reference frames and the current frame are different, the MVD may be scaled based on the POC distances for reference list 0 or reference list 1. Specifically, the distance between reference frame list 0 and the current frame may be denoted as td0, and the distance between reference frame list 1 and the current frame may be denoted as td1. If td0 is equal to or greater than td1, the joint_mvd may be directly used for reference list 0, and the MVD for reference list 1 may be derived based on Equation (1) according to the joint_mvd.

[0103]

[0104] Otherwise, if td1 is equal to or greater than td0, the joint_mvd is directly used for reference list 1, and the MVD for reference list 0 is derived based on Equation (2) according to the joint_mvd.

[0105]

[0106] In some example embodiments, another inter-frame coding mode named AMVDMV can be added to the single-reference case. When the AMVDMV mode is selected, it indicates that AMVD (Adaptive Motion Vector Difference) is applied to signal the MVD. A flag (e.g., named amvd_flag) can be added in the JOINT_NEWMV mode to indicate whether AMVD is applied to the joint MVD coding mode. When the adaptive MVD resolution is applied to the joint MVD coding mode (also named the joint AMVD coding mode), the MVDs of two reference frames are jointly signaled, and the precision of the MVD is implicitly determined by the MVD magnitude. Otherwise, the MVDs of two (or more than two) reference frames are jointly signaled, and traditional MVD coding is applied. Alternatively, a new inter-frame prediction mode named JOINT_AMVDNEWMV is added to indicate that AMVD is applied to the joint MVD coding mode.

[0107] Go to the composite inter-frame prediction mode, such as Figure 10 shown, these modes create the predicted value of the block in the current frame F i-1 by combining two hypothesized values of the motion vectors MV0 and MV1 from two different reference frames F i+1 and reference frame F i . Therefore, two motion information components (e.g., motion vectors) can be signaled for each block in the bitstream.

[0108] Alternatively, as Figure 11 shown, the information in two reference frames F i-1 and F i+1 can be combined, and can be projected to the same time instance as the current frame F i using an interpolation process to generate an interpolated frame. Multiple TIP modes can be supported. In one TIP mode, the interpolated frame can be used as an additional reference frame. The coded blocks of the current frame F i can directly refer to the interpolated frame, with only the overhead cost of the single inter-frame prediction mode, while leveraging the information from two different references. In another TIP mode, the interpolated frame can be directly specified as the output of the decoding process of the current frame F i , while skipping any other traditional encoding and decoding steps. This mode may have significant advantages in terms of coding and complexity, especially in low-bitrate applications.

[0109] In some embodiments, adaptive motion vector resolution (AMVR) can be supported. In a specific example embodiment, a total of 7 MV precisions (8, 4, 2, 1, 1 / 2, 1 / 4, 1 / 8) can be supported. For each prediction block, the AMVR encoder can search all supported precision values and signal the best precision to the decoder.

[0110] To reduce the encoder running time, two sets of MV precisions are supported. Each precision set can contain, for example, 4 predefined precisions. The precision set can be adaptively selected at the frame level based on the maximum precision value of the frame. This maximum precision can be signaled in the frame header. Table 5 summarizes the example supported precision values based on an example frame-level maximum precision.

[0111] Table 5: MV precisions supported in two sets

[0112] Frame - level Maximum Precision Supported MV Precision 1 / 8 1 / 8、1 / 2、1、4 1 / 4 1 / 4、1、4、8

[0113] In some example embodiments of AMVR, there can be a frame-level flag to indicate whether the MV of the frame contains sub-pixel precision. AMVR can be enabled only when the value of the cur_frame_force_integer_mv flag is 0. In AMVR, if the precision of a block is lower than the maximum precision, the motion model and interpolation filter may not be signaled. If the precision of a block is lower than the maximum precision, the motion mode is inferred to be translational motion, and the interpolation filter is inferred to be the REGULAR interpolation filter. Similarly, if the precision of a block is 4 pixels or 8 pixels, the inter-frame / intra-frame mode may not be signaled and can be inferred to be 0.

[0114] In some embodiments, when a block is encoded in the joint MVD coding mode JOINT_NEWMV or JOINT_AMVDNEWMV, a new syntax element named mvd_scaling_factor_idx can be signaled in the bitstream to explicitly indicate the MVD scaling factor between reference frame 0 and reference frame 1. As shown in Table 6 and Table 7, two predefined lookup tables can be used to store the scaling factors supported / allowed by JOINT_NEWMV or JOINT_AMVDNEWMV respectively. The associated entry index of the selected scaling factor in the lookup table can be signaled in the bitstream. For the JOINT_AMVDNEWMV mode, the same scaling factor is applied to both the vertical and horizontal components of the MVD of reference frame list 0 and / or reference frame list 1. For the JOINT_NEWMV mode, the scaling factor of one component (vertical or horizontal) of the MVD can be limited to 1, and the scaling factor of the other component of the MVD can be other values, such as 2 or 1 / 2. In one example, the MVD (mvd_ref0 or mvd_ref1) of reference frame list 0 or reference frame list 1 is calculated by the following equation.

[0115] mvd_ref0 = joint_mvd

[0116]

[0117] Table 6: Scaling factors for JOINT_AMVDNEWMV

[0118] Index Scaling Factors for x - axis and y - axis 0 1 1 2 2 1 / 2

[0119] Table 7: Scaling factors for JOINT_NEWMV

[0120] Index Scaling Factor for x - axis Scaling Factor for y - axis 0 1 1 1 1 2 2 2 1 3 1 1 / 2 4 1 / 2 1

[0121] In some example embodiments, in the Bi - directional prediction with CU - level weights (BCW) mode, a bi - directional prediction signal can be generated by averaging two prediction signals obtained from two different reference pictures, and / or using two different motion vectors. In some other embodiments, the bi - directional prediction mode can be extended beyond simple averaging to allow weighted averaging of the two prediction signals. For example, P bi-pred = ((8 - w)*P0+w*P1 + 4) >> 3. In weighted - average bi - directional prediction, five weights, w ∈ { - 2, 3, 4, 5, 10}, can be allowed. When w equals 4, equal weighting factors are used for weighted averaging of two prediction samples. For each bi - directional prediction CU, the weight w can be determined in one of the following two ways: 1) For non - merged CUs, after the motion vector difference, signal the weight index; and / or 2) For merged CUs, infer the weight index from adjacent blocks based on the merge candidate index. BCW can be applied only to CUs with 256 or more luma samples (i.e., CU width times CU height is greater than or equal to 256). For low - latency pictures, all 5 weights are used. For non - low - latency pictures, only 3 weights (w ∈ {3, 4, 5}) are used.

[0122] In another embodiment, the COMPOUND_AVERAGE mode can be extended to implement weighted averaging of the prediction signal. This extended mode is named Compound Weighted Prediction (CWP). Specifically, when the two reference frames are from different directions, five weighting factors w ∈ {8, 10, 6, 12, 4} are supported, while when the two reference frames are from the same direction, five different weighting factors w ∈ {8, 12, 4, 20, - 4} are supported. Thus, the above equation can be modified to P(x,y) = (W×P0(x,y)+(16 - W)×P1(x,y)+8) >> 4.

[0123] In some embodiments, the index of the selected weighting factor can be signaled when all of the following conditions are met: (i) the composite type is COMPOUND_AVERAGE; (ii) the inter-frame prediction mode is NEAR_NEARMV, JOINT_NEWMV, or JOINT_AMVDNEWMV; and (iii) the MVD scaling factor (1, 1) is used in the joint MVD mode. In some embodiments, the index of the selected weighting factor can be signaled when only a portion of the above conditions are met. In some embodiments, for the skip mode, the index of the selected weighting factor is the DRL index based on the motion vector prediction value, inferred from adjacent blocks.

[0124] In some embodiments, block adaptive weighted prediction (BAWP) can include block-level weighted prediction to model the local illumination change between the current block and its predicted block as a function of the local illumination change between the current block template (or causal samples of the current block) and the reference block template. The template of the current block (1210) (or referred to as the current template 1212) and the template of the reference block (1220) (or referred to as the reference template 1222) are shown in Figure 12 FIG. The reference block can be indicated or determined by a motion vector (MV, 1230). The current block can be in the current picture (or current frame), and the reference block can be in the reference picture (or reference frame). In some embodiments, the function can be a linear function. The parameters of the function can be represented by a scaling factor α and an offset β, which form a linear equation, i.e., α*p[x]+β, to compensate for the illumination change, where p[x] is the reference sample pointed to by the MV at position x on the reference picture. The scaling factor α can be the value (coefficient) of the multiplicative factor relative to the reference block, and the offset β can be the value (parameter) of the subtraction (or summation) offset relative to the reference block. In some embodiments, α and β can be derived based on the current block template and the reference block template, and no signaling overhead is required for α and β. In some embodiments, a BAWP flag can be signaled for a single inter-frame prediction mode to indicate the use of BAWP. BAWP can be applied to blocks larger than or equal to 8×8 and encoded in a single inter-frame prediction mode. In some embodiments, BAWP can be applied only to the luminance component.

[0125] In some embodiments, when a block is encoded in BAWP mode, another flag may be signaled to indicate whether explicit signaling or implicit signaling of the scaling factor is used. When implicit signaling of the BAWP scaling factor is employed, the BAWP scaling factor and the offset value may be derived based on a linear equation from the template of the current block and the template of the reference block pointed to by the motion vector of the current block. When explicit signaling of the BAWP scaling factor is used, another flag is signaled to indicate which scaling factor the current block uses, and the offset value is derived as zero. The supported scaling factors may be stored in one or more predefined look-up tables, and the index of the scaling factor in the look-up table may be signaled in the bitstream and parsed at the decoder side.

[0126] In some embodiments, there are some problems or challenges with the BAWP method, particularly the look-up table for storing the scaling factors. For example, when explicit signaling of BAWP is used, the encoding / decoding efficiency may not be optimal when only one fixed look-up table is employed. As another example, the signaling of the flag indicating explicit and implicit signaling does not consider the correlation between the current block and its neighboring blocks, which may not be optimal in terms of encoding / decoding efficiency. The present disclosure describes various embodiments for enhancing BAWP, solving at least one of the problems or challenges discussed above, improving encoding / decoding efficiency, and advancing video codec technology.

[0127] Each of the embodiments and / or implementations described in the present disclosure may be performed alone or in any order of combination. Additionally, each of these methods (or embodiments), encoders, and decoders may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). The one or more processors execute a program stored in a non-volatile computer-readable medium. In the present disclosure, the term "block" may be interpreted as a prediction block, a coded block, or a coding unit (CU).

[0128] Figure 13FIG. 1300 shows a flowchart of an exemplary method that follows the basic principles of the above-described embodiments and is used to enhance block adaptive weighted prediction (BAWP) to compensate for local illumination changes. The exemplary decoding method process starts at 1301 and may include some or all of the following steps: S1310, a decoding device receives an encoded video bitstream that includes a current block of a current frame and a first syntax element that indicates a prediction mode of the current block. The decoding device includes a memory that stores instructions and a processor that communicates with the memory. The memory of the decoding device further stores a plurality of scaling factor look-up tables, where the plurality of scaling factor look-up tables include scaling factors in different ranges, and in the plurality of scaling factor look-up tables, the step size of the scaling factors in each look-up table is the same, or the precision of the scaling factors is the same; S1320, the decoding device determines the prediction mode based on the value of the first syntax element, where the prediction mode is used to predict the current block based on a reference block of a reference frame; S1330, the decoding device determines a scaling factor from one of the plurality of scaling factor look-up tables; and / or S1340, the decoding device reconstructs the current block based on the reference block and the determined scaling factor according to a linear equation. The example method stops at S1399.

[0129] In any part or combination of the above-described embodiments, the plurality of scaling factor look-up tables include the same number of scaling factors.

[0130] In any part or combination of the above-described embodiments, determining a scaling factor from one of the plurality of scaling factor look-up tables includes: explicitly determining the scaling factor based on a second syntax element, or implicitly determining the scaling factor based on a local illumination value.

[0131] In some embodiments, local illumination changes may refer to intensity changes of a block (e.g., a luminance channel and / or a chrominance channel) caused by local illumination (e.g., illumination from a lamppost or a passing vehicle). The local illumination value may refer to at least one of the following: the local illumination change between a current block template and a reference block template, the local illumination change between causal samples of the current block and the reference block template; and / or a LIC parameter.

[0132] In any part or combination of the above-described embodiments, determining a scaling factor from one of the plurality of scaling factor look-up tables includes: selecting a scaling factor look-up table from the plurality of scaling factor look-up tables based on a derived scaling factor; and / or determining a scaling factor from the selected scaling factor look-up table.

[0133] In some embodiments, the range of the scaling factor lookup table may refer to the value between the smallest scaling factor and the largest scaling factor in the scaling factor lookup table. Taking {+1, -1, +3, -3, +6, -6} as a non-limiting exemplary scaling factor lookup table, its range may be 12, which is the value between -6 (the smallest scaling factor in the scaling factor lookup table) and +6 (the largest scaling factor in the scaling factor lookup table).

[0134] In some embodiments, the range of the scaling factor lookup table may refer to the value between the smallest absolute scaling factor and the largest absolute scaling factor in the scaling factor lookup table. Taking {+1, -1, +3, -3, +6, -6} as a non-limiting exemplary scaling factor lookup table, its range may be 5, which is the value between 1 (the smallest absolute scaling factor in the scaling factor lookup table) and 6 (the largest absolute scaling factor in the scaling factor lookup table).

[0135] Another exemplary decoding method may include some or all of the following steps: receiving an encoded video bitstream; determining, based on the encoded video bitstream, a prediction mode for predicting a current block based on a reference block of a reference frame; selecting, based on the derived scaling factor, one scaling factor lookup table from a plurality of scaling factor lookup tables, the plurality of scaling factor lookup tables including different ranges; determining, based on the encoded video bitstream, a scaling factor from the selected scaling factor lookup table; and / or reconstructing the current block based on the reference block and the determined scaling factor according to a linear equation.

[0136] In any part or combination of the above embodiments, the plurality of scaling factor lookup tables include the same number of scaling factors.

[0137] In any part or combination of the above embodiments, for each lookup table in the plurality of scaling factor lookup tables, the step size of the scaling factors in each lookup table is the same, or the precision of the scaling factors in each lookup table is the same.

[0138] In any part or combination of the above embodiments, among the plurality of scaling factor lookup tables, the step sizes of the scaling factors of different lookup tables are different, or the precisions of the scaling factors of different lookup tables are different.

[0139] In any part or combination of the above embodiments, among the plurality of scaling factor lookup tables, the step size of the scaling factors in one lookup table is the same, while the step size of the scaling factors in another lookup table is different.

[0140] In any part or combination of the above-described embodiments, the plurality of scaling factor lookup tables include a first lookup table and a second lookup table; based on the derived scaling factor, one of the plurality of scaling factor lookup tables is selected, including: determining whether the derived scaling factor is greater than a predefined threshold, when it is determined that the derived scaling factor is less than the predefined threshold, selecting the first lookup table, and / or when it is determined that the derived scaling factor is not less than the predefined threshold, selecting the second lookup table.

[0141] In any part or combination of the above-described embodiments, based on the encoded video bitstream, determining a prediction mode for predicting a current block based on a reference block of a reference frame includes: extracting a flag from the encoded video bitstream, where the flag indicates the prediction mode for predicting the current block based on the reference block of the reference frame.

[0142] In any part or combination of the above-described embodiments, the method may further include: deriving a derived scaling factor based on a current template of the current block and a reference template of the reference block.

[0143] In any part or combination of the above-described embodiments, the method may further include: determining a reference block of a reference frame based on a motion vector.

[0144] In any part or combination of the above-described embodiments, reconstructing the current block based on a reference block and a determined scaling factor according to a linear equation includes: calculating the pixel value of the current block as a*p + b, where a is the determined scaling factor, b is a determined offset, and p is the reference pixel value at a reference point determined by the motion vector.

[0145] In the present disclosure, the direction of a reference frame may be determined according to whether the display order of the reference frame is before the current frame or the display order is after the current frame.

[0146] In various embodiments, a block is encoded in the BAWP (or LIC, or Composite Weighted Prediction (CWP)) mode, hereinafter referred to as BAWP for simplicity of description. A flag called bawp_type can be signaled to indicate whether the BAWP scaling factor uses explicit signaling or implicit signaling. When explicit signaling of the BAWP scaling factor is adopted for the current block, the scaling factor or its corresponding index in the lookup table can be signaled in the bitstream and parsed on the decoder side; the offset value β, as the only parameter to be derived, can be derived from a linear equation between the template of the reference block and the template of the current block, for example, by the least squares fitting method as a non-limiting example. In some embodiments, when implicit signaling of the BAWP scaling factor is adopted for the current block, the scaling factor α and the offset value β, as two parameters to be derived, can be derived from a linear equation between the template of the reference block and the template of the current block, for example, by the least squares fitting method as a non-limiting example.

[0147] In various embodiments, the context for signaling bawp_type can depend on at least one of the following: the encoded information of the current block and adjacent blocks; and / or the number of adjacent blocks using the BAWP / LIC / CWP mode. As a non-limiting example, when none of the adjacent blocks are using the BAWP / LIC / CWP mode, the first context is used; when one of the adjacent blocks is using the BAWP / LIC / CWP mode, the second context is used; and / or when more than one adjacent block is using the BAWP / LIC / CWP mode, the third context is used.

[0148] In some embodiments, the context for signaling bawp_type can depend on the reference frame index of the current block and / or adjacent blocks. As a non-limiting example, when the reference frame index of the current block and / or adjacent blocks is the same, the first context is used; and / or when the reference frame index of the current block and / or adjacent blocks is different, the second context is used.

[0149] In some embodiments, the context for signaling bawp_type can depend on whether the current block is encoded in single-reference mode or composite prediction mode. As a non-limiting example, when the current block is encoded in single-reference mode, the first context is used; and / or when the current block is encoded in composite prediction mode, the second context is used.

[0150] In some embodiments, when BAWP is applied to a composite prediction block, the context for signaling bawp_type can depend on the syntax value associated with applying the adaptive composite prediction weights.

[0151] In various embodiments, when a flag in the bitstream indicates that the current block is using explicit signaling, another flag, called bawp_sign, is signaled in the bitstream to indicate whether the scaling factor of the current block is greater than or less than a threshold (TH), which can be predefined.

[0152] In some embodiments, the value of TH is signaled in the high-level syntax (HLS), and as a non-limiting example, it is signaled in the sequence parameter set (SPS), picture parameter set (PPS), picture header, slice header, tile header, or CTU header. In some embodiments, TH for all video sequences is set to a fixed value, such as 0 or 1. In some embodiments, TH is set to a derived ratio α (or a quantized value of the derived ratio α), which is calculated based on the local illumination change between the template of the current block (or the causal samples of the current block) and the template of the reference block. In some embodiments, the context for signaling bawp_sign can depend on the bawp_sign values of adjacent blocks. In some embodiments, after signaling bawp_sign, another flag can be signaled in the bitstream to indicate the magnitude of the scaling factor or the adjustment of the scaling factor.

[0153] In various embodiments, multiple scaling factor lookup tables are included, and the selection among different lookup tables can be implicitly determined based on the scaling factor values derived from the current block template and the reference block template. In some embodiments, there can be a predefined scaling factor offset and a predefined factor, and thus the actual scaling factor used to reconstruct the current block is calculated by adding the predefined scaling factor offset and then dividing by the predefined factor, i.e., ((scaling factor determined from multiple scaling factor lookup tables)+(predefined scaling factor offset)) / (predefined factor). As a non-limiting example, when the predefined scaling factor offset is +1 and the predefined factor is 16, the actual scaling factor = ((scaling factor determined from multiple scaling factor lookup tables)+1) / 16.

[0154] In some embodiments, the number of scaling factors in the multiple lookup tables is the same. As a non-limiting example, the number of scaling factors in the lookup table can be 2, 3, 4, 5, 6, 7, 8, or 10.

[0155] In some embodiments, the step size and / or precision of the scaling factor (or the absolute value of the scaling factor) in each look-up table are the same, but different in different look-up tables. As a non-limiting example, one look-up table is {+2, -2, +3, -3, +4, -4}, while another look-up table is {+2, -2, +4, -4, +6, -6}. As another non-limiting example, one look-up table is {+1, -1, +2, -2, +3, -3}, while another look-up table is {+2, -2, +4, -4, +6, -6}.

[0156] In some embodiments, in one look-up table, the step size and / or precision of the scaling factor (or the absolute value of the scaling factor) are the same, while in the remaining look-up tables, the step sizes are different. As a non-limiting example, two look-up tables are used. The first look-up table has a fixed step size and includes the values {+2, -2, +3, -3, +4, -4}. The second look-up table has different step sizes and consists of the values {+2, -2, +4, -4, +8, -8}.

[0157] In some embodiments, two scaling factor look-up tables are supported, and the choice between the two look-up tables can depend on the proximity of the derived scaling factor to a predefined value TH. As a non-limiting example, when the derived scaling factor is close to the predefined value TH (e.g., the derived scaling factor is less than the predefined value TH), the look-up table with a smaller step size is selected. Otherwise, the other look-up table is selected. For example, the TH value is set to 1.

[0158] In various embodiments, multiple scaling factor look-up tables are supported, and the selection among different look-up tables can be implicitly determined based on the picture order count (POC) distance between the reference frame and the current frame. In some embodiments, two look-up tables are supported. When the POC distance between the reference frame and the current frame is within a predefined value TH (e.g., less than TH), the look-up table with a smaller step size is selected. Otherwise, the other look-up table is selected. For example, the TH value is set to 4.

[0159] In various embodiments, only one predefined look-up table is supported, and in this look-up table, the step size of the absolute value of the scaling factor increases as the index of the scaling factor increases. In some embodiments, the absolute value of the scaling factor in the look-up table can only be a power of 2, such as 1, 2, 4, 8, 16. As a non-limiting example, the scaling factors in the look-up table are {+1, -1, +3, -3, +6, -6, +8, -8}. As another non-limiting example, the scaling factors in the look-up table are {+1, -1, +3, -3, +6, -6}.

[0160] Each embodiment in the present disclosure may include methods for encoding a current block into a video bitstream, which are performed by an encoder and include any part or all of the inverse processes of the processes described for a decoder.

[0161] As needed, the above operations may be combined or arranged in any quantity or order. Two or more of the steps and / or operations may be performed in parallel. The embodiments and implementations in the present disclosure may be used alone or in any combination. In addition, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored in a non-volatile computer-readable medium. The embodiments in the present disclosure may be applied to a luminance block or a chrominance block. The term block may be interpreted as a prediction block, an encoding block, or a coding unit (i.e., CU). The term block here may also be used to refer to a transform block. In the following terms, when referring to the block size, it may refer to the block width or block height, or the maximum value of the block width and height, or the minimum value of the block width and height, or the area size of the block (width * height), or the aspect ratio of the block (width:height or height:width).

[0162] The techniques described above may be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 14 FIG. shows a computer system (1800) suitable for implementing certain embodiments of the disclosed subject matter.

[0163] The computer software may be encoded using any suitable machine code or computer language, which may be subject to assembly, compilation, linking, or similar mechanisms to create code including instructions that may be executed directly or through interpretation, microcode execution, etc. by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.

[0164] The instructions may be executed on various types of computers or computer components, including, for example, personal computers, tablets, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0165] Figure 14 The components shown in FIG. for the computer system (1800) are exemplary in nature and are not intended to imply any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present application. Nor should the configuration of the components be construed as having any dependence or requirement on any one component or combination of components shown in the exemplary embodiments of the computer system (1800).

[0166] The computer system (1800) may include certain human-machine interface input devices. The input human-machine interface devices may include one or more of the following (only one of each is depicted): keyboard (1801), mouse (1802), trackpad (1803), touchscreen (1810), data glove (not shown), joystick (1805), microphone (1806), scanner (1807), camera (1808).

[0167] The computer system (1800) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate one or more human users' senses through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., the tactile feedback of the touchscreen (1810), data glove (not shown), or joystick (1805), but there may also be tactile feedback devices that do not act as input devices), audio output devices (e.g., speakers (1809), headphones (not depicted)), visual output devices (e.g., screens (1810), including cathode ray tube (CRT) screens, liquid crystal display (LCD) screens, plasma screens, organic light-emitting diode (OLED) screens, each with or without touchscreen input capabilities, each with or without tactile feedback capabilities - some of which are capable of outputting two-dimensional visual output or output greater than three-dimensional through, for example, stereoscopic flat painting output; virtual reality glasses (not depicted), holographic displays, and smoke boxes (not depicted), as well as printers (not depicted).

[0168] The computer system (1800) may also include human-accessible storage devices and the associated media of the storage devices, such as optical media, including CD / DVD ROM / RW (1820) with media such as CD / DVD (1821), thumb drives (1822), removable hard disk drives or solid-state drives (1823), legacy magnetic media such as tapes and floppy disks (not depicted), ROM / special application-specific integrated circuit (ASIC) / programmable logic device (PLD)-based special devices, such as security protection devices (not depicted), and so on.

[0169] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not cover transmission media, carrier waves, or other transient signals.

[0170] The computer system (1800) may also include an interface (1854) to one or more communication networks (1855). The network may be, for example, wireless, wired, optical. The network may also be local, wide area, metropolitan area, vehicular and industrial, real-time, delay-tolerant, and so on. Examples of networks include, for example, local area networks such as Ethernet, wireless LAN, cellular networks including Global System for Mobile Communications (GSM), Third Generation (3G), Fourth Generation (4G), Fifth Generation (5G), Long Term Evolution (LTE), etc., TV wired or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, vehicular networks and industrial networks including Controller Area Network Bus (CANBus), etc.

[0171] The above-mentioned human-machine interface device, human-accessible storage device, and network interface may be attached to the core (1840) of the computer system (1800).

[0172] The core (1840) may include one or more central processing units (CPUs) (1841), a graphics processing unit (GPU) (1842), a dedicated programmable processing unit in the form of a Field Programmable Gate Area (FPGA) (1843), a hardware accelerator (1844) for certain tasks, a graphics adapter (1850), and so on. These devices, together with a read-only memory (ROM) (1845), a random access memory (1846), and internal mass storage devices such as internal non-user-accessible hard disk drives, solid state drives (SSDs), etc. (1847), may be connected via a system bus (1848). In some computer systems, the system bus (1848) may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be attached directly or via a peripheral bus (1849) to the system bus (1848) of the core. In one embodiment, the screen (1810) may be connected to the graphics adapter (1850). Architectures for the peripheral bus include Peripheral Component Interconnect (PCI), USB, and so on.

[0173] Computer code for performing various computer-implemented operations may be present on a computer-readable medium. The medium and the computer code may be those designed and constructed specifically for the purposes of this application, or may be of the kind well-known and available to those skilled in the field of computer software.

[0174] Although this application describes several exemplary embodiments, within the scope of this application, there can be various modifications, permutations, and various alternative equivalents. Therefore, it should be understood that within the spirit and scope of the application, those skilled in the art can design various systems and methods that, although not explicitly shown or described herein, can embody the principles of this application.

Claims

1. A method for decoding a current block of a current frame in an encoded video bitstream, characterized in that, The method includes: A decoding device receives an encoded video bitstream, the encoded video bitstream including a current block of a current frame and a first syntax element that indicates a prediction mode of the current block. The decoding device includes a memory storing instructions and a processor communicating with the memory. The memory further stores a plurality of scaling factor lookup tables, where the plurality of scaling factor lookup tables include scaling factors in different ranges, and in the plurality of scaling factor lookup tables, the step size of the scaling factors in each lookup table is the same, or the precision of the scaling factors is the same; The decoding device determines the prediction mode based on the value of the first syntax element, where the prediction mode is used to predict the current block based on a reference block of a reference frame; The decoding device determines a scaling factor from one of the plurality of scaling factor lookup tables; The decoding device reconstructs the current block based on the reference block and the determined scaling factor according to a linear equation.

2. The method according to claim 1, wherein: The plurality of scaling factor lookup tables include the same number of scaling factors.

3. The method according to claim 1, wherein The determining a scaling factor from one of the plurality of scaling factor lookup tables includes: Explicitly determining the scaling factor based on a second syntax element, or implicitly determining the scaling factor based on a local illumination value.

4. The method according to claim 1, wherein The determining a scaling factor from one of the plurality of scaling factor lookup tables includes: Based on a derived scaling factor, selecting one scaling factor lookup table from the plurality of scaling factor lookup tables; and Determining the scaling factor from the selected scaling factor lookup table.

5. The method according to claim 4, wherein: The plurality of scaling factor lookup tables include a first lookup table and a second lookup table; The selecting one scaling factor lookup table from the plurality of scaling factor lookup tables based on a derived scaling factor includes: Determining whether the derived scaling factor is greater than a predefined threshold, When it is determined that the derived scaling factor is less than the predefined threshold, selecting the first lookup table, When it is determined that the derived scaling factor is not less than the predefined threshold, selecting the second lookup table.

6. The method according to claim 1, wherein: In the plurality of scaling factor lookup tables, the step size of the scaling factors in different scaling factor lookup tables is different, or the precision of the scaling factors in different scaling factor lookup tables is different.

7. The method according to claim 1, characterized in that Determining, based on the encoded video bitstream, a prediction mode for predicting the current block based on a reference block of a reference frame includes: Extracting a flag from the encoded video bitstream, where the flag indicates a prediction mode for predicting the current block based on a reference block of a reference frame.

8. The method according to claim 1, wherein The method further includes: The decoding device derives a derived scaling factor based on a current template of the current block and a reference template of the reference block.

9. The method according to claim 1, characterized in that, The method further includes: The decoding device determines a reference block of the reference frame based on a motion vector.

10. The method according to claim 1, wherein The reconstructing the current block based on the reference block and the determined scaling factor according to a linear equation includes: Calculate the pixel value of the current block as a*p + b, where a is a determined scaling factor, b is a determined offset, and p is the reference pixel value at the reference point determined according to the motion vector.

11. An apparatus for decoding a current block of a current frame in an encoded video bitstream, characterized in that, The apparatus comprises: a memory storing instructions; and a processor in communication with the memory, wherein when the processor executes the instructions, the processor is configured to cause the apparatus to perform the method according to any one of claims 1 to 10.

12. A non-volatile computer-readable storage medium for storing instructions, characterized in that, When the instructions are executed by the processor, the instructions are configured to cause the processor to perform the method according to any one of claims 1 to 10.