Method and apparatus for enhanced block adaptive weighted prediction

Through the enhanced block adaptive weighted prediction (BAWP) method, the reference block and scaling factor in video encoding technology are used to generate prediction blocks, which solves the impact of local lighting changes on encoding efficiency and quality, and achieves more efficient video encoding.

CN120380756APending Publication Date: 2025-07-25TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380083995.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-27
Filing Date
2023-11-28
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

Existing video encoding technology is difficult to effectively compensate for local lighting changes, resulting in limited encoding efficiency and quality.

Method used

The enhanced block adaptive weighted prediction (BAWP) method is used to receive the encoded video bit stream, identify the motion vector of the reference block, obtain the scaling factor and offset value, and generate the prediction block using a linear formula to reconstruct the current block.

Benefits of technology

The encoding efficiency and quality of video encoding are improved, especially in the case of severe local lighting changes, and the encoding robustness and accuracy are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120380756A_ABST
    Figure CN120380756A_ABST
Patent Text Reader

Abstract

The present disclosure generally relates to video encoding / decoding, and particularly to video encoding / decoding for enhanced block adaptive weighted prediction (BAWP). A method includes receiving an encoded video bitstream; identifying, from the encoded video bitstream, a motion vector corresponding to a reference block associated with a current block of the current frame; obtaining a scaling factor based on a syntax explicitly signed in the encoded video bitstream; determining a template for deriving an offset value; deriving an offset value based on the template; generating a prediction block based on the reference block according to a linear formula, wherein the linear formula is associated with a scaling factor and an offset value; and reconstructing, by the device, the current block based on the prediction block.
Need to check novelty before this filing date? Find Prior Art

Description

Cross - reference

[0001] This application claims the benefit of priority to U.S. Provisional Application No. 63 / 532,633, filed on August 14, 2023, which is hereby incorporated by reference in its entirety. This application also claims the benefit of priority to U.S. Non - Provisional Patent Application No. 18 / 519,859, filed on November 27, 2023, which is hereby incorporated by reference in its entirety. Technical Field

[0002] The present disclosure describes a set of advanced video / stream encoding / decoding techniques. More specifically, the disclosed techniques relate to enhancing Block Adaptive Weighted Prediction (BAWP) to compensate for local illumination variations. Background Art

[0003] Uncompressed digital video can include a sequence of pictures and can have specific bit - rate requirements for storage, data processing, and transmission bandwidth in streaming applications. One purpose of video encoding and video decoding can be to reduce redundancy in the uncompressed input video signal through various compression techniques. Summary of the Invention

[0004] The present disclosure describes various embodiments of methods, apparatuses, and computer - readable storage media for enhancing Block Adaptive Weighted Prediction (BAWP).

[0005] According to one aspect, embodiments of the present disclosure provide a method for decoding a current block of a current frame in an encoded video bitstream. The method includes receiving, by a device, the encoded video bitstream. The device includes a memory storing instructions and a processor communicatively coupled to the memory. The method further includes: identifying, by the device, a motion vector corresponding to a reference block associated with the current block of the current frame from the encoded video bitstream; obtaining, by the device, a scaling factor based on syntax explicitly signaled in the encoded video bitstream; determining, by the device, a template for deriving an offset value; deriving, by the device, the offset value based on the template; generating, by the device, a prediction block based on the reference block according to a linear formula associated with the scaling factor and the offset value; and reconstructing, by the device, the current block based on the prediction block.

[0006] According to another aspect, embodiments of the present disclosure provide an apparatus for processing a current block of a current frame in an encoded video bitstream. The apparatus includes: a memory storing instructions; and a processor communicatively coupled to the memory. When the processor executes the instructions, the processor is configured to cause the apparatus to perform the above - described method for video decoding and / or encoding.

[0007] In another aspect, embodiments of the present disclosure provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding and / or encoding, cause the computer to perform the above-described method for video decoding and / or encoding.

[0008] The above aspects and other aspects and their implementations are described in more detail in the accompanying drawings, the description, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Additional features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:

[0010] Figure 1 A schematic illustration showing a simplified block diagram of a communication system (100) according to an example embodiment;

[0011] Figure 2 A schematic illustration showing a simplified block diagram of a communication system (200) according to an example embodiment;

[0012] Figure 3 A schematic illustration showing a simplified block diagram of a video decoder according to an example embodiment;

[0013] Figure 4 A schematic illustration showing a simplified block diagram of a video encoder according to an example embodiment;

[0014] Figure 5 A block diagram showing a video encoder according to another example embodiment;

[0015] Figure 6 A block diagram showing a video decoder according to another example embodiment;

[0016] Figure 7 A scheme of coding block partitioning according to an example embodiment of the present disclosure is shown;

[0017] Figure 8 Another scheme of coding block partitioning according to an example embodiment of the present disclosure is shown;

[0018] Figure 9 Another scheme of coding block partitioning according to an example embodiment of the present disclosure is shown;

[0019] Figure 10 An example of a template of a current block and a reference block is shown;

[0020] Figure 11 An example of a reference line of a coding block is shown;

[0021] Figure 12Illustrates an example logical flow of the method in the present disclosure;

[0022] Figure 13 Illustrates an example of the upper regions of the current block and the reference block;

[0023] Figure 14 Illustrates an example of the left regions of the current block and the reference block;

[0024] Figure 15 Illustrates an example of the subsampling of the current block and / or the reference block; and

[0025] Figure 16 Illustrates a schematic diagram of a computer system according to an example embodiment of the present disclosure. Detailed Description

[0026] The present invention will now be described in detail below with reference to the accompanying drawings, which form a part of the present invention and illustrate specific examples of embodiments by way of illustration. However, note that the present invention can be implemented in various different forms, and thus, the subject matter covered or claimed is intended to be construed as not limited to any of the embodiments set forth below. Also note that the present invention can be implemented as a method, apparatus, component, or system. Therefore, embodiments of the present invention can take, for example, the form of hardware, software, firmware, or any combination thereof.

[0027] Throughout the specification and claims, terms may have nuanced meanings that are contextually presented or implied beyond the explicitly stated meanings. As used herein, the phrase "in one embodiment" or "in some embodiments" does not necessarily refer to the same embodiment, and the phrase "in another embodiment" or "in other embodiments" does not necessarily refer to different embodiments. Similarly, as used herein, the phrase "in one implementation" or "in some implementations" does not necessarily refer to the same implementation, and the phrase "in another implementation" or "in other implementations" does not necessarily refer to different implementations. For example, it is meant that the claimed subject matter includes combinations of all or part of the exemplary embodiments / implementations.

[0028] Typically, terms can be understood, at least in part, based on their usage in context. For example, terms such as "and", "or", or "and / or" as used herein can have a variety of meanings, which can depend, at least in part, on the context in which such terms are used. Generally, if "or" is used to associate a list, such as A, B, or C, then "or" is intended to mean: A, B, and C, used in an inclusive sense here; and A, B, or C, used in an exclusive sense here. Additionally, depending, at least in part, on the context, the terms "one or more" or "at least one" as used herein can be used to describe any feature, structure, or property in a singular sense, or can be used to describe a combination of features, structures, or properties in a plural sense. Similarly, terms such as "a", "an", or "the" can also be understood to convey a singular usage or to convey a plural usage, at least in part, depending on the context. Further, the terms "based on" or "determined by" can be understood to not necessarily convey an exclusive set of factors, and can alternatively allow for additional factors that are not necessarily explicitly described, again at least in part, depending on the context.

[0029] As Figure 1 shown, the terminal device can be implemented as a server, a personal computer, and a smart phone, but the applicability of the basic principles of the present disclosure is not limited thereto. Embodiments of the present disclosure can be implemented in a desktop computer, a laptop computer, a tablet computer, a media player, a wearable computer, dedicated video conferencing equipment, etc. The network (150) represents any number or type of network that conveys encoded video data between terminal devices, including, for example, wired (wired) and / or wireless communication networks. The communication network (150) can exchange data in circuit-switched channels, packet-switched channels, and / or other types of channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0030] As an example of an application of the disclosed subject matter, Figure 2 illustrates the placement of a video encoder and a video decoder in a video streaming environment. The disclosed subject matter can be equally applicable to other video applications, including, for example, video conferencing, digital television broadcasting, gaming, virtual reality, storage of compressed video on digital media including CD (Compact Disc), DVD (Digital Versatile Disc), memory sticks, etc.

[0031] As Figure 2As shown, a video streaming system may include a video capture subsystem (213), which may include a video source (201) such as a digital camera device for creating an uncompressed video picture or image stream (202). In an example, the video picture stream (202) includes samples recorded by the digital camera device of the video source (201). The video picture stream (202) is depicted as a thick line to emphasize the high data volume compared to the encoded video data (204) (or encoded video bitstream), and the video picture stream (202) may be processed by an electronic device (220) coupled to the video source (201) and including a video encoder (203). The video encoder (203) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded video data (204) (or encoded video bitstream (204)) is depicted as a thin line to emphasize the lower data volume compared to the uncompressed video picture stream (202), and the encoded video data (204) may be stored on a streaming server (205) for future use or directly stored to a downstream video device (not shown). One or more streaming client subsystems, such as Figure 2 the client subsystems (206) and (208) in

[0032] Figure 3 shows a block diagram of a video decoder (310) of an electronic device (330) according to any embodiment of the present disclosure below. The electronic device (330) may include a receiver (331) (e.g., receiving circuitry). The video decoder (310) may be used instead of Figure 2 the video decoder (210) in the example of

[0033] As Figure 3As shown, a receiver (331) may receive one or more encoded video sequences from a channel (301). To prevent network jitter and / or handle playback timing, a buffer memory (315) may be provided between the receiver (331) and an entropy decoder / parser (320) (hereinafter referred to as "parser (320)"). The parser (320) may reconstruct symbols (321) from the encoded video sequences. The categories of these symbols include information for managing the operation of a video decoder (310), and possibly information for controlling a rendering device such as a display (312) (e.g., a display screen). The parser (320) may parse / entropy decode the encoded video sequences. The parser (320) may extract from the encoded video sequences a set of subgroup parameters for at least one of a subgroup of pixels in the video decoder. The subgroup may include a Group of Pictures (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (320) may also extract information such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc. from the encoded video sequences. The reconstruction of the symbols (321) may involve multiple different processing or functional units. The units involved and how these units are involved may be controlled by subgroup control information parsed by the parser (320) from the encoded video sequences.

[0034] A first unit may include a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) may receive from the parser (320) quantized transform coefficients as symbols (321) and control information, including information indicating which inverse transform to use, block size, quantization factor / parameter, quantization scaling matrix, etc. The scaler / inverse transform unit (351) may output a block including sample values that may be input into an aggregator (355).

[0035] In some cases, the output samples of the scaler / inverse transform (351) can belong to an intra-coded block, i.e., a block that does not use predictive information from a previously reconstructed picture but can use predictive information from a previously reconstructed portion of the current picture. Such predictive information can be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) can use surrounding block information that has been reconstructed and stored in the current picture buffer (358) to generate a block having the same size and shape as the block being reconstructed. For example, the current picture buffer (358) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some implementations, the aggregator (355) can add the predictive information that the intra-prediction unit (352) has generated to the output sample information provided by the scaler / inverse transform unit (351) on a per-sample basis.

[0036] In other cases, the output samples of the scaler / inverse transform unit (351) can belong to an inter-coded and possibly motion-compensated block. In such a case, the motion compensation prediction unit (353) can access the reference picture memory (357) based on a motion vector to obtain samples for inter-picture prediction. After motion compensating the obtained reference samples according to the sign (321) belonging to the block, these samples can be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) (the output of unit 351 can be referred to as residual samples or a residual signal) to generate output sample information.

[0037] The output samples of the aggregator (355) can undergo various loop filtering techniques in a loop filter unit (356) that includes several types of loop filters. The output of the loop filter unit (356) can be a sample stream that can be output to a rendering device (312) and stored in the reference picture memory (357) for future inter-picture prediction.

[0038] Figure 4 A block diagram of a video encoder (403) according to an example embodiment of the present disclosure is shown. The video encoder (403) can be included in an electronic device (420). The electronic device (420) can also include a transmitter (440) (e.g., transmission circuitry). The video encoder (403) can be used instead of Figure 4 the video encoder (403) in the example of

[0039] The video encoder (403) may receive video samples from a video source (401). According to some example embodiments, the video encoder (403) may encode and compress pictures of the source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed constitutes a function of the encoding speed constitution controller (450). In some embodiments, the controller (450) may be functionally coupled to other functional units as described below and control the other functional units. The parameters set by the controller (450) may include rate control related parameters (picture skipping, quantizer, λ value of rate distortion optimization technique...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc.

[0040] In some example embodiments, the video encoder (403) may be configured to operate in an encoding loop. The encoding loop may include a source encoder (430) and an (in - loop) decoder (433) embedded in the video encoder (403). Even though the in - loop decoder 433 processes the encoded video stream of the source encoder 430 without entropy encoding, the decoder (433) reconstructs symbols in a manner similar to how a (remote) decoder would create sample data to create sample data (because in the video compression techniques contemplated in the disclosed subject matter, any compression between symbols and the encoded video bitstream in entropy encoding can be lossless). At this point, it can be observed that any decoder techniques that may exist only in the decoder other than parsing / entropy decoding may also necessarily exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter may sometimes focus on decoder operations related to the decoding part of the encoder. Thus, the description of encoder techniques can be simplified because the encoder techniques are the opposite of the fully described decoder techniques. A more detailed description of the encoder is provided only in certain areas or aspects below.

[0041] In some example implementations, during operation, the source encoder (430) may perform motion - compensated predictive encoding that predictively encodes an input picture by referring to one or more previously encoded pictures of the video sequence designated as "reference pictures".

[0042] The in - loop video decoder (433) may decode the encoded video data of pictures that may be designated as reference pictures. The in - loop video decoder (433) replicates the decoding process that may be performed by a video decoder on the reference pictures and may cause the reconstructed reference pictures to be stored in the reference picture buffer (434). In this way, the video encoder (403) may locally store a copy of the reconstructed reference pictures, which has the same content (in the absence of transmission errors) as the reconstructed reference pictures that would be obtained by a distal (remote) video decoder.

[0043] The predictor (435) can perform a prediction search for the coding engine (432). That is, for a new picture to be coded, the predictor (435) can search in the reference picture memory (434) for sample data (as candidate reference pixel blocks) that can be used as an appropriate prediction reference for the new picture or certain metadata such as reference picture motion vectors, block shapes, etc.

[0044] The controller (450) can manage the coding operations of the source encoder (430), including, for example, the setting of parameters and subgroup parameters for coding video data.

[0045] The outputs of all the above-mentioned functional units can undergo entropy coding in the entropy encoder (445). The transmitter (440) can buffer the coded video sequence created by the entropy encoder (445) to prepare for transmission via a communication channel (460), which can be a hardware / software link to a storage device storing the coded video data. The transmitter (440) can merge the coded video data from the video encoder (403) with other data to be transmitted, such as coded audio data and / or an auxiliary data stream (source not shown).

[0046] The controller (450) can manage the operations of the video encoder (403). During coding, the controller (450) can assign a certain coding picture type to each coded picture, which may affect the coding techniques that can be applied to the corresponding picture. For example, pictures can generally be assigned to one of the following picture types: intra picture (I picture), predictive picture (P picture), bi-predictive picture (B picture), multi-predictive picture. Source pictures can generally be spatially subdivided into multiple sample coding blocks, as described in further detail below.

[0047] Figure 5 A diagram of a video encoder (503) according to another example embodiment of the present disclosure is shown. The video encoder (503) is configured to receive the sample values of a processing block (e.g., a prediction block) within a current video picture in a sequence of video pictures and encode the processing block into a coded picture that is part of a coded video sequence. The example video encoder (503) can be used instead of Figure 4 the video encoder (403) in the example.

[0048] For example, the video encoder (503) receives a matrix of sample values of the processing block. The video encoder (503) then uses, for example, Rate-Distortion Optimization (RDO) to determine whether to best encode the processing block using an intra mode, an inter mode, or a bi-predictive mode.

[0049] In Figure 5 the example of, the video encoder (503) includes an inter-frame encoder (530), an intra-frame encoder (522), a residual calculator (523), a switch (526), a residual encoder (524), a general controller (521), and an entropy encoder (525) coupled together as shown in the example arrangement of Figure 5 .

[0050] The inter-frame encoder (530) is configured to: receive samples of a current block (e.g., a processing block); compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture in display order); generate inter-frame prediction information (e.g., a description of redundant information according to inter-frame coding techniques, a motion vector, merge mode information); and calculate an inter-frame prediction result (e.g., a predicted block) based on the inter-frame prediction information using any suitable technique.

[0051] The intra-frame encoder (522) is configured to: receive samples of a current block (e.g., a processing block); compare the block with blocks that have been encoded in the same picture; and generate quantized coefficients after transformation; and in some cases also generate intra-frame prediction information (e.g., intra-frame prediction direction information according to one or more intra-frame coding techniques).

[0052] The general controller (521) may be configured to determine general control data and control other components of the video encoder (503) based on the general control data to, for example, determine a prediction mode of a block and provide a control signal to the switch (526) based on the prediction mode.

[0053] The residual calculator (523) may be configured to calculate the difference (residual data) between the received block and the prediction result of a block selected from the intra-frame encoder (522) or the inter-frame encoder (530). The residual encoder (524) may be configured to encode the residual data to generate transform coefficients. Then, the transform coefficients are quantized to obtain quantized transform coefficients. In various example embodiments, the video encoder (503) further includes a residual decoder (528). The residual decoder (528) is configured to perform an inverse transform and generate decoded residual data. The entropy encoder (525) may be configured to format a bitstream to include the encoded block and perform entropy coding.

[0054] Figure 6 A diagram showing an example video decoder (610) according to another embodiment of the present disclosure is shown. The video decoder (610) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In the example, the video decoder (610) may be used instead of Figure 4The video decoder (410) in the example of

[0055] In Figure 6 the example of, the video decoder (610) includes an entropy decoder (671), an inter-frame decoder (680), a residual decoder (673), a reconstruction module (674), and an intra-frame decoder (672) coupled together as shown in the example arrangement of Figure 6

[0056] The entropy decoder (671) may be configured to reconstruct certain symbols representing the syntax elements that make up the coded picture based on the coded picture. The inter-frame decoder (680) may be configured to receive inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information. The intra-frame decoder (672) may be configured to receive intra-frame prediction information and generate a prediction result based on the intra-frame prediction information. The residual decoder (673) may be configured to perform inverse quantization to extract dequantized transform coefficients and process the dequantized transform coefficients to transform the residual from the frequency domain to the spatial domain. The reconstruction module (674) may be configured to combine the residual output by the residual decoder (673) with the prediction result (output by the inter-frame prediction module or the intra-frame prediction module, as appropriate) in the spatial domain to form a reconstructed block, and the reconstructed block forms a part of the reconstructed picture that is part of the reconstructed video.

[0057] Note that any suitable technology may be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610). In some example embodiments, one or more integrated circuits may be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610). In another embodiment, one or more processors executing software instructions may be used to implement the video encoders (203), (403), and (503) and the video decoders (210), (310), and (610).

[0058] ​Turning to block partitioning for encoding and decoding, the general partitioning can start from a basic block and can follow a predefined set of rules, a specific pattern, a partitioning tree, or any partitioning structure or scheme. The partitioning can be hierarchical and recursive. After splitting or partitioning the basic block following any of the example partitioning processes described below or other processes or a combination thereof, a final set of partitions or coding blocks can be obtained. Each of these partitions can be at one of the various partitioning levels in the partitioning hierarchy and can have various shapes. Each of the partitions can be referred to as a coding block (CB). For the various example partitioning implementations described further below, each resulting CB can have any allowed size and partitioning level. Such partitions are called coding blocks because they can form units for which some basic encoding / decoding decisions can be made and encoding / decoding parameters can be optimized, determined, and signaled in the encoded video bitstream. The highest or deepest level in the final partition represents the depth of the coding block partitioning structure of the tree. The coding block can be a luminance coding block or a chrominance coding block. The CB tree structure for each color can be referred to as a coding block tree (CBT). The coding blocks for all color channels can be collectively referred to as coding units (CUs). The hierarchical structures for all color channels can be collectively referred to as coding tree units (CTUs). The partitioning patterns or structures for the various color channels in a CTU can be the same or can be different.

[0059] In some implementations, the partitioning tree scheme or structure for the luminance channel and the chrominance channel may not need to be the same. In other words, the luminance channel and the chrominance channel can have separate coding tree structures or patterns. Additionally, whether the luminance channel and the chrominance channel use the same or different coding partitioning tree structures and the actual coding partitioning tree structure to be used can depend on whether the slice being encoded is a P slice, a B slice, or an I slice. For example, for an I slice, the chrominance channel and the luminance channel can have separate coding partitioning tree structures or coding partitioning tree structure patterns, while for a P slice or a B slice, the luminance channel and the chrominance channel can share the same coding partitioning tree scheme. When applying separate coding partitioning tree structures or patterns, the luminance channel can be partitioned into CBs by one coding partitioning tree structure, and the chrominance channel can be partitioned into chrominance CBs by another coding partitioning tree structure.

[0060] Figure 7 An example predefined 10-way partitioning structure / pattern that allows recursive partitioning to form a partitioning tree is shown. The root block can start at a predefined level (e.g., starting from a basic block at the 128×128 level or the 64×64 level). Figure 7The example partitioning structures include various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. In some example implementations, Figure 7 none of the rectangular partitions are allowed to be further subdivided. The coding tree depth can be further limited to indicate the splitting depth from the root node or root block. For example, the coding tree depth of the root node or root block can be set to 0, and after the root block is further split once according to Figure 7 , the coding tree depth increases by 1. In some implementations, only the all-square partitions in 710 are allowed to be recursively partitioned to the next level of the partition tree according to Figure 7 the pattern.

[0061] In some other example implementations for coding block partitioning, a quadtree structure can be used. Such quadtree splitting can be applied hierarchically and recursively to any square-shaped partition. Whether the basic block or intermediate block or partition is further quadtree split can be adapted to various local characteristics of the basic block or intermediate block / partition.

[0062] In still some other examples, a ternary partitioning scheme can be used to partition the basic block or any intermediate block, as Figure 8 shown. The ternary pattern can be implemented vertically as shown in 802, or horizontally as shown in 804. Although Figure 8 the example split ratio is shown as 1:2:1, other ratios can also be predefined. In some implementations, two or more different ratios can be predefined. In some implementations, the width and height of the partitions of the example ternary tree are always powers of 2 to avoid additional transformations.

[0063] The above partitioning schemes can be combined in any way at different partitioning levels. As an example, the above quadtree partitioning scheme and binary partitioning scheme can be combined to partition the basic block into a Quadtree-Binary-Tree (QTBT) structure. In such a scheme, according to a set of predefined conditions (if specified), the basic block or intermediate block / partition can be quadtree split or binary split. In Figure 9Specific examples are shown where a basic block is first split into four partitions by a quadtree as shown at 902, 904, 906, and 908. Thereafter, each of the resulting partitions is divided into four additional partitions (e.g., 908) by a quadtree at the next level, or is split into two additional partitions by a binary split (e.g., horizontally or vertically, e.g., 902 or 906, both of which are symmetric), or is not split (e.g., 904). For square-shaped partitions, a binary split or a quadtree split can be recursively allowed, as shown by the overall example partitioning pattern of 910 and the corresponding tree structure / representation in 920, where solid lines represent quadtree splits and dashed lines represent binary splits. A flag can be used for each binary split node (non-leaf binary partition) to indicate whether the binary split is horizontal or vertical. For example, as shown in 920 and consistent with the partitioning structure of 910, the flag "0" can represent a horizontal binary split, and the flag "1" can represent a vertical binary split. For quadtree split partitioning, there is no need to indicate the split type because a quadtree split always splits a block or partition horizontally and vertically to produce 4 sub-blocks / partitions of equal size. In some implementations, the flag "1" can represent a horizontal binary split, and the flag "0" can represent a vertical binary split.

[0064] In some example implementations of QTBT, the quadtree and binary split rule sets can be represented by the following predefined parameters and their associated corresponding functions: CTU size: The size of the root node of the quadtree (the size of the basic block) MinQTSize: The minimum allowed quadtree leaf node size MaxBTSize: The maximum allowed binary tree root node size MaxBTDepth: The maximum allowed binary tree depth MinBTSize: The minimum allowed binary tree leaf node size

[0065] In some example implementations of the QTBT partitioning structure, the CTU size can be set to 128×128 luma samples with two corresponding 64×64 chroma sample blocks (when considering and using example chroma subsampling), the MinQTSize can be set to 16×16, the MaxBTSize can be set to 64×64, the MinBTSize (for both width and height) can be set to 4×4, and the MaxBTDepth can be set to 4. Quadtree partitioning can first be applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes can have sizes ranging from their minimum allowed size of 16×16 (i.e., MinQTSize) to 128×128 (i.e., the CTU size). If a node is 128×128, then since the size exceeds MaxBTSize (i.e., 64×64), the node will not be first split by the binary tree. Otherwise, nodes not exceeding MaxBTSize can be partitioned by the binary tree. In Figure 9 the example of Figure 9 , the basic block is 128×128. According to the predefined rule set, the basic block can only be split by the quadtree. The basic block has a partitioning depth of 0. Each of the four resulting partitions is 64×64 - not exceeding MaxBTSize, and can be further split by the quadtree or binary tree at level 1. This process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), further splitting can be disregarded. When the width of a binary tree node equals MinBTSize (i.e., 4), further horizontal splitting can be disregarded. Similarly, when the height of a binary tree node equals MinBTSize, further vertical splitting is not considered.

[0066] In some example implementations, the above QTBT scheme can be configured to support the flexibility of having the same QTBT structure for luma and chroma or separate QTBT structures. For example, for P slices and B slices, the luma CTB and chroma CTB in a CTU can share the same QTBT structure. However, for I slices, the luma CTB can be partitioned into CUs by the QTBT structure, while the chroma CTB can be partitioned into chroma CUs by another QTBT structure. This means that a CU can be used to refer to different color channels in an I slice. For example, an I slice can consist of coded blocks of the luma component or coded blocks of the two chroma components, and a CU in a P slice or B slice can consist of coded blocks of all three color components.

[0067] The above various CU partitioning schemes and further partitioning of CUs to PBs can be combined in any way. The following specific implementations are provided as non - limiting examples.

[0068] Inter-frame prediction can be implemented, for example, in a single-reference mode or a composite-reference mode. In some implementations, a skip flag may first be included in the bitstream of the current block (or at a higher level) to indicate whether the current block is inter-frame encoded and not skipped. If the current block is inter-frame encoded, then another flag may be further included in the bitstream as a signal indicating whether the single-reference mode or the composite-reference mode is used for the prediction of the current block. For the single-reference mode, one reference block may be used to generate a prediction block for the current block. For the composite-reference mode, two or more reference blocks may be used to generate a prediction block, for example, by weighted averaging. One or more reference frame indices and additionally one or more motion vectors indicating the shift between the reference block and the current block in terms of the position relative to the frame (e.g., in horizontal pixels and vertical pixels) may be used to identify one or more reference blocks. For example, an inter-frame prediction block of the current block may be generated as a prediction block in the single-reference mode based on a single reference block identified by a motion vector in a reference frame, while for the composite-reference mode, a prediction block may be generated by weighted averaging of two reference blocks indicated by two reference frame indices and two corresponding motion vectors in two reference frames. The motion vectors may be encoded and included in the bitstream in various ways.

[0069] In some example implementations, one or more reference picture lists containing the identities of short-term reference frames and long-term reference frames for inter-frame prediction may be formed based on information in a Reference Picture Set (RPS). For example, a single picture reference list may be formed for uni-directional inter-frame prediction, and the single picture reference list is denoted as L0 reference (or reference list 0), while for bi-directional inter-frame prediction, two picture reference lists may be formed, and the two picture reference lists are denoted as L0 (or reference list 0) and L1 (or reference list 1) for each of the two prediction directions. The reference frames included in the L0 list and the L1 list may be sorted in various predetermined ways. The lengths of the L0 list and the L1 list may be signaled in the video bitstream. Uni-directional inter-frame prediction may be performed in the single-reference mode or in the composite-reference mode when multiple references used for generating a prediction block by weighted averaging are on the same side of the frame where the block to be predicted is located. Bi-directional inter-frame prediction may be performed only in the composite mode because bi-directional inter-frame prediction involves at least two reference blocks.

[0070] In some implementations, a Merge Mode (MM) for inter-frame prediction can be implemented. Generally, for the merge mode, one or more motion vectors among the motion vectors in the single-reference prediction of the current PB or the motion vectors in the composite-reference prediction can be derived from other motion vectors instead of being independently calculated and signaled. For example, in an encoding system, the current motion vector of the current PB can be represented by the difference between the current motion vector and one or more other already-encoded motion vectors (referred to as reference motion vectors). Such a difference of the motion vectors rather than the whole of the current motion vector can be encoded and included in the bitstream and can be linked to the reference motion vector. Correspondingly, in a decoding system, the motion vector corresponding to the current PB can be derived based on the decoded motion vector difference and the decoded reference motion vector linked thereto. As a specific form of the general Merge Mode (MM) inter-frame prediction, such an inter-frame prediction based on the motion vector difference can be referred to as a Merge Mode with Motion Vector Difference (MMVD). Thus, the general MM or the specific MMVD can be implemented to utilize the correlation between the motion vectors associated with different PBs to improve the encoding efficiency. For example, adjacent PBs may have similar motion vectors, so the MVD can be small and can be efficiently encoded. For another example, for blocks that are similarly positioned / located in space, the motion vectors can be temporally (between frames) correlated.

[0071] In some example implementations of MMVD, a list of reference motion vectors (RMVs) or MV predictor candidates for motion vector prediction can be formed for the block being predicted. The list of RMV candidates can include a predetermined number (e.g., 2) of MV predictor candidate blocks whose motion vectors can be used to predict the current motion vector. The RMV candidate blocks can include blocks selected from adjacent blocks and / or temporal blocks in the same frame (e.g., blocks at the same position in a previous or subsequent frame of the current frame). These options represent blocks at spatial or temporal positions relative to the current block that may have a motion vector similar to or the same as that of the current block. The size of the list of MV predictor candidates can be predetermined. For example, the list can include two or more candidates. In order for a candidate block to be on the list of RMV candidates, e.g., the candidate block may be required to have the same reference frame (or reference frames) as the current block, must exist (e.g., a boundary check may be required when the current block is close to the edge of the frame), and must have been encoded during the encoding process and / or decoded during the decoding process. In some implementations, if available and meeting the above conditions, the list of merge candidates can be first filled with spatially adjacent blocks (scanned in a specific predefined order), and then if space is still available in the list, the list of merge candidates can be filled with temporal blocks. For example, adjacent RMV candidate blocks can be selected from the left and top blocks of the current block. The list of RMV predictor candidates can be dynamically formed as a Dynamic Reference List (DRL) at various levels (sequence, picture, frame, slice, superblock, etc.). The DRL can be signaled in the bitstream.

[0072] In some implementations, the actual MV predictor candidates that are used as reference motion vectors for predicting the motion vector of the current block can be signaled. In the case where the RMV candidate list contains two candidates, a one-bit flag called the merge candidate flag can be used to indicate the selection of the reference merge candidate. For each of the multiple motion vectors predicted using the MV predictor for the current block being predicted in the composite mode, it can be associated with a reference motion vector from the merge candidate list. The encoder can determine which of the RMV candidates more closely predicts the MV of the current encoded block, and signal this selection as an index into the DRL.

[0073] In some example implementations of MMVD, after an RMV candidate is selected and used as a base motion vector predictor for a motion vector to be predicted, a motion vector difference (MVD (Motion Vector Difference, MVD), or delta MV, representing the difference between the motion vector to be predicted and a reference candidate motion vector) can be calculated in an encoding system. Such an MVD can include information representing both the magnitude of the MV difference and the direction of the MV difference, and both the magnitude of the MV difference and the direction of the MV difference can be signaled in the bitstream in various ways.

[0074] In some example implementations of MMVD, a distance index can be used to specify the magnitude information of the motion vector difference, and a distance index can be used to indicate one of a set of predefined offsets representing a predefined motion vector difference from a starting point (reference motion vector). Then, an MV offset according to the signaled index can be added to the horizontal or vertical component of the starting (reference) motion vector. An example predefined relationship between the distance index and the predefined offset is specified in Table 1. Table 1 - Example relationship between distance index and predefined MV offset

[0075] In some example implementations of MMVD, the direction index can be further signaled, and the direction index can be used to represent the direction of the MVD relative to the reference motion vector. In some implementations, the direction can be limited to either the horizontal direction or the vertical direction. Example 2-bit direction indexes are shown in Table 2. In the example of Table 2, the description of the MVD can vary according to the information of the start / reference MV. For example, when the start / reference MV corresponds to a uni-directionally predicted block or corresponds to a bi-directionally predicted block where both reference frame lists point to the same side of the current picture (i.e., the POCs of both reference pictures are greater than the POC of the current picture or both are less than the POC of the current picture), the signs in Table 2 can specify the sign (direction) of the MV offset added to the start / reference MV. When the start / reference MV corresponds to a bi-directionally predicted block with two reference pictures on different sides of the current picture (i.e., the POC of one reference picture is greater than the POC of the current picture and the POC of the other reference picture is less than the POC of the current picture), and the difference between the reference POC in picture reference list 0 and the current frame is greater than the difference between the reference POC in picture reference list 1 and the current frame, the signs in Table 2 can specify the sign of the MV offset added to the reference MV corresponding to the reference picture in picture reference list 0, and the sign of the offset of the MV corresponding to the reference picture in picture reference list 1 can have the opposite value (opposite sign of the offset). Otherwise, if the difference between the reference POC in picture reference list 1 and the current frame is greater than the difference between the reference POC in picture reference list 0 and the current frame, then the signs in Table 2 can specify the sign of the MV offset added to the reference MV associated with picture reference list 1, and the sign of the offset of the reference MV associated with picture reference list 0 has the opposite value. Table 2 - Example implementation of the signs of MV offsets specified by the direction index Direction IDX 00 01 10 11 x - axis (horizontal) + – N / A N / A y - axis (vertical) N / A N / A + –

[0076] In some example implementations, the MVD can be scaled according to the difference in POC in each direction. If the differences in POC in both lists are the same, no scaling is required. Otherwise, if the difference in POC in reference list 0 is greater than the difference in POC in reference list 1, scale the MVD of reference list 1. If the difference in POC in reference list 1 is greater than that in list 0, the MVD of list 0 can be scaled in the same way. If the start MV is uni-directionally predicted, then the MVD is added to the available or reference MV.

[0077] In some example implementations of MVD coding and signaling for bidirectional composite prediction, in addition to or as an alternative to separately coding and signaling two MVDS, symmetric MVD coding can be implemented such that only one MVD needs to be signaled and the other MVD can be derived from the signaled MVD. In such an implementation, the motion information including the reference picture indices of list 0 and list 1 is not all signaled. Specifically, at the slice level, a flag called "mvd_l1_zero_flag" can be included in the bitstream, which is used to indicate whether reference list 1 is not signaled in the bitstream. If this flag is 1, which indicates that reference list 1 is equal to zero (and thus not signaled), then the bidirectional prediction flag called "BiDirPredFlag" can be set to 0, which means that there is no bidirectional prediction. Otherwise, if mvd_l1_zero_flag is zero, if the nearest reference picture in list 0 and the nearest reference picture in list 1 form a forward and backward pair or a backward and forward pair of reference pictures, then BiDirPredFlag can be set to 1 and both the list 0 reference picture and the list 1 reference picture are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. BiDirPredFlag being 1 can indicate that a symmetric mode flag is additionally signaled in the bitstream. In the case where BiDirPredFlag is 1, the decoder can extract the symmetric mode flag from the bitstream. For example, the symmetric mode flag can be signaled at the CU level (if needed), and the symmetric mode flag can indicate whether the symmetric MVD coding mode is being used for the corresponding CU. When the symmetric mode flag is 1, it indicates that the symmetric MVD coding mode is used and only the reference picture indices of both list 0 and list 1 (called "mvp_l0_flag" and "mvp_l1_flag") and the MVD associated with list 0 (called "MVD0") are signaled, and the other motion vector difference "MVD1" is derived rather than signaled. For example, MVD1 can be derived as -MVD0. Thus, in the example symmetric MVD mode, only one MVD is signaled.

[0078] In some other example implementations of MV prediction, for both single-reference mode and combined-reference mode MV prediction, a coordination scheme can be used to implement the general merge mode MMVD and some other types of MV prediction. Various syntax elements can be used to signal the way to predict the MV of the current block. For example, for single-reference mode, the following MV prediction modes can be signaled: NEARMV – without any MVD, directly use one of the motion vector predictors (MVPs) indicated by the DRL (Dynamic Reference List) index in the list; NEWMV – use one of the motion vector predictors (MVPs) signaled by the DRL index in the list as a reference, and apply an increment to the MVP (e.g., use MVD); and GLOBALMV – use a motion vector based on frame-level global motion parameters.

[0079] Similarly, for the combined reference inter prediction mode using two reference frames corresponding to the two MVs to be predicted, the following MV prediction modes can be signaled: NEAR_NEARMV – for each of the two MVs to be predicted, without MVD, use one of the motion vector predictors (MVPs) signaled by the DRL index in the list. NEAR_NEWMV – for predicting the first of the two motion vectors, without MVD, use one of the motion vector predictors (MVPs) signaled by the DRL index in the list as the reference MV; for predicting the second of the two motion vectors, use one of the motion vector predictors (MVPs) signaled by the DRL index in the list as the reference MV combined with an additionally signaled incremental MV (MVD). NEW_NEARMV – for predicting the second of the two motion vectors, without MVD, use one of the motion vector predictors (MVPs) signaled by the DRL index in the list as the reference MV; for predicting the first of the two motion vectors, use one of the motion vector predictors (MVPs) signaled by the DRL index in the list as the reference MV combined with an additionally signaled incremental MV (MVD). NEW_NEWMV – use one of the motion vector predictors (MVPs) signaled by the DRL index in the list as the reference MV, and use this reference MV in combination with an additionally signaled incremental MV to predict each of the two MVs. GLOBAL_GLOBALMV – use the MV from each reference based on frame-level global motion parameters.

[0080] The term "NEAR" above refers to MV prediction that uses the reference MV as a general merge mode without any MVD, while the term "NEW" refers to MV prediction that involves using the reference MV in the MMVD mode and offsetting the reference MV by the MVD signaled or derived. For composite inter prediction, the above reference base motion vector and motion vector difference can generally be different or generally independent between two references or two MVDs, even if, for example, the two MVDs may be correlated and such correlation can be exploited to reduce the amount of information needed to signal the two motion vector differences. To exploit such correlation, joint signaling of the two MVDs can be implemented and indicated in the bitstream, as described in further detail below.

[0081] In some example implementations of the MVD, a predefined pixel resolution for the MVD can be allowed. For example, a motion vector precision (or accuracy) of 1 / 8 pixel can be allowed. The MVDs in the various MV prediction modes above can be constructed and signaled in various ways. In some implementations, various syntax elements can be used to signal one or more of the above motion vector differences in reference frame list 0 or list 1.

[0082] For example, a syntax element called "mv_joint" can specify which components of the associated motion vector difference are non-zero. For example, mv_joint has the following values: mv_joint with a value of 0 can indicate that there is no non-zero MVD along the horizontal or vertical direction; mv_joint with a value of 1 can indicate that there is only a non-zero MVD along the horizontal direction; mv_joint with a value of 2 can indicate that there is only a non-zero MVD along the vertical direction; and / or mv_joint with a value of 3 can indicate that there are non-zero MVDs along both the horizontal and vertical directions.

[0083] When the "mv_joint" syntax element of the MVD signals the absence of non-zero MVD components, no further MVD information is signaled. However, if the "mv_joint" syntax signals the presence of one or two non-zero components, additional syntax elements can be signaled for each of the non-zero MVD components as described below.

[0084] For example, a syntax element called "mv_sign" can be used to additionally specify whether the corresponding motion vector difference component is positive or negative.

[0085] For another example, a syntax element called "mv_class" can be used to specify a class of motion vector differences from a predefined set of classes for corresponding non-zero MVD components. For example, the predefined classes of motion vector differences can be used to partition the continuous magnitude space of motion vector differences into non-overlapping class ranges. Thus, the MVD class signaled indicates the magnitude range of the corresponding MVD component. In the example implementation shown in Table 3 below, higher classes correspond to motion vector differences with larger magnitude ranges. The notation (n,m] is used to denote a range of motion vector differences greater than n pixels and less than or equal to m pixels. Table 3: Magnitude Classes of Motion Vector Differences

[0086] In some other examples, a syntax element called "mv_bit" can be used to specify the integer part of the offset between a non-zero motion vector difference component and the starting magnitude of the corresponding signaled MV class magnitude range. In some other examples, a syntax element called "mv_fr" can be used to specify the first 2 fractional bits of the motion vector difference of the corresponding non-zero MVD component, while a syntax element called "mv_hp" can be used to specify the third fractional bit (high-resolution bit) of the motion vector difference of the corresponding non-zero MVD component. The two-bit "mv_fr" essentially provides an MVD resolution of 1 / 4 pixel, while the "mv_hp" bit can further provide a resolution of 1 / 8 pixel. In some other implementations, more than one "mv_hp" bit can be used to provide an MVD pixel resolution finer than 1 / 8 pixel. In some example implementations, an additional flag can be signaled at one or more of various levels to indicate whether 1 / 8 pixel or higher MVD resolution is supported. If the MVD resolution is not applied to a particular coding unit, the syntax elements for the corresponding unsupported MVD resolution may not be signaled.

[0087] In some example implementations, in Bi-prediction with CU-level Weight (BCW), a bi-predicted signal can be generated by averaging two predicted signals obtained from two different reference pictures and / or using two different motion vectors. In some other implementations, the bi-prediction mode can be extended beyond simple averaging to allow weighted averaging of the two predicted signals. For example, P 双向预测= ((8 - w) * P0 + w * P1 + 4) >> 3. Five weights are allowed in weighted bi - direction prediction, w ∈ {-2, 3, 4, 5, 10}. When w equals 4, equal weighting factors are used to perform weighted averaging of two prediction samples. For each bi - directionally predicted CU, the weight w can be determined in one of the following two ways: 1) For non - merged CUs, the weight index is signaled after the motion vector difference; and / or 2) For merged CUs, the weight index is inferred from neighboring blocks based on the merge candidate index. BCW can be applied only to CUs with 256 or more luma samples (i.e., CU width times CU height is greater than or equal to 256). For low - latency pictures, all 5 weights are used. For non - low - latency pictures, only 3 weights (w ∈ {3, 4, 5}) are used.

[0088] In some implementations, block - adaptive weighted prediction (BAWP) may include block - level weighted prediction to model the local illumination change between the current block and its predicted block as a function of the local illumination change between the current block template (or causal samples of the current block) and the reference block template. Figure 10 The template of the current block (1010) (or the current template, 1012) and the template of the reference block (1020) (or the reference template, 1022) are shown. The reference block can be indicated or determined by a motion vector (MV, 1030). The current block can be in the current picture (or current frame), and the reference block can be in the reference picture (or reference frame). In some implementations, the function can be a linear function. The parameters of the function can be represented by a scaling factor α and an offset β, which form a linear expression α * p[x]+β to compensate for the illumination change, where p[x] is the reference sample pointed to by the MV at position x on the reference picture. In some implementations, α and β can be derived based on the current block template and the reference block template, and they do not require signaling overhead. In some implementations, a BAWP flag can be signaled for the single - inter - frame prediction mode to indicate the use of BAWP. BAWP can be applied to blocks of size greater than or equal to 8×8 and encoded in the single - inter - frame prediction mode. In some implementations, BAWP can be applied only to the luma component. In some implementations, BAWP can be referred to as Local Illumination Compensation (LIC).

[0089] In some embodiments, if the current block includes more than one motion vector and each motion vector points to a block in a reference frame, multiple linear models can be used to describe the linear relationship between the current block and its multiple reference blocks in the BAWP / LIC mode. For each linear model, the linear function includes a scaling factor α and an offset β, and the scaling factor α and / or the offset β can be derived from the template of the current block and the template of each of the reference blocks, or can be signaled into the bitstream and parsed at the decoder side to reconstruct the predicted block. Each of the reference blocks is specified by a motion vector associated with its reference frame. In one embodiment, if the current block includes more than one motion vector and each motion vector points to a block in a reference block, at least one linear model is employed to describe the linear relationship between the current block and one of its reference blocks. In one embodiment, if the current block includes more than one motion vector and each motion vector points to a block in a reference frame, a separate linear model is employed for each motion vector. Specifically, according to the disclosed method, an exemplary decoder may receive an encoded video bitstream. Additionally, the decoder may identify from the encoded video bitstream a first motion vector corresponding to a first reference block and a second motion vector corresponding to a second reference block associated with the current block of the current frame. Then, the decoder may obtain a first scaling factor corresponding to the first motion vector and a second scaling factor corresponding to the second motion vector by parsing the encoded video bitstream. Additionally, the decoder may generate a first predicted block based on the first scaling factor and the first reference block according to a first linear equation associated with the first scaling factor and a first offset, and may generate a second predicted block based on the second reference block according to a second linear equation associated with the second scaling factor and a second offset. Then, the decoder may reconstruct the current block based on the first predicted block and the second predicted block.

[0090] In some implementations with explicit signaling of BAWP, in the case of predicting the current block according to its reference block using a linear function with a scaling factor α and an offset β, the selection / value of the scaling factor α and / or the offset β can be signaled into the bitstream and parsed at the decoder side to reconstruct the predicted block. The reference block is specified by a motion vector associated with the current block. All supported values for the scaling factor α are stored in a predefined look-up table, and the index of the scaling factor in the look-up table can be signaled in the bitstream and parsed at the decoder side. The offset value β can be derived according to the linear equation between the reference block and the current block. The offset value β is set to (cur_template_mean - α * ref_template_mean), where cur_template_mean indicates the average of the samples in the template of the current block, and ref_template_mean indicates the average of the samples in the template of the reference block.

[0091] In some implementations, one or more spatial motion vector predictors (SMVPs, both neighboring and non - neighboring SMVPs), one or more temporal motion vector predictors (TMVPs), one or more additional MV candidates and additionally derived MVPs, and one or more reference library MVPs are added. A stack with a fixed size can be generated on both the encoder and decoder sides to store MVPs, and this stack is referred to as the motion vector predictor list.

[0092] The MVP list can be constructed to hold a predetermined number of reconstructed MVP candidates (SMVPs, TMVPs, or other derived MVPs, or other types of MVP candidates) on both the encoder and decoder sides of the current coding block or super - block. When encoding the current prediction block in an inter - prediction mode, the encoder selects the MVP that provides the best coding efficiency from the candidates in the MVP candidate list as the predictor of the motion vector of the current prediction block. The index of the selected MVP in the MVP list can be signaled in the bitstream. The decoder will correspondingly update the MVP list of the current coding block or super - block when reconstructing the bitstream, extract the MVP index of the current inter - prediction prediction block, obtain the MVP from the MVP candidate list according to the MVP index extracted from the MVP list, and use the MVP as the predictor of the motion vector of the current prediction block to reconstruct the motion vector of the current prediction block (e.g., by combining the motion vector predictor extracted from the MVP list with the corresponding MVD). For example, the MVP list can represent a stack with a predefined fixed size.

[0093] The spatial motion vector predictor can be a neighboring SMVP or a non - neighboring SMVP. A neighboring SMVP can refer to the motion vector predictor belonging to a prediction block neighboring the current coding block or super - block. A non - neighboring SMVP can refer to the motion vector predictor belonging to a prediction block not immediately adjacent to the current coding block or super - block. Other types of MVP candidates can also be derived based on the reconstructed motion vectors. For another example, one or more additional MVP libraries can be maintained as one of the sources for building the MVP list, as described in further detail below.

[0094] In some implementations, the distribution of the supported scaling factors can vary depending on the coding information, so there may be room for further improvement in the signaling of the scaling factors. In the present disclosure, the lookup table can also be referred to as a list, and thus, the lookup table of the scaling factor candidates can be referred to as the list of scaling factor candidates.

[0095] In various embodiments of the present disclosure, for simplicity of description, Mode 1 may be referred to as an encoding mode that inherits the motion vector of an adjacent block; Mode 2 may be referred to as an encoding mode that signals a motion vector difference relative to a motion vector predictor selected from spatially or temporally adjacent blocks or a given derived motion vector (e.g., a global motion vector); Mode 3 may be referred to as an encoding mode that signals a motion vector difference relative to a motion vector predictor selected from spatially or temporally adjacent blocks or a given derived motion vector (e.g., a global motion vector), and implicitly determines the precision of the motion vector based on the magnitude of the motion vector.

[0096] In some implementations, for a current intra-coded block located at the boundary of a block (e.g., a superblock), a reference index indicates a neighboring reference line when it is zero; or a non-zero (or non-neighboring) reference line for intra prediction of the current coded block when it is a non-zero integer. Referring to Figure 11 the non-limiting example in, a coded block (also referred to as a decoded block or an encoded block) (1102) is positioned as the top boundary (1130) and the left boundary (1140) of a block (e.g., a superblock). The top boundary (1130) and the left boundary (1140) of the superblock may be indicated by thick lines, as Figure 11 shown.

[0097] In its top direction, the current coded block may have a top neighboring reference line with an index of zero (or referred to as the top closest neighboring reference line, or zero neighboring reference line) (1110), and one or more top non-neighboring reference lines (or referred to as top non-zero neighboring reference lines with non-zero indices) (1104, 1106, and 1108). For example, the first top non-neighboring reference line (1108) may have a reference index of 1, the second top non-neighboring reference line (1106) may have a reference index of 2, and / or the third top non-neighboring reference line (1104) may have a reference index of 3.

[0098] Similarly, in its left direction, the current coded block may have a left neighboring reference line with an index of zero (or referred to as the left closest neighboring reference line, or zero neighboring reference line) (1118), and one or more left non-neighboring reference lines (or referred to as left non-zero neighboring reference lines with non-zero indices) (1112, 1114, and 1116). For example, the first left non-neighboring reference line (1116) may have a reference index of 1, the second left non-neighboring reference line (1114) may have a reference index of 2, and / or the third left non-neighboring reference line (1112) may have a reference index of 3.

[0099] In some implementations, there are some problems or challenges associated with the BAWP method, particularly how to improve the flexibility and / or efficiency of determining the template for determining the offset value. For non-limiting examples, some implementations may lack the flexibility to determine the position of the template (e.g., only allowing the top template and the left template to be used together). For another non-limiting example, some implementations may lack the flexibility to determine the scope / size of the template (e.g., only allowing the top region and the left region of the current block to be used as the template). The present disclosure describes various implementations for enhancing BAWP, solving at least one of the above problems or challenges, improving the encoding / decoding efficiency, and advancing video codec technology.

[0100] The various implementations and / or realizations described in the present disclosure can be performed individually or in any order of combination. Additionally, each of the method (or implementation), the encoder, and the decoder can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). One or more processors execute a program stored in a non-transitory computer-readable medium. In the present disclosure, the term block can be interpreted as a prediction block, an encoding block, or a coding unit (CU). In the present disclosure, the direction of the reference frame can be determined by whether the reference frame is before the current frame in the display order or after the current frame in the display order.

[0101] Figure 12 FIG. 1200 is a flowchart of an exemplary method for decoding a current block of a current frame in an encoded video bitstream that follows the basic principles of the above implementation, and this exemplary method can be executed by an electronic device (e.g., a decoder). The exemplary decoding method flow starts at S1201 and can include some or all of the following steps: S1210, receiving the encoded video bitstream; S1220, identifying from the encoded video bitstream a motion vector corresponding to a reference block associated with the current block of the current frame; S1230, obtaining a scaling factor based on the syntax explicitly signaled in the encoded video bitstream; S1240, determining a template for deriving an offset value; S1250, deriving the offset value based on the template; S1260, generating a prediction block based on the reference block according to a linear formula associated with the scaling factor and the offset value; and / or S1270, reconstructing the current block by the device based on the prediction block. This exemplary method can stop at S1299.

[0103] In any part or combination of the above implementation, determining the template for deriving the offset value includes some or all of the following: obtaining a second syntax for indicating the position of the template for deriving the offset value; and / or determining the template for deriving the offset value based on the second syntax.

[0104] In any part or combination of the above implementation, the template for deriving the offset value based on the second syntax includes some or all of the following: The second syntax includes an integer between 0 and 2 (including 2); in response to the second syntax being 0, determining both the upper region and the left region of the current block as the template for deriving the offset value; in response to the second syntax being 1, determining only the upper region of the current block as the template for deriving the offset value; and / or in response to the second syntax being 2, determining only the left region of the current block as the template for deriving the offset value.

[0105] In any part or combination of the above implementation, the template for deriving the offset value based on the second syntax includes some or all of the following: The second syntax includes 0 or 1; in response to the second syntax being 0, determining only the upper region of the current block as the template for deriving the offset value; and / or in response to the second syntax being 1, determining only the left region of the current block as the template for deriving the offset value.

[0106] In any part or combination of the above implementation, the upper region of the current block includes both the upper sample of the current block and the upper-right sample of the current block.

[0107] In any part or combination of the above implementation, the upper sample of the current block and the upper-right sample of the current block have the same width.

[0108] In any part or combination of the above implementation, the left region of the current block includes both the left sample of the current block and the lower-left sample of the current block.

[0109] In any part or combination of the above implementation, the left sample of the current block and the lower-left sample of the current block have the same height.

[0110] In any part or combination of the above implementation, the second syntax is entropy-coded based on at least one of the following according to the context: the block shape of the current block, the block aspect ratio of the current block, the block size of the current block, or the position of the current block within the current tile, slice, sub-picture, or picture.

[0111] In any part or combination of the above implementation, the second syntax indicates whether a neighboring reference line or a non-neighboring reference line is used as the template for deriving the offset value.

[0112] In any part or combination of the above implementation, determining the template for deriving the offset value includes determining a subset of samples from the upper sample and the left sample of the current block as the template for deriving the offset value.

[0113] In any part or combination of the above implementation manners, determining a subset of samples from the upper samples and left samples of the current block as a template for deriving an offset value includes one or more of the following: determining only the upper samples of the current block as a template for deriving an offset value; and / or determining only the left samples of the current block as a template for deriving an offset value.

[0114] In any part or combination of the above implementation manners, determining a subset of samples from the upper samples and left samples of the current block as a template for deriving an offset value includes one or more of the following: in response to the block width of the current block being greater than the block height of the current block, determining only the upper samples of the current block as a template for deriving an offset value; in response to the block height of the current block being greater than the block width of the current block, determining only the left samples of the current block as a template for deriving an offset value; and / or in response to the block height of the current block being equal to the block width of the current block, determining both the upper samples and left samples of the current block as a template for deriving an offset value.

[0115] In any part or combination of the above implementation manners, determining a subset of samples from the upper samples and left samples of the current block as a template for deriving an offset value includes one or more of the following: in response to the current block being located at the top boundary of a tile, slice, subpicture, or picture, determining only the left samples of the current block as a template for deriving an offset value; and / or in response to the current block being located at the left boundary of a tile, slice, subpicture, or picture, determining only the upper samples of the current block as a template for deriving an offset value.

[0116] In any part or combination of the above implementation manners, determining a subset of samples from the upper samples and left samples of the current block as a template for deriving an offset value includes one or more of the following: in response to the current block being located at the right boundary of a tile, slice, subpicture, or picture, excluding the upper right sample of the current block as a template for deriving an offset value; and / or in response to the current block being located at the bottom boundary of a tile, slice, subpicture, or picture, excluding the lower left sample of the current block as a template for deriving an offset value.

[0117] In any part or combination of the above implementation manners, the number of samples in the sample subset is less than N, where N is a positive integer.

[0118] In any part or combination of the above implementation manners, N is determined based on the minimum or maximum value of the block width and block height of the current block.

[0119] In any part or combination of the above implementation manners, N is determined independently of the block size of the current block.

[0120] In any part or combination of the above implementation manners, the sample subset includes subsampling of the upper samples and left samples of the current block.

[0121] In various embodiments, when explicit signaling of a scaling factor is employed in block - adaptive weighted prediction, a flag, such as one called offset_region_idx, may be signaled to indicate which part of the template is used to derive the offset value β.

[0122] In some implementations, the value of offset_region_idx ranges from 0 to 2. When offset_region_idx is 0, it indicates that both the upper region and the left region of the template of the current block as well as the template of the reference block are used to derive the offset value. When offset_region_idx is 1, it indicates that the upper region of the template is used to derive the offset value. Otherwise, when offset_region_idx is 2, it indicates that the left region of the template is used to derive the offset value.

[0123] In some implementations, the value of offset_region_idx ranges from 0 to 1. When offset_region_idx is 0, it indicates that the upper region of the template is used to derive the offset value. Additionally, when offset_region_idx is 1, it indicates that the left region of the template is used to derive the offset value.

[0124] In some implementations, when the value of offset_region_idx indicates that the upper region of the current block is used to derive the offset value. Refer to Figure 13 , the upper region of the current block includes both the upper sample of the current block and the upper - right sample of the current block. In some implementations, the upper sample of the current block may have the same width as the upper - right sample of the current block, such that the upper region may have a width that is twice the width of the current block. Similarly, the upper region of the reference block includes both the upper sample of the reference block and the upper - right sample of the reference block. In some implementations, the upper sample of the reference block may have the same width as the upper - right sample of the reference block, such that the upper region may have a width that is twice the width of the reference block.

[0125] In some implementations, when the value of offset_region_idx indicates that the left region of the current block is used to derive the offset value. Refer to Figure 14 , the left region of the current block includes the left sample of the current block and the lower - left sample of the current block. In some implementations, the left sample of the current block may have the same height as the lower - left sample of the current block, such that the left region may have a height that is twice the height of the current block. Similarly, the left region of the reference block includes both the left sample of the reference block and the lower - left sample of the reference block. In some implementations, the left sample of the reference block may have the same height as the lower - left sample of the reference block, such that the left region may have a height that is twice the height of the reference block.

[0126] In some implementations, the context modeling for entropy encoding offset_region_idx depends on encoding information, but is not limited to: block shape, block aspect ratio, block size, and the position of the current block within the current tile, slice, sub-picture, or picture. For example, the position of the current block can refer to whether the current block is located at the top boundary and / or left boundary of the tile / slice / sub-picture / picture, and based on this, the syntax value of the encoding mode can be associated with the top neighboring block and / or left neighboring block.

[0127] In some implementations, when explicit signaling of a scaling factor for block-level adaptive weighted prediction is used for a block encoded with a block vector, a syntax element (e.g., a flag or an index) can be signaled to indicate which neighboring reference lines and non-neighboring reference lines are used to derive the offset value of the current block. For example, the flag can indicate whether neighboring reference lines or non-neighboring reference lines are used to derive the offset value; and / or the index can indicate whether neighboring reference lines or non-neighboring reference lines are used to derive the offset value, and when non-neighboring reference lines are used, the index can indicate which reference line among the non-neighboring reference lines is used to derive the offset value.

[0128] In some embodiments, when explicit signaling of a scaling factor is adopted in block adaptive weighted prediction, only a subset of samples in the upper region and / or left region is used to derive the offset value β. In some implementations, when explicit signaling of a scaling factor is adopted in block adaptive weighted prediction, only the upper sample or the left sample is used to derive the offset value β. In some implementations, when the block width is greater than the height, the upper sample is used to derive the offset value β. In some implementations, when the block height is greater than the width, the left sample is used to derive the offset value β. In some implementations, when the block height is equal to the block width, both the upper sample and the left sample are used to derive the offset value β. In some implementations, when the current block is located at the top boundary of the tile / slice / sub-picture / picture, the left sample is used to derive the offset value β. In some implementations, when the current block is located at the left boundary of the tile / slice / sub-picture / picture, the upper sample is used to derive the offset value β. In some implementations, when the current block is located at the right boundary of the tile / slice / sub-picture / picture, the upper-right sample is not used to derive the offset value β. In some implementations, when the current block is located at the bottom boundary of the tile / slice / sub-picture / picture, the lower-left sample is not used to derive the offset value β.

[0129] In some implementations, at most N samples in the upper region and / or left region of the template can be used to derive the offset value β. In some implementations, the value of N depends on the minimum or maximum of the block width and block height. For a non-limiting example, when the minimum of the block width is 4, N is set to 2; otherwise, N is set to 4. In some implementations, N is set to be equal to a fixed value such as 2 or 4, regardless of the block size.

[0130] In some implementations, the positions of the N samples are predefined as a subsampling of only the top sample, only the left sample, or both the top sample and the left sample. The subsampling may include even subsampling or non - even subsampling. Figure 15 A non - limiting example of forming a subsampling of the N samples for deriving an offset value for a current block is shown. The subsampling of the top sample (1510) may be performed the same as or differently from the subsampling of the left sample (1520) in terms of even or non - even subsampling, subsampling interval (or subsampling step).

[0131] Various embodiments in the present disclosure may include a method for encoding a current block into a video bitstream by an encoder, the method including an inverse process of any part or all of the processing described for a decoder. Various embodiments in the present disclosure may include a method for encoding a current block of a streamed video by one or more electronic devices (e.g., a streaming media player), the method including any part or all of the processing for a decoder and / or any part or all of the processing described for an encoder.

[0132] The above operations may be combined or arranged in any number or order as needed. Two or more of the steps and / or operations may be performed in parallel. The embodiments and implementations in the present disclosure may be used alone or in any order in combination. Additionally, each of the methods (or embodiments), the encoder, and the decoder may be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non - transitory computer - readable medium. The embodiments in the present disclosure may be applied to luminance blocks or chrominance blocks. The term "block" may be interpreted as a prediction block, an encoding block, or a coding unit, i.e., a CU. Here, the term "block" may also be used to refer to a transform block. In the following, when referring to the block size, it may refer to the block width or height, or the maximum of the width and height, or the minimum of the width and height, or the area size (width × height), the aspect ratio of the block (width∶height, or height∶width).

[0133] The techniques described above may be implemented as computer software using computer - readable instructions and physically stored on one or more computer - readable media. For example, Figure 16 A computer system (1800) suitable for implementing certain embodiments of the disclosed subject matter is shown.

[0134] Computer software can be encoded using any suitable machine code or computer language, which can be subject to mechanisms such as assembly, compilation, linking, etc. to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode execution, etc.

[0135] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.

[0136] Figure 16 The components shown in are exemplary in nature and are not intended to impose any limitations on the scope of use or functionality of the computer software implementing the present disclosure. The configuration of the components should also not be construed as having any dependencies or requirements related to any one or combination of the components shown in the exemplary embodiments of the computer system (1800).

[0137] The computer system (1800) may include certain human-machine interface input devices. The input human-machine interface devices may include one or more of the following (only one of each is shown): keyboard (1801), mouse (1802), touchpad (1803), touch screen (1810), data glove (not shown), joystick (1805), microphone (1806), scanner (1807), camera device (1808).

[0138] The computer system (1800) may also include certain human-machine interface output devices. Such human-machine interface output devices can stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include: tactile output devices (e.g., tactile feedback through the touch screen (1810), data glove (not shown), or joystick (1805), but there may also be tactile feedback devices that do not serve as input devices); audio output devices (e.g., speakers (1809), headphones (not depicted)); visual output devices (e.g., screen (1810), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of the screens may be able to output two-dimensional visual output or more than three-dimensional output through means such as stereoscopic output; virtual reality glasses (not depicted); holographic displays and fog machines (not depicted)); and printers (not depicted).

[0139] The computer system (1800) may also include human-accessible storage devices and their associated media, such as optical media including CD / DVD ROM (Read-Only Memory) / RW (1820) with media (1821) such as CD / DVD, thumb drives (1822), removable hard disk drives or solid state drives (1823), traditional magnetic media such as tapes and floppy disks (not depicted), dedicated ROM / ASIC / PLD-based devices such as security dongles (not depicted), etc.

[0140] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transient signals.

[0141] The computer system (1800) may also include an interface (1854) to one or more communication networks (1855). The network may be, for example, wireless, wired, optical. The network may also be local, wide area, metropolitan area, vehicular and industrial networks, real-time, delay-tolerant, etc. Examples of networks include: local area networks such as Ethernet; wireless LAN; cellular networks including GSM, 3G, 4G, 5G, LTE, etc.; television cable or wireless wide area digital networks including cable television, satellite television, and terrestrial broadcast television; vehicular and industrial networks including CAN bus, and so on.

[0142] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces may be attached to the core (1840) of the computer system (1800).

[0143] The core (1840) may include one or more central processing units (CPUs) (1841), a graphics processing unit (GPU) (1842), a dedicated programmable processing unit in the form of a field programmable gate area (FPGA) (1843), a hardware accelerator (1844) for certain tasks, a graphics adapter (1850), etc. These devices, together with a read-only memory (ROM) (1845), a random access memory (1846), an internal mass storage device such as an internal non-user-accessible hard disk drive, SSD, etc. (1847), may be connected via a system bus (1848). In some computer systems, the system bus (1848) may be accessed in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices may be directly attached to the system bus (1848) of the core or may be attached to the system bus (1848) of the core via a peripheral bus (1849). In an example, a screen (1810) may be connected to the graphics adapter (1850). The architecture of the peripheral bus includes PCI, USB, etc.

[0144] Computer-readable media may have computer code for performing various computer-implemented operations. The media and the computer code may be media and computer code specially designed and constructed for the purposes of this disclosure, or they may be of the type well-known and available to those of ordinary skill in the art of computer software.

[0145] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various equivalent substitutes that fall within the scope of the present disclosure. Accordingly, it will be recognized that those of ordinary skill in the art will be able to envision many systems and methods that, although not explicitly shown or described herein, implement the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.

Claims

1. A method for decoding a current block of a current frame in an encoded video bitstream, the method comprising: Receiving, by a device, the encoded video bitstream, the device comprising a memory storing instructions and a processor communicatively coupled to the memory; Identifying, by the device, from the encoded video bitstream a motion vector corresponding to a reference block associated with the current block of the current frame; Obtaining, by the device, a scaling factor based on syntax explicitly signified in the encoded video bitstream; Determining, by the device, a template for deriving an offset value; Deriving, by the device, the offset value based on the template; Generating, by the device, a prediction block based on the reference block according to a linear formula associated with the scaling factor and the offset value; And Reconstructing, by the device, the current block based on the prediction block.

2. The method according to claim 1, wherein Determining the template for deriving the offset value includes: Obtaining a second syntax for indicating a position of the template for deriving the offset value; and Determining the template for deriving the offset value based on the second syntax.

3. The method according to claim 2, wherein Determining the template for deriving the offset value based on the second syntax includes: The second syntax includes an integer between 0 and 2 inclusive; In response to the second syntax being 0, determining both an upper region and a left region of the current block as the template for deriving the offset value; In response to the second syntax being 1, determining only the upper region of the current block as the template for deriving the offset value; and In response to the second syntax being 2, determining only the left region of the current block as the template for deriving the offset value.

4. The method according to claim 2, wherein Determining the template for deriving the offset value based on the second syntax includes: The second syntax includes 0 or 1; In response to the second syntax being 0, determining only the upper region of the current block as the template for deriving the offset value; and In response to the second syntax being 1, determining only the left region of the current block as the template for deriving the offset value.

5. The method according to claim 3, wherein: The upper region of the current block includes both an upper sample of the current block and an upper-right sample of the current block.

6. The method according to claim 5, wherein: The upper sample of the current block and the upper-right sample of the current block have the same width.

7. The method according to claim 3, wherein: The left region of the current block includes both a left sample of the current block and a lower-left sample of the current block.

8. The method according to claim 7, wherein: The left sample of the current block and the lower-left sample of the current block have the same height.

9. The method according to claim 2, wherein: The second syntax is entropy-coded based on context according to at least one of: a block shape of the current block, a block aspect ratio of the current block, a block size of the current block, or a position of the current block within a current tile, slice, subpicture, or picture.

10. The method according to claim 2, wherein: The second syntax indication uses an adjacent reference line or a non - adjacent reference line as the template for obtaining the offset value.

11. The method according to claim 1, wherein Determining the template for obtaining the offset value includes: Determining a subset of samples in the upper samples and left samples of the current block as the template for obtaining the offset value.

12. The method according to claim 11, wherein, Determining the subset of samples in the upper samples and the left samples of the current block as the template for obtaining the offset value includes: Determining only the upper samples of the current block as the template for obtaining the offset value; or Determining only the left samples of the current block as the template for obtaining the offset value.

13. The method according to claim 11, wherein, Determining the subset of samples in the upper samples and the left samples of the current block as the template for obtaining the offset value includes: In response to the block width of the current block being greater than the block height of the current block, determining only the upper samples of the current block as the template for obtaining the offset value; In response to the block height of the current block being greater than the block width of the current block, determining only the left samples of the current block as the template for obtaining the offset value; or In response to the block height of the current block being equal to the block width of the current block, determining both the upper samples and the left samples of the current block as the template for obtaining the offset value.

14. The method according to claim 11, wherein, Determining the subset of samples in the upper samples and the left samples of the current block as the template for obtaining the offset value includes: In response to the current block being located at the top boundary of a tile, slice, sub - picture, or picture, determining only the left samples of the current block as the template for obtaining the offset value; or In response to the current block being located at the left boundary of a tile, slice, sub - picture, or picture, determining only the upper samples of the current block as the template for obtaining the offset value.

15. The method according to claim 11, wherein, Determining the subset of samples in the upper samples and the left samples of the current block as the template for obtaining the offset value includes: In response to the current block being located at the right boundary of a tile, slice, sub - picture, or picture, excluding the upper - right samples of the current block as the template for obtaining the offset value; or In response to the current block being located at the lower side of a tile, slice, sub - picture, or picture, excluding the lower - left samples of the current block as the template for obtaining the offset value.

16. The method according to claim 11, wherein: The number of samples in the sample subset is less than N, where N is a positive integer.

17. The method according to claim 16, wherein: N is determined based on the minimum or maximum of the block width and block height of the current block.

18. The method according to claim 16, wherein: N is determined independently of the block size of the current block.

19. The method according to claim 11, wherein: The sample subset includes sub - sampling of the upper samples and the left samples of the current block.

20. An apparatus for decoding a current block of a current frame in an encoded video bitstream, the apparatus comprising: A memory storing instructions; And A processor communicating with the memory, wherein when the processor executes the instructions, the processor is configured to cause the device to perform the method according to any one of claims 1 to 19.

21. A non-transitory computer-readable storage medium storing instructions, wherein, When the instructions are executed by the processor, the instructions are configured to cause the processor to perform the method according to any one of claims 1 to 19.