Method and apparatus for enhancing block adaptive weighted prediction using block vector

Through the block adaptive weighted prediction (BAWP) method, the prediction block is generated using block vectors and linear equations, which solves the problem of quality degradation caused by lighting changes in video encoding, and improves the encoding efficiency and quality.

CN120266477APending Publication Date: 2025-07-04TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380082021.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-27
Filing Date
2023-11-28
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing video encoding technology is inefficient when dealing with local lighting changes, and it is difficult to effectively compensate for the video quality decline caused by lighting changes.

Method used

Block adaptive weighted prediction (BAWP) method is used to identify the block vectors of the current block and the reference block, determine the scaling factor, and generate prediction blocks using linear equations to compensate for illumination changes.

Benefits of technology

Improve the efficiency and quality of video encoding, especially when dealing with local lighting changes, improve the reconstruction effect of video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120266477A_ABST
    Figure CN120266477A_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to video coding, and more particularly, to enhanced block adaptive weighted prediction (BAWP) using block vectors. A method includes: receiving an encoded video bitstream; identifying a block vector corresponding to a reference block associated with a current block of a current frame from the encoded video code stream; determining a scaling factor based on a syntax explicitly signaled in the encoded video bitstream; generating a prediction block based on the reference block according to a linear equation associated with the scaling factor; and reconstructing, by a device, the current block on the basis of the prediction block.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to U.S. Provisional Application No. 63 / 532,637, filed on Aug. 14, 2023, and U.S. Application No. 18 / 519,775, filed on Nov. 27, 2023, the entire contents of which are incorporated herein by reference. Technical Field

[0002] Embodiments of this application relate to a set of advanced video / stream encoding and decoding techniques. In particular, the disclosed techniques relate to using block vectors to enhance block adaptive weighted prediction (BAWP) to compensate for local illumination variations. Background Art

[0003] Uncompressed digital video can include a series of pictures and impose specific bitrate requirements on storage, data processing, and transmission bandwidth in streaming applications. One objective of video encoding and decoding can be to reduce redundancy in the uncompressed input video signal through various compression techniques. Summary of the Invention

[0004] Embodiments of this application relate to methods, apparatuses, and computer-readable storage media for enhancing block adaptive weighted prediction (BAWP).

[0005] According to one aspect, embodiments of this application provide a method for decoding a current block of a current frame in an encoded video bitstream, the method comprising:

[0006] receiving, by a device comprising a memory storing instructions and a processor in communication with the memory, the encoded video bitstream;

[0007] identifying, by the device, from the encoded video bitstream, a block vector corresponding to a reference block associated with the current block of the current frame;

[0008] determining, by the device, a scaling factor based on syntax explicitly signaled in the encoded video bitstream;

[0009] generating, by the device, based on the reference block, a prediction block according to a linear equation associated with the scaling factor; and,

[0010] reconstructing, by the device, the current block based on the prediction block.

[0011] According to another aspect, embodiments of this application further provide a decoding apparatus for a current block of a current frame in an encoded video bitstream, comprising a memory storing instructions; and a processor in communication with the memory, wherein when the processor executes the instructions, the processor is configured to cause the apparatus to perform the above video decoding and / or encoding method.

[0012] According to another aspect, an embodiment of the present application further provides a non-transitory computer-readable storage medium, on which instructions are stored, and when the instructions are executed by a computer for video decoding and / or encoding, the computer is caused to implement the above video decoding and / or encoding method.

[0013] The above and other aspects and their implementation manners are described in more detail in the accompanying drawings, the description, and the claims. Description of the Drawings

[0014] According to the following detailed description and the accompanying drawings, other features, properties, and various advantages of the disclosed subject matter will become further apparent, where:

[0015] Figure 1 is a schematic diagram of a simplified block diagram of a communication system (100) according to an embodiment;

[0016] Figure 2 is a schematic diagram of a simplified block diagram of a communication system (200) according to another embodiment;

[0017] Figure 3 is a schematic diagram of a simplified block diagram of a video decoder according to an embodiment;

[0018] Figure 4 is a schematic diagram of a simplified block diagram of a video encoder according to an embodiment;

[0019] Figure 5 shows a block diagram of a video encoder according to another embodiment;

[0020] Figure 6 shows a block diagram of a video decoder according to another embodiment;

[0021] Figure 7 shows a scheme for encoding block partitioning according to an embodiment of the present application;

[0022] Figure 8 shows another scheme for encoding block partitioning according to an embodiment of the present application;

[0023] Figure 9 shows yet another scheme for encoding block partitioning according to an embodiment of the present application;

[0024] Figure 10 shows a prediction mode using 4 predefined search regions according to an embodiment of the present application;

[0025] Figure 11A shows a first search example of subsampling based on a central sample with a 3x3 granularity;

[0026] Figure 11B shows an example of further correcting the search within a 3x3 window;

[0027] Figure 12 Shows an example of the template of the current block and the template of the reference block;

[0028] Figure 13 Shows an example of the reference line for encoding the block;

[0029] Figure 14 Shows an example of the logical flow of the method according to an embodiment of the present application;

[0030] Figure 15 Is a schematic diagram of a computer system according to an embodiment of the present application. Detailed implementation manners

[0031] The present invention will be described in detail below with reference to the accompanying drawings, which form a part of the present invention and show specific examples of embodiments by way of illustration. However, it should be noted that the present invention can be embodied in various different forms, and thus, the subject matter covered or claimed is intended to be construed as not limited to any of the embodiments described below. It should also be noted that the present invention can be embodied as a method, an apparatus, a component, or a system. Therefore, the embodiments of the present invention can, for example, take the form of hardware, software, firmware, or any combination thereof.

[0032] Throughout the specification and claims, terms may have nuanced meanings that are implied or suggested in the context rather than being explicit. As used herein, the phrase "in one embodiment" or "in some embodiments" does not necessarily refer to the same embodiment, and the phrase "in another embodiment" or "in other embodiments" as used herein does not necessarily refer to different embodiments. Similarly, the phrase "in one implementation" or "in some implementations" as used herein does not necessarily refer to the same implementation, and "in another implementation" or "in other implementations" as used herein does not necessarily refer to different implementations. For example, the claimed subject matter includes all or part of the combination of exemplary embodiments / implementations.

[0033] In general, terms can be understood, at least in part, from their use in context. For example, terms used herein, such as "and," "or," and "and / or," can have multiple meanings, which can depend, at least in part, on the context in which these terms are used. Generally, "or" when used in a list such as A, B, or C, means A, B, and C, used herein in an inclusive sense, as well as A, B, or C, used herein only in an exclusive sense. Additionally, the terms "at least one" or "one or more" used herein can, at least in part, depend on context, be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, and characteristics in a plural sense. Similarly, terms such as "a," "an," or "the" can also be understood to convey a singular or plural usage, at least in part, depending on context. Additionally, the terms "based on" or "comprising" can be understood to not necessarily be intended to convey a set of exclusive factors, and instead, can allow for the presence of other factors that are not necessarily explicitly described, again, at least in part, depending on context.

[0034] As Figure 1 shown, the terminal device can be implemented as a server, a personal computer, and a smartphone, but the applicability of the basic principles of this application may not be limited thereto. Embodiments of this application can be implemented in a desktop computer, a laptop computer, a tablet computer, a media player, a wearable computer, a dedicated video conferencing device, and the like. The network (130) represents any number or type of network that conveys encoded video data between terminal devices, including, for example, a wired (wired) and / or wireless communication network. The communication network (130) can exchange data in circuit-switched, packet-switched, and / or other types of channels. Representative networks include a telecommunications network, a local area network, a wide area network, and / or the Internet.

[0035] As an example, Figure 2 illustrates the placement of a video encoder and a video decoder in a streaming environment. The subject matter disclosed in this application can be equally applicable to other video-supported applications, including, for example, video conferencing, digital TV, storing compressed video on digital media including CDs, DVDs, memory sticks, and the like.

[0036] As Figure 2As shown, the streaming system may include an acquisition subsystem (213), and the acquisition subsystem may include a video source (201), such as a digital camera, for creating an uncompressed video picture stream (202). In an embodiment, the video picture stream (202) includes samples captured by the digital camera of the video source (201). Compared with the encoded video data (204) (or encoded video stream), the video picture stream (202) is depicted as a thick line to emphasize the high data volume of the video picture stream. The video picture stream (202) may be processed by an electronic device (420), and the electronic device (420) includes a video encoder (203) coupled to the video source (201). The video encoder (203) may include hardware, software, or a combination of both to implement or carry out aspects of the disclosed subject matter described in more detail below. Compared with the video picture stream (202), the encoded video data (204) (or encoded video stream (204)) is depicted as a thin line to emphasize the lower data volume of the encoded video data (204) (or encoded video stream (204)), which may be stored on the streaming server (205) for future use. At least one streaming client subsystem, such as Figure 3 the client subsystem (206) and the client subsystem (208) in

[0037] Figure 3 is a block diagram of a video decoder (310) according to an embodiment disclosed in the present application. The video decoder (310) may be provided in the electronic device (330). The electronic device (330) may include a receiver (331) (such as a receiving circuit). The video decoder (310) may be used to replace Figure 2 the video decoder (210) in the

[0038] A receiver (331) may receive at least one encoded video sequence to be decoded by a video decoder (310). To prevent network jitter, a buffer memory (315) may be coupled between the receiver (331) and an entropy decoder / parser (320) (hereinafter referred to as "parser (320)"). The parser (320) reconstructs symbols (321) from the encoded video sequence. The categories of these symbols include information for managing the operation of the video decoder (310), and potential information for controlling a display device such as a display device (312) (e.g., a display screen). The parser (320) may extract a set of subgroup parameters for at least one subgroup of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to a group. Subgroups may include Group of Pictures (GOP), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs), and so on. The parser (320) may also extract information from the encoded video sequence, such as transform coefficients, quantizer parameter values, motion vectors, and so on. The reconstruction of the symbols (321) may involve multiple different units. Which units are involved and how they are involved may be controlled by subgroup control information parsed by the parser (320) from the encoded video sequence. For the sake of brevity, such subgroup control information flows between the parser (320) and multiple units below are not described.

[0039] The first unit is a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) receives quantized transform coefficients as symbols (321) and control information from the parser (320), including which transform mode to use, block size, quantization factor, quantization scaling matrix, and so on. The scaler / inverse transform unit (351) may output a block including sample values, and the sample values may be input into an aggregator (355).

[0040] In some cases, the output samples of the scaler / inverse transform unit (351) may belong to intra-coded blocks; that is: blocks that do not use predictive information from previously reconstructed pictures, but may use predictive information from previously reconstructed parts of the current picture. Such predictive information may be provided by an intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) generates surrounding blocks of the same size and shape as the block being reconstructed using the reconstructed information extracted from the current picture buffer (358). For example, the current picture buffer (358) buffers the partially reconstructed current picture and / or the fully reconstructed current picture. In some cases, the aggregator (355) adds the prediction information generated by the intra-picture prediction unit (352) to the output sample information provided by the scaler / inverse transform unit (351) based on each sample.

[0041] In other cases, the output samples of the scaler / inverse transform unit (351) may belong to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit (353) may access the reference picture memory (357) to extract samples for prediction. After motion compensation of the extracted samples according to the sign (321), these samples may be added by the aggregator (355) to the output of the scaler / inverse transform unit (which is referred to as residual samples or a residual signal in this case), thereby generating output sample information.

[0042] The output samples of the aggregator (355) may be employed by various loop filtering techniques in the loop filter unit (356). The output of the loop filter unit (356) may be a sample stream, which may be output to the display device (312) and stored in the reference picture memory (357) for subsequent inter-picture prediction.

[0043] Figure 4 is a block diagram of a video encoder (403) according to an embodiment disclosed in the present application. The video encoder (403) is provided in an electronic device (420). The electronic device (420) includes a transmitter (440) (e.g., a transmission circuit). The video encoder (403) can be used to replace Figure 4 the video encoder (403) in the embodiment.

[0044] The video encoder (403) may receive video samples from a video source (401). According to an embodiment, the video encoder (403) may encode and compress pictures of the source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by an application. Enforcing an appropriate encoding speed is a function of the controller (450). In some embodiments, the controller (450) controls and is functionally coupled to other functional units as described below. Parameters set by the controller (450) may include rate control related parameters (picture skipping, quantizer, λ value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc.

[0045] In some embodiments, the video encoder (403) operates in an encoding loop. As a simple description, in an embodiment, the encoding loop may include a source encoder (430) (e.g., responsible for creating symbols, such as a symbol stream, based on an input picture to be encoded and reference pictures) and a (local) decoder (433) embedded in the video encoder (403). At this time, it can be observed that any decoder technology other than parsing / entropy decoding existing in the decoder must also exist in the corresponding encoder in substantially the same functional form. For this reason, the present application focuses on decoder operations. The description of encoder technology can be simplified because encoder technology is the reverse of the decoder technology described comprehensively. More detailed descriptions are needed only in certain areas and are provided below.

[0046] During operation, in some embodiments, the source encoder (430) may perform motion-compensated predictive coding. Referring to at least one previously encoded picture designated as a "reference picture" in the video sequence, the motion-compensated predictive coding performs predictive coding on the input picture.

[0047] The local video decoder (433) may decode the encoded video data that can be designated as a reference picture. The local video decoder (433) replicates the decoding process that can be performed by the video decoder on the reference picture and may store the reconstructed reference picture in the reference picture cache (434). In this way, the video encoder (403) can locally store a copy of the reconstructed reference picture, which has the same content (in the absence of transmission errors) as the reconstructed reference picture to be obtained by the remote video decoder.

[0048] The predictor (435) may perform a prediction search for the encoding engine (432). That is, for a new picture to be encoded, the predictor (435) may search in the reference picture memory (434) for sample data (as a candidate reference pixel block) or some metadata, such as a reference picture motion vector, block shape, etc., that can be used as an appropriate prediction reference for the new picture.

[0049] The controller (450) may manage the encoding operations of the source encoder (430), including, for example, setting parameters and subgroup parameters for encoding the video data.

[0050] The outputs of all the above functional units can be entropy encoded in the entropy encoder (445). The transmitter (440) can buffer the encoded video sequence created by the entropy encoder (445) to prepare for transmission over a communication channel (460), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (440) can merge the encoded video data from the video encoder (403) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).

[0051] The controller (450) can manage the operation of the video encoder (403). During encoding, the controller (450) can assign a certain encoded picture type to each encoded picture, but this may affect the encoding techniques applicable to the corresponding picture. For example, a picture can generally be assigned to any of the following picture types: intra picture (I picture), predictive picture (P picture), bi - predictive picture (B picture), multiple predictive pictures. The source picture can generally be spatially subdivided into a plurality of sample - coded blocks, as described in further detail below.

[0052] Figure 5 is a diagram of a video encoder (503) according to another embodiment disclosed in the present application. The video encoder (503) is configured to receive sample values in a processing block (e.g., a prediction block) within a current video picture in a sequence of video pictures, and encode the processing block into an encoded picture that is part of an encoded video sequence. In this embodiment, the video encoder (503) is used to replace Figure 4 the video encoder (403) in the embodiment.

[0053] For example, the video encoder (503) receives a matrix of sample values for the processing block. The video encoder (503) uses, for example, rate - distortion (RD) optimization to determine whether to use an intra mode, an inter mode, or a bi - predictive mode to encode the processing block.

[0054] In Figure 5 the embodiment, the video encoder (503) includes an inter - frame encoder (530), an intra - frame encoder (522), a residual calculator (523), a switch (526), a residual encoder (524), a general controller (521), and an entropy encoder (525) coupled together as shown in Figure 5 the figure.

[0055] The inter-frame encoder (530) is configured to receive samples of a current block (e.g., a processing block), compare the block with at least one reference block in a reference picture (e.g., blocks in a previous picture and a subsequent picture), generate inter-frame prediction information (e.g., a description of redundancy information according to inter-frame coding techniques, a motion vector, merge mode information), and calculate an inter-frame prediction result (e.g., a predicted block) based on the inter-frame prediction information using any suitable technique. In some embodiments, the reference picture is a decoded reference picture decoded based on the encoded video information.

[0056] The intra-frame encoder (522) is configured to receive samples of a current block (e.g., a processing block), compare the block with encoded blocks in the same picture in some cases, generate quantization coefficients after transformation, and also generate intra-frame prediction information in some cases (e.g., intra-frame prediction direction information according to at least one intra-frame coding technique). In an embodiment, the intra-frame encoder (522) further calculates an intra-frame prediction result (e.g., a predicted block) based on the intra-frame prediction information and reference blocks in the same picture.

[0057] The general controller (521) is configured to determine general control data and control other components of the video encoder (503) based on the general control data. For example, it determines the prediction mode of a block and provides a control signal to the switch (526) based on the prediction mode.

[0058] The residual calculator (523) is configured to calculate the difference (residual data) between the received block and a prediction result selected from the intra-frame encoder (522) or the inter-frame encoder (530). The residual encoder (524) is configured to operate based on the residual data to encode the residual data to generate transform coefficients. The transform coefficients are then subjected to quantization processing to obtain quantized transform coefficients. In various embodiments, the video encoder (503) further includes a residual decoder (528). The residual decoder (528) is configured to perform an inverse transformation and generate decoded residual data. The decoded residual data can be appropriately used by the intra-frame encoder (722) and the inter-frame encoder (730). The entropy encoder (525) is configured to format the bitstream to produce an encoded block.

[0059] Figure 6 is a diagram of a video decoder (610) according to another embodiment disclosed in the present application. The video decoder (610) is configured to receive an encoded image as part of an encoded video sequence and decode the encoded image to generate a reconstructed picture. In an embodiment, the video decoder (610) is used to replace Figure 4 the video decoder (410) in the embodiment.

[0060] In Figure 6 the embodiment, the video decoder (610) includes as Figure 6The entropy decoder (671), inter - frame decoder (680), residual decoder (673), reconstruction module (674) and intra - frame decoder (672) coupled together as shown.

[0061] The entropy decoder (671) can be used to reconstruct certain symbols from an encoded picture, where these symbols represent the syntax elements that make up the encoded picture. The inter - frame decoder (680) is used to receive inter - frame prediction information and generate an inter - frame prediction result based on the inter - frame prediction information. The intra - frame decoder (672) is used to receive intra - frame prediction information and generate a prediction result based on the intra - frame prediction information. The residual decoder (673) is used to perform inverse quantization to extract the de - quantized transform coefficients and process the de - quantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The reconstruction module (674) is used to combine in the spatial domain the residual output by the residual decoder (673) with the prediction result (which can be output by an inter - frame prediction module or an intra - frame prediction module) to form a reconstructed block, and the reconstructed block can be part of a reconstructed picture, and the reconstructed picture in turn can be part of a reconstructed video. It should be noted that other suitable operations such as de - blocking operations can be performed to improve the visual quality.

[0062] It should be noted that any suitable technology can be used to implement the video encoders (203), (403) and (503) and the video decoders (210), (310) and (610). In an embodiment, at least one integrated circuit can be used to implement the video encoders (203), (403) and (503) and the video decoders (210), (310) and (610). In another embodiment, at least one processor executing software instructions can be used to implement the video encoders (203), (403) and (503) and the video decoders (210), (310) and (610).

[0063] Turning to block partitioning for encoding and decoding, the general partitioning can start from a basic block and can follow a predefined set of rules, a specific pattern, a partitioning tree, or any partitioning structure or scheme. The partitioning can be hierarchical and recursive. After dividing or partitioning the basic block following any example partitioning procedure or other procedures described below or a combination thereof, a final set of partitions or encoded blocks can be obtained. Each of these partitions can be at one of the various partitioning levels in a partitioning hierarchy and can have various shapes. Each of the partitions can be referred to as an encoded block (CB). For the various example partitioning implementations described further below, each resulting CB can have any allowed size and partitioning levels. Such partitions are called encoded blocks because they can form units on which some basic encoding / decoding decisions can be made and the encoding / decoding parameters can be optimized, determined, and signaled in the encoded video bitstream. The highest or deepest level in the final partitioning represents the depth of the encoded block partitioning structure of the tree. The encoded blocks can be luminance encoded blocks or chrominance encoded blocks. The CB tree structure for each color can be referred to as an encoded block tree (CBT). The encoded blocks for all color channels can be collectively referred to as a coding unit (CU). The hierarchical structures for all color channels can be collectively referred to as a coding tree unit (CTU). The partitioning patterns or structures for the various color channels in a CTU can be the same or can be different.

[0064] In some embodiments, the partitioning tree scheme or structure for the luminance and chrominance channels may not have to be the same. In other words, the luminance and chrominance channels can have independent coding tree structures or patterns. Additionally, whether the luminance and chrominance channels use the same or different coding partitioning tree structures and the actual coding partitioning tree structure to be used can depend on whether the slice being encoded is a P-slice, a B-slice, or an I-slice. For example, for an I-slice, the chrominance channel and the luminance channel can have independent coding partitioning tree structures or coding partitioning tree structure patterns, while for a P-slice or a B-slice, the luminance and chrominance channels can share the same coding partitioning tree scheme. When applying an independent coding partitioning tree structure or pattern, the luminance channel can be partitioned into CBs by one coding partitioning tree structure and the chrominance channel can be partitioned into chrominance CBs by another coding partitioning tree structure.

[0065] Figure 7 An example of a predefined 10-way partitioning structure / pattern is shown, which allows recursive partitioning to form a partitioning tree. The root block can start at a predefined level (e.g., starting from a basic block at the 128×128 or 64×64 level). Figure 7 Example partitioning structures include various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. In some example embodiments, further subdivision is not allowed Figure 7Rectangular partitions. The coding tree depth can be further defined to indicate the depth of the split from the root node or root block. For example, the coding tree depth of the root node or root block can be set to 0, and after the root block is further split once, as Figure 7 shown, the coding tree depth is increased by 1. In some embodiments, only all square partitions in 710 can be allowed to follow Figure 7 the pattern and be recursively partitioned to the next level of the partition tree.

[0066] In some other example embodiments for coding block partitioning, a quadtree structure can be used. This quadtree splitting can be applied hierarchically and recursively to any square-shaped partition. Whether a basic block or an intermediate block or partition is further quadtree split can be applied to various local characteristics of the basic block or intermediate block / partition.

[0067] In yet some other examples, a ternary partitioning scheme can be used to partition a basic block or any intermediate block, as Figure 8 shown. The ternary pattern can be implemented vertically, as shown in 802, or horizontally, as shown in 804. Although Figure 8 the example split ratio in is shown as 1:2:1, other ratios can be predefined. In some embodiments, two or more different ratios can be predefined. In some embodiments, the width and height of the partitions of the example ternary tree are always powers of 2 to avoid additional transforms.

[0068] The above partitioning schemes can be combined in any way at different partitioning levels. As an example, the above quadtree and binary partitioning schemes can be combined to partition a basic block into a quadtree-binary tree (QTBT) structure. In such a scheme, a basic block or an intermediate block / partition can be quadtree split or binary split, if specified, subject to a predefined set of conditions. Figure 9Specific examples are illustrated in which a basic block is first divided into four partitions by a quadtree, as shown at 902, 904, 906, and 908. Thereafter, each of the resulting partitions is divided into four additional partitions (such as 908) by a quadtree at the next level, or is divided binary into two additional partitions (horizontally or vertically, such as 902 or 906, for example both are symmetric), or is non - divided (such as 904). For square - shaped partitions, binary or quadtree division can be allowed recursively, as shown by the entire example partition pattern of 910 and the corresponding tree structure / representation in 920, where solid lines represent quadtree division and dashed lines represent binary division. A flag can be used for each binary - division node (non - leaf binary partition) to indicate whether the binary division is horizontal or vertical. For example, as shown in 920 and consistent with the partition structure of 910, the flag "0" can represent a horizontal binary division, and the flag "1" can represent a vertical binary division. For quadtree - divided partitions, no indication of the division type is needed because a quadtree always divides a block or partition horizontally and vertically to produce 4 sub - blocks / partitions of equal size. In some embodiments, the flag "1" can represent a horizontal binary division, and the flag "0" can represent a vertical binary division.

[0069] In some example embodiments of QTBT, the quadtree and binary - division rule sets can be represented by the following predefined parameters and their associated corresponding functions:

[0070] CTU size: The size of the root node of the quadtree (the size of the basic block)

[0071] MinQTSize: The minimum allowed quadtree leaf - node size

[0072] MaxBTSize: The maximum allowed binary - tree root - node size

[0073] MaxBTDepth: The maximum allowed binary - tree depth

[0074] MinBTSize: The minimum allowed binary - tree leaf - node size

[0075] In some example embodiments of the QTBT partition structure, the CTU size can be set to 128×128 luma samples with two corresponding 64×64 chroma sample blocks (when considering and using example chroma subsampling), MinQTSize can be set to 16×16, MaxBTSize can be set to 64×64, MinBTSize (for both width and height) can be set to 4×4, and MaxBTDepth can be set to 4. Quadtree partitioning can be first applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes can have sizes ranging from the minimum size allowed for them, which is 16×16 (i.e., MinQTSize), to 128×128 (i.e., the CTU size). If a node is 128×128, it will not be first split by the binary tree because that size exceeds MaxBTSize (i.e., 64×64). Otherwise, nodes that do not exceed MaxBTSize can be partitioned by the binary tree. In Figure 9 the example of Figure 9 , the basic block is 128×128. According to a predefined set of rules, the basic block can only be split by the quadtree. The basic block has a partition depth of 0. Each of the resulting four partitions is 64×64, which does not exceed MaxBTSize, and can be further split by the quadtree or the binary tree at level 1. This process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), further splitting can be disregarded. When a binary tree node has a width equal to MinBTSize (i.e., 4), further horizontal splitting can be disregarded. Similarly, when a binary tree node has a height equal to MinBTSize, further vertical splitting is not considered.

[0076] In some example embodiments, the above QTBT scheme can be configured to support the flexibility of having the same QTBT structure for luma and chroma or independent QTBT structures. For example, for P slices and B slices, the luma and chroma CTBs in a CTU can share the same QTBT structure. However, for I slices, the luma CTB can be partitioned into CUs by one QTBT structure, and the chroma CTB can be partitioned into chroma CUs by another QTBT structure. This means that CUs can be used to refer to different color channels in I slices. For example, an I slice can be composed of coded blocks of the luma component or coded blocks of two chroma components, and a CU in a P slice or B slice can be composed of coded blocks of all three color components.

[0077] The above various CU partitioning schemes and the further partitioning of CUs to PBs can be combined in any way. The following specific embodiments are provided as non - limiting examples.

[0078] Inter-frame prediction can be implemented, for example, in a single-reference mode or a combined-reference mode. In some embodiments, a skip flag may first be included in the bitstream of the current block (or at a higher level) to indicate whether the current block is inter-frame encoded and not skipped. If the current block is inter-frame encoded, another flag may be further included in the bitstream to signal whether the single-reference mode or the combined-reference mode is used for prediction of the current block. For the single-reference mode, one reference block may be used to generate a prediction block for the current block. For the combined-reference mode, two or more reference blocks may be used to generate a prediction block (e.g., by weighted averaging). At least one reference frame index and additionally at least one corresponding motion vector may be used to identify at least one reference block, the at least one corresponding motion vector indicating at least one shift in the position relative to the frame (e.g., in horizontal and vertical pixels) between the at least one reference block and the current block. For example, an inter-frame prediction block for the current block may be generated from a single reference block identified by a motion vector in one of the reference frames as a prediction block in the single-reference mode, while for the combined-reference mode, the prediction block may be generated by weighted averaging of two reference blocks in two reference frames indicated by two reference frame indices and two corresponding motion vectors. The at least one motion vector may be encoded in various ways and included in the bitstream.

[0079] The embodiments and implementations in this application can be used alone or in any order combination. In addition, each of the method (or embodiment), encoder, and decoder can be implemented by a processing circuit (e.g., at least one processor or at least one integrated circuit). One or more processors execute a program stored in a non-transitory computer-readable medium. In this application, the term block may be interpreted as a prediction block, an encoded block, or a coding unit (CU). In this application, the direction of a reference frame can be determined by whether the reference frame is before or after the current frame in the display order.

[0080] In some example implementations of motion vector (MV) prediction, a coordination scheme can be used to implement a common merge mode, MMVD, and some other types of MV prediction for both single-reference mode and combined-reference mode MV prediction. Various syntax elements can be used to signal the way to predict the MV of the current block. For example, for the single-reference mode, the following MV prediction modes can be signaled: NEARMV - directly use one of the motion vector prediction values (MVPs) in the list indicated by the dynamic reference list (DRL) index without using any MVD; NEWMV - use one of the motion vector prediction values (MVPs) in the list signaled by the DRL index as a reference and apply an increment to the MVP (e.g., use MVD); and GLOBALMV - use a motion vector based on frame-level global motion parameters.

[0081] Similarly, for the composite reference frame - based inter - prediction mode using two reference frames corresponding to two MVs to be predicted, the following MV prediction modes can be signaled: NEAR_NEARMV - For each of the two MVs to be predicted, use one of the motion vector prediction values (MVPs) in the list signaled by the DRL without using the MVD. NEAR_NEWMV - For predicting the first of the two motion vectors, use one of the motion vector prediction values (MVPs) in the list signaled by the DRL as the reference MV without using the MVD; for predicting the second of the two motion vectors, use one of the motion vector prediction values (MVPs) in the list signaled by the DRL as the reference MV combined with an additionally signaled delta MV (MVD). NEW_NEARMV - For predicting the second of the two motion vectors, use one of the motion vector prediction values (MVPs) in the list signaled by the DRL as the reference MV without using the MVD; for predicting the first of the two motion vectors, use one of the motion vector prediction values (MVPs) in the list signaled by the DRL as the reference MV combined with an additionally signaled delta MV (MVD). NEW_NEWMV - Use one of the motion vector prediction values (MVPs) in the list signaled by the DRL as the reference MV and use it in combination with an additionally signaled delta MV to predict each of the two MVs. GLOBAL_GLOBALMV - Use the MV from each reference based on the frame - level global motion parameters of each reference.

[0082] Thus, the term "NEAR" above refers to MV prediction using a reference MV without any MVD as a general merge mode, while the term "NEW" refers to MV prediction involving using a reference MV and using a signaled or derived MVD to offset the reference MV as in the MMVD mode. For composite inter - prediction, both the above - mentioned reference base motion vectors and motion vector deltas can generally be different or independent between the two references or two MVDs, even though the two MVDs can be related, for example, and this correlation can be exploited to reduce the amount of information needed to signal the two motion vector deltas. To exploit this correlation, joint signaling of the two MVDs can be implemented and indicated in the bitstream, as described in further detail below.

[0083] In some example implementations of MVD, a predefined pixel resolution of the MVD may be allowed. For example, 1 / 8 pixel motion vector precision (or accuracy) may be allowed. The MVDs described above in the various MV prediction modes may be constructed and signaled in various ways. In some implementations, various syntax elements may be used to signal one or more motion vector differences in reference frame list 0 or list 1 as described above.

[0084] For example, a syntax element called "mv_joint" may specify which components of the associated motion vector difference are non-zero. For example, mv_joint has the following values: 0 may indicate that there is no non-zero MVD along the horizontal or vertical direction; 1 may indicate that there is a non-zero MVD only along the horizontal direction; 2 may indicate that there is a non-zero MVD only along the vertical direction; and / or 3 may indicate that there is a non-zero MVD along both the horizontal and vertical directions.

[0085] When the "mv_joint" syntax element for MVD signals that there are no non-zero MVD components, further MVD information may not be signaled. However, if the "mv_joint" syntax signals the presence of one or two non-zero components, additional syntax elements may be signaled for each of the non-zero MVD components, as described below.

[0086] For example, a syntax element called "mv_sign" may be used to further specify whether the corresponding motion vector difference component is positive or negative.

[0087] For another example, a syntax element called "mv_class" may be used to specify the class of the motion vector difference within a predefined set of classes for the corresponding non-zero MVD component. For example, the predefined classes for the motion vector difference may be used to divide the continuous magnitude space of the motion vector difference into non-overlapping class ranges. Thus, the signaled MVD class indicates the magnitude range of the corresponding MVD component. In some implementations, higher classes may correspond to motion vector differences with larger magnitude ranges.

[0088] In some other examples, a syntax element called "mv_bit" can be further used to specify the integer part of the offset between a non-zero motion vector difference component and the starting amplitude of a correspondingly signaled MV class amplitude range. In some other examples, a syntax element called "mv_fr" can be further used to specify the first 2 fractional bits of the motion vector difference of the corresponding non-zero MVD component, while a syntax element called "mv_hp" can be used to specify the third fractional bit (high-resolution bit) of the motion vector difference of the corresponding non-zero MVD component. The two-bit "mv_fr" basically provides 1 / 4 pixel MVD resolution, and the "mv_hp" bit can further provide 1 / 8 pixel resolution. In some other implementations, more than one "mv_hp" bit can be used to provide a finer MVD pixel resolution than 1 / 8 pixel. In some example implementations, additional flags can be signaled at one or more levels in each hierarchy to indicate whether 1 / 8 pixel or higher MVD resolution is supported. If the MVD resolution is not applied to a particular coding unit, the above syntax elements for the corresponding unsupported MVD resolution may not be signaled.

[0089] In some example implementations, in bi-directional prediction with CU-level weighting (BCW), a bi-directional prediction signal can be generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In some other implementations, the bi-directional prediction mode can be extended beyond simple averaging to allow weighted averaging of the two prediction signals. For example, five weights can be allowed in weighted-average bi-directional prediction, and when w equals 4, equal weighting factors are used to perform weighted averaging on the two prediction samples. For each bi-directionally predicted CU, the weight w can be determined in one of two ways: 1) for non-merged CUs, the weight index is signaled after the motion vector difference; and / or 2) for merged CUs, the weight index is inferred from adjacent blocks based on the merge candidate index. BCW can only be applied to CUs with 256 or more luma samples (i.e., CU width multiplied by CU height is greater than or equal to 256). For low-delay pictures, all 5 weights are used. For non-low-delay pictures, only 3 weights (w ∈ {3, 4, 5}) are used.

[0090] In some implementations, the prediction mode (which may be referred to as the first prediction mode in various embodiments) can be defined as copying the best prediction block from the reconstructed portion of the current frame, whose template matches the template of the current block (or is referred to as the current template). The template consists of a set of adjacent reconstructed samples. For example, the top and left adjacent reconstructed samples forming an L shape, or only the top adjacent reconstructed samples, or only the left adjacent reconstructed samples. For a predefined search range, the encoder searches the reconstructed portion of the current frame for the template that is most similar to the current template and uses the corresponding block as the prediction block. The similar template is identified by measuring the cost value between the reference template (the template of the candidate prediction block) and the template of the current block (e.g., by using the sum of absolute differences (SAD) error or the sum of squared errors (SSE) cost). Then, the encoder signals the use of this mode, and the same prediction operation is performed on the decoder side.

[0091] In some implementations, the reference Figure 10 , a prediction signal can be generated by matching the template of the current block (1050) with another block (e.g., the matching block or the prediction block, 1060) in a predefined search region. The search region can include: R1(1010): the current coding tree unit (CTU); R2(1020): the upper left CTU; R3(1030): the upper CTU; and R4(1040): the left CTU. In some implementations, the sum of absolute differences (SAD) is used as the cost function. In each region, the decoder searches for the template with the minimum SAD or minimum SSE relative to the current template and uses the corresponding block as the prediction block. In some implementations, the vector from the current block to the prediction block can be referred to as the block vector (BV). In some implementations, the BV can be different from the generally defined motion vector: the block vector corresponds to a prediction block within the same frame as the current block, and / or the block vector corresponds to a prediction block that is relatively farther from the current block than the prediction block of the motion vector.

[0092] In some implementations, the dimensions of all regions (SearchRange_w, SearchRange_h) are set to be proportional to the block dimensions (BlkW, BlkH) to have a fixed number of SAD comparisons per pixel. SearchRange_w is the width of the search range, and SearchRange_h is the height of the search range. BlkW is the width of the block, and BlkH is the height of the block. In some implementations, SearchRange_w = a * BlkW, and SearchRange_h = a * BlkH, where a is a constant that controls the gain / complexity trade-off. A non-limiting example value of a is equal to 5.

[0093] In some implementations, the first prediction mode may have some variations. In one variation, a candidate list of the best N block vectors (BVs) is maintained for the current coding block. In some implementations, N is a non - negative value, such as 15. Each BV_i corresponds to a predictor (Pred_i), where the range of i is [0, N - 1]. In some implementations, the final predictor may be a fusion of multiple predictors Pred_i. In some implementations, the weight of each predictor is derived based on a function, where the larger the error (e.g., SAD or SSE) generated by the predictor, the smaller the weight of the predictor.

[0094] In another variation, the search range may have a minimum search range. For example, SearchRange_w is defined as max(a*BlkW, 64), and SearchRange_h is defined as max(a*BlkH, 64). The search range can also be determined / limited based on the CTU size.

[0095] In another variation, the search for the most similar template is performed in several stages: in the first stage, subsampling (e.g., the central sample 1110 within a 3x3 sample) is used to perform the search at the granularity of 3x3 samples (1120), and the block vector with the best cost function is stored as the starting point for the next stage, as Figure 11A shown; in the second stage, a refined search is performed based on the best BV (e.g., 1160) in the first stage, and its surrounding samples are searched within a 3x3 window, as Figure 11B shown; and / or for an optional stage, a further refined search can be performed at the sub - pixel level.

[0096] In some implementations, another prediction mode (which may be referred to as the second prediction mode in various implementations) may use a vector to indicate the prediction from a previously decoded region of the same frame, where the vector may be referred to as a block vector (BV). The BV can be signaled in the bitstream with integer or sub - pixel sample precision. The prediction process in the second prediction mode is similar to the prediction process in the inter - frame prediction mode, with the main difference being that in the second prediction mode, a prediction block is formed from the current frame, while in inter - frame prediction, after applying the loop filter, a prediction block is formed from the reconstructed samples of the previously encoded frame.

[0097] In some implementations, when encoding the current block, a flag indicating whether to use the second prediction mode is signaled first. When the flag indicates that the second prediction mode is used to predict the current block, the BV difference is calculated by subtracting the value of the predicted BV from the value of the current BV, and the predicted BV is derived from the BV used by the encoded block.

[0098] This application describes various embodiments based on linear (or non - linear) block - level adaptive weighted prediction. Part or all of the proposed adaptive weighted prediction design can be used in many existing codecs mentioned in the background section.

[0099] In some implementations, block - adaptive weighted prediction (BAWP) can include block - level weighted prediction to model the local illumination change between a current block and its predicted block based on the local illumination change between the current block template (or causal samples of the current block) and the reference block template. Figure 12 The template of the current block (1210) (or the current template, 1212) and the template of the reference block (1220) (or the reference template, 1222) are illustrated. The reference block can be indicated or determined by a motion vector (MV, 1230). The current block can be in the current picture (or current frame), and the reference block can be in the reference picture (or reference frame). In some implementations, the function can be a linear function. The parameters of the function can be represented by a scaling factor α and an offset β, which form a linear equation, i.e., α*p[x]+β to compensate for the illumination change, where p[x] is the reference sample pointed to by the MV at position x on the reference picture. In some implementations, α and β can be derived based on the current block template and the reference block template, and they do not require signaling overhead. In some implementations, a BAWP flag can be signaled for a single inter - prediction mode to indicate the use of BAWP. BAWP can be applied to blocks with a size greater than or equal to 8x8 and encoded in a single inter - prediction mode. In some implementations, BAWP can be applied only to the luminance component. In some implementations, BAWP can be referred to as local illumination compensation (LIC).

[0100] In some implementations, when a block is encoded in the block - adaptive weighted prediction mode, a flag can be signaled to indicate whether to use explicit signaling or implicit signaling for the weighted - prediction scaling factor. In some implementations, when a block is encoded in the block - adaptive weighted prediction mode, explicit signaling of the weighted - prediction scaling factor can be directly adopted. In some implementations, when implicit signaling of the BAWP scaling factor is adopted, the weighted - prediction scaling factor and the offset value are derived from the linear equation based on the template of the current block and the template of the reference block pointed to by the motion vector of the current block.

[0101] In some implementations of explicit signaling with BAWP, when predicting a current block from its reference block using a linear function with a scaling factor α and an offset β, the selection / value of the scaling factor α and / or the offset β can be signaled into the bitstream and parsed at the decoder side to reconstruct the predicted block. The reference block is specified by a motion vector associated with the current block. All supported values of the scaling factor α are stored in a predefined lookup table, and the index of the scaling factor in the lookup table can be signaled in the bitstream and parsed at the decoder side. The offset value β can be derived from a linear equation between the reference block and the current block. The offset value β is set to (cur_template_mean – α * ref_template_mean), where cur_template_mean indicates the average of the samples in the template of the current block, and ref_template_mean indicates the average of the samples in the template of the reference block.

[0102] In some implementations, another prediction mode, which can be referred to as the third prediction mode, can be an encoding mode that inherits the motion vectors of adjacent blocks. In some implementations, another prediction mode, which can be referred to as the fourth prediction mode, can be an encoding mode that signals the motion vector difference relative to a motion vector prediction value selected from spatial or temporal adjacent blocks. In some implementations, another prediction mode, which can be referred to as the fifth prediction mode, can be an encoding mode that signals the motion vector difference relative to a motion vector prediction value selected from spatial or temporal adjacent blocks and implicitly determines the precision of the motion vector based on the magnitude of the motion vector.

[0103] In some implementations, the distribution of the supported scaling factors may vary depending on the encoded information, and thus there may be room for further improving the signaling of the scaling factors. In this application, the lookup table can also be referred to as a list, and thus, the lookup table of scaling factor candidates can be referred to as the list of scaling factor candidates.

[0104] In various embodiments of this application, for simplicity of description, mode 1 can be referred to as an encoding mode that inherits the motion vectors of adjacent blocks; mode 2 can be referred to as an encoding mode that signals the motion vector difference relative to a motion vector prediction value selected from spatial or temporal adjacent blocks or a given derived motion vector (such as a global motion vector); mode 3 can be referred to as an encoding mode that signals the motion vector difference relative to a motion vector prediction value selected from spatial or temporal adjacent blocks or a given derived motion vector (such as a global motion vector) and implicitly determines the precision of the motion vector based on the magnitude of the motion vector.

[0105] In some implementations, for a current block (as a non-limiting example, located at the boundary of a block (e.g., a superblock)), a reference cue index indicates an adjacent reference line when it is zero; or a non-zero (or non-adjacent) reference line for intra prediction of the current encoded block when it is a non-zero integer. Refer to Figure 13 In a non-limiting example of , an encoded block (also referred to as an encoded block or a coded block) (1302) is positioned at the top boundary (1330) and the left boundary (1340) of a block (e.g., a superblock). The top boundary (1330) and the left boundary (1340) of the superblock may be indicated by thick lines, as Figure 13 shown.

[0106] In some implementations, in its top direction, the current encoded block may have a top adjacent reference line with an index of zero (or referred to as the top closest adjacent reference line, or zero adjacent reference line) (1310), and one or more top non-adjacent reference lines (or referred to as top non-zero adjacent reference lines with non-zero indices) (1304, 1306, and 1308). For example, the first top non-adjacent reference line (1308) may have a reference cue index of 1, the second top non-adjacent reference line (1306) may have a reference cue index of 2, and / or the third top non-adjacent reference line (1304) may have a reference cue index of 3.

[0107] In some implementations, similarly, in its left direction, the current encoded block may have a left adjacent reference line with an index of zero (or referred to as the left closest adjacent reference line, or zero adjacent reference line) (1318), and one or more left non-adjacent reference lines (or referred to as left non-zero adjacent reference lines with non-zero indices) (1312, 1314, and 1316). For example, the first left non-adjacent reference line (1316) may have a reference cue index of 1, the second left non-adjacent reference line (1314) may have a reference cue index of 2, and / or the third left non-adjacent reference line (1312) may have a reference cue index of 3.

[0108] In some implementations, there are some problems or challenges associated with the BAWP method, particularly how to improve the flexibility and / or efficiency of determining the scaling factor for BAWP that utilizes a block vector. This application describes various embodiments for enhancing BAWP using a block vector, solving at least one of the problems or challenges discussed above, improving encoding / decoding efficiency, and advancing video codec technology.

[0109] The various embodiments and / or implementations described in this application can be performed individually or combined in any order. Additionally, each method (or embodiment), encoder, and decoder can be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). One or more processors execute a program stored in a non-volatile computer-readable medium. In this application, the term block can be interpreted as a prediction block, a coding block, or a coding unit (CU). In this application, the direction of a reference frame can be determined by whether the reference frame is before or after the current frame in the display order.

[0110] Figure 14 FIG. 1400 is a flowchart of an exemplary method that follows the above implementation principle for decoding a current block of a current frame in an encoded video bitstream, and this decoding can be performed by an electronic device (e.g., a decoder). The exemplary decoding method flow starts from S1401 and can include some or all of the following steps: S1410, receiving the encoded video bitstream; S1420, identifying, from the encoded video bitstream, a block vector corresponding to a reference block associated with the current block of the current frame; S1430, determining a scaling factor based on the syntax explicitly signaled in the encoded video bitstream; S1440, generating a prediction block based on the reference block according to a linear equation associated with the scaling factor; and / or S1450, reconstructing the current block by the device based on the prediction block. The example method can stop at S1499.

[0111] In any part or combination of the above implementations, the reference block is pointed to by a block vector relative to the current block within the current frame.

[0112] In any part or combination of the above implementations, determining the scaling factor includes: determining the scaling factor based on a predefined look-up table, where the predefined look-up table for the reference block within the current frame in which the current block is located is different from a second look-up table for the reference block in a frame different from the current frame.

[0113] In any part or combination of the above implementations, the predefined look-up table and the second look-up table differ in at least one of the following aspects: the size of the predefined look-up table is larger than the size of the second look-up table, and / or, the predefined look-up table has higher precision than the second look-up table.

[0114] In any part or combination of the above implementations, determining the scaling factor includes: determining the scaling factor based on a predefined look-up table, where the predefined look-up table for the reference block within the current frame in which the current block is located is the same as a third look-up table for an inter-frame prediction mode.

[0115] In any part or combination of the above implementations, the current block has a first prediction mode, in which the best prediction block is copied from the reconstructed part of the current frame, and the template of the best prediction block matches the template of the current block.

[0116] In any part or combination of the above implementations, the current block has a second prediction mode, in which a block vector is used to indicate a prediction block from a previously decoded region of the current frame.

[0117] In any part or combination of the above implementations, the method further includes extracting a first flag from the encoded video bitstream, the first flag indicating whether block-level weighted prediction for the current block is explicitly signaled, wherein the context for signaling the first flag is based on whether the current block is encoded with the block vector.

[0118] In any part or combination of the above implementations, determining a scaling factor includes: extracting a second flag from the encoded video bitstream, the second flag indicating whether the scaling factor for the current block is explicitly signaled; and / or, in response to the second flag indicating that the scaling factor is explicitly signaled, determining the scaling factor based on the syntax explicitly signaled in the encoded video bitstream.

[0119] In any part or combination of the above implementations, the current block has multiple linear models associated with multiple scaling factors; and / or, in the encoded video bitstream, at least one of the multiple scaling factors is explicitly signaled.

[0120] In any part or combination of the above implementations, the method may further include extracting a third flag from the encoded video bitstream, the third flag indicating whether block-level weighted prediction for the current block is explicitly signaled, wherein the context for signaling the third flag is based on whether encoding of one or more adjacent blocks of the current block is by explicitly signaling block-level weighted prediction for the one or more adjacent blocks.

[0121] In any part or combination of the above implementations, the method may further include, in response to explicitly signaling the scaling factor for the current block and the block vector, determining an adjacent reference line based on a second syntax explicitly signaled in the encoded video bitstream, the adjacent reference line being used to derive an offset value for the current block.

[0122] In any part or combination of the above implementations, the method may further include, in response to explicitly signaling the scaling factor for the current block and the block vector, determining adjacent samples among multiple adjacent reference lines for deriving an offset value for the current block.

[0123] In any part or combination of the above implementations, the method may further include determining an offset value of the current block as a predefined value in response to explicitly signaling the scaling factor for the current block and the block vector.

[0124] In any part or combination of the above implementations, the method may further include determining a predefined component of the scaling factor for the current block in response to explicitly signaling the scaling factor for the current block and the block vector.

[0125] In any part or combination of the above implementations, the method may further include extracting high-level syntax from the encoded video bitstream, the high-level syntax indicating whether to explicitly signal block-level weighted prediction for the current block, wherein the high-level syntax includes one of the following: sequence-level flag, picture-level flag, slice-level flag, sub-picture-level flag, or tile-level flag.

[0126] This application describes some non-limiting examples in the following various embodiments, which are used as exemplary embodiments and do not impose any restrictions on this application. When the reference block of the current block is pointed to by a block vector within the current frame / slice / sub-picture / picture, explicit signaling of the scaling factor in block-level weighted prediction is used to generate the prediction samples of the current block.

[0127] In some embodiments, a flag may be signaled, which indicates whether multiple scaling factors at the explicitly signaled block level are used to generate the prediction samples of the current block.

[0128] In some embodiments, when the reference block of the current block is pointed to by a block vector within the current frame / slice / picture, different predefined lookup tables may be used to store the scaling factor, compared to the lookup table used to store the scaling factor when the reference block of the current block comes from a different picture or frame (i.e., inter-frame prediction).

[0129] In some implementations, when the reference block of the current block is pointed to by a block vector within the current frame / slice / picture, the size of the lookup table used to store the scaling factor is larger than the lookup table used to store the scaling factor when the reference block of the current block comes from a different picture.

[0130] For a non-limiting example, the size of the scaling factor lookup table for block-adaptive weighted prediction with a block vector is twice the size of the scaling factor lookup table for block-adaptive weighted prediction with a motion vector pointing to a block in another picture.

[0131] In some implementations, when the reference block of the current block is pointed to by a block vector within the current frame / slice / picture, a lookup table with a higher-precision scaling factor is used, compared to the lookup table used to store the scaling factor when the reference block of the current block comes from a different picture or frame (i.e., inter-frame prediction).

[0132] For a non - limiting example, when the reference block of the current block is pointed to by a block vector within the current frame / slice / picture, a lookup table with a scaling factor of 1 / 32 precision is used, as compared to the 1 / 16 scaling factor lookup table used when the reference block of the current block comes from a different picture (i.e., inter - frame prediction).

[0133] For another non - limiting example, when the reference block of the current block is pointed to by a block vector within the current frame / slice / picture, an example of the scaling factor lookup table is {30 / 32, 31 / 32, 33 / 32, 34 / 32}.

[0134] In some embodiments, when a block vector within the current frame points to the reference block of the current block, the lookup table used to store the scaling factor is the same as the lookup table for one of the inter - frame prediction modes.

[0135] In some implementations, when a block vector within the current frame points to the reference block of the current block, the lookup table used to store the scaling factor is the same as the lookup table for the third prediction mode.

[0136] In some embodiments, the first prediction mode and the second prediction mode are two examples of using a block vector to point to a reference block within the current frame.

[0137] In some implementations, a set of the best N BVs is maintained for the current coded block. These BVs correspond to different templates, so different predictors for BAWP Pred_i can be derived, where i ranges from [0, N - 1]. The final prediction value of BAWP is obtained based on the fusion of some or all of the BAWP prediction values Pred_i, and the fusion can be a simple average or a weighted average. For a non - limiting example, a set of the best 15 BVs is maintained, and 5 of the 15 BVs are used to generate the final BAWP prediction value (e.g., by averaging).

[0138] In some embodiments, when multiple linear models are used in the current block, the scaling factor of at least one linear model is signaled explicitly.

[0139] In certain implementations, which linear model's scaling factor is signaled explicitly depends on the encoded / available information, including but not limited to: the number of samples assigned to each group, the block size, the value or magnitude of the BV.

[0140] In some implementations, the samples in the current block are divided into multiple groups according to a predefined criterion, and a separate scaling factor can be applied to each group. The group with the least number of samples is determined to use the explicitly signaled scaling factor.

[0141] In some implementations, when the block size is greater than a threshold, the parameters of the linear model can be implicitly derived; otherwise, when the block size is not greater than the threshold, the parameters of the linear model are signaled explicitly.

[0142] In some implementations, multiple linear models can be applied to the samples in the current block. A weighted average of the predictors generated by each linear model is obtained to get the final predicted sample of the current block.

[0143] In some embodiments, the context for signaling a flag depends on whether a block is coded with a block vector, and this flag indicates explicit signaling of block-level weighted prediction.

[0144] In some implementations, a separate context is used for signaling a flag that indicates whether explicit signaling of block-level weighted prediction is performed when a block is coded with a block vector.

[0145] In some embodiments, when the current block is coded by the first method or the second method, the context for signaling a flag depends on whether multiple adjacent blocks are coded with explicitly signaled block-level weighted prediction, and this flag indicates explicit signaling of block-level weighted prediction.

[0146] In some embodiments, when an explicit signaling of a scaling factor for block-level weighted prediction is used for a block coded with a block vector, a flag can be used to indicate which adjacent reference line is used to derive the offset value of the current block.

[0147] In some embodiments, when an explicit signaling of a scaling factor for block-level weighted prediction is used for a block coded with a block vector, adjacent samples among multiple adjacent reference lines can be used to derive the offset value of the current block.

[0148] In some embodiments, when an explicit signaling of a scaling factor for block-level weighted prediction is used for a block coded with a block vector, the offset value β is set to a predefined value, such as zero.

[0149] In some embodiments, when an explicit signaling of a scaling factor for block-level weighted prediction is used for a block coded with a block vector, the method is only applied to a specific component of the current block, such as the luminance component.

[0150] In some embodiments, it can be signaled with high-level syntax whether to explicitly signal block-level weighted prediction and / or implicitly derive linear model parameters, including but not limited to sequence / picture / slice / sub-picture / tile-level flags.

[0151] Various embodiments in the present application may include methods for encoding a current block into a video bitstream, which are executed by an encoder and include inverse processes of any part or all of the processes described as a decoder. Various embodiments in the present application may include methods for encoding a current block of a streaming video, which are executed by one or more electronic devices (such as a streaming media player) and include any part or all of the processes of a decoder and / or any part or all of the processes described as an encoder.

[0152] The above operations can be combined or arranged in any quantity or order as needed. Two or more steps and / or operations can be executed in parallel. The embodiments and implementations in the present application can be used alone or in any combination. The steps in one embodiment / method can be split to form multiple sub-methods, and each of the sub-methods can be independent of other steps in the embodiment and can form an independent solution. In addition, each of the method (or embodiment), encoder, and decoder can be implemented by a processing circuit (such as at least one processor or at least one integrated circuit). In one example, at least one processor executes a program stored in a non-volatile computer-readable medium. The embodiments in the present application can be applied to luminance blocks or chrominance blocks. The term "block" can be interpreted as a prediction block, a coding block, or a coding unit, i.e., a CU. The term "block" can also be used here to refer to a transform block. In the following items, when referring to the block size, it can refer to the block width or height, or the maximum of the width and height, or the minimum of the width and height, or the area size (width * height), or the aspect ratio of the block (width: height, or height: width).

[0153] The above technologies can be implemented as computer software by computer-readable instructions and physically stored in at least one computer-readable storage medium. For example, Figure 15 FIG. shows a computer system (1800) suitable for implementing certain embodiments of the disclosed subject matter.

[0154] The computer software can be encoded by any suitable machine code or computer language, and code including instructions is created through mechanisms such as assembly, compilation, and linking. The instructions can be directly executed by at least one computer central processing unit (CPU), graphics processing unit (GPU), etc., or executed through methods such as decoding and microcode.

[0155] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablets, servers, smartphones, gaming devices, Internet of Things devices, etc.

[0156] Figure 15The components shown for the computer system (1800) are exemplary in nature and are not intended to impose any limitation on the scope of use or functionality of the computer software implementing the embodiments of the present application. Nor should the configuration of the components be construed as having any dependency or requirement on any one component or combination thereof shown in the exemplary embodiments of the computer system (1800).

[0157] The computer system (1800) may include certain human-machine interface input devices. The human-machine interface input devices may include at least one of the following (only one is shown): keyboard (1801), mouse (1802), touchpad (1803), touch screen (1810), data glove (not shown), joystick (1805), microphone (1806), scanner (1807), camera (1808).

[0158] The computer system (1800) may also include certain human-machine interface output devices. Such human-machine interface output devices may stimulate the senses of at least one human user through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include tactile output devices (e.g., tactile feedback through the touch screen (1810), data glove (not shown), or joystick (1805), but there may also be tactile feedback devices that do not serve as input devices), audio output devices (e.g., speakers (1809), headphones (not shown)), visual output devices (e.g., screens (1810) including cathode ray tube screens, liquid crystal screens, plasma screens, organic light-emitting diode screens, each of which may or may not have touch screen input functionality, each of which may or may not have tactile feedback functionality - some of which may output two-dimensional visual output or output above three dimensions through means such as stereoscopic picture output; virtual reality glasses (not shown), holographic displays, and smoke boxes (not shown)), and printers (not shown).

[0159] The computer system (1800) may also include human-accessible storage devices and their associated media, such as optical media including high-density read-only / rewritable optical discs with CD / DVD (CD / DVD ROM / RW) (1820) or similar media (1821), thumb drives (1822), removable hard disk drives or solid state drives (1823), traditional magnetic media such as tapes and floppy disks (not shown), dedicated devices based on ROM / ASIC / PLD such as security software protectors (not shown), and so on.

[0160] Those skilled in the art should also understand that the term "computer-readable storage medium" used in connection with the disclosed subject matter does not include transmission media, carrier waves, or other transient signals.

[0161] The computer system (1800) may also include an interface (1854) to at least one communication network (1855). For example, the network may be wireless, wired, or optical. The network may also be a local area network, a wide area network, a metropolitan area network, a vehicular network, and an industrial network, a real-time network, a delay-tolerant network, and so on. The network also includes local area networks such as Ethernet, wireless local area network, cellular networks (GSM, 3G, 4G, 5G, LTE, etc.), television cable or wireless wide area digital networks (including cable television, satellite television, and terrestrial broadcast television), vehicular and industrial networks (including CANBus), and so on. Some networks typically require an external network interface adapter for connection to certain general-purpose data ports or peripheral buses (1849) (e.g., the USB port of the computer system (1800)); other systems are typically integrated into the core of the computer system (1800) by connecting to the system bus as described below (e.g., an Ethernet interface is integrated into a PC computer system or a cellular network interface is integrated into a smart phone computer system). By using any of these networks, the computer system (1800) can communicate with other entities. The communication may be unidirectional, only for receiving (e.g., wireless television), unidirectional only for sending (e.g., CAN bus to certain CAN bus devices), or bidirectional, e.g., via a local or wide area digital network to other computer systems. Each of the above networks and network interfaces may use certain protocols and protocol stacks.

[0162] The above-mentioned human-machine interface device, human-accessible storage device, and network interface may be connected to the core (1820) of the computer system (1800).

[0163] The core (1820) may include at least one central processing unit (CPU) (1821), a graphics processing unit (GPU) (1842), a dedicated programmable processing unit in the form of a field-programmable gate array (FPGA) (1843), a hardware accelerator (1844) for specific tasks, a graphics adapter (1830), and so on. These devices, as well as read-only memory (ROM) (1845), random access memory (1846), internal mass storage (e.g., internal non-user-accessible hard disk drive, solid-state drive, etc.) (1847), and so on, may be connected via a system bus (1848). In some computer systems, the system bus (1848) may be accessed in the form of at least one physical plug for expansion by an additional central processing unit, graphics processing unit, and so on. Peripherals may be directly attached to the system bus (1848) of the core or connected via a peripheral bus (1849). In one example, the screen (1810) may be connected to the graphics adapter (1830). The architecture of the peripheral bus includes an external controller interface PCI, a universal serial bus USB, and so on.

[0164] Computer code may be present on the computer-readable storage medium for performing various computer-implemented operations. The medium and the computer code may be those specially designed and constructed for the purposes of this application, or they may be of the kind well-known and available to those having skill in the art of computer software.

[0165] Although the present application has described a number of exemplary embodiments, various changes, permutations and various equivalent substitutions of the embodiments are within the scope of the present application. It should thus be understood that those skilled in the art can design various systems and methods that, although not explicitly shown or described herein, embody the principles of the present application and are thus within the spirit and scope of the present application.

Claims

1. A decoding method for a current block of a current frame in an encoded video bitstream, characterized in that, The method includes: receiving, by a device, the encoded video bitstream, the device including a memory storing instructions and a processor in communication with the memory; identifying, by the device, from the encoded video bitstream, a block vector corresponding to a reference block associated with the current block of the current frame; determining, by the device, a scaling factor based on syntax explicitly signaled in the encoded video bitstream; generating, by the device, a prediction block based on the reference block according to a linear equation associated with the scaling factor; and reconstructing, by the device, the current block based on the prediction block.

2. The method according to claim 1, wherein The reference block is pointed to by a block vector relative to the current block within the current frame.

3. The method according to claim 1, characterized in that, The determining the scaling factor includes: determining the scaling factor based on a predefined look-up table, wherein a predefined look-up table for a reference block within the current frame in which the current block is located is different from a second look-up table for a reference block in a frame different from the current frame.

4. The method according to claim 3, wherein The predefined look-up table and the second look-up table differ in at least one of the following aspects: the size of the predefined look-up table is larger than the size of the second look-up table, and / or the predefined look-up table has a higher accuracy than the second look-up table.

5. The method according to claim 1, wherein The determining the scaling factor includes: determining the scaling factor based on a predefined look-up table, wherein a predefined look-up table for a reference block within the current frame in which the current block is located is the same as a third look-up table for an inter prediction mode.

6. The method according to claim 1, wherein The current block has a first prediction mode, wherein the best prediction block is copied from a reconstructed portion of the current frame, and a template of the best prediction block matches a template of the current block.

7. The method according to claim 1, characterized in that, The current block has a second prediction mode, wherein the block vector is used to indicate a prediction block from a previously decoded region of the current frame.

8. The method according to claim 1, further comprising: extracting, from the encoded video bitstream, a first flag indicating whether block-level weighted prediction of the current block is explicitly signaled, wherein a context for signaling the first flag is based on whether the current block is encoded with the block vector.

9. The method according to claim 1, characterized in that, The determining the scaling factor includes: extracting, from the encoded video bitstream, a second flag indicating whether the scaling factor of the current block is explicitly signaled; in response to the second flag indicating that the scaling factor is explicitly signaled, determining the scaling factor based on syntax explicitly signaled in the encoded video bitstream.

10. The method according to claim 1, characterized in that, The current block has multiple linear models associated with multiple scaling factors; at least one of the multiple scaling factors is explicitly signaled in the encoded video bitstream.

11. The method according to claim 1, further comprising: Extract a third flag from the encoded video bitstream, the third flag indicating whether block-level weighted prediction for the current block is signaled explicitly, wherein the context for signaling the third flag is based on whether block-level weighted prediction for one or more neighboring blocks of the current block is signaled explicitly during encoding of the one or more neighboring blocks.

12. The method according to claim 1, further comprising: In response to explicitly signaling the scaling factor for the current block and the block vector, determine a neighboring reference line based on a second syntax signaled explicitly in the encoded video bitstream, the neighboring reference line being used to derive an offset value for the current block.

13. The method according to claim 1, further comprising: In response to explicitly signaling the scaling factor for the current block and the block vector, determine neighboring samples among a plurality of neighboring reference lines, the neighboring samples being used to derive an offset value for the current block.

14. The method according to claim 1, further comprising: In response to explicitly signaling the scaling factor for the current block and the block vector, determine the offset value of the current block as a predefined value.

15. The method according to claim 1, further comprising: In response to explicitly signaling the scaling factor for the current block and the block vector, determine a predefined component of the scaling factor for the current block.

16. The method according to claim 1, further comprising: Extract high-level syntax from the encoded video bitstream, the high-level syntax indicating whether block-level weighted prediction for the current block is signaled explicitly, wherein the high-level syntax includes one of the following: a sequence-level flag, a picture-level flag, a slice-level flag, a sub-picture-level flag, or a tile-level flag.

17. A decoding apparatus for a current block of a current frame in an encoded video bitstream, characterized in that, Comprising A memory storing instructions; and A processor in communication with the memory, wherein when the processor executes the instructions, the processor is configured to cause the apparatus to perform the method according to any one of claims 1-16.

18. A non-volatile storage medium, characterized in that, For storing instructions which, when executed by a processor, are configured to cause the processor to perform the method according to any one of claims 1-16.