Block-level adaptive weighted prediction using multiple scaling factors
By adopting the block adaptive weighted prediction (BAWP) method in video encoding technology, the problem of inefficiency of block-level prediction and local lighting compensation is solved, and more efficient and accurate video encoding is achieved.
Patent Information
- Application Number
- CN202480004512.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-01-19
- Publication Date
- 2025-06-03
AI Technical Summary
Existing video encoding technologies have problems with inefficiency in block-level prediction and local lighting compensation.
Block Adaptive Weighted Prediction (BAWP) method is used to improve block-level prediction and local lighting compensation by determining scaling factors and offset parameters.
Improves the efficiency and accuracy of video encoding, reduces signaling costs, and enhances encoding gain.
Smart Images

Figure CN120092443A_ABST
Abstract
Description
Incorporation by reference
[0001] This application claims the benefit of priority to U.S. Non - Provisional Application No. 18 / 400,223, filed on December 29, 2023, which claims the benefit of priority to U.S. Provisional Application No. 63 / 540,009, filed on September 22, 2023. Each of the above - mentioned applications is incorporated herein by reference in its entirety. Technical Field
[0002] This disclosure describes a set of advanced video / stream encoding / decoding techniques. More specifically, the disclosed techniques relate to enhancements to block - level prediction. Background Art
[0003] Uncompressed digital video can include a sequence of pictures and may have specific bit - rate requirements for storage, data processing, and transmission bandwidth in streaming applications. One purpose of video encoding and video decoding can be to reduce redundancy in the uncompressed input video signal through various compression techniques. Summary of the Invention
[0004] This disclosure describes various embodiments of methods, apparatuses, and computer - readable storage media for block - level prediction enhancement and block - adaptive weighted prediction (BAWP) for modeling local illumination compensation (LIC).
[0005] According to one aspect, embodiments of the present disclosure provide a method for decoding a current block of a current frame in an encoded video bitstream. The method includes: receiving a video bitstream that includes a current block and a reference block, the reference block being used to predict the current block by using a prediction function; receiving, from the video bitstream, a first syntax element that indicates how to determine prediction parameters of the prediction function, the prediction parameters including a scaling factor, and the first syntax element indicating at least one of: whether the scaling factor is explicitly signaled in the video bitstream via a syntax element, whether the scaling factor needs to be derived from the video bitstream, and whether a first subset of the scaling factor is signaled in the video bitstream and a second subset of the scaling factor is derived from the video bitstream, the second subset and the first subset of the scaling factor being non - overlapping; determining the scaling factor based on the first syntax element; predicting a current sample in the current block based on at least one of: a prediction function including the determined scaling parameter; local illumination variation; current block template and reference block template; and reconstructing the current block based on the predicted current sample.
[0006] According to another aspect, an embodiment of the present disclosure provides an apparatus / decoder for decoding a video. The apparatus / decoder includes: a memory storing instructions; and a processor communicatively coupled to the memory. When the processor executes the instructions, the processor is configured to cause the apparatus to perform the above-described method for video decoding and / or encoding.
[0007] In another aspect, an embodiment of the present disclosure provides a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding and / or encoding, cause the computer to perform the above-described method for video decoding and / or encoding.
[0008] The above aspects and other aspects and their implementations are described in more detail in the accompanying drawings, the description, and the claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Additional features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings, in which:
[0010] Figure 1 A schematic diagram showing a simplified block diagram of a communication system (100) according to an example embodiment;
[0011] Figure 2 A schematic diagram showing a simplified block diagram of a communication system (200) according to an example embodiment;
[0012] Figure 3 A schematic diagram showing a simplified block diagram of a video decoder according to an example embodiment;
[0013] Figure 4 A schematic diagram showing a simplified block diagram of a video encoder according to an example embodiment;
[0014] Figure 5 A block diagram showing a video encoder according to another example embodiment;
[0015] Figure 6 A block diagram showing a video decoder according to another example embodiment;
[0016] Figure 7 A scheme of coding block partitioning according to an example embodiment of the present disclosure;
[0017] Figure 8 Another scheme of coding block partitioning according to an example embodiment of the present disclosure;
[0018] Figure 9 Another scheme of coding block partitioning according to an example embodiment of the present disclosure;
[0019] Figure 10 Examples of the template of the current block and the template of the reference block are shown.
[0020] Figure 11 Examples of a sample and its neighboring samples are shown.
[0021] Figure 12 An example logical flow of the method in the present disclosure is shown.
[0022] Figure 13 A schematic diagram of a computer system according to an example embodiment of the present disclosure is shown. Detailed Description
[0023] The present invention will now be described in detail below with reference to the accompanying drawings, which form a part of the present invention and illustrate specific examples of the embodiments by way of illustration. However, note that the present invention can be implemented in various different forms, and thus, the subject matter covered or claimed is intended to be construed as not limited to any of the embodiments to be described below. Also note that the present invention can be implemented as a method, device, component, or system. Therefore, the embodiments of the present invention can take, for example, the form of hardware, software, firmware, or any combination thereof.
[0024] Throughout the specification and claims, terms may have nuanced meanings that are implicit or implied in the context beyond the explicitly stated meaning. The phrase "in one embodiment" or "in some embodiments" used herein does not necessarily refer to the same embodiment, and the phrase "in another embodiment" or "in other embodiments" used herein does not necessarily refer to different embodiments. Similarly, the phrase "in one implementation" or "in some implementations" used herein does not necessarily refer to the same implementation, and the phrase "in another implementation" or "in other implementations" used herein does not necessarily refer to different implementations. For example, it is intended that the claimed subject matter includes combinations of all or part of the exemplary embodiments / implementations.
[0025] Typically, terms can be understood, at least in part, from their usage in context. For example, terms such as "and," "or," or "and / or" used herein can include a variety of meanings, at least in part, depending on the context in which such terms are used. Generally, "or" (if used in an associative list, e.g., A, B, or C) is intended to mean: A, B, and C, used herein in an inclusive sense; and A, B, or C, used herein in an exclusive sense. Additionally, at least in part depending on the context, the terms "one or more" or "at least one" used herein can be used to describe any feature, structure, or property in a singular sense, or can be used to describe a combination of features, structures, or properties in a plural sense. Similarly, terms such as "a," "an," or "the" can likewise be understood to convey a singular usage or to convey a plural usage, at least in part depending on the context. Additionally, the terms "based on" or "determined by" can be understood to not necessarily be intended to convey an exclusive set of factors, but can allow for the existence of additional factors that are not necessarily explicitly described, again at least in part depending on the context.
[0026] As Figure 1 shown, the terminal device can be implemented as a server, a personal computer, and a smart phone, but the applicability of the basic principles of the present disclosure can not be so limited. Embodiments of the present disclosure can be implemented in a desktop computer, a laptop computer, a tablet computer, a media player, a wearable computer, a dedicated video conferencing device, etc. The network (150) represents any number or type of network for transmitting encoded video data between terminal devices, including, for example, wired (wired) and / or wireless communication networks. The communication network (150) can exchange data in circuit-switched, packet-switched, and / or other types of channels. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.
[0027] As an example of an application for the disclosed subject matter, Figure 2 illustrates the placement of a video encoder and a video decoder in a video streaming environment. The disclosed subject matter can equally apply to other video applications, including, for example, video conferencing, digital television broadcasting, gaming, virtual reality, storing compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0028] As Figure 2As shown, a video streaming system may include a video capture subsystem (213), which may include a video source (201) for creating an uncompressed video picture or image stream (202), such as a digital camera device. In an example, the video picture stream (202) includes samples recorded by the digital camera device of the video source (201). The video picture stream (202) is depicted as a thick line to emphasize the high data volume when compared to the encoded video data (204) (or encoded video bitstream). The video picture stream (202) may be processed by an electronic device (220) coupled to the video source (201) and including a video encoder (203). The video encoder (203) may include hardware, software, or a combination of hardware and software to implement or carry out aspects of the disclosed subject matter described in more detail below. The encoded video data (204) (or encoded video bitstream (204)) is depicted as a thin line to emphasize the lower data volume when compared to the uncompressed video picture stream (202). The encoded video data (204) may be stored on a streaming server (205) for future use or directly stored to a downstream video device (not shown). One or more streaming client subsystems such as Figure 2 the client subsystems (206) and (208) in
[0029] Figure 3 may access the streaming server (205) to retrieve copies (207) and (209) of the encoded video data (204). The client subsystem (206) may include, for example, a video decoder (210) in an electronic device (230). The video decoder (210) decodes an incoming copy (207) of the encoded video data and creates an outgoing video picture stream (211) that is uncompressed and may be presented on a display (212) (e.g., a display screen) or other presentation device (not depicted). Figure 2 shows a block diagram of a video decoder (310) of an electronic device (330) according to any embodiment of the present disclosure below. The electronic device (330) may include a receiver (331) (e.g., receiving circuitry). The video decoder (310) may be used in place of
[0030] In the example of Figure 3As shown, a receiver (331) may receive one or more encoded video sequences from a channel (301). To prevent network jitter and / or handle playback timing, a buffer memory (315) may be provided between the receiver (331) and an entropy decoder / parser (320) (hereinafter referred to as "parser (320)"). The parser (320) may reconstruct symbols (321) from the encoded video sequences. The categories of these symbols include: information for managing the operation of a video decoder (310), and information that may be used to control a rendering device such as a display (312) (e.g., a display screen). The parser (320) may parse / entropy decode the encoded video sequences. The parser (320) may extract a set of subgroup parameters of at least one in a subgroup of pixels in the video decoder from the encoded video sequences. The subgroup may include a Group of Picture (GOP), a picture, a tile, a slice, a macroblock, a Coding Unit (CU), a block, a Transform Unit (TU), a Prediction Unit (PU), etc. The parser (320) may also extract information such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, motion vectors, etc. from the encoded video sequences. The reconstruction of the symbols (321) may involve multiple different processing or functional units. The units involved and how they are involved may be controlled by subgroup control information parsed by the parser (320) from the encoded video sequences.
[0031] The first unit may include a scaler / inverse transform unit (351). The scaler / inverse transform unit (351) may receive quantized transform coefficients as (one or more) symbols (321) and control information from the parser (320), including information indicating which inverse transform to use, block size, quantization factor / parameter, quantization scaling matrix, etc. The scaler / inverse transform unit (351) may output blocks including sample values, and these blocks may be input into an aggregator (355).
[0032] In some cases, the output samples of the scaler / inverse transform (351) may belong to an intra-coded block, i.e., a block that does not use predictive information from a previously reconstructed picture but may use predictive information from a previously reconstructed portion of the current picture. Such predictive information may be provided by the intra-picture prediction unit (352). In some cases, the intra-picture prediction unit (352) may use surrounding block information that has been reconstructed and stored in the current picture buffer (358) to generate a block having the same size and shape as the block being reconstructed. For example, the current picture buffer (358) buffers a partially reconstructed current picture and / or a fully reconstructed current picture. In some implementations, the aggregator (355) may add the prediction information that the intra-prediction unit (352) has generated to the output sample information provided by the scaler / inverse transform unit (351) on a per-sample basis.
[0033] In other cases, the output samples of the scaler / inverse transform unit (351) may belong to an inter-coded and potentially motion-compensated block. In such a case, the motion compensation prediction unit (353) may access the reference picture memory (357) based on a motion vector to obtain samples for inter-picture prediction. After motion compensating the obtained reference samples according to the sign (321) belonging to the block, these samples may be added by the aggregator (355) to the output of the scaler / inverse transform unit (351) (the output of unit 351 may be referred to as residual samples or a residual signal) to generate output sample information.
[0034] The output samples of the aggregator (355) may undergo various loop filtering techniques in a loop filter unit (356) that includes several types of loop filters. The output of the loop filter unit (356) may be a sample stream that may be output to a rendering device (312) and stored in the reference picture memory (357) for future inter-picture prediction.
[0035] Figure 4 A block diagram of a video encoder (403) according to an example embodiment of the present disclosure is shown. The video encoder (403) may be included in an electronic device (420). The electronic device (420) may also include a transmitter (440) (e.g., transmission circuitry). The video encoder (403) may be used to replace Figure 4 the video encoder (403) in the example of
[0036] The video encoder (403) may receive video samples from a video source (401). According to some example embodiments, the video encoder (403) may encode and compress pictures of the source video sequence into an encoded video sequence (443) in real time or under any other time constraints required by the application. Implementing an appropriate encoding speed constitutes a function of the rate controller (450). In some embodiments, the controller (450) may be functionally coupled to other functional units as described below and control the other functional units. The parameters set by the controller (450) may include rate control related parameters (picture skipping, quantizer, λ value of rate-distortion optimization techniques...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc.
[0037] In some example embodiments, the video encoder (403) may be configured to operate in an encoding loop. The encoding loop may include a source encoder (430) and an (in - loop) decoder (433) embedded in the video encoder (403). The decoder (433) reconstructs symbols in a manner similar to how a (remote) decoder would create sample data to create sample data, although the embedded decoder 433 processes the encoded video stream of the source encoder 430 without entropy encoding (since in the video compression techniques considered in the disclosed subject matter, any compression between the symbols and the encoded video bit stream in entropy encoding can be lossless). At this point, it can be observed that, except for parsing / entropy decoding which may only exist in the decoder, any decoder technique may also necessarily exist in the corresponding encoder in substantially the same functional form. For this reason, the disclosed subject matter may sometimes focus on decoder operations, which also applies to the decoding part of the encoder. Thus, the description of encoder techniques can be simplified because encoder techniques are the inverse process of the fully described decoder techniques. A more detailed description of the encoder is provided only in certain areas or aspects below.
[0038] In some example implementations, during operation, the source encoder (430) may perform motion - compensated predictive coding that predictively encodes an input picture by referring to one or more previously encoded pictures of the video sequence designated as "reference pictures".
[0039] The local video decoder (433) may decode the encoded video data of pictures that may be designated as reference pictures. The local video decoder (433) replicates the decoding process that a video decoder may perform on the reference pictures and may store the reconstructed reference pictures in the reference picture cache (434). In this way, the video encoder (403) may locally store a copy of the reconstructed reference pictures, which has the same content (in the absence of transmission errors) as the reconstructed reference pictures that would be obtained by a distal (remote) video decoder.
[0040] The predictor (435) can perform a prediction search for the encoding engine (432). That is, for a new picture to be encoded, the predictor (435) can search the reference picture memory (434) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, etc., which can be used as a suitable prediction reference for the new picture.
[0041] The controller (450) can manage the encoding operations of the source encoder (430), including, for example, setting parameters and subgroup parameters for encoding video data.
[0042] The outputs of all the above functional units can be entropy encoded in the entropy encoder (445). The transmitter (440) can buffer the encoded video sequence(s) created by the entropy encoder (445) to prepare for transmission via the communication channel (460), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (440) can merge the encoded video data from the video encoder (403) with other data to be transmitted, such as encoded audio data and / or auxiliary data streams (source not shown).
[0043] The controller (450) can manage the operations of the video encoder (403). During encoding, the controller (450) can assign a specific encoded picture type to each encoded picture, which may affect the encoding techniques that can be applied to the corresponding picture. For example, pictures can generally be designated as one of the following picture types: intra picture (I picture), predictive picture (P picture), bi-predictive picture (B picture), multi-predictive picture. Source pictures can generally be spatially subdivided into multiple sample encoding blocks, as described in further detail below.
[0044] Figure 5 A diagram of a video encoder (503) according to another example embodiment of the present disclosure is shown. The video encoder (503) is configured to receive a processing block (e.g., a prediction block) of sample values within a current video picture in a sequence of video pictures and encode the processing block into an encoded picture that is part of an encoded video sequence. An example video encoder (503) can be used in place of Figure 4 the video encoder (403) in the example of
[0045] For example, the video encoder (503) receives a matrix of sample values of the processing block. The video encoder (503) then uses, for example, Rate-Distortion Optimization (RDO) to determine whether to best encode the processing block using an intra mode, an inter mode, or a bi-predictive mode.
[0046] In Figure 5 the example of, the video encoder (503) includes an inter-frame encoder (530), an intra-frame encoder (522), a residual calculator (523), a switch (526), a residual encoder (524), a general controller (521), and an entropy encoder (525) coupled together as shown in the example arrangement of Figure 5 .
[0047] The inter-frame encoder (530) is configured to: receive samples of a current block (e.g., a processing block); compare the block with one or more reference blocks in a reference picture (e.g., blocks in a previous picture and a subsequent picture in display order); generate inter-frame prediction information (e.g., a motion vector, merge mode information, a description of redundant information according to an inter-frame coding technique); and calculate an inter-frame prediction result (e.g., a predicted block) based on the inter-frame prediction information using any suitable technique.
[0048] The intra-frame encoder (522) is configured to: receive samples of a current block (e.g., a processing block); compare the block with blocks that have been encoded in the same picture; generate quantized coefficients after transformation; and in some cases also generate intra-frame prediction information (e.g., intra-frame prediction direction information according to one or more intra-frame coding techniques).
[0049] The general controller (521) may be configured to determine general control data and control other components of the video encoder (503) based on the general control data to, for example, determine a prediction mode of a block and provide a control signal to the switch (526) based on the prediction mode.
[0050] The residual calculator (523) may be configured to calculate a difference (residual data) between a received block and a prediction result of a block selected from the intra-frame encoder (522) or the inter-frame encoder (530). The residual encoder (524) may be configured to encode the residual data to generate transform coefficients. Then, the transform coefficients are quantized to obtain quantized transform coefficients. In various example embodiments, the video encoder (503) further includes a residual decoder (528). The residual decoder (528) is configured to perform an inverse transform and generate decoded residual data. The entropy encoder (525) may be configured to format a bitstream to include encoded blocks and perform entropy coding.
[0051] Figure 6 FIG. shows an example video decoder (610) according to another embodiment of the present disclosure. The video decoder (610) is configured to receive an encoded picture as part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In the example, the video decoder (610) may be used instead of Figure 4A video decoder (410) in an example of FIG.
[0052] exist Figure 6 In the example of FIG. 6 , the video decoder ( 610 ) includes Figure 6 The example arrangement shown is an entropy decoder (671), an inter-frame decoder (680), a residual decoder (673), a reconstruction module (674) and an intra-frame decoder (672) coupled together.
[0053] The entropy decoder (671) may be configured to reconstruct certain symbols from the coded picture, which represent the syntax elements constituting the coded picture. The inter-frame decoder (680) may be configured to receive inter-frame prediction information and generate an inter-frame prediction result based on the inter-frame prediction information. The intra-frame decoder (672) may be configured to receive intra-frame prediction information and generate a prediction result based on the intra-frame prediction information. The residual decoder (673) may be configured to perform inverse quantization to extract dequantized transform coefficients, and process the dequantized transform coefficients to convert the residual from the frequency domain to the spatial domain. The reconstruction module (674) may be configured to combine the residual output by the residual decoder (673) with the prediction result (output by the inter-frame prediction module or the intra-frame prediction module, as the case may be) in the spatial domain to form a reconstructed block, which forms a part of a reconstructed picture as a part of the reconstructed video.
[0054] Note that the video encoders (203), (403) and (503) and the video decoders (210), (310) and (610) may be implemented using any suitable technology. In some example embodiments, the video encoders (203), (403) and (503) and the video decoders (210), (310) and (610) may be implemented using one or more integrated circuits. In another embodiment, the video encoders (203), (403) and (503) and the video decoders (210), (310) and (610) may be implemented using one or more processors executing software instructions.
[0055] Turning to block partitioning for encoding and decoding, the general partitioning can start from a base block and can follow a predefined set of rules, a specific pattern, a partitioning tree, or any partitioning structure or scheme. The partitioning can be hierarchical and recursive. After chunking or partitioning the base block following any of the example partitioning processes or other processes or combinations thereof described below, a final set of partitions or coding blocks can be obtained. Each of these partitions can be at one of the various partitioning levels in the partitioning hierarchy and can have various shapes. Each partition can be referred to as a Coding Block (CB). For the various example partitioning implementations described further below, each resulting CB can have any allowed size and partitioning level. Such partitions are called coding blocks because they can form the units for which some basic encoding / decoding decisions can be made and encoding / decoding parameters can be optimized, determined, and signaled in the encoded video bitstream. The highest or deepest level in the final partition represents the depth of the coding block partitioning structure of the tree. The coding blocks can be luminance coding blocks or chrominance coding blocks. The CB tree structure for each color can be called a CodingBlock Tree (CBT). The coding blocks for all color channels can be collectively referred to as Coding Units (CUs). The hierarchical structures for all color channels can be collectively referred to as Coding Tree Units (CTUs). The partitioning patterns or structures for the various color channels in a CTU can be the same or can be different.
[0056] In some implementations, the partitioning tree scheme or structure for the luminance channel and the chrominance channel may not need to be the same. In other words, the luminance channel and the chrominance channel can have separate coding tree structures or patterns. Additionally, whether the luminance channel and the chrominance channel use the same or different coding partitioning tree structures and the actual coding partitioning tree structure to be used can depend on whether the slice being encoded is a P slice, a B slice, or an I slice. For example, for an I slice, the chrominance channel and the luminance channel can have separate coding partitioning tree structures or coding partitioning tree structure patterns, while for a P slice or a B slice, the luminance channel and the chrominance channel can share the same coding partitioning tree scheme. When applying separate coding partitioning tree structures or patterns, the luminance channel can be partitioned into CBs by one coding partitioning tree structure, and the chrominance channel can be partitioned into chrominance CBs by another coding partitioning tree structure.
[0057] Figure 7 An example predefined 10-way partitioning structure / pattern that allows recursive partitioning to form a partitioning tree is shown. The root block can start at a predefined level (e.g., starting from a base block at the 128×128 level or the 64×64 level). Figure 7 The example partitioning structure includes various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. In some example implementations,Figure 7 None of the rectangular partitions are allowed to be further subdivided. The coding tree depth can be further limited to indicate the depth of the split starting from the root node or root block. For example, the coding tree depth of the root node or root block can be set to 0, and after the root block is split further once according to Figure 7 , the coding tree depth increases by 1. In some implementations, only the full square partitions in 710 are allowed to be recursively divided into the next level of the partitioning tree according to Figure 7 the pattern.
[0058] In some other example implementations for coding block partitioning, a quadtree structure can be used. Such quadtree splitting can be applied hierarchically and recursively to any square-shaped partition. Whether a base block or an intermediate block or partition is further quadtree split can be adjusted according to various local characteristics of the base block or intermediate block / partition.
[0059] In still some other examples, as Figure 8 shown, a ternary partitioning scheme can be used to partition a base block or any intermediate block. The ternary pattern can be implemented vertically as shown in 802 or horizontally as shown in 804. Although Figure 8 the example split ratio in is shown as 1:2:1, other ratios can also be predefined. In some implementations, two or more different ratios can be predefined. In some implementations, the width and height of the partitions of the example ternary tree are always powers of 2 to avoid additional transforms.
[0060] The above partitioning schemes can be combined in any way at different partitioning levels. As an example, the above quadtree partitioning scheme and the binary partitioning scheme can be combined to partition a base block into a Quadtree-Binary-Tree (QTBT) structure. In such a scheme, a base block or an intermediate block / partition can be quadtree split or binary split, subject to a set of predefined conditions (if specified). Figure 9 A specific example is shown in Figure 9In the example of, the base block is first divided into four partitions by a quadtree, as shown by 902, 904, 906, and 908. Thereafter, each of the resulting partitions is divided into four further partitions (e.g., 908) by the quadtree at the next level, or is divided into two additional partitions by binary partitioning (horizontally or vertically, e.g., 902 or 906, e.g., both are symmetric), or is not divided (e.g., 904). Binary or quadtree partitioning can be allowed to be used recursively for square-shaped partitions, as shown by the overall example partitioning pattern of 910 and the corresponding tree structure / representation in 920, where solid lines represent quadtree partitioning and dashed lines represent binary partitioning. A flag can be used for each binary partitioning node (non-leaf binary partition) to indicate whether the binary partitioning is horizontal or vertical. For example, as shown in 920 and consistent with the partitioning structure of 910, the flag "0" can represent a horizontal binary partition, and the flag "1" can represent a vertical binary partition. For quadtree-partitioned partitions, there is no need to indicate the partitioning type because quadtree partitioning always divides the block or partition horizontally and vertically to produce 4 sub-blocks / partitions of equal size. In some implementations, the flag "1" can represent a horizontal binary partition, and the flag "0" can represent a vertical binary partition.
[0061] In some example implementations of QTBT, the quadtree and binary partitioning rule sets can be represented by the following predefined parameters and their associated corresponding functions: – CTU size: the size of the root node of the quadtree (the size of the base block) – MinQTSize: the minimum allowable quadtree leaf node size – MaxBTSize: the maximum allowable binary tree root node size – MaxBTDepth: the maximum allowable binary tree depth – MinBTSize: the minimum allowable binary tree leaf node size
[0062] In some example implementations of the QTBT partitioning structure, the CTU size can be set to 128×128 luma samples having two corresponding 64×64 chroma sample blocks (when considering and using example chroma subsampling), MinQTSize can be set to 16×16, MaxBTSize can be set to 64×64, MinBTSize (for both width and height) can be set to 4×4, and MaxBTDepth can be set to 4. Quadtree partitioning can be first applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes can have sizes ranging from their minimum allowed size of 16×16 (i.e., MinQTSize) to 128×128 (i.e., the CTU size). If a node is 128×128, since the size exceeds MaxBTSize (i.e., 64×64), the node will not be first split by the binary tree. Otherwise, nodes not exceeding MaxBTSize can be partitioned by the binary tree. In Figure 9 the example of Figure 9 , the base block is 128×128. According to a predefined rule set, the base block can only be quadtree partitioned. The base block has a partitioning depth of 0. Each of the four resulting partitions is 64×64 - not exceeding MaxBTSize and can be further quadtree or binary tree partitioned at level 1. The process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), further partitioning can be disregarded. When the width of a binary tree node equals MinBTSize (i.e., 4), further horizontal partitioning can be disregarded. Similarly, when the height of a binary tree node equals MinBTSize, further vertical partitioning is disregarded.
[0063] In some example implementations, the above QTBT scheme can be configured to support the flexibility of having the same QTBT structure for luma and chroma or separate QTBT structures. For example, for P slices and B slices, the luma CTB and chroma CTB in a CTU can share the same QTBT structure. However, for I slices, the luma CTB can be partitioned into CUs by the QTBT structure, and the chroma CTB can be partitioned into chroma CUs by another QTBT structure. This means that in I slices, a CU can refer to different color channels. For example, an I slice can consist of coded blocks of the luma component or coded blocks of the two chroma components, and a CU in a P slice or B slice can consist of coded blocks of all three color components.
[0064] The above various CU partitioning schemes and further partitioning of CUs into PUs can be combined in any way. The following specific implementations are provided as non - limiting examples.
[0065] Inter-frame prediction can be implemented, for example, in a single-reference mode or a combined-reference mode. In some implementations, a skip flag may first be included in the bitstream of the current block (or elsewhere at a higher level) to indicate whether the current block is inter-frame encoded and not skipped. If the current block is inter-frame encoded, then another flag may be further included in the bitstream as a signal indicating whether the single-reference mode or the combined-reference mode is used for predicting the current block. For the single-reference mode, one reference block may be used to generate a predicted block of the current block. For the combined-reference mode, two or more reference blocks may be used to generate a predicted block, for example, by weighted averaging. One or more reference frame indices and additionally one or more corresponding motion vectors may be used to identify one or more reference blocks, and these motion vectors indicate the shift in the position relative to the frame (e.g., in horizontal pixels and vertical pixels) between the (one or more) reference blocks and the current block. For example, in the single-reference mode, an inter-frame predicted block of the current block may be generated as a predicted block based on a single reference block identified by one motion vector in a reference frame, while for the combined-reference mode, a predicted block may be generated by weighted averaging of two reference blocks indicated by two reference frame indices and two corresponding motion vectors in two reference frames. The (one or more) motion vectors may be encoded and included in the bitstream in various ways.
[0066] In some example implementations, one or more reference picture lists containing identifications of short-term reference frames and long-term reference frames for inter-frame prediction may be formed based on information in a Reference Picture Set (RPS). For example, a single picture reference list may be formed for uni-directional inter-frame prediction, and the single picture reference list is denoted as L0 reference (or reference list 0), while two picture reference lists may be formed for bi-directional inter-frame prediction, denoted as L0 (or reference list 0) and L1 (or reference list 1) for each of the two prediction directions. The reference frames included in the L0 list and the L1 list may be sorted in various predetermined ways. The lengths of the L0 list and the L1 list may be signaled in the video bitstream. Uni-directional inter-frame prediction may be in the single-reference mode, or may be in the combined-reference mode when multiple references used for generating a predicted block by weighted averaging in the combined prediction mode are on the same side of the frame where the block to be predicted is located. Bi-directional inter-frame prediction may only be in the combined mode because bi-directional inter-frame prediction involves at least two reference blocks.
[0067] In some implementations, a Merge Mode (MM) for inter - frame prediction can be implemented. Generally, for the Merge Mode, one or more motion vectors in the single - reference prediction of the current PB or in the composite - reference prediction can be derived from one or more other motion vectors, rather than being independently computed and signaled. For example, in an encoding system, one or more current motion vectors of the current PB can be represented by one or more differences between one or more current motion vectors and one or more other already - encoded motion vectors (referred to as reference motion vectors). Such one or more differences in one or more motion vectors, rather than the entirety of one or more current motion vectors, can be encoded and included in the bitstream, and can be linked to one or more reference motion vectors. Correspondingly, in a decoding system, one or more motion vectors corresponding to the current PB can be derived based on one or more decoded motion - vector differences and one or more decoded reference motion vectors linked thereto. As a specific form of general Merge - Mode (MM) inter - frame prediction, such inter - frame prediction based on one or more motion - vector differences can be referred to as Merge Mode with Motion Vector Difference (MMVD). Thus, general MM or specific MMVD can be implemented to utilize the correlation between motion vectors associated with different PBs to improve the encoding efficiency. For example, neighboring PBs can have similar motion vectors, and thus the Motion Vector Difference (MVD) can be small and can be efficiently encoded. For another example, for blocks similarly positioned / placed in space, the motion vectors can be temporally (between frames) correlated.
[0068] In some example implementations of MMVD, a list of reference motion vectors (RMVs) or motion vector (MV) predictor candidates for motion vector prediction can be formed for the block being predicted. The list of RMV candidates can include a predetermined number (e.g., 2) of MV predictor candidate blocks, and the motion vectors of these MV predictor candidate blocks can be used to predict the current motion vector. The RMV candidate blocks can include blocks selected from neighboring blocks and / or temporal blocks in the same frame (e.g., blocks at the same position in a previous frame or a subsequent frame of the current frame). These options represent blocks at spatial or temporal positions relative to the current block that may have a motion vector similar to or the same as that of the current block. The size of the list of MV predictor candidates can be predetermined. For example, the list can include two or more candidates. In order to be on the list of RMV candidates, for example, a candidate block may be required to have the same reference frame (or multiple reference frames) as the current block, must exist (e.g., a boundary check needs to be performed when the current block is close to the edge of the frame), and must have been encoded during the encoding process and / or decoded during the decoding process. In some implementations, if available and meeting the above conditions, the list of merge candidates can be first filled with spatially neighboring blocks (scanned in a specific predefined order), and then, if space is still available in the list, the list of merge candidates can be filled with temporal blocks. For example, neighboring RMV candidate blocks can be selected from the left block and the top block of the current block. The list of RMV predictor candidates can be dynamically formed as a dynamic reference list (DRL) at various levels (sequence, picture, frame, slice, superblock, etc.). The DRL can be signaled in the bitstream.
[0069] In some implementations, the actual MV predictor candidates that are used as reference motion vectors for predicting the motion vector of the current block can be signaled. In the case where the RMV candidate list contains two candidates, a one-bit flag called the merge candidate flag can be used to indicate the selection of the reference merge candidate. For a current block predicted in the composite mode, each of the multiple motion vectors predicted using the MV predictor can be associated with a reference motion vector from the merge candidate list. The encoder can determine which of the RMV candidates more closely predicts the MV of the current encoded block and signal this selection as an index into the DRL.
[0070] In some example implementations of MMVD, after an RMV candidate is selected and used as a base motion vector predictor for the motion vector to be predicted, a motion vector difference (MVD or ΔMV, representing the difference between the motion vector to be predicted and the reference candidate motion vector) can be calculated in the coding system. Such an MVD can include information representing both the magnitude of the MV difference and the direction of the MV difference, both of which can be signaled in the bitstream in various ways.
[0071] In some example implementations of MMVD, a distance index can be used to specify the magnitude information of the motion vector difference and to indicate one of a set of predefined offsets that represent predefined motion vector differences from a starting point (reference motion vector). The MV offset according to the signaled index can then be added to the horizontal or vertical component of the starting (reference) motion vector. An example predefined relationship between the distance index and the predefined offsets is specified in Table 1. Table 1 - Example relationship between distance index and predefined MV offsets
[0072] In some example implementations of MMVD, the direction index can be further signaled, and the direction index is used to represent the direction of the MVD relative to the reference motion vector. In some implementations, the direction can be restricted to either the horizontal direction or the vertical direction. Table 2 shows example 2-bit direction indexes. In the example of Table 2, the description of the MVD can vary according to the information of the start / reference MV. For example, when the start / reference MV corresponds to a uni-predicted block or corresponds to a bi-predicted block where both reference frame lists point to the same side of the current picture (i.e., the POCs of both reference pictures are greater than the POC of the current picture or both are less than the POC of the current picture), the signs in Table 2 can specify the sign (direction) of the MV offset added to the start / reference MV. When the start / reference MV corresponds to a bi-predicted block where the two reference pictures are at different sides of the current picture (i.e., the POC of one reference picture is greater than the POC of the current picture and the POC of the other reference picture is less than the POC of the current picture), and the difference between the reference POC in picture reference list 0 and the current frame is greater than the difference between the reference POC in picture reference list 1 and the current frame, the signs in Table 2 can specify the sign of the MV offset added to the reference MV corresponding to the reference picture in picture reference list 0, and the sign of the offset of the MV corresponding to the reference picture in picture reference list 1 can have an opposite value (opposite sign of the offset). Otherwise, if the difference between the reference POC in picture reference list 1 and the current frame is greater than the difference between the reference POC in picture reference list 0 and the current frame, then the signs in Table 2 can specify the sign of the MV offset added to the reference MV associated with picture reference list 1, and the sign of the offset of the reference MV associated with picture reference list 0 has an opposite value. Table 2 – Example implementation of the signs of MV offsets specified by the direction index Direction Index 00 01 10 11 x - axis (Horizontal) + – N / A N / A y - axis (Vertical) N / A N / A + –
[0073] In some example implementations, the MVD can be scaled according to the difference in POC in each direction. If the differences in POC in both lists are the same, no scaling is required. Otherwise, if the difference in POC in reference list 0 is greater than the difference in POC in reference list 1, scale the MVD of reference list 1. If the POC difference in reference list 1 is greater than that in list 0, the MVD of list 0 can be scaled in the same way. If the start MV is uni-predicted, then the MVD is added to the available or reference MV.
[0074] In some example implementations of MVD coding and signaling for bidirectional composite prediction, in addition to or as an alternative to separately coding and signaling two MVDs, symmetric MVD coding can be implemented such that only one MVD needs to be signaled and the other MVD can be derived from the signaled MVD. In such an implementation, not all of the motion information including the reference picture indices of list 0 and list 1 is signaled. Specifically, at the slice level, a flag called "mvd_l1_zero_flag" can be included in the bitstream, which is used to indicate whether reference list 1 is not signaled in the bitstream. If this flag is 1, indicating that reference list 1 is equal to zero (and thus not signaled), then the bidirectional prediction flag called "BiDirPredFlag" can be set to 0, which means that there is no bidirectional prediction. Otherwise, if mvd_l1_zero_flag is zero, if the nearest reference picture in list 0 and the nearest reference picture in list 1 form a forward and backward pair of reference pictures or a backward and forward pair of reference pictures, then BiDirPredFlag can be set to 1 and both the list 0 reference picture and the list 1 reference picture are short-term reference pictures. Otherwise, BiDirPredFlag is set to 0. BiDirPredFlag being 1 can indicate that a symmetric mode flag is additionally signaled in the bitstream. When BiDirPredFlag is 1, the decoder can extract the symmetric mode flag from the bitstream. For example, the symmetric mode flag can be signaled at the CU level (if needed), and the symmetric mode flag can indicate whether the symmetric MVD coding mode is being used for the corresponding CU. When the symmetric mode flag is 1, it indicates that the symmetric MVD coding mode is used, and only the reference picture indices of both list 0 and list 1 (called "mvp_l0_flag" and "mvp_l1_flag") and the MVD associated with list 0 (called "MVD0") are signaled, and the other motion vector difference "MVD1" is derived rather than signaled. For example, MVD1 can be derived as -MVD0. Thus, in the example symmetric MVD mode, only one MVD is signaled.
[0075] In some other example implementations of MV prediction, for both single-reference mode and composite-reference mode MV prediction, a coordination scheme can be used to implement the general merge mode MMVD and some other types of MV prediction. Various syntax elements can be used to signal the way to predict the MV of the current block. For example, for the single-reference mode, the following MV prediction modes can be signaled:
[0076] NEARMV – Use one of the motion vector predictors (MVPs) in the list indicated by DRL (Dynamic Reference List) indexing
[0077] NEWMV – Use one of the motion vector predictors (MVPs) in the list signaled by DRL indexing as a reference and apply an increment to the MVP
[0078] GLOBALMV – Use a motion vector based on frame-level global motion parameters
[0079] Similarly, for the composite reference inter-prediction mode that uses two reference frames corresponding to the two MVs to be predicted, the following MV prediction modes can be signaled:
[0080] NEAR_NEARMV – Use one of the motion vector predictors (MVPs) in the list signaled by DRL indexing
[0081] NEAR_NEWMV – Use one of the motion vector predictors (MVPs) in the list signaled by DRL indexing as a reference and send an incremental MV for the second MV
[0082] NEW_NEARMV – Use one of the motion vector predictors (MVPs) in the list signaled by DRL indexing as a reference and send an incremental MV for the first MV
[0083] NEW_NEWMV – Use one of the motion vector predictors (MVPs) in the list signaled by DRL indexing as a reference and send incremental MVs for both MVs
[0084] GLOBAL_GLOBALMV – Use the MV from each reference based on the frame-level global motion parameters for each reference
[0085] The term "NEAR" above refers to using the reference MV as an MV prediction in the general merge mode without MVD, while the term "NEW" refers to an MV prediction that involves using the reference MV as in the MMVD mode and offsetting the reference MV using the signaled motion vector difference (MVD). For composite inter-prediction, both the above reference basic motion vectors and motion vector increments can generally be different or independent between the two references, even though the two references may be related and this correlation can be exploited to reduce the amount of information needed to signal the two motion vector increments. In such a case, joint signaling of the two MVDs can be achieved and indicated in the bitstream.
[0086] The above dynamic reference list (DRL) can be used to save a set of indexed motion vectors, which are dynamically maintained and regarded as candidate motion vector predictors.
[0087] Motion vector difference coding
[0088] In some example implementations, in coding techniques such as AV1, motion vector precision (or accuracy) of fractions such as 1 / 8 pixel (i.e., one - eighth pixel) is allowed / supported, and the following syntax is used to signal the motion vector difference in reference frame list 0 (L0) or reference frame list 1 (L1). ● mv_joint specifies which components of the motion vector difference are non - zero o0 indicates that there is no non - zero MVD along the horizontal or vertical direction o1 indicates that there is a non - zero MVD only along the horizontal direction o2 indicates that there is a non - zero MVD only along the vertical direction o3 indicates that there are non - zero MVDs along both the horizontal and vertical directions · mv_sign specifies whether the motion vector difference is positive or negative · mv_class specifies the class of the motion vector difference. As shown in Table 3, the higher the class, the larger the magnitude of the motion vector difference. Table 3: Magnitude classes of motion vector differences MV Category Magnitude of MVD MV_CLASS_0 (0,2] MV_CLASS_1 (2,4] MV_CLASS_2 (4,8] MV_CLASS_3 (8,16] MV_CLASS_4 (16,32] MV_CLASS_5 (32,64] MV_CLASS_6 (64,128] MV_CLASS_7 (128,256] MV_CLASS_8 (256,512] MV_CLASS_9 (512,1024] MV_CLASS_10 (1024,2048] · mv_bit specifies the integer part of the offset between the motion vector difference and the starting magnitude of each MV class · mv_fr specifies the first 2 fractional bits of the motion vector difference · mv_hp specifies the third fractional bit of the motion vector difference
[0089] Adaptive MVD resolution
[0090] In some example implementations, the resolution of MVDs in various MVD magnitude classes can be differentiated. For example, high - resolution MVDs for large MVD magnitudes in higher MVD classes may not provide a statistically significant improvement in compression efficiency. Thus, for larger MVD magnitude ranges corresponding to higher MVD magnitude classes, the MVD can be encoded at a reduced resolution (integer - pixel resolution or fractional - pixel resolution). Similarly, for generally larger MVD values, the MVD can be encoded at a reduced resolution (integer - pixel resolution or fractional - pixel resolution). Such MVD class - related or MVD magnitude - related MVD resolution can generally be referred to as adaptive MVD resolution.
[0091] Since statistically observing that processing the MVD resolution of a large - magnitude or high - class MVD at a level similar to that of a low - magnitude or low - class MVD in a non - adaptive manner may not significantly increase the inter - frame prediction residual coding efficiency of the blocks of the large - magnitude or high - class MVD, using an adaptive MVD resolution, the number of signaling bits reduced by targeting a less precise MVD may be greater than the additional bits required to encode the inter - frame prediction residual due to such a less precise MVD. In other words, using a higher MVD resolution for a large - magnitude or high - class MVD may not result in much coding gain compared to using a lower MVD resolution.
[0092] In some example implementations, further constraints may be imposed on composite reference modes (such as the NEW_NEARMV and NEAR_NEWMV modes as described above). Specifically, the precision of the motion vector difference (MVD) depends on the associated class and the magnitude of the MVD.
[0093] In some example implementations, fractional MVD is allowed only when the MVD magnitude is equal to or less than one pixel. Alternatively or additionally, in some example implementations, when the value of the associated MV class is equal to or greater than MV_CLASS_1, only one MVD value is allowed, and for MV classes 1 (MV_CLASS_1), 2 (MV_CLASS_2), 3 (MV_CLASS_3), 4 (MV_CLASS_4), or 5 (MV_CLASS_5), the MVD values in each MV class are derived as 4, 8, 16, 32, 64 respectively.
[0094] Exemplarily, Table 4 shows the allowed MVD values in each MV class. Table 4: Adaptive MVD in Each MV Magnitude Class · Additionally, if the current block is encoded in the NEW_NEARMV or NEAR_NEWMV mode, one context is used to signal mv_joint or mv_class. Otherwise, a different context is used to signal mv_joint or mv_class. Note that in AV1, mv_class specifies the class of the motion vector difference. A higher class means a larger update represented by the motion vector difference; mv_joint specifies which components of the motion vector difference are non - zero.
[0095] Improvement of Adaptive MVD Resolution
[0096] In some example implementations, the adaptive MVD implementation can be further improved. A new inter-frame coding mode (named AMVDMV) can be added to the single-reference case. When the AMVDMV mode is selected and / or flagged, it indicates that AMVD is applied to the signal MVD, thereby adopting an adaptive MVD resolution.
[0097] In one solution, a flag (named amvd_flag) is added in the JOINT_NEWMV mode to indicate whether AMVD is applied to the joint MVD coding mode. When the adaptive MVD resolution is applied to the joint MVD coding mode (which can also be named the joint AMVD coding mode), the MVDs of two reference frames are jointly signaled, and the accuracy of the MVD is implicitly determined by, for example, the MVD magnitude. Otherwise, the MVDs of two (or more than two) reference frames are jointly signaled, and conventional MVD coding is applied.
[0098] Alternatively, the MVDs of two (or more than two) reference frames can be jointly signaled, and conventional MVD coding is applied. In this case, instead of adding the amvd_flag as discussed above, a new inter-frame prediction mode (named JOINT_AMVDNEWMV) is added to indicate that AMVD is applied to the joint MVD coding mode.
[0099] Adaptive Motion Vector Resolution (AMVR)
[0100] In some example implementations, AMVR can be implemented with various coding techniques (such as AV1). For example, a total of 7 MV precisions are supported (e.g., 8, 4, 2, 1, 1 / 2, 1 / 4, 1 / 8). For each prediction block, the encoder (e.g., AVM, AV1 encoder) searches all supported precision values and signals the best precision to the decoder.
[0101] To reduce the complexity during encoder runtime, two precision sets are supported. Each precision set can contain, for example, 4 predefined precisions. Based on the maximum precision value of the frame, one of the precision sets is adaptively selected at the frame level. Exemplarily, the maximum precision can be signaled in the frame header. Table 5 summarizes the precision values supported based on the frame-level maximum precision. Table 5: MV Precisions Supported in Two Sets Frame - level Maximum Precision Supported MV Precision 1 / 8 1 / 8,1 / 2,1,4 1 / 4 1 / 4,1,4,8
[0102] In some example implementations, there is a frame-level flag to indicate whether the MV of a frame includes sub-pel (i.e., sub-pixel) precision. AMVR is enabled only when the value of the cur_frame_force_integer_mv flag is 0. Under AMVR, if the precision of a block is lower than the maximum precision, the motion model and interpolation filter are not signaled. If the precision of a block is lower than the maximum precision, the motion mode can be inferred as translational motion and the interpolation filter can be inferred as a REGULAR interpolation filter. Similarly, if the precision of a block is 4 pixels or 8 pixels, the inter-intra mode is not signaled and is inferred as 0.
[0103] Joint MVD Coding (JMVD)
[0104] In some example implementations, in an encoding technique such as AV1, an inter-frame coding mode (named JOINT_NEWMV) is applied to indicate whether the MVDs of two reference lists are jointly signaled. If the inter-frame prediction mode is equal to the JOINT_NEWMV mode, the MVDs of reference list 0 and reference list 1 are jointly signaled. In this case, only one MVD (named joint_mvd) is signaled and transmitted to the decoder, and the incremental MVs of reference list 0 and reference list 1 are derived from joint_mvd.
[0105] In some example implementations, the JOINT_NEWMV mode is signaled together with the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes. No additional context needs to be added.
[0106] When the JOINT_NEWMV mode is signaled and the Picture Order Count (POC) distances between the two reference frames and the current frame are different, the MVD is scaled for reference list 0 or reference list 1 based on the POC distance. Specifically, the distance between reference frame list 0 and the current frame is denoted as td0, and the distance between reference frame list 1 and the current frame is denoted as td1. If td0 is equal to or greater than td1, joint_mvd is the jointly signaled MVD and is directly used for reference list 0, and the mvd of reference list 1 is derived from joint_mvd based on Equation (1):
[0107] Otherwise, if td1 is equal to or greater than td0, joint_mvd is directly used for reference list 1, and the mvd of reference list 0 is derived from joint_mvd based on Equation (2):
[0108] Improved Joint MVD Coding
[0109] In some example implementations, in an encoding technique such as AV1, if a block is encoded in a joint MVD coding mode (e.g., JOINT_NEWMV or JOINT_AMVDNEWMV), a new syntax (named mvd_scaling_factor_idx) is signaled into the bitstream to explicitly indicate the scaling factor of the MVD between reference frame 0 and reference frame 1.
[0110] Exemplarily, as shown in Tables 6 and 7 below, two predefined lookup tables can be used to store the scaling factors supported / allowed by JOINT_NEWMV or JOINT_AMVDNEWMV respectively. The associated entry index of the selected scaling factor in the lookup table is signaled in the bitstream. For the JOINT_AMVDNEWMV mode, the same scaling factor is applied to both the vertical and horizontal components of the MVD of reference frame list 0 and / or 1. For the JOINT_NEWMV mode, the scaling factor of one component (vertical or horizontal component) of the MVD is restricted to 1, while the scaling factor of the other component of the MVD can be other values, such as 2 or 1 / 2. In one example, the MVDs of reference frame lists 0 and 1 are calculated in the following equations: mvd_ref0 = joint_mvd (3)
[0111] Here, mvd_ref0 and mvd_ref1 represent the MVD of reference frame list 1 and the MVD of reference frame list 2 respectively. The distance between reference frame 0 and the current frame is denoted as td0, and the distance between reference frame 1 and the current frame is denoted as td1; joint_mvd represents the jointly signaled MVD; and jmvd_scale represents the scaling factor. Table 6: Scaling Factors for JOINT_AMVDNEWMV Index Scaling Factor for both x - axis and y - axis 0 1 1 2 2 1 / 2 Table 7: Scaling Factors for JOINT_NEWMV
[0112] Bi-prediction with CU-level Weight (BCW)
[0113] In some example implementations, in video coding techniques such as HEVC (High Efficiency Video Coding), a bi-prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different motion vectors. In VVC, the bi-prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. For example, the bi-prediction denoted as P bi-pred can be calculated using Equation 5 below: P bi-pred = ((8 - w) * P 0 + w * P 1 + 4) >> 3 (5)
[0114] Exemplarily, five weights are allowed in weighted-average bi-prediction, w ∈ {-2, 3, 4, 5, 10}. When w equals 4, equal weight factors are used for weighted averaging of the two prediction samples. For each bi-predicted CU, the weight w is determined in one of the following two ways: 1) For non-merged CUs, the weight index is signaled after the motion vector difference; 2) For merged CUs, the weight index is inferred from neighboring blocks based on the merge candidate index.
[0115] In some example implementations, BCW is only applied to CUs with 256 or more luma samples (i.e., CU width multiplied by CU height is greater than or equal to 256). For low-delay pictures, all 5 weights are used. For non-low-delay pictures, only 3 weights (w ∈ {3, 4, 5}) are used.
[0116] Local Illumination Compensation (LIC)
[0117] LIC is a video coding tool that can be utilized by video encoders and video decoders. In some example implementations, LIC can be applied based on a linear model to compensate for illumination changes between one or more temporal reference pictures and the current picture (e.g., in the motion compensation stage). The linear model is based on LIC parameters including a scaling factor α and an offset β, and is described in more detail in the block adaptive weighted prediction section below. LIC can be enabled or disabled via, for example, high-level signaling at various levels.
[0118] In some example implementations, a bidirectional prediction reference template can be generated for a template sample associated with a current block. One or more motion vectors associated with the current block can be used to identify the reference template sample. For example, the reference template sample can include neighboring temporal reference CUs of the current CU and can correspond to the template sample of the current CU. In LIC parameter derivation, the reference template samples can be jointly considered (e.g., averaged).
[0119] In some example implementations, the Least Mean-Squared-Error (LMSE) algorithm can be applied to derive the LIC parameter. For example, LMSE-based calculations can be performed to determine the LIC parameter such that the difference between the reference template reference samples for bidirectional prediction and the template samples of the current CU can be minimized.
[0120] In some example implementations, a similar method can be used under unidirectional prediction. In this case, the LIC parameter can be determined such that the difference between the reference template reference samples for unidirectional prediction and the template samples of the current CU can be minimized.
[0121] Note that the LMSE algorithm described herein is merely an example for deriving the LIC parameter. One or more other methods / algorithms can be used.
[0122] Block - Adaptive Weighted Prediction (BAWP)
[0123] In some example implementations, BAWP is used to model local illumination changes.
[0124] Referring to Figure 10 , for example, BAWP can be block-level weighted prediction to model the local illumination change between the current block and its predicted block according to the local illumination change between the current block template (or the causal sample of the current block) and the reference block template. Figure 10 Shows the template of the current block (1012, or the predicted block, i.e., the block currently being encoded / decoded) (or the current template, current block template, 1010) and the template of the reference block (1022) (or the reference template, 1020). The template of the current block can be referred to as the current block template or the current template. The template of the reference block can be referred to as the reference block template or the reference template. Each template can include an upper part and / or a left part. For example, the current block 1012 includes an upper part 1014 and / or a left part 1016. The reference block can be indicated or determined by a motion vector (MV, 1030). The current block can be in the current picture (or current frame). The reference block can be in a different picture such as a reference picture (or reference frame) of the current picture, or the reference block can be at a different position in the current picture.
[0125] In some example implementations, the current template and the reference template may be extended to the upper right and lower left areas of the corresponding blocks.
[0126] In some implementations, the function may be a linear function. The parameters of the function may be represented by a scaling factor α and an offset β, which form a linear equation. The scaling factor may also be referred to as scaling, α factor, or α value.
[0127] When a block is encoded in BAWP mode, an exemplary linear function for compensating for illumination changes is listed below: p′(x′)= α*p(x)+β (6)
[0128] Where: p′(x′) is the prediction sample at position x′ in the current block (or the prediction sample in the prediction unit (PU) in the current block), p(x) is the sample at position x in the reference block corresponding to p′(x′), α is a scaling factor (or scaling), and β is an offset value. Note that the reference block can be identified or derived from the MV associated with the current block, and p(x) is the reference sample at position x on the reference picture pointed to by the MV. Note that in Equation 6, the reference sample and the prediction sample can have the same coordinates (i, j) in their respective blocks (i.e., the reference block and the current block, or the reference block and the prediction unit). Alternatively, the coordinates of the reference sample in the reference block can be based on the coordinates of the prediction sample in the current block. For example, the coordinates of the prediction sample can be adjusted by a delta value to obtain the coordinates of the corresponding reference sample (in the reference picture). That is, the prediction sample has coordinates of (i, j) (relative to the prediction block). Instead of using the coordinates (i, j) of the reference sample in the reference block, the coordinates (i, j) may be adjusted by the increment value in each dimension to obtain updated coordinates (i+Δi, j+Δj) in the reference block.
[0129] In the present disclosure, the prediction samples in the prediction block (current block) as described above and its reference samples may have a co-location relationship (i.e., they have the same spatial position). The block vector is used to identify the reference block of the current block. In some example implementations, the reference block is restricted to a picture (or frame) different from the current picture (or current frame) to which the current block belongs. In some example implementations, the reference block is restricted to the same picture as the current block. In some example implementations, the reference block identified by the motion vector is within a predetermined distance from the current block.
[0130] In the present disclosure, unless otherwise specified, a video bitstream includes payload data and signaling data. The signaling data is different from the payload data and may include, for example, syntax elements. When a parameter is signaled, this means that the parameter is transmitted via, for example, syntax elements rather than in the payload data. Syntax elements are explicitly signaled in the bitstream, and the decoder can directly extract the content of the syntax elements without using a derivation process. In a derivation process, decoding can use existing information extracted from, for example, the payload data to derive other information. For example, explicit signaling can be used to directly indicate an encoding mode, or the decoder can derive the encoding mode based on existing information (or already reconstructed information). Explicit signaling may increase the video bitstream overhead, but can improve the encoding / decoding efficiency. On the other hand, implicit indication (requiring further derivation) may reduce the bitstream overhead.
[0131] In the present disclosure, various embodiments for improving video encoding / decoding techniques in BAWP and / or LIC modes are disclosed, aiming to improve the prediction accuracy with a minimum signaling cost overhead. Specifically, various methods for signaling and / or deriving the scaling factors and offsets (α and β) specified in Equation 6 are described.
[0132] Based on investigations and statistical observations, high-precision, high-accuracy scaling factors can help improve the encoding efficiency and increase the encoding gain. High precision is particularly beneficial when the scaling factor has a low magnitude. Therefore, instead of deriving the scaling factor at the decoder side, signaling the scaling factor in the video bitstream can have some advantages. In addition, when the decoder stores the supported (candidate) scaling factors, sorting them in a certain way can help improve the entropy coding efficiency. In the present disclosure, various embodiments for achieving such goals are described below.
[0133] In the present disclosure, the term block may refer to a transform block, an encoded block, a prediction block, an encoding block, a coding unit (CU), etc. The term chroma block may refer to a block in any chroma (color) channel. The direction of a reference frame is determined by whether the reference frame is before the current frame in the display order or after the current frame in the display order.
[0134] In the present disclosure, a sample may be interpreted as the pixel value of a pixel. It generally may refer to any component (luminance or chroma).
[0135] In the present disclosure, the terms x-axis and y-axis refer to the horizontal and vertical components of 2-D values. They may also be replaced by two other axes along two predefined directions perpendicular to each other, and the same embodiments apply. That is, the x-axis and y-axis may be rotated by a certain degree. For example, the x-axis and y-axis may be replaced by the 45-degree axis and the 135-degree axis.
[0136] To improve the modeling of local illumination changes between the current block and its predicted block, it has been observed that when predicting a predicted sample (or the current sample in the current block), more accurate predictions can be achieved when considering the co-located reference sample plus one or more neighboring samples of the co-located reference sample compared to Equation 6 that uses only one reference sample. This observation applies to both the LIC mode and the BAWP mode. In this improved model, the local illumination changes between the current block and its predicted block can be simulated according to the following: a) the local illumination changes between the current block template (or the causal samples of the current block) and / or the current block; and b) the reference block template and / or the current block reference block. As previously described, Figure 10 An example template 1010 of the current block and an example template 1020 of the reference block are shown. The reference block can be located in a different picture such as a reference picture of the current picture, or the reference block can be located at different positions in the same picture of the current picture.
[0137] In one embodiment, for each sample in the current block template and the current block, a set of reference samples in the reference block template and / or the reference block is determined, and this set of reference samples is used to predict the sample. One or more samples from the reference block or the reference block template can be selected. Figure 11 An example of candidate samples in the reference block for predicting the current sample C' in the current block is shown. C is the co-located sample of the current sample located in the reference block, and NW, N, NE, W, E, SW, S, SE are the neighboring samples of C. Relative to C, NW is the northwest direction (upper left), N is the north direction (up), NW is the northwest direction (upper left), NE is the northeast direction (upper right), E is the east direction (right), SE is the southeast direction (lower right), SE is the southeast direction (lower right), S is the south direction (down), SW is the southwest direction (lower left), and W is the west direction (left).
[0138] In some example implementations, the predicted value of the current sample in the current block can be derived from the following equation:
[0139] predVal = c0C + c1N + c2S + c3W + c4E + c5NW + c6NE + c7SW + c8SE + c9P + c10B(7)
[0140] where c0 to c10 are scaling factors, which are natural numbers; C, N, S, W, E, NW, NE, SW, and SE are Figure 11 the samples in the reference block as shown. P is a non-linear term, and B is an offset. P and B are used for compensation purposes. For example, they can be used to compensate for the average value of the predicted value.
[0141] P is a non - linear term. Exemplarily, P=(C * C+medianOrAverage_value)>>(bit_depth), where ">>" is a bit - shift - right operation; medianOrAverage is the median or average of all samples in the current block; and bit_depth is the bit - depth of the current block. Alternatively, medianOrAverage is the median or average of all samples in the reference block; and bit_depth is the bit - depth of the reference block.
[0142] B is an offset, which can be the median or average of the current block or the reference block (e.g., depending on the choice of calculating P).
[0143] In some example implementations, the prediction uses only the samples defined in Equation 7 and not other samples. However, in some other example implementations, additional parameters can be added.
[0144] In some example implementations, only a subset of neighboring samples is used for prediction.
[0145] In some example implementations, only the samples located above (N), to the left (W), below (S), and to the right (E) of C are used. The prediction of the current block is derived from the following equation, and the parameters follow the same definitions as in Equation 7.
[0146] predVal = c0C + c1N + c2S + c3W + c4E + C5P + C6B (8)
[0147] In some example implementations, only the samples located in the upper - left (NW), upper - right (NE), lower - left (SW), and lower - right (SE) directions of C are used. The prediction of the current block is derived from the following equation, and the parameters follow the same definitions as in Equation 7.
[0148] predVal = c0C + c1NW + c2NE + c3SW + C4SE + C5P + C6B (9)
[0149] In some example implementations, P can be removed from the above Equations 7 to 9.
[0150] In equations 7 to 9 for obtaining the predicted value (predVal), illumination changes are compensated. Since all parameters are derived from the templates of the reference block and the current block and the information is already available (already reconstructed), the derivation does not require signaling overhead. A flag can be signaled to indicate enabling of the multiple scaling factor feature (or more specifically, the multiple scaling factor weighted prediction feature at the block level), referred to as the multiple scaling factor flag. In some example implementations, the multiple scaling factor feature only applies to coded blocks predicted in the single inter-frame prediction mode, so signaling of the multiple scaling factor is only required when predicting coded blocks in the single inter-frame prediction mode. Additionally or alternatively, the multiple scaling factor feature may only apply to blocks having a size greater than a predefined threshold (e.g., 8×8 (in pixels)).
[0151] In some example implementations, when encoding a block using the multiple scaling factor feature, a flag can be signaled to indicate whether to use explicit signaling or implicit signaling for the weighted prediction scaling factor. Alternatively, when encoding a block using the multiple scaling factor feature, explicit signaling of the weighted prediction scaling factor can be directly adopted. That is, turning on or enabling the multiple scaling factor feature will automatically indicate explicit signaling for weighted prediction.
[0152] In some example implementations, all scaling factors (such as those used in equations 7 to 9) are explicitly signaled via, for example, syntax elements. In this case, neighboring templates are not used for scaling factor derivation.
[0153] In some example implementations, some of the scaling factors are signaled and the rest are derived from neighboring templates.
[0154] In some example implementations, a flag can be signaled to indicate whether to use a linear model with a single scaling factor (e.g., using equation 6) or multiple scaling factors (e.g., using equations 7 to 9).
[0155] In some example implementations, only some of the scaling factors are explicitly signaled or no scaling factors are signaled explicitly, and the remaining scaling factors need to be derived. The derivation can be based on, for example, using the minimum mean square estimation (e.g., without square root Cholesky decomposition) based on the templates of the current block and the reference block. Note that the reference block is pointed to / identified by the motion vector of the current block.
[0156] In some example implementations, when using explicit signaling of weighted prediction scaling factors, another flag is signaled to indicate which scaling factor(s) are used for the current block. Exemplarily, the supported scaling factors can be stored in one or more predefined look-up tables, and the indices of the scaling factors in the look-up table are signaled into the bitstream and parsed at the decoder side. In some example implementations, an offset value can be derived based on the signaled scaling factor and the templates of the current block and the reference block. In some example implementations, other parameters than c0 are derived by (cur_template_sample - c0 * ref_template_C), where cur_template_C indicates the co-located sample of the current sample in the current block template, and ref_template_C indicates the current sample in the reference block template. In the present disclosure, the look-up table can be implemented in various data structures such as lists, sets, maps (e.g., hash maps), tables (e.g., hash tables), trees, etc.
[0157] In some example implementations, it has been observed that the distribution of the supported scaling factors has a strong correlation with the coding information. Therefore, a further improvement to the signaling of the scaling factors has been proposed and implemented.
[0158] In some example implementations, the context for signaling a flag indicating whether to enable or apply the multiple scaling factor feature can depend on whether the current block is coded using a block vector pointing to a reference block. That is, if the current block is predicted using a block vector, context A can be selected; otherwise, a different context B is selected and the use of context is not allowed.
[0159] In some example implementations, when the multiple scaling factor feature is enabled or applied for a current block coded using a block vector, another flag is signaled to indicate how many reference lines and which reference lines are used to derive the parameters (e.g., including scaling factors as shown in equations 7, 8, or 9) for predicting the current block via, for example, equations 7, 8, or 9.
[0160] In some example implementations, when the multiple scaling factor feature is enabled or applied for a current block coded using a block vector, neighboring samples in more than one neighboring reference line can be used to derive the parameters (e.g., including scaling factors as shown in equations 7, 8, or 9) for predicting the current block via, for example, equations 7, 8, or 9.
[0161] In some example implementations, when the multiple scaling factor feature is enabled or applied for a current block coded using a block vector, the multiple scaling factor feature is only allowed to be applied to a specific component of the current block, such as the luminance component.
[0162] In some example implementations, whether to use explicit signaling of block-level weighted prediction and / or implicit derivation of linear model parameters can be signaled in a high-level syntax that includes, but is not limited to, sequence / picture / slice / sub-picture / tile-level flags.
[0163] In some example implementations, when multiple scaling factor features are enabled or applied for a current block, a flag can be signaled to indicate whether the scaling factor is explicitly signaled or needs to be derived. The flag can be signaled via a high-level syntax that can be at least one of the following levels: Sequence Parameter Set (SPS) level; Picture Parameter Set (PPS) level; picture level; slice level; or tile level.
[0164] In the present disclosure, unless otherwise specified, signaling can include one or more sub-signalings. The one or more sub-signalings can be transmitted together or separately.
[0165] For the embodiments described below, the coding block or the coded block can be coded in the BAWP (or LIC, or Compound Weighted Prediction (CWP)) mode (hereinafter referred to as BAWP for simplicity of description). This embodiment can be implemented in a decoder and / or an encoder.
[0166] In the present disclosure, signaling (e.g., syntax elements) can indicate values, such as scaling factors and / or offset β, in an explicit manner or an implicit manner. Explicit indication can be performed by sending an index to a lookup table to look up the value, or by directly sending the value, which can be the result of entropy decoding of raw bits from the bitstream - once entropy decoded, the decoder can use the value without further derivation. When the value is sent in an implicit manner, the decoder may need to perform further derivation based on the signal syntax elements. In the case of implicit indication, the syntax element can also be referred to as being "associated with" the value to be derived.
[0167] The various embodiments and / or implementations described in the present disclosure can be performed individually or in any order in combination. In addition, each of the method (or embodiment), encoder, and decoder can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). One or more processors execute a program stored in a non-transitory computer-readable medium. In the present disclosure, the term block can be interpreted as a prediction block, a coding block, or a coding unit (CU).
[0168] Figure 12FIG. 1200 shows a flow chart of an exemplary method that follows the underlying principles of the above-described implementation for decoding a video bitstream based on a prediction function using multiple scaling factors. The exemplary decoding method flow may include some or all of the following steps: S1210, receiving a video bitstream that includes a current block and a reference block, the reference block being used to predict the current block by using a prediction function; S1220, receiving a first syntax element from the video bitstream, the first syntax element indicating how to determine prediction parameters of the prediction function, the prediction parameters including scaling factors, and the first syntax element indicating at least one of the following: whether the scaling factors are explicitly signaled in the video bitstream via a syntax element; whether the scaling factors need to be derived from the video bitstream; and whether a first subset of the scaling factors is signaled in the video bitstream and a second subset of the scaling factors is derived from the video bitstream, the second subset of the scaling factors and the first subset of the scaling factors being non-overlapping; S1230, determining the scaling factors based on the first syntax element; S1240, determining a local illumination change between the current block and a predicted block corresponding to the current block based on the prediction function including the determined scaling parameters, a current block template, and a reference block template; S1250, predicting a current sample in the current block based at least in part on the local illumination change and the prediction function; and S1260, reconstructing the current block based on the predicted current sample.
[0169] In any part or combination of the above implementation, the prediction function is a weighted sum of at least one of the following: a co-located sample of a current sample located in a reference block; a set of neighboring samples of the co-located sample; a compensation value based on the co-located sample; or an offset; and the parameters include at least one of the following: a scaling factor corresponding to the co-located sample; a scaling factor corresponding to each of the set of neighboring samples; a scaling factor corresponding to the compensation value; or a scaling factor corresponding to the offset.
[0170] In any part or combination of the above implementation, the prediction function may include any one of Equations 7 to 9.
[0171] In any part or combination of the above implementation, the prediction parameters include scaling factors, such as c0 to c10 in Equation 7; c0 to c6 in Equation 8 or 9.
[0172] In any part or combination of the above implementation, the prediction function may use two or more scaling factors, each scaling factor being applied to a different sample.
[0173] In any part or combination of the above implementation manners, the compensation value is determined using the following formula: (C * C + median_value) >> bit_depth, where: C is the value of the co-located sample; median_value is the median or average value of all samples in the reference block; and bit_depth is the bit depth of the reference block.
[0174] Embodiments in the present disclosure can be applied to blocks using the BAWP or LIC mode.
[0175] In the present disclosure, the direction of the reference frame can be determined by whether the reference frame is before or after the current frame in the display order.
[0176] The above operations can be combined or arranged in any number or order as needed. Two or more of the steps and / or operations can be executed in parallel. Embodiments and implementation manners in the present disclosure can be used individually or in combination in any order. The steps in one embodiment / method can be split to form multiple sub-methods, and each of the sub-methods can be independent of the other steps in the embodiment and can form an independent solution. Additionally, each of the method (or embodiment), encoder, and decoder can be implemented by processing circuitry (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium. Embodiments in the present disclosure can be applied to luminance blocks or chrominance blocks. The term block can be interpreted as a prediction block, a coding block, or a coding unit, i.e., a CU. The term block herein can also be used to refer to a transform block. In the following items, when referring to the block size, it can refer to the block width or height, or the maximum of the width and height, or the minimum of the width and height, or the area size (width * height), or the aspect ratio of the block (width:height or height:width).
[0177] The above technologies can be implemented as computer software using computer-readable instructions and physically stored in one or more computer-readable media. For example, Figure 13 A computer system (1800) suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0178] Computer software can be encoded using any suitable machine code or computer language, and any such suitable machine code or computer language can be subject to mechanisms such as assembly, compilation, linking, etc. to create code including instructions that can be directly executed by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc., or executed through interpretation, microcode execution, etc.
[0179] The instructions can be executed on various types of computers or their components, and the various types of computers or their components include, for example, personal computers, tablet computers, servers, smart phones, gaming devices, Internet of Things devices, etc.
[0180] Figure 13 The components shown for the computer system (1800) are exemplary in nature and are not intended to impose any limitations on the scope of use or functionality of the computer software implementing the present disclosure. The configuration of the components should also not be construed as having any dependencies or requirements related to any one or combination of the components shown in the exemplary embodiments of the computer system (1800).
[0181] The computer system (1800) may include certain human-machine interface input devices. The input human-machine interface devices may include one or more of the following (each is depicted only one): keyboard (1801), mouse (1802), touchpad (1803), touch screen (1810), data glove (not shown), joystick (1805), microphone (1806), scanner (1807), camera device (1808).
[0182] The computer system (1800) may also include certain human-machine interface output devices. Such human-machine interface output devices can stimulate the senses of one or more human users through, for example, tactile output, sound, light, and smell / taste. Such human-machine interface output devices may include: tactile output devices (e.g., tactile feedback through a touch screen (1810), a data glove (not shown), or a joystick (1805), but there may also be tactile feedback devices that do not serve as input devices), audio output devices (such as: speakers (1809), headphones (not depicted)), visual output devices (such as a screen (1810), including a CRT screen, an LCD screen, a plasma screen, an OLED screen, each with or without touch screen input capabilities, each with or without tactile feedback capabilities - some of which may be able to output two-dimensional visual output or more than three-dimensional output through means such as stereoscopic output; virtual reality glasses (not depicted); holographic displays and smoke generators (not depicted)); and printers (not depicted).
[0183] The computer system (1800) may also include human-accessible storage devices and their associated media, such as optical media including a CD / DVD ROM / RW (1820) with media such as CD / DVD (1821), thumb drives (1822), removable hard disk drives or solid state drives (1823), traditional magnetic media such as tapes and floppy disks (not depicted), devices based on dedicated ROM / ASIC / PLD such as security dongles (not depicted), etc.
[0184] Those skilled in the art should also understand that the term "computer-readable medium" used in connection with the presently disclosed subject matter does not encompass transmission media, carrier waves, or other transient signals.
[0185] The computer system (1800) may also include an interface (1854) to one or more communication networks (1855). The network can be, for example, wireless, wired, or optical. The network can also be local, wide area, metropolitan area, vehicular, and industrial, real-time, delay-tolerant, etc. Examples of networks include: local area networks, such as Ethernet, wireless LAN; cellular networks, including GSM, 3G, 4G, 5G, LTE, etc.; television cable or wireless wide area digital networks, including cable television, satellite television, and terrestrial broadcast television; vehicular and industrial networks, including CAN bus, etc.
[0186] The above-mentioned human-machine interface devices, human-accessible storage devices, and network interfaces can be attached to the core (1840) of the computer system (1800).
[0187] The core (1840) may include one or more central processing units (CPUs) (1841), a graphics processing unit (GPU) (1842), a dedicated programmable processing unit in the form of a field programmable gate area (FPGA) (1843), a hardware accelerator for specific tasks (1844), a graphics adapter (1850), etc. These devices, together with a read-only memory (ROM) (1845), a random access memory (1846), an internal mass storage device such as an internal non-user-accessible hard disk drive, SSD, etc. (1847), may be connected via a system bus (1848). In some computer systems, the system bus (1848) may be accessed in the form of one or more physical plugs to enable expansion via additional CPUs, GPUs, etc. Peripheral devices may be attached directly or via a peripheral bus (1849) to the system bus (1848) of the core. In an example, a screen (1810) may be connected to the graphics adapter (1850). The architecture of the peripheral bus includes PCI, USB, etc.
[0188] A computer-readable medium may have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be those specially designed and constructed for the purposes of the present disclosure, or the medium and the computer code may be of the type well known and available to those of skill in the computer software art.
[0189] Although the present disclosure has described several exemplary embodiments, there are changes, permutations, and various equivalent substitutes that fall within the scope of the present disclosure. Accordingly, it will be understood that those skilled in the art will be able to envision numerous systems and methods that, although not explicitly shown or described herein, embody the principles of the present disclosure and are thus within the spirit and scope of the present disclosure.
Claims
1. A method for decoding a video bitstream performed by a decoder in a decoder, the method comprising: receiving the video bitstream, the video bitstream comprising a current block and a reference block, the reference block being used to predict the current block by using a prediction function; A first syntax element is received from the video bitstream, the first syntax element indicating how to determine prediction parameters of a prediction function, the prediction parameters comprising a scaling factor, and the first syntax element indicating at least one of: whether the scaling factor is explicitly signaled in the video bitstream via a syntax element; whether the scaling factor needs to be derived from the video bitstream; as well as whether a first subset of scaling factors is signaled in the video bitstream and a second subset of scaling factors is derived from the video bitstream, the second subset of scaling factors and the first subset of scaling factors not overlapping; determining the scaling factor based on the first syntax element; determining a local illumination variation between the current block and a prediction block corresponding to the current block based on the prediction function including the determined scaling parameter, a current block template, and a reference block template; predicting a current sample in the current block based at least in part on the local illumination variation and the prediction function; and The current block is reconstructed based on the predicted current samples.
2. The method according to claim 1, wherein: The prediction function is a weighted sum of at least one of: A co-located sample of the current sample located in the reference block; a set of neighboring samples of the co-located sample; a compensation value based on the co-located sample; or offset; and The scaling factor includes at least one of the following: a scaling factor corresponding to the co-located sample; a scaling factor corresponding to each of the set of neighboring samples; a scaling factor corresponding to the compensation value; or A scaling factor corresponding to the offset.
3. The method according to claim 2, wherein: The compensation value is determined using the following formula: (C*C+median_value)>>bit_depth Where: C is the value of the co-located sample; median_value is the median or average of all samples in the reference block; and bit_depth is the bit depth of the reference block.
4. The method according to claim 2, wherein: The set of neighboring samples only includes neighboring samples in the following directions of the co-located sample: an upper direction; a left direction; a lower direction; and a right direction.
5. The method according to claim 2, wherein: The set of neighboring samples only includes neighboring samples in the following directions of the co-located sample: a left direction; an upper right direction; a lower left direction; and a lower right direction. 6 . The method according to claim 1 , further comprising receiving from the video bitstream whether the current block is to be predicted using the prediction function.
7. The method according to any one of claims 1 to 5, wherein: The prediction function uses at least two scaling factors, and the current block is only allowed to be encoded in a single inter-frame prediction mode.
8. The method according to claim 7, wherein: The size of the current block is only allowed to be equal to or larger than 8×8 pixels.
9. The method according to any one of claims 1 to 5, further comprising: In response to the first syntax element indicating that the second subset of the scaling factors needs to be derived, the subset of the scaling factors is derived based on a template of the current block and a template of the reference block.
10. The method according to claim 9, wherein: Deriving the subset of the scaling factors includes deriving the subset of the scaling factors based on a template of the current block and a template of the reference block by using least mean square estimation.
11. The method according to any one of claims 1 to 5, wherein: At least one of the scaling factors is explicitly signaled in the video bitstream via a syntax element, the method further comprising: A second syntax element is received from the video bitstream, the second syntax element indicating the at least one of the scaling factors, the second syntax element carrying an index to a scaling factor in a lookup table storing candidate scaling factors.
12. The method according to claim 11, wherein: The at least one of the scaling factors comprises a scaling factor for one of: a co-located sample of the current sample located in the reference block; or a neighboring sample of the co-located sample.
13. The method according to any one of claims 1 to 5, wherein: A context for entropy encoding the first syntax element depends on whether a block vector is used to determine the reference block for the current block.
14. The method according to claim 13, wherein: A context in which the first syntax element is entropy encoded when the block vector is used is different from a context in which the first syntax element is entropy encoded when the block vector is not used.
15. The method according to any one of claims 1 to 5, wherein: The current block is predicted using a block vector and at least one of the scaling factors needs to be derived, the method further comprising: A third syntax element is received from the video bitstream, the third syntax element indicating which reference lines are to be used to derive the at least one of the scaling factors, wherein the reference lines are in at least one of: a template of the current block; or a template of the reference block.
16. The method according to any one of claims 1 to 5, wherein: The current block is predicted using a block vector and at least one of the scaling factors needs to be derived, the method further comprising: The at least one of the scaling factors is derived based on more than one neighboring reference lines selected from at least one of a template of the current block or a template of the reference block.
17. The method according to any one of claims 1 to 5, wherein: The current block is predicted using a block vector, and the prediction function is applied only to the luminance component of the current block.
18. The method according to any one of claims 1 to 5, wherein: The first syntax element is signaled via a high-level syntax, the high-level syntax being in at least one of the following levels: Sequence Parameter Set (SPS) level; Picture Parameter Set (PPS) level; Picture level; Slice level; or Tile level.
19. An apparatus for decoding a video bitstream, the apparatus comprising a memory for storing computer instructions and a processor in communication with the memory, wherein: When the processor executes the computer instructions, the processor is configured to cause the device to: receiving the video bitstream, the video bitstream comprising a current block and a reference block, the reference block being used to predict the current block by using a prediction function; A first syntax element is received from the video bitstream, the first syntax element indicating how to determine prediction parameters of a prediction function, the prediction parameters comprising a scaling factor, and the first syntax element indicating at least one of: whether the scaling factor is explicitly signaled in the video bitstream via a syntax element; whether the scaling factor needs to be derived from the video bitstream; as well as whether a first subset of scaling factors is signaled in the video bitstream and a second subset of scaling factors is derived from the video bitstream, the second subset of scaling factors and the first subset of scaling factors not overlapping; determining the scaling factor based on the first syntax element; determining a local illumination variation between the current block and a prediction block corresponding to the current block based on the prediction function including the determined scaling parameter, a current block template, and a reference block template; predicting a current sample in the current block based at least in part on the local illumination variation and the prediction function; and The current block is reconstructed based on the predicted current samples.
20. A non-transitory storage medium storing computer-readable instructions that, when executed by a processor in a decoder for decoding a video bitstream, cause the processor to: receiving the video bitstream, the video bitstream comprising a current block and a reference block, the reference block being used to predict the current block by using a prediction function; A first syntax element is received from the video bitstream, the first syntax element indicating how to determine prediction parameters of a prediction function, the prediction parameters comprising a scaling factor, and the first syntax element indicating at least one of: whether the scaling factor is explicitly signaled in the video bitstream via a syntax element; whether the scaling factor needs to be derived from the video bitstream; as well as whether a first subset of scaling factors is signaled in the video bitstream and a second subset of scaling factors is derived from the video bitstream, the second subset of scaling factors and the first subset of scaling factors not overlapping; determining the scaling factor based on the first syntax element; Determine a local illumination change between the current block and a prediction block corresponding to the current block based on the prediction function including the determined scaling parameter, a current block template, and a reference block template; predict a current sample in the current block based at least in part on the local illumination change and the prediction function; and The current block is reconstructed based on the predicted current samples.