Improvements to explicit signaling for block-based adaptive weighted predictions

Enhanced BAWP and LIC techniques address redundancy in video coding by using scale factors and lookup tables for improved luminance compensation, enhancing bitrate management and video quality in streaming applications.

JP2026515574APending Publication Date: 2026-05-19TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
TENCENT AMERICA LLC
Filing Date
2023-09-14
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing video coding and decoding technologies face challenges in efficiently reducing redundancy in uncompressed video signals, particularly in handling local luminance variations and bitrate requirements for streaming applications.

Method used

Enhanced Block Adaptive Weighted Prediction (BAWP) and Local Illumination Compensation (LIC) techniques are employed to improve video decoding and encoding processes by using scale factors and lookup tables to predict and reconstruct video blocks, incorporating motion compensation and loop filtering.

Benefits of technology

These techniques enhance video compression efficiency by reducing redundancy and improving luminance compensation, resulting in more effective bitrate management and improved video quality in streaming applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026515574000001_ABST
    Figure 2026515574000001_ABST
Patent Text Reader

Abstract

This disclosure relates in general to video coding / decoding, and more particularly to video coding / decoding for augmenting BAWP. One method includes the steps of: receiving a video bitstream including a current block and a reference block, wherein the reference block is used to predict the current block and is identified by a motion vector associated with the current block; receiving a first syntax element indicating a scale factor, wherein the scale factor is stored in one of two or more lookup tables maintained by the decoder to store candidate scale factors or candidate scale factor differences, and the candidate scale factor difference is the difference between a candidate scale factor and a threshold; selecting a lookup table; determining a scale factor based on the first syntax element and the selected lookup table; predicting the current block based on the reference block, the scale factor, and the offset; and reconstructing the current block based on the predicted current block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Incorporation by Reference

[0001] This application claims the benefit of priority based on U.S. Provisional Application No. 63 / 462,482, filed on April 27, 2023, and claims the benefit of priority based on U.S. Non-Provisional Application No. 18 / 461,759, filed on September 6, 2023, the entire contents of which are hereby incorporated by reference.

[0002]

[0002] This disclosure describes a set of advanced video / streaming coding / decoding techniques. More particularly, the disclosed techniques involve enhancing Block Adaptive Weighted Prediction (BAWP) and Local Illumination Compensation (LIC) to compensate for local luminance variations.

Background Art

[0003]

[0003] Uncompressed digital video can include a series of pictures and can have specific bitrate requirements for memory, data processing, and transmission bandwidth in streaming applications. One goal of video coding and decoding can be to reduce redundancy in the uncompressed input video signal through various compression techniques.

Summary of the Invention

Means for Solving the Problems

[0004]

[0004] This disclosure describes various embodiments of methods, apparatuses, and computer-readable storage media for enhancing Block Adaptive Weighted Prediction (BAWP) and modeling Local Illumination Compensation (LIC).

[0005]

[0005] According to one aspect, one embodiment of the present disclosure provides a method for decoding the current block of the current frame in a coded video bitstream. The method includes the steps of: receiving a video bitstream including a current block and a reference block, wherein the reference block is used to predict the current block and is identified by a motion vector associated with the current block; receiving a first syntax element from the video bitstream indicating a scale factor (α), wherein the scale factor is stored in one of two or more lookup tables maintained by the decoder to store candidate scale factors or candidate scale factor differences, and the candidate scale factor difference is the difference between a candidate scale factor and a threshold; selecting a lookup table to store the scale factor; determining the scale factor based on the value of the first syntax element and the selected lookup table; predicting the current block based on the reference block, the scale factor, and the offset; and reconstructing the current block based on the predicted current block.

[0006]

[0006] In another aspect, one embodiment of the present disclosure provides an apparatus for processing the current block of the current frame in a coded video bitstream. The apparatus includes a memory for storing instructions and a processor for communicating with the memory. When the processor executes an instruction, the processor is configured to cause the apparatus to perform the above method for video decoding and / or encoding.

[0007]

[0007] In another aspect, one embodiment of the present disclosure provides a non-temporary computer-readable medium that, when executed by a computer for video decoding and / or encoding, stores instructions causing the computer to perform the above method for video decoding and / or encoding.

[0008]

[0008] The above and other embodiments and their implementations will be described in more detail in the drawings, description and claims.

[0009]

[0009] Further features, properties, and various advantages of the disclosed subject matter will become clearer from the following detailed description and accompanying drawings. [Brief explanation of the drawing]

[0010] [Figure 1]

[0010] This is a schematic diagram of a simplified block diagram of a communication system (100) according to an exemplary embodiment. [Figure 2]

[0011] This is a schematic diagram of a simplified block diagram of a communication system (200) according to an exemplary embodiment. [Figure 3]

[0012] This is a schematic diagram of a simplified block diagram of a video decoder according to an exemplary embodiment. [Figure 4]

[0013] This is a schematic diagram of a simplified block diagram of a video encoder according to an exemplary embodiment. [Figure 5]

[0014] This is a block diagram of a video encoder according to another exemplary embodiment. [Figure 6]

[0015] This is a block diagram of a video decoder according to another exemplary embodiment. [Figure 7]

[0016] This figure shows a coding block partitioning scheme according to an exemplary embodiment of the present disclosure. [Figure 8]

[0017] This figure shows another scheme for coding block partitioning according to an exemplary embodiment of the present disclosure. [Figure 9]

[0018] This figure shows another scheme for coding block partitioning according to an exemplary embodiment of the present disclosure. [Figure 10]

[0019] This is a diagram showing the combined motion compensation. [Figure 11]

[0020] This figure shows an example of an interpolated reference frame for motion compensation. [Figure 12]

[0021] It is a diagram showing an example of a template of a current block and a template of a reference block. [Figure 13]

[0022] It is a diagram showing an exemplary logic flow for the method in the present disclosure. [Figure 14]

[0023] It is a schematic diagram of a computer system according to an exemplary embodiment of the present disclosure.

Mode for Carrying Out the Invention

[0011]

[0024] Here, a part of the present invention will be described in detail below with reference to the accompanying drawings that form a part of the present invention and show specific examples of embodiments as examples. However, it should be noted that the present invention may be embodied in various different forms, and thus, the covered or claimed subject matter is not intended to be construed as limited to any of the embodiments described below. Also, it should be noted that the present invention can be embodied as a method, a device, a component, or a system. Therefore, the embodiments of the present invention may take the form of, for example, hardware, software, firmware, or any combination thereof.

[0012]

[0025] Throughout this specification and the claims, terms may have meanings that are implied or suggested in context beyond their expressly stated meaning. Where used herein, the phrases “in one embodiment” or “in some embodiments” do not necessarily refer to the same embodiment, and where used herein, the phrases “in another embodiment” or “in other embodiments” do not necessarily refer to different embodiments. Similarly, where used herein, the phrases “in one implementation” or “in some implementations” do not necessarily refer to the same implementation, and where used herein, the phrases “in another implementation” or “in other implementations” do not necessarily refer to different implementations. For example, the claimed subject matter is intended to include, in whole or in part, exemplary combinations of embodiments / implementations.

[0013]

[0026] Generally, technical terms may be understood, at least in part, from their usage in context. For example, as used herein, terms such as "and," "or," or "and / or" may include various meanings that may depend, at least in part, on the context in which such terms are used. Typically, "or" when used to associate a list such as A, B, or C herein is intended to mean A, B, and C in an inclusive sense here, as well as A, B, or C in an exclusive sense here. Also, as used herein, the terms "one or more" or "at least one" may be used, at least in part depending on the context, to describe any feature, structure, or property in a singular sense, or to describe a combination of features, structures, or properties in a plural sense. Similarly, terms such as "a," "an," or "the" may also be understood, at least in part depending on the context, to convey a singular usage or a plural usage. Further, the terms "based on" or "determined by" may be understood not necessarily to convey an exclusive set of factors, but rather, again, at least in part depending on the context, to allow for the presence of additional factors that are not necessarily explicitly recited.

[0014]

[0027] As shown in Figure 1, terminal devices may be implemented as servers, personal computers, and smartphones, but the applicability of the basic principles of this disclosure is not limited to these. Embodiments of this disclosure may be implemented in desktop computers, laptop computers, tablet computers, media players, wearable computers, dedicated video conferencing equipment, etc. Network (150) represents any number or type of network that transmits coded video data between terminal devices, including, for example, wireline and / or wireless communication networks. The communication network (150) may exchange data on circuit-switched, packet-switched, and / or other types of channels. Typical networks include telecommunications networks, local area networks, wide area networks, and / or the Internet.

[0015]

[0028] Figure 2 shows the arrangement of a video encoder and video decoder in a video streaming environment as an example of the application of the disclosed subject matter. The disclosed subject matter may be equally applicable to other video applications, such as video conferencing, digital television broadcasting, games, virtual reality, and storage of compressed video on digital media, including CDs, DVDs, memory sticks, etc.

[0016]

[0029] As shown in Figure 2, the video streaming system may include a video capture subsystem (213) which may include a video source (201), for example, a digital camera, for creating a stream (202) of uncompressed video pictures or images. In one example, the stream (202) of video pictures includes samples recorded by the digital camera of the video source (201). The stream (202) of video pictures is drawn as a thick line to emphasize that it has a larger data volume compared to encoded video data (204) (or encoded video bitstream) and is processable by an electronic device (220) which includes a video encoder (203) coupled to the video source (201). The video encoder (203) may include hardware, software, or a combination thereof to enable or implement embodiments of the subject matter disclosed, as will be described in more detail below. The encoded video data (204) (or encoded video bitstream (204)) is drawn as a thin line to emphasize its smaller data volume compared to the stream of uncompressed video pictures (202), and can be stored in a streaming server (205) for future use or directly in a downstream video device (not shown). One or more streaming client subsystems, such as client subsystems (206) and (208) in Figure 2, can access the streaming server (205) to retrieve copies (207) and (209) of the encoded video data (204). The client subsystem (206) may include a video decoder (210) within, for example, an electronic device (230). The video decoder (210) decodes the incoming copy of the encoded video data (207) to create a stream (211) of outgoing video pictures that is uncompressed and can be rendered to a display (212) (e.g., a display screen) or other rendering device (not shown).

[0017]

[0030] Figure 3 shows a block diagram of a video decoder (310) of an electronic device (330) according to any embodiment of the present disclosure described below. The electronic device (330) may include a receiver (331) (e.g., a receiving circuit). The video decoder (310) can be used instead of the video decoder (210) in the example of Figure 2.

[0018]

[0031] As shown in Figure 3, the receiver (331) may receive one or more coded video sequences from the channel (301). To deal with network jitter and / or handle playback timing, a buffer memory (315) may be placed between the receiver (331) and the entropy decoder / analyzer (320) (hereinafter "analyzer (320)"). The analyzer (320) may reconstruct symbols (321) from the coded video sequences. Categories of such symbols include information used to manage the operation of the video decoder (310) and, optionally, information for controlling rendering devices such as a display (312) (e.g., a display screen). The analyzer (320) may analyze / entropy decode the coded video sequences. The analyzer (320) may extract from the coded video sequences a set of subgroup parameters for at least one of the pixel subgroups in the video decoder. Subgroups may include groups of pictures (GOP), pictures, tiles, slices, macroblocks, coding units (CU), blocks, transform units (TU), and prediction units (PU). The analyzer (320) may also extract information from the coded video sequence, such as transform coefficients (e.g., Fourier transform coefficients), quantizer parameter values, and motion vectors. Multiple different processing or function units may be involved in the reconstruction of the symbol (321). The units involved, and how they are involved, may be controlled by subgroup control information analyzed from the coded video sequence by the analyzer (320).

[0019]

[0032] The first unit may include a scaler / inverse unit (351). The scaler / inverse unit (351) may receive control information from the analyzer (320) as a symbol (321), which includes the quantized transformation coefficients, as well as information indicating which type of inverse transformation to use, the block size, quantization factors / parameters, the quantization scaling matrix, etc. The scaler / inverse unit (351) can output a block having sample values ​​that can be input to the aggregator (355).

[0020]

[0033] In some cases, the output samples from the scaler / inverse transform (351) may relate to intracoded blocks, i.e., blocks that do not use prediction information from previously reconstructed pictures but may use prediction information from parts reconstructed before the current picture. Such prediction information may be provided by an intrapicture prediction unit (352). In some cases, the intrapicture prediction unit (352) may generate a block of the same size and shape as the block being reconstructed, using surrounding block information that has already been reconstructed and is stored in the current picture buffer (358). The current picture buffer (358) buffers, for example, partially reconstructed current pictures and / or fully reconstructed current pictures. In some implementations, the aggregator (355) may add the prediction information generated by the intra prediction unit (352) to the output sample information, such as that provided by the scaler / inverse transform unit (351), on a sample-by-sample basis.

[0021]

[0034] In other cases, the samples output by the scaler / inverse unit (351) may relate to an intercoded and possibly motion-compensated block. In such cases, the motion-compensated prediction unit (353) can access the reference picture memory (357) based on the motion vector to fetch samples to be used for interpicture prediction. After motion-compensating the reference samples fetched according to the symbols (321) related to the block, these samples can be added by the aggregator (355) to the output of the scaler / inverse unit (351) (the output of unit 351 may also be called residual samples or residual signals) to generate output sample information.

[0022]

[0035] The samples output by the aggregator (355) can undergo various loop filtering techniques in the loop filter unit (356), including several types of loop filters. The output of the loop filter unit (356) can be output to the rendering device (312) and can be stored in the reference picture memory (357) for use in future interpicture prediction as a sample stream.

[0023]

[0036] Figure 4 shows a block diagram of a video encoder (403) according to an exemplary embodiment of the present disclosure. The video encoder (403) may be included in an electronic device (420). The electronic device (420) may further include a transmitter (440) (e.g., a transmitting circuit). The video encoder (403) can be used instead of the video encoder (403) in the example of Figure 4.

[0024]

[0037] The video encoder (403) may receive video samples from the video source (401). According to some exemplary embodiments, the video encoder (403) may code and compress the pictures of the source video sequence into a coded video sequence (443) in real time or under any other arbitrary time constraints, as required for each application. Enforcing an appropriate coding speed constitutes one function of the controller (450). In some embodiments, the controller (450) may be functionally coupled to and control other functional units, as described below. Parameters set by the controller (450) may include rate control-related parameters (picture skip, quantizer, lambda value of rate distortion optimization technique, ...), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc.

[0025]

[0038] In some exemplary embodiments, the video encoder (403) may be configured to operate in a coding loop. The coding loop may include a source coder (430) and a (local) decoder (433) built into the video encoder (403). Even if the built-in decoder 433 processes the video stream coded by the source coder 430 without entropy coding, the decoder (433) reconstructs the symbols to create sample data in a manner similar to that created by a (remote) decoder (because in entropy coding, any compression between the symbols and the coded video bitstream can be reversible in the video compression techniques considered in the disclosed subject). An observation that can be made at this point is that any decoder technique other than parsing / entropy decoding that may exist only in the decoder may necessarily exist in the corresponding encoder in a nearly identical functional form. For this reason, the disclosed subject may occasionally focus on the decoder operation, which is tied to the decoding portion of the encoder. Thus, the description of the encoder technique can be simplified, as it is the inverse of the more comprehensive description of the decoder technique. The following provides a more detailed description of the encoder, but only in specific areas or embodiments.

[0026]

[0039] In operation in some exemplary implementations, the source coder (430) may perform motion-compensated predictive coding, which predictively codes an input picture by referencing one or more previously coded pictures from a video sequence designated as “reference pictures”.

[0027]

[0040] The local video decoder (433) may decode the coded video data of a picture that may be designated as a reference picture. The local video decoder (433) may replicate the decoding process that may be performed on the reference picture by the video decoder and store the reconstructed reference picture in the reference picture cache (434). Thus, the video encoder (403) may locally store a copy of the reconstructed reference picture that has the same content as the reconstructed reference picture that will be obtained by the (transmission error-free) far-end (remote) video decoder.

[0028]

[0041] The predictor (435) may perform a predictive search for the coding engine (432). That is, for a new picture to be coded, the predictor (435) may search the reference picture memory (434) for sample data (as candidate reference pixel blocks) or specific metadata such as reference picture motion vectors, block shapes, etc., which may serve as appropriate predictive references for the new picture.

[0029]

[0042] The controller (450) may manage the coding operations of the source coder (430), including, for example, setting parameters and subgroup parameters used to encode video data.

[0030]

[0043] The outputs of all the aforementioned functional units may undergo entropy coding in the entropy coder (445). The transmitter (440) may buffer the coded video sequence, such as that created by the entropy coder (445), in preparation for transmission over the communication channel (460), which may be a hardware / software link to a storage device that stores the coded video data. The transmitter (440) may merge the coded video data from the video coder (403) with other data being transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).

[0031]

[0044] The controller (450) may manage the operation of the video encoder (403). During coding, the controller (450) may assign a specific coded picture type to each coded picture, which may influence the coding technique that can be applied to each picture. For example, a picture may often be assigned as one of the following picture types: intra picture (I picture), predictive picture (P picture), bidirectional predictive picture (B picture), or multiple predictive picture. As will be described in more detail below, a source picture may generally be spatially subdivided into multiple sample coding blocks.

[0032]

[0045] Figure 5 shows a diagram of a video encoder (503) according to another exemplary embodiment of the present disclosure. The video encoder (503) is configured to receive a processing block (e.g., a prediction block) of sample values ​​in the current video picture within a sequence of video pictures, and to encode the processing block into a coded picture which is part of a coded video sequence. The exemplary video encoder (503) may be used instead of the video encoder (403) in the example of Figure 4.

[0033]

[0046] For example, the video encoder (503) receives a matrix of sample values ​​for a processing block. The video encoder (503) then uses, for example, rate-distortion optimization (RDO) to determine whether it is best for the processing block to be coded using intra-mode, inter-mode, or bi-predictive mode.

[0034]

[0047] In the example shown in Figure 5, the video encoder (503) includes an interencoder (530), an intraencoder (522), a residual computer (523), a switch (526), ​​a residual encoder (524), a master controller (521), and an entropy encoder (525), all coupled together as shown in the illustrative arrangement of Figure 5.

[0035]

[0048] The interencoder (530) is configured to receive a sample of the current block (e.g., a processing block), compare the block with one or more reference blocks in the reference picture (e.g., blocks in the previous and subsequent pictures in the display order), generate interprediction information (e.g., a description of redundant information by an intercoding technique, motion vectors, merge mode information), and compute an interprediction result (e.g., a predicted block) based on the interprediction information using any preferred technique.

[0036]

[0049] The intra encoder (522) is configured to receive a sample of the current block (e.g., a processing block), compare the block to a block already coded within the same picture, generate transformed quantized coefficients, and optionally also generate intra prediction information (e.g., intra prediction direction information by one or more intra coding techniques).

[0037]

[0050] The master controller (521) may be configured to determine master control data, control other components of the video encoder (503) based on the master control data, for example, determine the prediction mode of a block, and provide control signals to the switch (526) based on the prediction mode.

[0038]

[0051] The residual computer (523) may be configured to calculate the difference (residual data) between the received block and the prediction result for a block selected from the intra encoder (522) or interencoder (530). The residual encoder (524) may be configured to encode the residual data to generate transformation coefficients. The transformation coefficients are then subjected to a quantization process to obtain quantized transformation coefficients. In various exemplary embodiments, the video encoder (503) also includes a residual decoder (528). The residual decoder (528) is configured to perform an inverse transformation to generate decoded residual data. The entropy encoder (525) may be configured to format the bitstream to include the encoded blocks and to perform entropy coding.

[0039]

[0052] Figure 6 shows a diagram of an illustrative video decoder (610) according to another embodiment of the present disclosure. The video decoder (610) is configured to receive a coded picture which is part of a coded video sequence, and to decode the coded picture to produce a reconstructed picture. In one example, the video decoder (610) may be used instead of the video decoder (410) in the example of Figure 4.

[0040]

[0053] In the example in Figure 6, the video decoder (610) includes an entropy decoder (671), an interdecoder (680), a residual decoder (673), a reconstruction module (674), and an intradecoder (672) coupled together as shown in the illustrative arrangement in Figure 6.

[0041]

[0054] An entropy decoder (671) may be configured to reconstruct specific symbols representing the syntax elements that make up a coded picture from the coded picture. An interdecoder (680) may be configured to receive interprediction information and generate interprediction results based on the interprediction information. An intradecoder (672) may be configured to receive intraprediction information and generate prediction results based on the intraprediction information. A residual decoder (673) may be configured to perform inverse quantization to extract dequantized conversion coefficients and process the dequantized conversion coefficients to convert the residuals from the frequency domain to the spatial domain. A reconstruction module (674) may be configured to combine the residuals output by the residual decoder (673) and the prediction results (optionally output by the inter or intraprediction module) in the spatial domain to form reconstructed blocks as part of the reconstructed video that form part of the reconstructed picture.

[0042]

[0055] It should be noted that the video encoders (203), (403), and (503), as well as the video decoders (210), (310), and (610), may be implemented using any preferred technique. In some exemplary embodiments, the video encoders (203), (403), and (503), as well as the video decoders (210), (310), and (610), may be implemented using one or more integrated circuits. In another embodiment, the video encoders (203), (403), and (503), as well as the video decoders (210), (310), and (610), may be implemented using one or more processors that execute software instructions.

[0043]

[0056] Moving on to coding and decoding block partitioning, general partitioning may begin with a base block and follow a predefined set of rules, a specific pattern, a partition tree, or any partition structure or scheme. Partitioning may be hierarchical and recursive. After partitioning or dividing the base block according to one of the example partitioning procedures described below, or other procedures, or a combination thereof, a final set of partitions or coding blocks may be obtained. Each of these partitions may be one of the various partitioning levels in the partitioning hierarchy and may be of various shapes. Each partition may also be called a coding block (CB). In the various example partitioning implementations described further below, each resulting CB may be of any of the allowed sizes and partitioning levels. Such partitions are called coding blocks because they may form units on which several basic coding / decoding decisions may be made, and on which coding / decoding parameters can be optimized, determined, and signaled within the encoded video bitstream. The highest or deepest level in the final partition represents the depth of the coding block partitioning structure of the tree. A coding block may be a lumen coding block or a chroma coding block. The CB tree structure for each color may be called a coding block tree (CBT). The coding blocks for all color channels may be collectively called coding units (CU). The hierarchical structure for all color channels may be collectively called coding tree units (CTU). The partitioning patterns or structures for the various color channels within a CTU may be the same or different.

[0044]

[0057] In some implementations, the partition tree scheme or structure used for lumern and chroma channels may not be the same. In other words, lumern and chroma channels may have separate coding tree structures or patterns. Furthermore, whether lumern and chroma channels use the same or different coding partition tree structures, and the actual coding partition tree structures used, may depend on whether the slice being coded is a P-slice, a B-slice, or an I-slice. For example, in an I-slice, chroma and lumern channels may have separate coding partition tree structures or coding partition tree structure modes, while in a P-slice or B-slice, lumern and chroma channels may share the same coding partition tree scheme. When separate coding partition tree structures or modes are applied, a lumern channel may be partitioned into CBs by one coding partition tree structure, and a chroma channel may be partitioned into chroma CBs by another coding partition tree structure.

[0045]

[0058] Figure 7 shows 10 exemplary predefined partitioning structures / patterns that allow for the formation of a partition tree by recursive partitioning. The root block may start at a predefined level (e.g., from a base block at the 128x128 or 64x64 level). The exemplary partitioning structures in Figure 7 include various 2:1 / 1:2 and 4:1 / 1:4 rectangular partitions. In some exemplary implementations, none of the rectangular partitions in Figure 7 can be further subdivided. A coding tree depth may be further defined to indicate the partitioning depth from the root node or root block. For example, the coding tree depth of the root node or root block may be set to 0, and if the root block is partitioned one more time according to Figure 7, the coding tree depth increases by 1. In some implementations, only all square partitions in 710 may be able to recursively partition to the next level of the partitioning tree according to the patterns in Figure 7.

[0046]

[0059] In some other exemplary implementations of coding block partitioning, a quadtree structure may be used. Such quadtree partitioning may be applied hierarchically and recursively to partitions of any square shape. Whether the base block or intermediate block or partition is further quadtree partitioned may be adapted to various local characteristics of the base block or intermediate block / partition.

[0047]

[0060] In some further examples, a ternary partitioning scheme may be used to partition a base block or any intermediate block, as shown in Figure 8. The ternary pattern may be implemented vertically, as shown in 802, or horizontally, as shown in 804. In the example partition ratio in Figure 8, it is shown as 1:2:1, but other ratios may be predefined. In some implementations, two or more different ratios may be predefined. In some implementations, the width and height of the partitions in the example ternary tree are always powers of 2 to avoid additional transformations.

[0048]

[0061] The above partitioning schemes may be combined in any way at different partitioning levels. For example, the base block can be a quadtree-binary (QTBT). To partition into a -tree structure, the quadtree partitioning scheme and the binary partitioning scheme described above may be combined. In such a scheme, the base block or intermediate block / partition may be quadtree partitioned or binary partitioned, if specified, subject to a set of predefined conditions. A particular example is shown in Figure 9, in which the base block is initially quadtree partitioned into four partitions, as indicated by 902, 904, 906, and 908. Each of the resulting partitions may then be quadtree partitioned into four further partitions (e.g., 908) or binary partitioned into two further partitions (e.g., 902 or 906, which are either horizontal or vertical and both symmetric), or remain unpartitioned (e.g., 904). Binary or quadtree partitioning may be recursively allowed for square-shaped partitions, as shown by the overall example partitioning pattern in 910 and the corresponding tree structure / representation in 920, where solid lines represent quadtree partitioning and dashed lines represent binary partitioning. A flag may be used for each binary node (non-leaf binary partition) to indicate whether the binary is horizontal or vertical. For example, as shown in 920 and consistent with the partitioning structure in 910, flag "0" may represent horizontal binary and flag "1" may represent vertical binary. In quadtree partitioning, there is no need to indicate the partition type, as quadtree partitioning always divides a block or partition both horizontally and vertically to produce four subblocks / partitions of equal size. In some implementations, flag "1" may represent horizontal binary and flag "0" may represent vertical binary.

[0049]

[0062] In some exemplary implementations of QTBT, the rule sets for quadtree partitioning and binary partitioning may be represented by the following predefined parameters and their associated corresponding functions. - CTU size: The size of the root node of the quadtree (the size of the base block). - MinQTSize: Minimum allowable quadtree leaf node size - MaxBTSize: Maximum allowable binary tree root node size - MaxBTDepth: Maximum allowable binary tree depth - MinBTSize: Minimum allowable binary tree leaf node size

[0050]

[0063] In some exemplary implementations of the QTBT partitioning structure, the CTU size may be set as a 128x128 chroma sample with two corresponding 64x64 blocks of chroma samples (when exemplary chroma subsampling is considered and used), MinQTSize may be set as 16x16, MaxBTSize may be set as 64x64, MinBTSize (for both width and height) may be set as 4x4, and MaxBTDepth may be set as 4. Quadratic partitioning may be applied to the CTU first to generate quadratic leaf nodes. Quadratic leaf nodes may have sizes ranging from their minimum allowable size of 16x16 (i.e., MinQTSize) to 128x128 (i.e., CTU size). If a node is 128x128, the node will not be initially partitioned by a binary tree because its size exceeds MaxBTSize (i.e., 64x64). Otherwise, nodes that do not exceed MaxBTSize may be partitioned by a binary tree. In the example in Figure 9, the base block is 128x128. The base block can only be quadruped according to a predefined set of rules. The base block has a partitioning depth of 0. Each of the resulting four partitions is 64x64 and does not exceed MaxBTSize, and may be further quadruped or binary-tree partitioned at level 1. The process continues. When the binary tree depth reaches MaxBTDepth (i.e., 4), further partitioning may not be considered. When a binary tree node has a width equal to MinBTSize (i.e., 4), further horizontal partitioning may not be considered. Similarly, when a binary tree node has a height equal to MinBTSize, further vertical partitioning is not considered.

[0051]

[0064] In some exemplary implementations, the above QTBT scheme may be configured to support flexibility for lumens and chromens by having the same QTBT structure or separate QTBT structures. For example, in the case of P-slice and B-slice, the lumens CTB and chromens CTB in one CTU may share the same QTBT structure. However, in the case of I-slice, the lumens CTB may be partitioned into CBs by a QTBT structure, and the chromens CTB may be partitioned into chromens CBs by another QTBT structure. This means that CUs can be used to point to different color channels in an I-slice; for example, an I-slice may consist of a coding block for the lumens component or a coding block for two chromens, and a CU in a P-slice or B-slice may consist of a coding block for all three color components.

[0052]

[0065] The various CB partitioning schemes described above, and further partitioning of CBs into PBs, may be combined in any manner. The following specific implementations are provided as non-limiting examples.

[0053]

[0066] Interpretation may be implemented, for example, in single-reference mode or composite-reference mode. In some implementations, a skip flag may be initially included in the bitstream of the current block (or at a higher level) to indicate whether the current block is being intercoded and should not be skipped. If the current block is being intercoded, another flag may be further included in the bitstream as a signal to indicate whether single-reference mode or composite-reference mode is used for predicting the current block. In single-reference mode, one reference block may be used to generate a prediction block about the current block. In composite-reference mode, two or more reference blocks may be used, for example, by a weighted average, to generate a prediction block. One or more reference blocks may be identified using one or more reference frame indices, and further, using one or more corresponding motion vectors indicating the shift between the reference block and the current block in location relative to the frame, for example, in horizontal and vertical pixels. For example, in single-reference mode, the prediction block for the current block may be generated from a single-reference block identified by a single motion vector in the reference frame, while in composite-reference mode, the prediction block may be generated by a weighted average of two reference blocks in two reference frames, indicated by two reference frame indices and two corresponding motion vectors. The motion vectors may be coded in various ways and included in the bitstream.

[0054]

[0067] In some exemplary implementations, one or more reference picture lists, including the identification of short-term and long-term reference frames for inter-prediction, may be formed based on information in a Reference Picture Set (RPS). For example, a single picture reference list may be formed for unidirectional inter-prediction and denoted as L0 reference (or reference list 0), while two picture reference lists may be formed for bidirectional inter-prediction and denoted as L0 (or reference list 0) and L1 (or reference list 1) for each of the two prediction directions. The reference frames included in the L0 and L1 lists may be ordered in various predetermined ways. The lengths of the L0 and L1 lists may be signaled in the video bitstream. Unidirectional inter-prediction may be in single-reference mode, or in composite-reference mode when multiple references for generating prediction blocks by weighted averaging in composite-prediction mode are on the same side of the frame in which the predicted block lies. Bidirectional inter-prediction can only be in composite mode, as it involves at least two reference blocks in bidirectional inter-prediction.

[0055]

[0068] In some implementations, a merge mode (MM) for interpretation may be implemented. Generally, in merge mode, one or more motion vectors in a single reference prediction, or in a composite reference prediction for the current PB, may be derived from other motion vectors rather than being independently computed and signaled. For example, in an encoding system, the current motion vector of the current PB may be represented by the difference between the current motion vector and one or more other already encoded motion vectors (called reference motion vectors). Such a difference in the motion vector, rather than the entire current motion vector, may be encoded and included in the bitstream and linked to the reference motion vector. Correspondingly, in a decoding system, the motion vector corresponding to the current PB may be derived based on the decoded difference motion vector and the decoded reference motion vector linked to it. As a unique form of general merge mode (MM) interpretation, such interpretation based on difference motion vectors is sometimes called Merge Mode with Motion Vector Difference (MMVD). Therefore, a general MM, or in particular an MMVD, may be implemented to improve coding efficiency by leveraging the correlation between motion vectors associated with different PBs. For example, adjacent PBs may have similar motion vectors, and thus the MVD may be small and can be coded efficiently. As another example, motion vectors may also be temporally correlated (between frames) with respect to blocks that are similarly in / placed in space.

[0056]

[0069] In some exemplary implementations of MMVD, a list of reference motion vectors (RMVs), or MV predictor candidates for motion vector prediction, may be formed for the block being predicted. The list of RMV candidates may contain a predetermined number (e.g., two) of MV predictor candidate blocks, the motion vectors of which may be used to predict the current motion vector. The RMV candidate blocks may include blocks selected from adjacent blocks in the same frame and / or from temporal blocks (e.g., blocks that are in exactly the same position in frames preceding or following the current frame). These choices represent blocks that are spatially or temporally located relative to the current block and are likely to have a similar or identical motion vector to the current block. The size of the list of MV predictor candidates may be predetermined. For example, the list may contain two or more candidates. To be on the list of RMV candidates, candidate blocks may need to have the same reference frame (or multiple reference frames) as the current block, must exist (for example, boundary checking must be performed when the current block is near the edge of a frame), must have already been encoded during the encoding process, and / or must have already been decoded during the decoding process. In some implementations, the list of merge candidates may first populate spatially adjacent blocks (scanned in a specific predefined order) if available and satisfying the above conditions, and then temporal blocks if space is still available in the list. Adjacent RMV candidate blocks may be selected from, for example, blocks to the left and above the current block. The list of RMV predictor candidates may be dynamically formed as a Dynamic Reference List (DRL) at various levels (sequence, picture, frame, slice, superblock, etc.). The DRL may be signaled with a bitstream.

[0057]

[0070] In some implementations, the actual MV predictor candidates used as reference motion vectors to predict the motion vector of the currently coded block may be signaled. If the RMV candidate list contains two candidates, a one-bit flag called a merge candidate flag may be used to indicate the selection of the reference merge candidate. If the currently coded block is predicted in composite mode, each of the multiple motion vectors predicted using the MV predictors may be associated with a reference motion vector from the merge candidate list. The encoder may determine which of the RMV candidates more faithfully predicts the MV of the currently coded block and signal the selection as an index to the DRL.

[0058]

[0071] In some exemplary implementations of MMVD, after an RMV candidate is selected and used as a base motion vector predictor for the predicted motion vector, a differential motion vector (MVD or delta MV representing the difference between the predicted motion vector and the reference candidate motion vector) may be computed in the encoding system. Such an MVD may include information representing the magnitude and direction of the differential MV, and both the magnitude and direction of the differential MV may be signaled in the bitstream in various ways.

[0059]

[0072] In some exemplary implementations of MMVD, the distance index may be used to specify the magnitude information of the difference motion vector and to indicate one of a set of predefined offsets that represent a predefined difference motion vector from the starting point (reference motion vector). The MV offset by the signaled index may then be added to either the horizontal or vertical component of the starting (reference) motion vector. Exemplary predefined relationships between the distance index and the predefined offsets are shown in Table 1.

[0060] [Table 1]

[0061]

[0073] In some exemplary implementations of MMVD, a direction index may be further signaled and used to represent the direction of the MVD relative to the reference motion vector. In some implementations, the direction may be constrained to either the horizontal or vertical direction. Exemplary 2-bit direction indices are shown in Table 2. In the examples in Table 2, the interpretation of the MVD may differ depending on the information of the start / reference MV. For example, when the start / reference MV corresponds to a single predictive block, or to a double predictive block where both reference frame lists point to the same side of the current picture (i.e., the POCs of both reference pictures are either greater than or less than the POC of the current picture), the sign in Table 2 may specify the sign (direction) of the MV offset added to the start / reference MV. If the start / reference MV corresponds to a biprediction block where two reference pictures are on opposite sides of the current picture (i.e., the POC of one reference picture is greater than the POC of the current picture, and the POC of the other reference picture is smaller than the POC of the current picture), and the difference between the reference POC in picture reference list 0 and the current frame is greater than the difference between the reference POC in picture reference list 1 and the current frame, then the sign in Table 2 may specify the sign of the MV offset added to the reference MV corresponding to the reference picture in picture reference list 0, and the sign for the offset of the MV corresponding to the reference picture in picture reference list 1 may have the opposite value (the opposite sign for the offset). Otherwise, if the difference between the reference POC in picture reference list 1 and the current frame is greater than the difference between the reference POC in picture reference list 0 and the current frame, then the sign in Table 2 may specify the sign of the MV offset added to the reference MV associated with picture reference list 1, and the sign for the offset for the reference MV associated with picture reference list 0 may have the opposite value.

[0062] [Table 2]

[0063]

[0074] In some exemplary implementations, MVD may be scaled according to the difference in POCs in each direction. If the difference in POCs in both lists is the same, scaling is not necessary. If, instead, the difference in POCs in reference list 0 is greater than that in reference list 1, the MVD for reference list 1 is scaled. If the difference in POCs in reference list 1 is greater than that of list 0, the MVD for list 0 may be scaled in the same way. If the starting MV is single-predicted, the MVD is added to the available MV or reference MV.

[0064]

[0075] In some exemplary implementations of MVD coding and signaling for bidirectional composite prediction, in addition to coding and signaling two MVDs separately, or as an alternative, symmetric MVD coding may be implemented such that only one MVD requires signaling, and the other MVD can be derived from the signaled MVD. In such implementations, motion information, including the reference picture indices of List-0 and List-1, is not signaled together. Specifically, at the slice level, a flag called "mvd_l1_zero_flag" may be included in the bitstream to indicate whether reference list-1 is not signaled in the bitstream. If this flag is 1, indicating that reference list-1 is equal to zero (and therefore not signaled), a bidirectional prediction flag called "BiDirPredFlag" may be set to 0, which means there is no bidirectional prediction. Instead, if mvd_l1_zero_flag is zero and the nearest neighbor reference picture in list-0 and the nearest neighbor reference picture in list-1 form a forward-to-backward or backward-to-forward reference picture pair, then BiDirPredFlag may be set to 1, and both reference pictures in list-0 and list-1 become short-circuit reference pictures. Otherwise, BiDirPredFlag is set to 0. A BiDirPredFlag of 1 may indicate that a symmetric mode flag is signaled in the bitstream as an additional. The decoder may extract the symmetric mode flag from the bitstream when BiDirPredFlag is 1. The symmetric mode flag may be signaled, for example, at the CU level (if necessary) to indicate whether a symmetric MVD coding mode is being used for the corresponding CU.When the symmetric mode flag is 1, it indicates the use of the symmetric MVD coding mode, and that only the reference picture indices in both List-0 and List-1 (called "mvp_l0_flag" and "mvp_l1_flag") are signaled along with the MVD associated with List-0 (called "MVD0"), while the other differential motion vector "MVD1" is derived rather than signaled. For example, MVD1 may be derived as -MVD0. Thus, in the exemplary symmetric MVD mode, only one MVD is signaled.

[0065]

[0076] In some other exemplary implementations of MV prediction, a cooperative scheme may be used to implement a common merge mode, i.e., MMVD, and several other types of MV prediction for both single-reference mode and compound-reference mode MV prediction. Various syntactic elements may be used to signal the way in which the MV of the current block is predicted. For example, in single-reference mode, the following MV prediction modes may be signaled:

[0066]

[0077] NEARMV uses one of the motion vector predictors (MVPs) in a list indicated by a DRL (Dynamic Reference List) index.

[0067]

[0078] NEWMV - Uses one of the motion vector predictors (MVPs) in a list signaled by a DRL index as a reference, and applies the delta to the MVP.

[0068]

[0079] GLOBALMV - Uses motion vectors based on global motion parameters at the frame level.

[0069]

[0080] Similarly, if the composite reference interpretation prediction mode uses two reference frames corresponding to the two predicted MVs, the following MV prediction modes may be signaled:

[0070]

[0081] NEAR_NEARMV - Uses one of the motion vector predictors (MVPs) in a list signaled by the DRL index.

[0071]

[0082] NEAR_NEWMV - Uses one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference and sends a delta MV to the second MV.

[0072]

[0083] NEW_NEARMV - Uses one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference and sends a delta MV to the first MV.

[0073]

[0084] NEW_NEWMV - Uses one of the motion vector predictors (MVPs) in the list signaled by the DRL index as a reference and sends a delta MV for both MVs.

[0074]

[0085] GLOBAL_GLOBALMV - Uses MV from each reference based on its frame-level global motion parameters.

[0075]

[0086] The term "NEAR" above refers to MV prediction using a reference MV without MVD as a general merge mode, while the term "NEW" refers to MV prediction involving the use of a reference MV and offsetting it using a signaled differential motion vector (MVD), as in MMVD mode. In the case of composite interpretation, the reference base motion vector and the motion vector delta above may be correlated, and such correlation may be utilized to reduce the amount of information required to signal the two motion vector deltas, although generally they may be different or independent between the two references. In such situations, joint signaling of the two MVDs may be implemented and represented in a bitstream.

[0076]

[0087] The above dynamic reference list (DRL) may be used to hold a set of indexed motion vectors that are dynamically maintained and considered as candidate motion vector predictors.

[0077]

[0088] Differential motion vector coding

[0078]

[0089] In some example implementations, coding techniques such as AV1 allow / support fractional motion vector precision (or accuracy), such as 1 / 8 of a pixel (i.e., one-eighth of a pixel), and the following syntax is used to signal the differential motion vector in reference frame list 0 (L0) or list 1 (L1). • mv_joint specifies which components of the difference motion vector are non-zero. ○ 0 indicates the absence of non-zero MVD along either the horizontal or vertical direction. ○ 1 indicates that there is a non-zero MVD only along the horizontal direction. ○ 2 indicates that there is a non-zero MVD only along the vertical direction. ○ 3 indicates that there is a non-zero MVD along both the horizontal and vertical directions. • mv_sign specifies whether the difference motion vector is positive or negative. • `mv_class` specifies the class of the differential motion vector. As shown in Table 3, a higher class means a larger differential motion vector.

[0079] [Table 3] • mv_bit specifies the integer part of the offset between the difference motion vector and the starting magnitude for each MV class. • mv_fr specifies the first two fractional bits of the difference motion vector. • mv_hp specifies the third fractional bit of the differential motion vector.

[0080]

[0090] Adaptive MVD resolution

[0081]

[0091] In some exemplary implementations, the resolution of MVDs may be differentiated across different MVD size classes. For example, a high-resolution MVD for larger MVDs in higher MVD classes may not result in a statistically significant improvement in compression efficiency. Therefore, MVDs may be coded with reduced resolution (integer pixel resolution or fractional pixel resolution) for a wider range of MVD size corresponding to higher MVD size classes. Similarly, MVDs may generally be coded with reduced resolution (integer pixel resolution or fractional pixel resolution) for larger MVD values. Such MVD class-dependent or MVD size-dependent MVD resolutions can generally be called adaptive MVD resolutions.

[0082]

[0092] Statistical observations suggest that treating the MVD resolution of large or high-class MVDs at the same level as that of small or low-class MVDs in a non-adaptive manner may not significantly improve inter-predictive residual coding efficiency for blocks using large or high-class MVDs. This suggests that the reduction in signaling bits achieved by aiming for a less accurate MVD using adaptive MVD resolution may outweigh the additional bits required to code the inter-predictive residuals resulting from such a less accurate MVD. In other words, using a higher MVD resolution for large or high-class MVDs may not yield more coding gain than using a lower MVD resolution.

[0083]

[0093] In some exemplary implementations, further constraints may be imposed on the composite reference mode, such as NEW_NEARMV and NEAR_NEWMV, as described above. Specifically, the precision of the differential motion vector (MVD) depends on the associated class and the size of the MVD.

[0084]

[0094] In some exemplary implementations, a decimal MVD is allowed only when the size of the MVD is one pixel or less. Alternatively or additionally, in some exemplary implementations, only one MVD value is allowed when the value of the relevant MV class is MV_CLASS_1 or greater, and the MVD values ​​for each MV class are derived as 4, 8, 16, 32, and 64 for MV classes 1 (MV_CLASS1), 2 (MV_CLASS2), 3 (MV_CLASS_3), 4 (MV_CLASS_4), or 5 (MV_CLASS_5), respectively.

[0085]

[0095] As an example, the acceptable MVD values ​​for each MV class are shown in Table 4.

[0086] [Table 4] In addition, if the block is currently coded as NEW_NEARMV or NEAR_NEWMV mode, a certain context is used to signal mv_joint or mv_class. Otherwise, a different context is used to signal mv_joint or mv_class. Note that in AV1, mv_class specifies the class of the differential motion vector. A higher class means that the differential motion vector represents a larger update, and mv_joint specifies which components of the differential motion vector are non-zero.

[0087]

[0096] Improvement of Adaptive MVD Resolution

[0088]

[0097] In some exemplary implementations, further improvements to adaptive MVD may be implemented. For single references, a new intercoded mode named AMVDMV may be added. When the AMVDMV mode is selected and / or flagged, this indicates that AMVD is applied to signal MVD in such a way that adaptive MVD resolution is adopted.

[0089]

[0098] One solution involves adding a flag named amvd_flag under JOINT_NEWMV mode to indicate whether AMVD is applied to Joint MVD coding mode. When Adaptive MVD Resolution is applied to Joint MVD coding mode, which may also be called Joint AMVD coding mode, the MVDs of two reference frames are joined and signaled, and the precision of the MVD is implicitly determined, for example, by the size of the MVD. Otherwise, the MVDs of two (or more than three) reference frames are joined and signaled, and conventional MVD coding is applied.

[0090]

[0099] Alternatively, MVD for two (or more) reference frames may be joined and signaled, and conventional MVD coding will apply. In this case, instead of adding an amvd_flag as described above, a new inter-prediction mode named JOINT_AMVDNEWMV is added to indicate that AMVD is applying to the joint MVD coding mode.

[0091]

[0100] Adaptive Motion Vector Resolution (AMVR)

[0092]

[0101] In some exemplary implementations, AMVR may be implemented using various coding techniques such as AV1. For example, a total of seven MV accuracies (e.g., 8, 4, 2, 1, 1 / 2, 1 / 4, 1 / 8) are supported. For each prediction block, an encoder (e.g., AVM, AV1 encoder) searches all supported accuracy values ​​and signals the decoder to the best accuracy.

[0093]

[0102] To reduce the complexity of encoder execution time, two precision sets are supported. Each precision set may include, for example, four predefined precisions. One of the precision sets is adaptively selected at the frame level based on the maximum precision value of the frame. For example, the maximum precision may be signaled in the frame header. Table 5 summarizes the supported precision values ​​based on the maximum precision at the frame level.

[0094] [Table 5]

[0095]

[0103] In some exemplary implementations, there is a frame-level flag to indicate whether the frame's MV includes sub-per- (i.e., sub-pixel) precision. AMVR is only enabled when the value of the cur_frame_force_integer_mv flag is 0. In AMVR, if the block precision is lower than the maximum precision, the motion model and interpolation filter are not signaled. If the block precision is lower than the maximum precision, the motion mode may be inferred to be translational motion, and the interpolation filter is inferred to be a REGULAR interpolation filter. Similarly, if the block precision is 4-per- or 8-per-, the inter-intra mode is not signaled and is inferred to be 0.

[0096]

[0104] Joint MVD Coding (JMVD)

[0097]

[0105] In some exemplary implementations, coding techniques such as AV1 apply an intercoded mode named JOINT_NEWMV to indicate whether the MVDs of two reference lists are joined and signaled. When the interprediction mode is equal to the JOINT_NEWMV mode, the MVDs of reference list 0 and reference list 1 are joined and signaled. In this case, only one MVD named joint_mvd is signaled and sent to the decoder, and the delta MVs of reference list 0 and reference list 1 are derived from joint_mvd.

[0098]

[0106] In some exemplary implementations, the JOINT_NEWMV mode is signaled along with the NEAR_NEARMV, NEAR_NEWMV, NEW_NEARMV, NEW_NEWMV, and GLOBAL_GLOBALMV modes. No further context is required.

[0099]

[0107] When the JOINT_NEWMV mode is signaled and the Picture Order Count (POC) distance between the two reference frames and the current frame is different, the MVD is scaled for reference list 0 or reference list 1 based on the POC distance. Specifically, the distance between reference frame list 0 and the current frame can be denoted as td0, and the distance between reference frame list 1 and the current frame can be denoted as td1. If td0 is greater than or equal to td1, joint_mvd becomes the MVD that is signaled as a joint and is used directly for reference list 0, while the mvd for reference list 1 is derived from joint_mvd based on equation (1).

[0100]

number

[0101]

[0108] Instead, if td1 is greater than or equal to td0, joint_mvd is used directly for reference list 1, and the mvd for reference list 0 is derived from joint_mvd based on equation (2).

[0102]

number

[0103]

[0109] Improved joint MVD coding

[0104]

[0110] In some exemplary implementations, when a block is coded in a coding technique such as AV1 as a joint MVD coding mode, e.g., JOINT_NEWMV or JOINT_AMVDNEWMV, a new syntax named mvd_scaling_factor_idx is signaled to the bitstream to explicitly indicate the MVD scaling factor between reference frame 0 and reference frame 1.

[0105]

[0111] As an example, two predefined lookup tables may be used to store separately supported / allowed scale factors for JOINT_NEWMV or JOINT_AMVDNEWMV, as shown in Tables 6 and 7 below. The relevant entry index for the selected scale factor in the lookup table is signaled in the bitstream. In JOINT_AMVDNEWMV mode, the same scale factor is applied to both the vertical and horizontal components of the MVD for reference framelists 0 and / or 1. In JOINT_NEWMV mode, the scale factor for one component of the MVD (either the vertical or horizontal component) is constrained to 1, while the scale factor for the other component of the MVD can be 2 or other values ​​such as 1 / 2. In one example, the MVD for reference framelists 0 and 1 is calculated using the following formula: mvd_ref0=joint_mvd (3)

[0106]

number

[0107]

[0112] Here, mvd_ref0 and mvd_ref1 represent the MVD of reference frame list 1 and reference frame list 2, respectively. The distance between reference frame list 0 and the current frame is denoted as td0, and the distance between reference frame list 1 and the current frame is denoted as td1. joint_mvd represents the MVD that is signaled by the joint, and jmvd_scale represents the scale factor.

[0108] [Table 6]

[0109] [Table 7]

[0110]

[0113] Biprediction with CU-level weights (BCW)

[0111]

[0114] In some exemplary implementations, in video coding techniques such as HEVC, the dual-prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or by using two different motion vectors. In VVC, the dual-prediction mode is extended beyond simple averaging to allow weighted averaging of the two prediction signals. For example, P bi-pred The bidirectional prediction, as indicated by the notation, may be calculated using Equation 5 below. P bi-pred =((8-w)*P0+w*P1+4>>3 (5)

[0112]

[0115] As an example, in weighted averaging biprediction, five weights are allowed, i.e., w ∈ {-2, 3, 4, 5, 10}. When w is equal to 4, equal weighting factors are used to perform a weighted average of the two predicted samples. For each CU being bipredicted, the weight w is determined in one of two ways: 1) In non-merged CUs, the weight index is signaled after the difference motion vector. 2) In merged CUs, the weight index is inferred from adjacent blocks based on the merge candidate index.

[0113]

[0116] In some example implementations, BCW is applied only to CUs that have 256 or more lumens (i.e., CU width × CU height is 256 or greater). For low-latency pictures, all five weights are used. For non-low-latency pictures, only three weights (w ∈ {3, 4, 5}) are used.

[0114]

[0117] Local Luminance Compensation (LIC)

[0115]

[0118] LIC is a video coding tool that can be used by video encoders and video decoders. In some exemplary implementations, LIC may be applied based on a linear model for compensating for luminance changes (e.g., during motion compensation) between one or more temporal reference pictures and the current picture. The linear model is based on LIC parameters including a scale factor α and an offset β, which are described in detail below in the section on block adaptive weighting prediction. LIC can be enabled or disabled, for example, through high-level signaling at various levels.

[0116]

[0119] In some exemplary implementations, a bipredictive reference template may be generated for template samples associated with the current block. Reference template samples may be identified based on one or more motion vectors associated with the current block. For example, a reference template sample may include a temporal reference CU that is a neighbor of the current CU, and may correspond to a template sample for the current CU. Reference template samples may be considered jointly (e.g., averaged) in the LIC parameter derivation.

[0117]

[0120] In some exemplary implementations, the least mean-squared error (LMSE) algorithm may be applied to derive the LIC parameters. For example, an LMSE-based calculation may be performed to determine the LIC parameters so that the difference between the bipredicted reference template reference sample and the current CU template sample is minimized.

[0118]

[0121] In some exemplary implementations, a similar approach may be used for single prediction. In this case, the LIC parameters may be determined so as to minimize the difference between the reference template reference sample being single predicted and the template sample of the current CU.

[0119]

[0122] It should be noted that the LMSE algorithm described herein is merely one example of how to derive LIC parameters. One or more other approaches / algorithms may be used.

[0120]

[0123] Block-Adaptive Weighted Prediction (BAWP)

[0121]

[0124] In some example implementations, BAWP is used to model local luminance fluctuations.

[0122]

[0125] Referring to Figure 12, for example, BAWP may be a block-level weighted prediction to model the local luminance variation between the current block and its predicted block as a function of the local luminance variation between the current block template (or causally related sample of the current block) and the reference block template. Figure 12 shows the template for the current block (1212) (or referred to as current template 1210) and the template for the reference block (1222) (or referred to as reference template 1220). Each template may include an upper and left portion. For example, current block 1212 includes an upper portion 1214 and a left portion 1216. The reference block may be indicated or determined by a motion vector (MV1230). The current block may be in the current picture (or current frame), and the reference block may be in the reference picture (or reference frame). In some implementations, the function may be a linear function. The parameters of the function may be expressed by a scale factor α and an offset β, which form a linear equation. The scale factor is sometimes called the scale, alpha factor, or alpha value.

[0123]

[0126] The following is an example of a linear function used to compensate for brightness changes when a block is coded in BAWP mode. p'(x') = α*p(x) + β (6)

[0124]

[0127] Here, p'(x') is the predicted sample at location x' in the current block (or the predicted sample at the predicted unit (PU) in the current block), p(x) is the sample corresponding to p'(x') at location x in the reference block, α is the scale factor (or scale), and β is the offset value. Note that the reference block may be identified or derived from the MV associated with the current block, and p(x) is the reference sample pointed to by the MV at location x in the reference picture. Note that in Equation 6, the reference sample and the predicted sample may have the same coordinates (i,j) in their respective blocks (i.e., the reference block and the current block, or the reference block and the predicted unit). Alternatively, the coordinates of the reference sample in the reference block may be based on the coordinates of the predicted sample in the current block. For example, the coordinates of the predicted sample may be adjusted by a delta value to obtain the coordinates (in the reference picture) of the corresponding reference sample.

[0125]

[0128] In some exemplary implementations, α and β may be derived based on the current block template and the reference block template, and therefore no signaling overhead is required for them, except that the BAWP flag is signaled for a single inter-prediction mode to indicate the use of BAWP. In some exemplary implementations, the BWAP method is applied only to blocks that are 8x8 or larger in size and coded in a single inter-prediction mode. In some exemplary implementations, the BAWP method is applied only to the ruma component.

[0126]

[0129] Deriving both scale factors and offsets (α and β) saves some signaling overhead, but in certain scenarios, there may be some potential drawbacks. For example, the effectiveness and / or accuracy of scale factors depend on the similarity between the current block and its template in the reference frame. When the correlation between the current block and its template in the reference frame is high, the derived scale factors can significantly improve the accuracy of predictions. However, if the current block and its template in the reference frame are significantly different, the scale factors may be inaccurate and not even useful, potentially impairing the accuracy of predictions.

[0127]

[0130] This disclosure discloses various embodiments for improving video coding / decoding techniques in BAWP mode and / or LIC mode, with the aim of enhancing prediction accuracy while minimizing overhead for signaling costs. Specifically, various methods for signaling and / or deriving scale factors and offsets (α and β) as specified in Equation 6 are described.

[0128]

[0131] Based on research and statistical observations, high-precision and high-accuracy scale factors can help improve coding efficiency and increase coding gain. High precision is particularly beneficial when the scale factor is small. Therefore, signaling the scale factor in the video bitstream instead of deriving it on the decoder side may have several advantages. Furthermore, sorting the supported (candidate) scale factors according to a specific scheme when the decoder stores them may help improve entropy coding efficiency. Various embodiments for achieving these goals are described below in this disclosure.

[0129]

[0132] In this disclosure, the term "block" may refer to a transformed block, a coded block, a predicted block, a coded block, a coding unit (CU), etc. The term "chroma block" may refer to a block in any of the chromaticity (color) channels. The orientation of a reference frame is determined by whether the reference frame is before or after the current frame in the display order.

[0130]

[0133] In this disclosure, a sample can be interpreted as the pixel value of a pixel. This can generally refer to any component (luma, or chroma).

[0131]

[0134] In this disclosure, the terms x-axis and y-axis refer to the horizontal and vertical components of a 2D value. These may also be replaced by two other axes along two predefined directions perpendicular to each other, and the same embodiments apply similarly. That is, the x-axis and y-axis may be rotated by a certain angle. For example, the x-axis and y-axis may be replaced by a 45-degree axis and a 135-degree axis.

[0132]

[0135] In this disclosure, "conventional JMVD" may refer to a JMVD having normal full MV resolution or a JMVD having AMVR.

[0133]

[0136] In this disclosure, unless otherwise specified, a signaling may include one or more subsignalings. These subsignalings may be transmitted together or separately.

[0134]

[0137] In the embodiments described below, coding blocks or coded blocks may be coded in BAWP (or LIC or compound weighted prediction (CWP)) mode, and the BAWP mode will be referred to as BAWP hereafter for simplicity of explanation. These embodiments may be implemented in decoders and / or encoders.

[0135]

[0138] In one embodiment, when the current block is predicted from its reference block using a linear function having a scale factor and an offset β, such as the linear equation 6 shown above, the selection of the scale factor and / or offset β may be signaled to the bitstream and analyzed on the decoder side in order to reconstruct the predicted block. The reference block may be specified, for example, by a motion vector associated with the current block. That is, the scale factor and / or offset β may be explicitly signaled rather than derived by the decoder. In some exemplary implementations, only the scale factor may be signaled and the offset β may be derived based on the scale factor, or vice versa. In this disclosure, the scale factor may also be referred to as the scale, scale factor, scale factor α, or alpha factor.

[0136]

[0139] In one embodiment, all supported values ​​of the scale factor may be stored in a predefined lookup table, and the index of the scale factor in the lookup table is signaled in the bitstream and parsed on the decoder side. In this case, by signaling the index in the lookup table, the signaling explicitly indicates or identifies the scale factor.

[0137]

[0140] In one embodiment, the supported values ​​in a predefined lookup table are distributed symmetrically along a single threshold value (TH). For example, the supported values ​​stored in the lookup table may be [TH-d0, TH-d1, TH-d2, ...TH, ..., TH+d2, TH+d1, TH+d0], where TH, d0, d1, and d2 are natural numbers. For example, TH=1, d0=0.1, d1=0.2, and d2=0.3. In another example, TH itself is not included in the predefined lookup table (the values ​​in the lookup table are still distributed symmetrically along TH).

[0138]

[0141] In some exemplary implementations, the TH value may be signaled in High Level Signaling (HLS), and the syntax may be signaled through at least one of the following levels: Sequence Parameter Set (SPS) level, Picture Parameter Set (SPS) level, Picture level, Slice level, Tile level, and Coding Tree Unit (CTU) level. For example, the signaling may be carried in headers corresponding to these levels, such as a CTU header.

[0139]

[0142] In some example implementations, TH may be set to 0 or 1.

[0140]

[0143] In one embodiment, instead of the symmetric distribution described above, the order of the numerical values ​​(scale factors) in the lookup table may depend on the magnitude of each scale factor, with smaller values ​​for these scale factors resulting in smaller indices. In other words, the lookup table may be sorted based on the magnitude of the scale factors.

[0141]

[0144] For example, the order of numbers in a lookup table may depend on the absolute difference between the magnitude of the scale factor and the TH. Such a difference may be called the distance between the scale factor and the TH.

[0142]

[0145] In one embodiment, in order to improve coding efficiency and increase coding gain, the context for entropy coding / decoding (signaling) the index or value of the scale factor and / or the index or value of the offset (β) value may depend on already decoded information associated with the current block and / or already decoded information associated with adjacent blocks, which may include, but are not limited to, the block size, the interprediction mode, the reference frame of the current block, the distance between the reference frame associated with the reference block and the current frame to which the current block belongs, the temporal level of the current frame in a group of pictures (GOP) structure, the temporal level of the reference frame in a GOP structure, an index that identifies the scale factor stored in a lookup table, the index used to decode the adjacent block of the current block, the value of a selected scale factor used to decode the adjacent block of the current block, and so on.

[0143]

[0146] In one embodiment, a plurality of predefined lookup tables may be supported, and the decoder may be configured to have and maintain a plurality of predefined lookup tables. The selection of a particular lookup table for each block may be implicitly determined / derived based on already decoded information associated with the current block and / or already decoded information associated with adjacent blocks, which includes, but is not limited to, block size, interpretation mode, reference frame of the current block, distance between the reference frame associated with the reference block and the current frame to which the current block belongs, temporal level of the current frame in the GOP structure, temporal level of the reference frame in the GOP structure, index identifying a scale factor stored in the lookup table used to decode adjacent blocks of the current block, and a value of a selected scale factor used to decode adjacent blocks of the current block.

[0144]

[0147] In one embodiment, multiple predefined lookup tables are supported, and the selection of a lookup table for each block is explicitly signaled to the bitstream and parsed on the decoder side.

[0145]

[0148] In one embodiment, the context for signaling the lookup table index may depend on already decoded information associated with the current block and / or adjacent blocks, which may include, but are not limited to, block size, interpretation mode, reference frame of the current block, distance between the reference frame associated with the reference block and the current frame to which the current block belongs, temporal level of the current frame in the GOP structure, temporal level of the reference frame in the GOP structure, an index identifying a scale factor stored in the lookup table used to decode the adjacent blocks of the current block, and a value of a selected scale factor used to decode the adjacent blocks of the current block.

[0146]

[0149] In one embodiment, instead of directly storing the scale factors, the difference between the magnitude of all supported scale factors and the threshold TH (this difference is called the scale factor difference) may be stored in a predefined lookup table, and the index of the scale factor difference in the lookup table is signaled in a bitstream and analyzed on the decoder side.

[0147]

[0150] In some exemplary implementations, the values ​​or magnitudes of scale factor differences stored in a predefined lookup table are distributed symmetrically along a single threshold TH, such as 0.

[0148]

[0151] In some exemplary implementations, the threshold TH may be predefined or signaled in the video bitstream.

[0149]

[0152] In some exemplary implementations, the decoder may derive a scale factor α (or a quantized value of the scale factor α), which may be used as a basis for determining the final, more accurate scale factor α. In this case, TH may be set to the derived scale factor α (or a quantized value of the scale factor α), which may be calculated, for example, based on local luminance fluctuations between the current block template (or causally related samples of the current block) and a reference block template.

[0150]

[0153] In some exemplary implementations, once a new TH value is derived as described above, the lookup table may be updated and resorted based on the new TH value.

[0151]

[0154] In some example implementations, zero is not included in the predefined lookup table.

[0152]

[0155] In some exemplary implementations, the TH value may be signaled in High Level Signaling (HLS), and the syntax may be signaled through at least one of the following levels: Sequence Parameter Set (SPS) level, Picture Parameter Set (SPS) level, Picture level, Slice level, Tile level, and Coding Tree Unit (CTU) level. For example, the signaling may be carried in headers corresponding to these levels, such as a CTU header.

[0153]

[0156] In some example implementations, TH may be set to 0 or 1.

[0154]

[0157] In one embodiment, instead of the symmetric distribution described above, the order of the scale factor differences in the lookup table may depend on their respective magnitudes, with smaller values ​​for these scale factor differences resulting in smaller indices. That is, the lookup table may be sorted based on the magnitude of the scale factor differences.

[0155]

[0158] For example, the sequence of values ​​in a lookup table (i.e., scale factor differences) may depend on the absolute difference between the scale factor's magnitude and the TH. Such a difference may be called the distance between the scale factor and the TH.

[0156]

[0159] In one embodiment, the offset value β may be derived from a linear equation between the reference block and the current block.

[0157]

[0160] In some example implementations, the offset value β may be set to (cur_template_mean - α * ref_template_mean), where cur_template_mean is the mean of the samples in the current block's template and ref_template_mean is the mean of the samples in the referenced block's template.

[0158]

[0161] In one embodiment, either the scale factor α or the offset β may be derived from adjacent reconstructed samples of the current block and the reference block, while the other is signaled in the bitstream.

[0159]

[0162] In one embodiment, the precision (e.g., accuracy or step size) of the scale factor α and the offset β may be different. In another embodiment, the scale factor α and the offset β may have the same precision.

[0160]

[0163] In some exemplary implementations, the offset β may have higher precision because it is derived. On the other hand, the scaling factor may have lower precision to cover a larger range, taking into account signaling overhead constraints such as the number of bits allocated for signaling. Lower precision allows for fitting a larger range with the same number of signaling bits.

[0161]

[0164] In one embodiment, the predicted sample for the current block is obtained by subtracting the mean value of the corresponding reference block from the reference sample, multiplying the result by a signaled scaling factor α for the block, and adding an offset value β.

[0162]

[0165] For example, this prediction process can be represented by the following equation 7. Cur pred =α*(ref_sample-ref_mean)+offset (7)

[0163]

[0166] Here, ref_sample is the reference sample (in the reference block) corresponding to the predicted sample in the current block, ref_mean is the arithmetic mean (or mean) of the reference block, and Cur pred This is the prediction result for the predicted samples within the current block.

[0164]

[0167] In some exemplary implementations, the offset value β in Equation 7 may be set to the arithmetic mean (or average) of the reference blocks.

[0165]

[0168] In one embodiment, a scale factor (either the value of the scale factor or an index that identifies the scale factor in a lookup table) may be signaled to the bitstream, and an offset β (either the value of the offset β or an index that identifies the offset β in a lookup table) may be signaled or derived. Furthermore, whether β is signaled or derived can be determined by, for example, the coding mode of the current block, the coding mode of the adjacent block, the number of adjacent blocks coded in BAWP mode, explicitly signaled flags, etc.

[0166]

[0169] In this disclosure, signaling (e.g., syntax elements) may indicate values ​​such as scale factors and / or offset β, either explicitly or implicitly. Explicit indication may be performed by sending an index to a lookup table to find the value, or by sending the value directly. When the value is sent implicitly, the decoder may need to perform further derivation based on the signaled syntax element. In the case of implicit indication, the syntax element may also be called the one that "associates" the value to be derived.

[0167]

[0170] The various embodiments and / or implementations described herein may be implemented separately or in combination in any order. Furthermore, each of the methods (or embodiments), encoders, and decoders may be implemented by processing circuits (e.g., one or more processors or one or more integrated circuits). One or more processors execute a program stored on a non-temporary computer-readable medium. In this disclosure, the term "block" may be interpreted as a prediction block, coding block, or coding unit (CU).

[0168]

[0171] Figure 13 shows a flowchart 1300 of an exemplary method following the principles based on the above implementation for indicating scale factors and / or offset values. An exemplary decoding method flow may include part or all of the following steps: S1310 receiving a video bitstream containing a current block and a reference block, wherein the reference block is used to predict the current block and is identified by a motion vector associated with the current block; S1320 receiving a first syntax element from the video bitstream indicating a scale factor (α), wherein the scale factor is stored in one of two or more lookup tables maintained by the decoder to store candidate scale factors or candidate scale factor differences, and the candidate scale factor difference is the difference between a candidate scale factor and a threshold; S1330 selecting a lookup table to store the scale factor; S1340 determining the scale factor based on the value of the first syntax element and the selected lookup table; S1350 predicting the current block based on the reference block, the scale factor, and the offset; and S1360 reconstructing the current block based on the predicted current block.

[0169]

[0172] In any part or combination of the above implementation configurations, the current block may be located within the current frame, and the reference block may be located within the reference frame.

[0170]

[0173] In any part or combination of the above implementation, the scaling factor may be the coefficient of the first-order model p'(x') = α*p(x) + β, where p'(x') is the sample in the current block at location x', and p(x) is the reference sample corresponding to p'(x') at location x in the reference block. p(x) may also be the variation or adjustment of the reference sample, and β is the offset.

[0171]

[0174] In any part or combination of the above implementations, the variation of the reference sample is obtained by subtracting the arithmetic mean of the reference block from the reference sample, and the implementation may further include determining that the offset is the arithmetic mean of the reference block.

[0172]

[0175] In any part or combination of the above implementation forms, the lookup table storing the scale factor may be explicitly signaled (for example, as an index to the lookup table or as a value of the scale factor) or implicitly signaled (the signaled syntax element is used to derive the scale factor).

[0173]

[0176] In this disclosure, the orientation of a reference frame may be determined by whether the reference frame is before or after the current frame in the display order.

[0174]

[0177] The operations described above may be combined or arranged in any quantity or order as needed. Two or more of the processes and / or operations may be performed in parallel. The embodiments and implementations in this disclosure may be used separately or in combination in any order. The processes in one embodiment / method may be divided to form a plurality of sub-methods, each of which may be independent of the other processes in the embodiment or may form a standalone solution. Furthermore, each of the methods (or embodiments), encoders, and decoders may be implemented by processing circuits (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-temporary computer-readable medium. Embodiments of this disclosure may be applied to rumor blocks or chroma blocks. The term block may be interpreted as a prediction block, coding block, or coding unit, i.e., CU. The term block as herein may also be used to refer to a transformation block. In the following sections, when referring to block size, block size may refer to the block width or height, the maximum or minimum width and height of the block, the area (width * height), or the aspect ratio (width:height, or height:width).

[0175]

[0178] The techniques described above can be implemented as computer software, using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 14 shows a computer system (1800) suitable for implementing a particular embodiment of the subject matter disclosed.

[0176]

[0179] Computer software can be coded using any suitable machine code or computer language, which may undergo mechanisms such as assembly, compilation, and linking to create code with executable instructions, either directly or through interpretation, microcode execution, etc., by one or more computer central processing units (CPUs), graphics processing units (GPUs), etc.

[0177]

[0180] The instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, and Internet of Things devices.

[0178]

[0181] The components shown in Figure 14 for the computer system (1800) are essentially illustrative and are not intended to imply any limitation on the scope of use or functionality of computer software implementing embodiments of this disclosure. Furthermore, the configuration of the components should not be construed as having any dependency or requirement on any one or any combination thereof of the components shown in the exemplary embodiments of the computer system (1800).

[0179]

[0182] The computer system (1800) may include certain human interface input devices. The input human interface devices may include one or more of the following (only one of each is shown): a keyboard (1801), a mouse (1802), a trackpad (1803), a touchscreen (1810), a data glove (not shown), a joystick (1805), a microphone (1806), a scanner (1807), and a camera (1808).

[0180]

[0183] The computer system (1800) may also include certain human interface output devices. Such human interface output devices may stimulate the senses of one or more human users, for example, through tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback via a touchscreen (1810), data glove (not shown), or joystick (1805), but which may also be tactile feedback devices that do not function as input devices), audio output devices (e.g., speakers (1809), headphones (not shown)), visual output devices (screens (1810), including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capability, each with or without tactile feedback capability, some of which may be capable of outputting two-dimensional visual output or three-dimensional or more output through means such as stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).

[0181]

[0184] The computer system (1800) may also include human-accessible storage devices and associated media, such as CD / DVD ROM / RW (1820) media including CD / DVD media (1821), thumb drives (1822), removable hard drives or solid-state drives (1823), legacy magnetic media such as tapes and floppy disks (not shown), and dedicated ROM / ASIC / PLD-based devices such as security dongles (not shown).

[0182]

[0185] Those skilled in the art will understand that the term “computer-readable medium” as used in relation to the subject matter of this disclosure does not include transmission media, carrier waves, or other transient signals.

[0183]

[0186] The computer system (1800) may also include an interface (1854) to one or more communication networks (1855). The networks may be, for example, wireless networks, wireline networks, or optical networks. Furthermore, the networks may be local networks, wide area networks, metropolitan networks, vehicle and industrial networks, real-time networks, delay-tolerant networks, etc. Examples of networks include local area networks such as Ethernet, wireless LANs, cellular networks including GSM, 3G, 4G, 5G, LTE, etc., TV wireline or wireless wide area digital networks including cable TV, satellite TV, and terrestrial broadcast TV, and vehicle and industrial networks including CAN bus.

[0184]

[0187] The aforementioned human interface devices, human-accessible memory devices, and network interfaces can be attached to the core (1840) of the computer system (1800).

[0185]

[0188] The core (1840) may include one or more central processing units (CPUs) (1841), graphics processing units (GPUs) (1842), dedicated programmable processing units in the form of field-programmable gate areas (FPGAs) (1843), hardware accelerators for specific tasks (1844), graphics adapters (1850), etc. These devices may be connected via a system bus (1848) along with read-only memory (ROM) (1845), random access memory (1846), internal mass storage such as internal user-inaccessible hard drives (1847), SSDs, etc. In some computer systems, the system bus (1848) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices can be connected directly to the core's system bus (1848) or via a peripheral bus (1849). For example, a screen (1810) may be connected to a graphics adapter (1850). The architecture for peripheral buses includes PCI, USB, etc.

[0186]

[0189] Computer-readable media can have computer code on them for performing operations implemented on various computers. The media and computer code can be specifically designed and constructed for the purposes of this disclosure, or they can be of a type that is well known and available to those skilled in the computer software field.

[0187]

[0190] While this disclosure has described several exemplary embodiments, there are many modifications, substitutions, and equivalents that fall within the scope of this disclosure. Therefore, it will be understood that a number of systems and methods not expressly shown or described herein, but embodying the principles of this disclosure and thus falling within the spirit and scope of this disclosure, can be devised by those skilled in the art.

Claims

1. A method for processing video data in a decoder, A step of receiving a video bitstream including a current block and a reference block, wherein the reference block is used to predict the current block and is identified by a motion vector associated with the current block; A step of receiving a first syntax element representing a scale factor (α) from the video bitstream, wherein the scale factor is stored in one of two or more lookup tables maintained by the decoder for storing candidate scale factors or candidate scale factor differences, and the candidate scale factor difference is the difference between the candidate scale factor and a threshold. A step of selecting the lookup table that stores the scale factor, A step of determining the scale factor based on the value of the first syntax element and the selected lookup table, A step of predicting the current block based on the reference block, the scale factor, and the offset, A step of reconstructing the current block based on the predicted current block, Methods that include...

2. The step of determining the lookup table includes the step of determining the lookup table based on decoded information associated with the current block or the reference block, and the decoded information is Block size, Interpretation mode, The reference frame of the current block, The distance between the reference frame associated with the aforementioned reference block and the current frame to which the current block belongs, The temporal level of the current frame in the Group of Pictures (GOP) structure, The temporal level of the reference frame in the GOP structure, An index that identifies a scale factor stored in one of the two or more lookup tables, wherein the scale factor is used to decode the adjacent block of the current block, or The value of the scale factor selected to decode the adjacent block of the current block. The method according to claim 1, comprising at least one of the following.

3. The step of determining the lookup table is, The steps include receiving a second syntax element from the video bitstream that indicates one of the two or more lookup tables maintained by the decoder, A step of determining the lookup table based on the value of the second syntax element, The method according to claim 1, including the method described in claim 1.

4. The context used to entropy encode the second syntax element is based on the decoded information associated with the current block or the reference block, and the decoded information is Block size, Interpretation mode, The reference frame of the current block, The distance between the reference frame associated with the aforementioned reference block and the current frame to which the current block belongs, The temporal level of the current frame in the GOP structure, The temporal level of the reference frame in the GOP structure, An index that identifies a scale factor stored in one of the two or more lookup tables, wherein the scale factor is used to decode the adjacent block of the current block, or The value of the scale factor selected to decode the adjacent block of the current block. The method according to claim 3, comprising at least one of the following.

5. The method according to claim 1, wherein the candidate scale factor differences in each of the two or more lookup tables are sorted.

6. The method according to claim 5, wherein the candidate scale factor differences in each of the two or more lookup tables are distributed symmetrically along the threshold based on the magnitude of each of the candidate scale factor differences.

7. A step of deriving a scale factor predicted based on local luminance fluctuations between the template of the current block and the template of the reference block, The steps of setting the threshold to the predicted scale factor and The method according to claim 6, further comprising:

8. A step of receiving a high-level syntax indicating the threshold from the video bitstream, wherein the high-level syntax is Sequence parameter set (SPS) level, Picture Parameter Set (PPS) level, Picture level, Slice level, Tile level, or Coding Tree Unit (CTU) level The method according to claim 5, further comprising the step of signaling at least one of the following levels.

9. The method according to claim 5, wherein the threshold is predetermined and includes 0 or 1.

10. The method according to claim 5, wherein the candidate scale factor differences in each of the two or more lookup tables are sorted based on the magnitude of each of the candidate scale factor differences.

11. The method according to claim 5, wherein none of the two or more lookup tables contain a value of 0.

12. The process further includes deriving the offset using the following formula: β=cur_template_mean-α*ref_template_mean The method according to claim 1, wherein cur_template_mean is the average of the samples in the template of the current block, and ref_template_mean is the average of the samples in the template of the reference block.

13. The method according to claim 1, further comprising the step of deriving the offset from adjacent reconstructed samples of the current block and adjacent reconstructed samples of the reference block.

14. The method according to claim 1, wherein the accuracy of the offset is higher than the accuracy of the scale factor.

15. Explicitly signaled flags, The coding mode of the current block mentioned above, The coding mode of the adjacent block to the current block, or The number of adjacent blocks to the current block, coded in BWAP mode. A step of determining whether to derive or obtain the offset using a third syntax element carried in the video bitstream based on one of the following: The decision to obtain the offset using the third syntax element includes the step of determining the offset based on the value of the third syntax element. The method according to claim 1, further comprising:

16. The method according to claim 1, wherein the current block is coded in either a block adaptive weighting prediction (BAWP) mode or a local luminance compensation (LIC) mode.

17. The first syntax element includes an index that identifies the scale factor in the lookup table, The aforementioned scaling factor (α) functions as the slope of the linear equation, and the linear equation takes the following form: p'(x')=α*p(x)+β Here, p'(x') is a sample in the current block at location x', p(x) is a reference sample corresponding to p'(x'), which is either at location x in the reference block or a variation of the reference sample, and β is the offset. The method according to claim 1, wherein the step of predicting the current block includes the step of predicting the current block based on the reference block and the linear equation.

18. The variation of the reference sample is obtained by subtracting the arithmetic mean of the reference block from the reference sample. The method according to claim 17, further comprising the step of determining that the offset is the arithmetic mean of the reference block.

19. A program that causes at least one processor in a computer to perform the method according to any one of claims 1 to 18.

20. A device for processing video data, wherein the device comprises a memory for storing computer instructions and a processor for communicating with the memory, and when the processor executes the computer instruction, the processor provides the device with Receiving a video bitstream including a current block and a reference block, wherein the reference block is used to predict the current block and is identified by the motion vector associated with the current block. Receiving a first syntax element from the video bitstream that indicates a scale factor (α), wherein the scale factor is stored in one of two or more lookup tables maintained by the decoder for storing candidate scale factors or candidate scale factor differences, and the candidate scale factor difference is the difference between the candidate scale factor and a threshold. Selecting the lookup table that stores the scale factor, The scaling factor is determined based on the value of the first syntax element and the selected lookup table, Predicting the current block based on the aforementioned reference block, the scale factor, and the offset, Reconstructing the current block based on the predicted current block and A device configured to perform a certain action.

21. A non-temporary storage medium for storing computer-readable instructions, wherein when the computer-readable instructions are executed by a processor, the processor receives the instructions. Receiving a video bitstream including a current block and a reference block, wherein the reference block is used to predict the current block and is identified by the motion vector associated with the current block. Receiving a first syntax element from the video bitstream that indicates a scale factor (α), wherein the scale factor is stored in one of two or more lookup tables maintained by the decoder for storing candidate scale factors or candidate scale factor differences, and the candidate scale factor difference is the difference between the candidate scale factor and a threshold. Selecting the lookup table that stores the scale factor, The scaling factor is determined based on the value of the first syntax element and the selected lookup table, Predicting the current block based on the aforementioned reference block, the scale factor, and the offset, Reconstructing the current block based on the predicted current block and A non-temporary storage medium that enables the operation of [a certain function].