Index reordering of bi-prediction with CU-level weight (BCW) by using template-matching

JP2025090674A5Pending Publication Date: 2025-09-08TENCENT AMERICA LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025035849
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-10-21
Filing Date
2025-03-06
Publication Date
2025-09-08

AI Technical Summary

Technical Problem

Existing video encoding technologies face challenges in efficiently encoding and decoding videos, particularly in reducing redundancy and improving compression efficiency, especially with the increasing complexity of intra prediction directions and motion vector prediction mechanisms.

Method used

The proposed solution involves an apparatus and method for video decoding that uses template matching (TM) to reorder bi-prediction with coding unit (CU)-level weights (BCW) by determining a respective TM cost for each BCW candidate weight, based on a current template and bi-prediction sub-templates, to select the optimal BCW weight for reconstructing the current block.

Benefits of technology

This approach enhances video encoding and decoding efficiency by optimizing the selection of BCW weights through template matching, thereby improving compression performance and reducing computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a method and apparatus for video decoding.SOLUTION: A video decoding apparatus includes processing circuitry that decodes prediction information indicating bi-prediction with coding unit (CU)-level weights (BCW) for a current block in a current picture. The processing circuitry performs template matching (TM) on BCW candidate weights by determining a respective TM cost corresponding to each BCW candidate weight. Each TM cost is determined based on a portion or all of a current template of the current block and a respective bi-predictor template. The bi-predictor template is based on the respective BCW candidate weight, a portion or all of a first reference template in a first reference picture, and a portion or all of a second reference template in a second reference picture. The processing circuitry reorders the BCW candidate weights based on the respectively determined TM costs.SELECTED DRAWING: Figure 26
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] [Incorporation by Reference] This application claims the benefit of priority to U.S. Patent Application No. 17 / 971,255, filed October 21, 2022, entitled "INDEX REORDERING OF BI-PREDICTION WITH CU-LEVEL WEIGHT (BCW) BY USING TEMPLATE-MATCHING", which claims the benefit of priority to U.S. Provisional Application No. 63 / 274,286, filed November 1, 2021, entitled "INDEX REORDERING OF BI-PREDICTION WITH CU-LEVEL WEIGHT (BCW) BY USING TEMPLATE-MATCHING" and U.S. Provisional Application No. 63 / 289,135, filed December 13, 2021, entitled "INDEX REORDERING OF BI-PREDICTION WITH CU-LEVEL WEIGHT (BCW) BY USING TEMPLATE-MATCHING". The entire disclosure of the prior applications is incorporated by reference.

[0002] [Technical Field] This disclosure generally describes embodiments related to video encoding (coding).

Background Art

[0003] The background description provided herein is for the purpose of generally presenting the context of the present disclosure. Aspects of the description that are the work of the inventors named in this application and that are within the scope of this background section, as well as aspects of the description that may not be eligible as prior art at the time of filing for other reasons, are not admitted as prior art to the present disclosure, either expressly or implicitly.

[0004] Uncompressed digital images and / or videos can include a series of pictures, each picture having spatial dimensions of, for example, 1920×1080 luminance samples and associated chrominance samples. The series of pictures can have a fixed or variable picture rate (informally also known as the frame rate), for example, a picture rate of 60 pictures per second or 60 Hz. Uncompressed images and / or videos have specific bitrate requirements. For example, 1080p60 4:2:0 video with 8 bits per sample (1920×1080 luminance sample resolution at a frame rate of 60 Hz) requires a bandwidth close to 1.5 Gbit / s. One hour of such video requires a storage space of more than 600 GB.

[0005] One purpose of image and / or video encoding and decoding can be the reduction of redundancy in the input image and / or video signal by compression. Compression can help reduce the aforementioned bandwidth and / or storage space requirements, sometimes by more than two orders of magnitude. The description herein uses video encoding / decoding as an illustrative example, but the same techniques can be equally applied to image encoding / decoding without departing from the spirit of the present disclosure. Both reversible compression and irreversible compression, as well as combinations thereof, can be used. Reversible compression refers to a technique where an exact copy of the original signal can be reconstructed from the compressed original signal. When using irreversible compression, the reconstructed signal may not be identical to the original signal, but the distortion between the original signal and the reconstructed signal is small enough to make the reconstructed signal useful for its intended purpose. In the case of video, irreversible compression is widely used. The amount of acceptable distortion depends on the application. For example, users of certain consumer streaming applications may tolerate higher distortion than users of television distribution applications. The achievable compression ratio can reflect that higher acceptable / tolerable distortion can result in a higher compression ratio.

[0006] Video encoders and decoders can utilize techniques from several broad categories, including, for example, motion compensation, transform processing, quantization, and entropy coding.

[0007] Video codec techniques can include techniques known as intra coding. In intra coding, sample values are represented without reference to samples from previously reconstructed reference pictures or other data. In some video codecs, a picture is spatially divided into blocks of samples. If all blocks of samples are encoded in an intra mode, that picture can be an intra picture. Intra pictures and their derivatives, such as independent decoder refresh pictures, can be used to reset the decoder state and thus can be used as the first picture in an encoded video bitstream and video session or as a still image. The samples of an intra block can be subjected to a transform, and the transform coefficients can be quantized prior to entropy coding. Intra prediction can be a technique that minimizes the sample values in the pre-transform region. In some cases, the smaller the post-transform DC value and the smaller the AC coefficients, the fewer bits are required at a given quantization step size to represent the block after entropy coding.

[0008] For example, traditional intra coding used in MPEG-2 generation coding techniques does not use intra prediction. However, some newer video compression techniques include techniques that attempt to perform prediction based on, for example, surrounding sample data and / or metadata obtained during encoding / decoding of a block of data. Such techniques are hereinafter referred to as "intra prediction" techniques. Note that in at least some cases, intra prediction uses only reference data from the currently reconstructed picture and does not use reference data from reference pictures.

[0009] There can be various forms of intra prediction. In a given video coding technology, if two or more such technologies can be used, the specific technology used can be encoded as a specific intra prediction mode that uses the specific technology. In certain cases, the intra prediction mode can have sub - modes and / or parameters, and the sub - modes and / or parameters can be encoded individually, or can be included in the mode codeword that defines the prediction mode. Which codeword to use for a given combination of mode, sub - mode and / or parameter can affect the coding efficiency gain through intra prediction, and can similarly affect the entropy coding technology used to convert the codeword into the bitstream.

[0010] A certain mode of intra prediction was introduced in H.264, refined in H.265, and further refined in newer coding technologies such as the Joint Exploration Model (JEM), Versatile Video Coding (VVC), and Benchmark Set (BMS). The predictor block can be formed using the neighboring sample values of already available samples. The sample values of the neighboring samples are copied into the predictor block according to a certain direction. The reference to the direction used can be encoded in the bitstream, or can be predicted itself.

[0011] Referring to Figure 1A, in the lower right, a subset of 9 known predictor directions out of 33 possible predictor directions (corresponding to 33 of the 35 intra - modes defined in H.265) is depicted. The point (101) where the arrows converge represents the sample to be predicted. The arrows represent the direction in which the sample is predicted. For example, arrow (102) indicates that sample (101) is predicted from the sample(s) in the upper right at an angle of 45 degrees from the horizontal. Similarly, arrow (103) indicates that sample (101) is predicted from the sample(s) in the lower left of sample (101) at an angle of 22.5 degrees from the horizontal.

[0012] Continuing to refer to FIG. 1A, in the upper left, a square block (104) of 4×4 samples is depicted (shown by the thick dashed line). The square block (104) contains 16 samples, and each sample is labeled with "S" and its position in the Y dimension (e.g., row index) and its position in the X dimension (e.g., column index). For example, sample S21 is the second sample from the top in the Y dimension and the first sample from the left in the X dimension. Similarly, sample S44 is the fourth sample within block (104) in both the Y and X dimensions. Since the block is of size 4×4 samples, S44 is in the lower right. Further, reference samples following a similar numbering scheme are shown. The reference samples are labeled with "R" and its Y position (e.g., row index) and X position (column index) relative to block (104). In both H.264 and H.265, the predicted samples are in the vicinity of the block being reconstructed, and thus there is no need to use negative values.

[0013] Intra-picture prediction can function by copying the reference sample value from neighboring samples indicated by the signaling predicted direction. For example, assume that the encoded video bitstream includes signaling indicating the predicted direction that aligns with arrow (102) for this block. That is, the samples are predicted from the upper right sample at an angle of 45 degrees from the horizontal. In that case, samples S41, S32, S23, and S14 are predicted from the same reference sample R05. Then, sample S44 is predicted from reference sample R08.

[0014] In certain cases, especially when the direction is not divisible by 45 degrees, the values of multiple reference samples can be combined, for example, by interpolation, to calculate the reference sample.

[0015] With the development of video coding technology, the number of possible directions has been increasing. In H.264 (2003), nine different directions could be represented. This increased to 33 in H.265 (2013). Currently, JEM / VVC / BMS can support up to 65 directions. Experiments are conducted to identify the most likely directions, and certain techniques in entropy coding are used to represent those likely directions with a few bits while accepting a penalty for some of the less likely directions. Furthermore, the direction itself may be predicted from neighboring directions already decoded in neighboring blocks.

[0016] FIG. 1B shows a schematic diagram (110) depicting 65 intra prediction directions by JEM to show the number of prediction directions increasing over time.

[0017] The mapping of intra prediction direction bits in the encoded video bitstream representing the direction can be different for each video coding technology. Such mappings can range from a simple direct mapping to codewords, complex adaptive schemes related to the most probable mode, and similar techniques. However, in most cases, in video content, there may be certain directions that are statistically less likely to occur than other specific directions. Since the goal of video compression is to reduce redundancy, in a well-functioning video coding technology, such less likely methods are represented by a larger number of bits than the more likely directions.

[0018] Image and / or video encoding and decoding can be performed using inter-picture prediction with motion compensation. Motion compensation can be an irreversible compression technique, and a block of sample data from a previously reconstructed picture or a part thereof (reference picture) is spatially shifted in the direction indicated by a motion vector (hereinafter, MV) and then used for prediction of a newly reconstructed picture or a part thereof. In some cases, the reference picture can be the same as the picture currently being reconstructed. The MV can have two dimensions of X and Y, or three dimensions, and the third dimension is an indication of the reference picture used (which can be indirectly the temporal dimension).

[0019] In some video compression techniques, the MV applicable to a certain region of sample data can be predicted from other MVs, for example, from an MV related to another region of sample data that is spatially adjacent to the region being reconstructed and that precedes that MV in decoding order. By doing so, the amount of data required for encoding the MV can be significantly reduced, thereby removing redundancy and increasing compression. MV prediction can function effectively, for example, when encoding an input video signal derived from a camera (known as natural video), because a larger region than the region to which a single MV is applicable moves in a similar direction, and thus, in certain cases, there is a statistical likelihood that a similar motion vector derived from the MVs of neighboring regions can be used for prediction. As a result, the MV found for a given region will be similar or identical to the MV predicted from the surrounding MVs, and it can be represented with fewer bits than would be used if the MV were directly encoded after entropy encoding. In some cases, MV prediction can be an example of lossless compression of a signal (i.e., an MV) derived from the original signal (i.e., the sample stream). In other cases, MV prediction itself can be lossy, for example, due to rounding errors when calculating predictors from some surrounding MVs.

[0020] H.265 / HEVC (ITU-T Rec. H.265, "High Efficiency Video Coding", December 2016) describes various MV prediction mechanisms. Among the many MV prediction mechanisms provided by H.265, the one described with reference to FIG. 2 is hereinafter a technique called "spatial merge".

[0021] Referring to FIG. 2, the current block (201) includes samples found by the encoder during the motion search process that the current block can be predicted from a previous block of the same size that has been spatially shifted. Instead of directly encoding the MV, the MV can be derived from metadata associated with one or more reference pictures, for example, from the latest reference picture (in decoding order), using an MV associated with any of five surrounding samples denoted as A0, A1, and B0, B1, B2 (202 to 206 respectively). In H.265, MV prediction can use predictors from the same reference picture that neighboring blocks are using. SUMMARY OF THE INVENTION

[0022] Aspects of the present disclosure provide methods and apparatuses for video encoding and decoding. In some examples, an apparatus for video decoding includes a processing circuit. The processing circuit is configured to decode prediction information of a current block within a current picture from a coded video bitstream. The prediction information indicates that the current block is predicted by bi-prediction with coding unit (CU)-level weights (BCW). The processing circuit can perform template matching (TM) for each BCW candidate weight by determining a respective TM cost for each of the respective BCW candidate weights. Each TM cost can be determined based at least in part on a current template of the current block and part or all of each of the bi-prediction sub-templates. The bi-prediction sub-templates can be determined based on each BCW candidate weight, part or all of a first reference template in a first reference picture, and part or all of a second reference template in a second reference picture. The first reference template and the second reference template correspond to the current template. A part of the first reference template and a part of the second reference template correspond to a part of the current template. The processing circuit can perform TM for each BCW candidate weight by selecting, from each BCW candidate weight, a BCW candidate weight that becomes the BCW weight used to reconstruct the current block based on the respective determined TM costs. The processing circuit can reconstruct the current block based on the selected BCW weight.

[0023] In one embodiment, the processing circuit sorts the BCW candidate weights based on the respective determined TM costs and selects a BCW candidate weight that becomes the BCW weight from the sorted BCW candidate weights.

[0024] In one embodiment, all of the current templates are used to determine each TM cost. For each BCW candidate weight, all of the first reference templates determined based on the first motion vector (MV) of the current block are used to calculate the bidirectional predictor template, and all of the second reference templates determined based on the second MV of the current block are used to calculate the bidirectional predictor template.

[0025] In one example, for each BCW candidate weight, the bidirectional predictor template is a weighted average of all of the first reference templates and all of the second reference templates, and the weights of the weighted average are based on the respective BCW candidate weights.

[0026] In one example, the prediction information indicates that the current block is predicted in an affine adaptive motion vector prediction (AMVP) mode having a plurality of control points. The first MV and the second MV are associated with a certain control point among the plurality of control points.

[0027] In one embodiment, the shape of the current template is based on one or more of (i) the reconstructed samples of the adjacent blocks of the current block, (ii) the decoding order of the current block, or (iii) the size of the current block.

[0028] In one example, the current template includes one or more reconstructed regions that are adjacent regions of the current block.

[0029] In one example, one or more reconstructed regions that are adjacent regions of the current block are one of (i) the left adjacent region and the upper adjacent region, (ii) the left adjacent region, the upper adjacent region, and the upper left adjacent region, (iii) the upper adjacent region, or (iv) the left adjacent region.

[0030] In one embodiment, the prediction information indicates that the current block is predicted in affine mode. The current template includes the current sub-block, and a part of the current template used to determine each TM cost is one of the current sub-blocks. For each BCW candidate weight, the first reference template includes a first reference sub-block corresponding to the current sub-block respectively, and a part of the first reference template used to calculate the bidirectional predictor template is one of the first reference sub-blocks. The second reference template includes a second reference sub-block corresponding to the current sub-block respectively, and a part of the second reference template used to calculate the bidirectional predictor template is one of the second reference sub-blocks. The bidirectional predictor template is based on each BCW candidate weight, one of the first reference sub-blocks, and one of the second reference sub-blocks.

[0031] In one example, for each BCW candidate weight, the bidirectional predictor template is a weighted average of one of the first reference sub-blocks and one of the second reference sub-blocks, and the weights of the weighted average are based on the respective BCW candidate weights.

[0032] In one example, the BCW candidate weights are normalized to 8, 16, or 32.

[0033] In one embodiment, the processing circuit decodes prediction information of a current block in a coded video bitstream. The processing circuit determines that the prediction information indicates that (1) the current block is predicted by bidirectional prediction and (2) bi-prediction CU-level weights at the coding unit (CU) level are valid for the current block. The processing circuit performs template matching (TM) for each BCW candidate weight by determining a respective TM cost for each of the respective BCW candidate weights. Each TM cost is determined based at least in part on a current template of the current block and a part or all of each of the respective bidirectional predictor templates. The bidirectional predictor templates are based on each BCW candidate weight, a part or all of a first reference template in a first reference picture, and a part or all of a second reference template in a second reference picture. The first reference template and the second reference template correspond to the current template. The processing circuit sorts the BCW candidate weights based on the respectively determined TM costs. The processing circuit reconstructs the current block based on the sorted BCW candidate weights.

[0034] Aspects of the present disclosure also provide a non-transitory computer-readable storage medium storing a program executable by at least one processor to perform a method for video decoding. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and the accompanying drawings.

Figure 1A

Figure 1B

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13A

Figure 13B

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20A

Figure 20B

Figure 20C

Figure 20D

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

DETAILED DESCRIPTION OF THE INVENTION

[0036] FIG. 3 shows an exemplary block diagram of a communication system (300). The communication system (300) includes a plurality of terminal devices that can communicate with each other, for example, via a network (350). For example, the communication system (300) includes a first pair of terminal devices (310) and (320) interconnected via the network (350). In the example of FIG. 3, the first pair of terminal devices (310) and (320) perform unidirectional transmission of data. For example, the terminal device (310) may encode video data (e.g., a stream of video pictures captured by the terminal device (310)) for transmission to the other terminal device (320) via the network (350). The encoded video data can be transmitted in the form of one or more encoded video bitstreams. The terminal device (320) may receive the encoded video data from the network (350), decode the encoded video data to restore the video pictures, and display the video pictures according to the restored video data. Unidirectional data transmission may be common in media service applications and the like.

[0037] In another example, the communication system (300) includes a second pair of terminal devices (330) and (340) that perform bidirectional transmission of encoded video data, for example, during a video conference. For bidirectional data transmission, in one example, each of the terminal devices (330) and (340) may encode video data (e.g., a stream of video pictures captured by the terminal device) for transmission to the other of the terminal devices (330) and (340) via the network (350). Each of the terminal devices (330) and (340) may receive the encoded video data transmitted by the other of the terminal devices (330) and (340), decode the encoded video data to restore the video pictures, and display the video pictures on an accessible display device according to the restored video data.

[0038] In the example of FIG. 3, the terminal devices (310), (320), (330), and (340) are shown as a server, a personal computer, and smartphones, respectively, but the principles of the present disclosure may not be limited thereto. Embodiments of the present disclosure find application in laptop computers, tablet computers, media players, and / or dedicated video conferencing facilities. The network (350) represents any number of networks that transfer encoded video data between the terminal devices (310), (320), (330), and (340), including, for example, a wired (wired) and / or wireless communication network. The communication network (350) may exchange data in a circuit-switched and / or packet-switched channel. Representative networks include telecommunications networks, local area networks, wide area networks, and / or the Internet. For the purposes of the discussion herein, the architecture and topology of the network (350) may not be important for the operation of the present disclosure, unless otherwise described below.

[0039] FIG. 4 shows a video encoder and a video decoder in a streaming environment as an example of an application for the disclosed subject matter. The disclosed subject matter may be equally applicable to other video-enabled applications, including, for example, video conferencing, digital TV, streaming services, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.

[0040] The streaming system can include a video source (401), such as a digital camera, and may include a capture subsystem (413) that generates a stream (402) of, for example, uncompressed video pictures. In one example, the stream of video pictures (402) includes samples captured by a digital camera. The stream of video pictures (402), depicted as a thick line to emphasize the high data volume when compared to the encoded video data (404) (or encoded video bitstream), can be processed by an electronic device (420) that includes a video encoder (403) coupled to the video source (401). The video encoder (403) can include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as will be described in more detail below. The encoded video data (404) (or encoded video bitstream), depicted as a thin line to emphasize the lower data volume when compared to the stream of video pictures (402), can be stored in a streaming server (405) for future use. One or more streaming client subsystems, such as client subsystems (406) and (408) of FIG. 4, can access the streaming server (405) to retrieve copies (407) and (409) of the encoded video data (404). The client subsystem (406) can include, for example, a video decoder (410) within an electronic device (430). The video decoder (410) decodes an input copy (407) of the encoded video data and generates an output stream (411) of video pictures that can be rendered on a display (412) (such as a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data (404), (407), and (409) (e.g., video bitstreams) can be encoded according to a particular video encoding / compression standard. Examples of these standards include ITU-T Recommendation H.265.In one example, the video coding standard under development is informally known as Versatile Video Coding (VVC). The disclosed subject matter may be used in the context of VVC.

[0041] Note that the electronic devices (420) and (430) can include other components (not shown). For example, the electronic device (420) can include a video decoder (not shown), and the electronic device (430) can also include a video encoder (not shown).

[0042] FIG. 5 shows an exemplary block diagram of a video decoder (510). The video decoder (510) can be included in an electronic device (530). The electronic device (530) can include a receiver (531) (e.g., a receiving circuit). The video decoder (510) can be used in place of the video decoder (310) in the example of FIG. 4.

[0043] The receiver (531) may receive one or more encoded video sequences to be decoded by the video decoder (510). In certain embodiments, one encoded video sequence is received at a time, and the decoding of each encoded video sequence is independent of the decoding of other encoded video sequences. The encoded video sequences may be received from a channel (501), which may be a hardware / software link to a storage device storing the encoded video data. The receiver (531) may receive the encoded video data together with other data, such as encoded audio data and / or auxiliary data streams, and these data may be transferred to their respective usage entities (not shown). The receiver (531) can separate the encoded video sequences from the other data. As a network jitter countermeasure, a buffer memory (515) may be coupled between the receiver (531) and the entropy decoder / parser (520) (hereinafter “parser”). In certain applications, the buffer memory (515) is part of the video decoder (510). In other applications, it can be external to the video decoder (510) (not shown). In yet other applications, for example to counter network jitter, there may be a buffer memory (not shown) external to the video decoder (510), and further, for example to handle playback timing, there may be another buffer memory (515) inside the video decoder (510). If the receiver (531) is receiving data from a storage / transfer device with sufficient bandwidth and controllability, or from an isochronous network, the buffer memory (515) may not be needed, or may be small. For use in a best-effort packet network such as the Internet, a buffer memory (515) may be required, may be relatively large, and advantageously may be of an adaptable size and may be implemented, at least in part, in an operating system or similar element (not shown) external to the video decoder (510).

[0044] Video decoder (510) may include a parser (520) for reconstructing symbols (521) from an encoded video sequence. The categories of these symbols include information used to manage the operation of the video decoder (510) and potentially information for controlling a rendering device such as a rendering device (512) (e.g., a display screen). The rendering device can be coupled to the electronic device (530) rather than being an integral part of the electronic device (530) as shown in FIG. 5. The control information for the rendering device(s) may be in the form of Supplementary Enhancement Information (SEI message) or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (520) can parse / entropy decode the received encoded video sequence. The encoding of the encoded video sequence can follow video encoding techniques or standards and can follow various principles including variable length encoding, Huffman encoding, arithmetic encoding with or without context sensitivity, etc. The parser (520) can extract a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder from the encoded video sequence based on at least one parameter corresponding to the group. The subgroups can include Group of Pictures (GOP), picture, tile, slice, macroblock, Coding Unit (CU), block, Transform Unit (TU), Prediction Unit (PU), etc. The parser (520) can also extract information such as transform coefficients, quantizer parameter values, motion vectors, etc. from the encoded video sequence.

[0045] The parser (520) can perform an entropy decoding / parsing operation on the video sequence received from the buffer memory (515), thereby generating symbols (521).

[0046] The reconstruction of the symbols (521) can involve multiple different units depending on the type of the encoded video picture or its parts (e.g., inter and intra pictures, inter and intra blocks) and other factors. How each unit is involved can be controlled by subgroup control information parsed by the parser (520) from the encoded video sequence. Such a flow of subgroup control information between the parser (520) and the multiple units below is not depicted for clarity.

[0047] In addition to the functional blocks already described, the video decoder (510) can conceptually be divided into several functional units as described below. In a practical implementation operating under commercial constraints, many of these units can interact closely with each other and can at least partially be integrated with each other. However, for the purpose of describing the disclosed subject matter, a conceptual subdivision into the following functional units is appropriate.

[0048] The first unit is the scaler / inverse transform unit (551). The scaler / inverse transform unit (551) receives from the parser (520) the quantized transform coefficients and control information as symbols (singular or plural) (521). The control information includes which transform to use, block size, quantization coefficient, quantization scaling matrix, etc. The scaler / inverse transform unit (551) can output a block containing sample values that can be input to the aggregator (555).

[0049] In some cases, the output samples of the scaler / inverse transform (551) can relate to intra-coded blocks. An intra-coded block is a block that does not use prediction information from a previously reconstructed picture, but can use prediction information from previously reconstructed parts of the current picture. Such prediction information can be provided by the intra-picture prediction unit (552). In some cases, the intra-picture prediction unit (552) can use surrounding already reconstructed information taken from the current picture buffer (558) to generate a block of the same size and shape as the block being reconstructed. The current picture buffer (558) buffers, for example, a partially reconstructed current picture and / or a fully reconstructed current picture. The aggregator (555) can, in some cases, add, for each sample, the prediction information generated by the intra prediction unit (552) to the output sample information provided by the scaler / inverse transform unit (551).

[0050] In other cases, the output samples of the scaler / inverse transform unit (551) can relate to inter-coded and potentially motion-compensated blocks. In such cases, the motion compensation prediction unit (553) can access the reference picture memory (557) to fetch the samples used for prediction. After motion-compensating the fetched samples according to the symbols (521) related to the block, these samples can be added by the aggregator (555) to the output of the scaler / inverse transform unit (in this case called the residual samples or residual signal), thereby generating the output sample information. The address in the reference picture memory (557) from which the motion compensation unit (553) fetches the prediction samples can be controlled by the motion vectors available to the motion compensation unit (553) in the form of symbols (521). The symbols can have, for example, X, Y, and reference picture components. Motion compensation can include interpolation of the sample values fetched from the reference picture memory (557) when exact motion vectors below the sample level are used, a motion vector prediction mechanism, etc.

[0051] The output samples of the aggregator (555) can be subjected to various loop filtering techniques within the loop filter unit (556). Video compression techniques can include in-loop filtering techniques. The in-loop filtering techniques are controlled by parameters included in the encoded video sequence (also referred to as the encoded video bitstream) and are made available to the loop filter unit (556) as symbols (521) from the parser (520). Video compression can also respond to meta information obtained during the decoding of the previous part (in decoding order) of the encoded picture or encoded video sequence and can respond to previously reconstructed and loop filtered sample values.

[0052] The output of the loop filter unit (556) can be a sample stream, which can be output to the render device (512) and can also be stored in the reference picture memory (557) for use in future inter-picture prediction.

[0053] Once an encoded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, when the encoded picture corresponding to the current picture is fully reconstructed and the encoded picture is identified as a reference picture (e.g., by the parser (520)), the current picture buffer (558) can become part of the reference picture memory (557), and a fresh current picture buffer can be reallocated before starting the reconstruction of subsequent encoded pictures.

[0054] The video decoder (510) can perform a decoding operation according to a predetermined video compression technology or standard such as ITU-T Recommendation H.265. The encoded video sequence can conform to the syntax of the video compression technology or standard and the profile documented in the video compression technology or standard, meaning that the encoded video sequence can comply with the syntax defined by the video compression technology or standard being used. Specifically, the profile can select specific tools from all the tools available in the video compression technology or standard as the tools that are only available for use under that profile. For compliance, it may also be necessary that the complexity of the encoded video sequence be within the range defined by the level of the video compression technology or standard. In some cases, the level restricts the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level can, in some cases, be further restricted through the virtual reference decoder (Hypothetical Reference Decoder, HRD) specifications and metadata for HRD buffer management signaled in the encoded video sequence.

[0055] In one embodiment, the receiver (531) may receive additional (redundant) data along with the encoded video. The additional data may be included as part of the encoded video sequence(s). The additional data may be used by the video decoder (510) for proper decoding of the data and / or for more accurately reconstructing the original video data. The additional data can be in the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.

[0056] FIG. 6 shows an exemplary block diagram of a video encoder (603). The video encoder (603) is included in an electronic device (620). The electronic device (620) includes a transmitter (640) (e.g., a transmission circuit). The video encoder (603) can be used in place of the video encoder (403) in the example of FIG. 4.

[0057] The video encoder (603) can receive video samples from a video source (601) (which is not part of the electronic device (620) in the example of FIG. 6) that can capture a video image to be encoded by the video encoder (603). In another example, the video source (601) is part of the electronic device (620).

[0058] The video source (601) can provide a source video sequence to be encoded by the video encoder (603) in the form of a digital video sample stream that can be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, …), any color space (e.g., BT.601 YCrCB, RGB, …), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (601) may be a storage device storing pre-prepared videos. In a video conferencing system, the video source (601) may be a camera that locally captures image information as a video sequence. The video data may be provided as a plurality of individual pictures that impart motion when viewed in sequence. Each picture itself may be organized as a spatial array of pixels, and each pixel can include one or more samples depending on the sampling structure, color space, etc. in use. Those skilled in the art can easily understand the relationship between pixels and samples. The following description focuses on samples.

[0059] According to one embodiment, a video encoder (603) can encode and compress pictures of a source video sequence in real time or under any other required temporal constraints to produce an encoded video sequence (643). Enforcing an appropriate encoding rate is one function of a controller (650). In some embodiments, the controller (650) controls and is functionally coupled to other functional units as described below. Such couplings are not drawn for clarity. Parameters set by the controller (650) can include parameters related to rate control (picture skip, quantizer, lambda value of rate-distortion optimization techniques, …), picture size, group of pictures (GOP) layout, maximum motion vector search range, etc. The controller (650) can be configured to have other suitable functions regarding the video encoder (603) optimized for a particular system design.

[0060] In some embodiments, the video encoder (603) is configured to operate in an encoding loop. As a radically simplified explanation, in one example, the encoding loop can include a source encoder (630) (e.g., responsible for generating symbols such as a symbol stream based on an input picture and reference picture(s) to be encoded) and a (local) decoder (633) embedded in the video encoder (603). The decoder (633) reconstructs the symbols to generate sample data in the same way as a (remote) decoder would also generate. The reconstructed sample stream (sample data) is input into the reference picture memory (634). Since the decoding of the symbol stream yields bit-exact results regardless of the decoder position (local or remote), the content of the reference picture memory (634) is also bit-exact between the local encoder and the remote encoder. In other words, the prediction unit of the encoder "sees" the same sample values as reference picture samples that the decoder "sees" when using prediction during decoding. This basic principle of reference picture synchronization (and the resulting drift if synchronization cannot be maintained, e.g., due to channel errors) is also used in some related arts.

[0061] The operation of the "local" decoder (633) may be the same as that of the "remote" decoder, e.g., the video decoder (410), already described in detail above in connection with FIG. 5. However, briefly referring also to FIG. 5, since symbols are available and the encoding / decoding of the symbols into an encoded video sequence by the entropy encoder (645) and the parser (420) can be reversible, the entropy decoding part of the video decoder (410) including the buffer memory (415) and the parser (420) may not be fully implemented in the local decoder (633).

[0062] In some embodiments, decoder techniques, excluding parse / entropy decoding that exists within the decoder, exist in the same or substantially the same functional form within the corresponding encoder. Thus, the disclosed subject matter focuses on decoder operation. The description of encoder techniques can be omitted since it is the reverse of the decoder techniques described comprehensively. In certain areas, more detailed explanations are provided below.

[0063] During operation, in some examples, the source coder (630) can perform motion-compensated predictive coding that predictively codes an input picture by referring to one or more previously coded pictures from the video sequence designated as "reference pictures". In this way, the coding engine (632) codes the difference between a pixel block of the input picture and a pixel block of the reference picture(s) that can be selected as a predictive reference for the input picture.

[0064] The local video decoder (633) can decode the coded video data of a picture that can be designated as a reference picture based on the symbols generated by the source coder (630). The operation of the coding engine (632) can advantageously be a lossy process. When the coded video data can be decoded by a video decoder (not shown in FIG. 6), the reconstructed video sequence can typically be a copy of the source video sequence with some errors. The local video decoder (633) can replicate the decoding process that can be performed on the reference picture by the video decoder and cause the reconstructed reference picture to be stored in the reference picture memory (634). In this way, the video encoder (603) can locally store a copy of the reconstructed reference picture that has (in the absence of transmission errors) the common content as the reconstructed reference picture that would be obtained by a remote video decoder.

[0065] The predictor (635) can perform a prediction search for the encoding engine (632). That is, for a new picture to be encoded, the predictor (635) can search the reference picture memory (634) to obtain sample data (as candidate reference pixel blocks) or specific metadata, such as reference picture motion vectors, block shapes, etc., that can function as an appropriate prediction reference for the new picture. The predictor (635) can operate on a sample block-by-pixel block basis to find an appropriate prediction reference. In some cases, depending on what is determined by the search results obtained by the predictor (635), the input picture can have a prediction reference drawn from a plurality of reference pictures stored in the reference picture memory (634).

[0066] The controller (650) may manage the encoding operation of the source encoder (630), including, for example, setting parameters and subgroup parameters used for encoding video data.

[0067] The outputs of all the above functional units can undergo entropy encoding in the entropy encoder (645). The entropy encoder (645) converts the symbols generated by various functional units into an encoded video sequence by applying lossless compression to the symbols according to techniques such as Huffman coding, variable-length coding, arithmetic coding, etc.

[0068] The transmitter (640) can buffer the encoded video sequence generated by the entropy encoder (645) and prepare it for transmission via the communication channel (660). The communication channel (660) may be a hardware / software link to a storage device that stores the encoded video data. The transmitter (640) can merge the encoded video data from the video encoder (630) with other data to be transmitted, such as encoded audio data and / or an auxiliary data stream (source not shown).

[0069] The controller (650) may manage the operation of the video encoder (603). During encoding, the controller (650) can assign a certain encoded picture type to each encoded picture. The encoded picture type can affect the encoding technique that can be applied to each picture. For example, a picture may often be assigned as one of the following picture types.

[0070] An intra picture (I picture) can be encoded and decoded without using other pictures in the sequence as a prediction source. Some video codecs allow different types of intra pictures, including, for example, an Independent Decoder Refresh (IDR) picture. Those skilled in the art will recognize these variations of I pictures, as well as their respective uses and characteristics.

[0071] A predicted picture (P picture) can be encoded and decoded using intra prediction or inter prediction that uses at most one motion vector and a reference index to predict the sample values of each block.

[0072] Bidirectional prediction pictures (B pictures) can be encoded and decoded using intra prediction or inter prediction that uses up to two motion vectors and reference indices to predict the sample values of each block. Similarly, multi-prediction pictures can use three or more reference pictures and associated metadata for the reconstruction of a single block.

[0073] The source picture can usually be spatially divided into a plurality of sample blocks (e.g., blocks of 4×4, 8×8, 4×8, or 16×16 samples each) and encoded block by block. The blocks can be predictively encoded by referring to other (already encoded) blocks as determined by the encoding assignment applied to each picture of the block. For example, the blocks of an I picture may be encoded non-predictively or predictively by referring to already encoded blocks of the same picture (spatial prediction or intra prediction). The pixel blocks of a P picture may be predictively encoded via spatial prediction or temporal prediction by referring to one previously encoded reference picture. The blocks of a B picture may be predictively encoded via spatial prediction or temporal prediction by referring to one or two previously encoded reference pictures.

[0074] The video encoder (603) can perform an encoding operation according to a predetermined video encoding technology or standard such as ITU-T Recommendation H.265. In that operation, the video encoder (603) can perform various compression operations including a predictive encoding operation that utilizes the temporal and spatial redundancy in the input video sequence. Thus, the encoded video data can conform to the syntax specified by the video encoding technology or standard used.

[0075] In one embodiment, the transmitter (640) may transmit additional data along with the encoded video. The source encoder (630) may include such data as part of the encoded video sequence. The additional data may include other forms of redundant data such as temporal / spatial / SNR enhancement layers, redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.

[0076] Video may be captured as a plurality of source pictures (video pictures) in a temporal sequence. Intra picture prediction (often abbreviated as intra prediction) utilizes the spatial correlation in a given picture, and inter picture prediction utilizes the (temporal or other) correlation between pictures. In one example, a particular picture to be encoded / decoded, called the current picture, is divided into blocks. If a block within the current picture is similar to a reference block within a reference picture that has been previously encoded and is still in the buffer in the video, that block within the current picture can be encoded by a vector called a motion vector. The motion vector points to the reference block within the reference picture and can have a third dimension that identifies the reference picture if multiple reference pictures are used.

[0077] In some embodiments, dual prediction techniques can be used in inter picture prediction. According to the dual prediction technique, two reference pictures such as a first reference picture and a second reference picture that both precede the current picture in decoding order (although in display order, they may be past and future respectively) in the video are used. A block within the current picture can be encoded by a first motion vector pointing to a first reference block within the first reference picture and a second motion vector pointing to a second reference block within the second reference picture. The block can be predicted by a combination of the first reference block and the second reference block.

[0078] Furthermore, to improve encoding efficiency, merge mode techniques can be used in inter picture prediction.

[0079] According to some embodiments of the present disclosure, predictions such as inter-picture prediction and intra-picture prediction are performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are divided into coding tree units (CTUs) for compression, and those CTUs in a picture have the same size such as 64×64 pixels, 32×32 pixels, or 16×16 pixels. Generally, a CTU includes three coding tree blocks (CTBs) which are one luma CTB and two chroma CTBs. Each CTU can be recursively quad-tree divided into one or more coding units (CUs). For example, a 64×64 pixel CTU can be divided into one 64×64 pixel CU, or four 32×32 pixel CUs, or sixteen 16×16 pixel CUs. In one example, each CU is analyzed to determine a prediction type for that CU, such as an inter prediction type or an intra prediction type. The CU is divided into one or more prediction units (PUs) depending on its temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB) and two chroma PBs. In some embodiments, the prediction operation in coding (encoding / decoding) is performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of values (e.g., luma values) for pixels such as 8×8 pixels, 16×16 pixels, 8×16 pixels, 16×8 pixels, etc.

[0080] FIG. 7 shows an exemplary diagram of a video encoder (703). The video encoder (703) receives a processing block (e.g., a prediction block) of sample values in a current video picture within a sequence of video pictures and is configured to encode the processing block into an encoded picture that is part of an encoded video sequence. In one example, the video encoder (703) is used in place of the video encoder (403) in the example of FIG. 4.

[0081] In an example of HEVC, a video encoder (703) receives a matrix of sample values for a processing block, such as a prediction block of 8×8 samples. The video encoder (703) determines, for example using rate-distortion optimization, which of an intra mode, an inter mode, or a bi-prediction mode the processing block is best encoded using. If the processing block is encoded in the intra mode, the video encoder (703) may use intra prediction techniques to encode the processing block into the encoded picture. If the processing block is encoded in the inter mode or the bi-prediction mode, the video encoder (703) may use inter prediction techniques or bi-prediction techniques, respectively, to encode the processing block into the encoded picture. In certain video encoding techniques, a merge mode may be an inter-picture prediction sub-mode in which motion vectors are derived from one or more motion vector predictors but there is no benefit of the encoded motion vector components outside of the predictors. In certain other video encoding techniques, there may be motion vector components applicable to the target block. In one example, the video encoder (703) includes other components, such as a mode decision module (not shown) for determining the mode of the processing block.

[0082] In the example of FIG. 7, the video encoder (703) includes an inter encoder (730), an intra encoder (722), a residual calculator (723), a switch (726), a residual encoder (724), a general controller (721), and an entropy encoder (725) coupled together as shown in FIG. 7.

[0083] The inter-encoder (730) receives samples of a current block (e.g., a processing block), compares the block with one or more reference blocks (e.g., blocks in a previous picture and a subsequent picture) in a reference picture, generates inter-prediction information (e.g., a description of redundant information by an inter-coding technique, a motion vector, merge mode information), and is configured to calculate an inter-prediction result (e.g., a predicted block) using any suitable technique based on the inter-prediction information. In some examples, the reference picture is a decoded reference picture decoded based on encoded video information.

[0084] The intra-encoder (722) receives samples of a current block (e.g., a processing block), optionally compares the block with blocks already encoded within the same picture, generates quantized coefficients after transformation, and is optionally also configured to generate intra-prediction information (e.g., intra-prediction direction information by one or more intra-coding techniques). In one example, the intra-encoder (722) also calculates an intra-prediction result (e.g., a predicted block) based on the intra-prediction information and reference blocks within the same picture.

[0085] The overall controller (721) is configured to determine overall control data and control other components of the video encoder (703) based on the overall control data. In one example, the overall controller (721) determines the mode of a block and provides a control signal to the switch (726) based on that mode. For example, when the mode is the intra mode, the overall controller (721) controls the switch (726) to select the intra mode result for use by the residual calculator (723), selects the intra prediction information, and controls the entropy encoder (725) to include the intra prediction information in the bitstream. When the mode is the inter mode, the overall controller (721) controls the switch (726) to select the inter prediction result for use by the residual calculator (723), selects the inter prediction information, and controls the entropy encoder (725) to include the inter prediction information in the bitstream.

[0086] The residual calculator (723) is configured to calculate the difference (residual data) between the received block and the prediction result selected from the intra encoder (722) or the inter encoder (730). The residual encoder (724) is configured to encode the residual data based on the residual data to generate transformation coefficients. In one example, the residual encoder (724) is configured to convert the residual data from the spatial domain to the frequency domain and generate transformation coefficients. The transformation coefficients are then subjected to quantization processing to obtain quantized transformation coefficients. In various embodiments, the video encoder (703) also includes a residual decoder (728). The residual decoder (728) is configured to perform inverse transformation to generate decoded residual data. The decoded residual data can be suitably used by the intra encoder (722) and the inter encoder (730). For example, the inter encoder (730) can generate a decoded block based on the decoded residual data and the inter prediction information, and the intra encoder (722) can generate a decoded block based on the decoded residual data and the intra prediction information. The decoded block is suitably processed to generate a decoded picture, and the decoded picture is buffered in a memory circuit (not shown) and can be used as a reference picture in some examples.

[0087] The entropy encoder (725) is configured to format the bitstream to include the encoded block. The entropy encoder (725) is configured to include various information in the bitstream according to a suitable standard such as the HEVC standard. In one example, the entropy encoder (725) is configured to include general control data, selected prediction information (e.g., intra prediction information or inter prediction information), residual information, and other suitable information in the bitstream. Note that according to the disclosed subject matter, there is no residual information when encoding a block in either the merge submode of the inter mode or the bi-prediction mode.

[0088] FIG. 8 shows an exemplary diagram of a video decoder (810). The video decoder (810) is configured to receive an encoded picture that is part of an encoded video sequence and decode the encoded picture to generate a reconstructed picture. In one example, the video decoder (810) is used in place of the video decoder (410) in the example of FIG. 4.

[0089] In the example of FIG. 8, the video decoder (810) includes an entropy decoder (871), an inter decoder (880), a residual decoder (873), a reconstruction module (874), and an intra decoder (872) coupled together as shown in FIG. 8.

[0090] The entropy decoder (871) can be configured to reconstruct from the encoded picture specific symbols that represent the syntax elements that the encoded picture is composed of. Such symbols can include, for example, the mode in which a block is encoded (e.g., the latter two in intra mode, inter mode, bi-prediction mode, merge sub-mode or another sub-mode), and prediction information (e.g., intra prediction information or inter prediction information, etc.) that can identify specific samples or metadata used for prediction by the intra decoder (872) or the inter decoder (880) respectively. The symbols can also include, for example, residual information in the form of quantized transform coefficients. In one example, when the prediction mode is inter or bi-prediction mode, the inter prediction information is provided to the inter decoder (880). When the prediction type is intra prediction type, the intra prediction information is provided to the intra decoder (872). The residual information can undergo inverse quantization and is provided to the residual decoder (873).

[0091] The inter decoder (880) is configured to receive inter prediction information and generate an inter prediction result based on the inter prediction information.

[0092] The intra decoder (872) is configured to receive intra prediction information and generate a prediction result based on the intra prediction information.

[0093] The residual decoder (873) is configured to perform inverse quantization to extract dequantized transform coefficients, process the dequantized transform coefficients, and convert the residual information from the frequency domain to the spatial domain. The residual decoder (873) may also require certain control information (including quantization parameter (QP)), and such information may be provided by the entropy decoder (871) (since this is only low-volume control information, the data path is not depicted).

[0094] The reconstruction module (874) is configured to combine, in the spatial domain, the residual information output by the residual decoder (873) and the prediction result (output by the intra or inter prediction module as the case may be) to form a reconstructed block, and the reconstructed block may be part of a reconstructed picture, and the reconstructed picture may be part of a reconstructed video. Note that other suitable operations such as a deblocking operation can be performed to improve visual quality.

[0095] Note that the video encoders (403), (603), (703), and the video decoders (410), (510), (810) can be implemented using any suitable technology. In one embodiment, the video encoders (403), (603), (703) and the video decoders (410), (510), (810) can be implemented using one or more integrated circuits. In another embodiment, the video encoders (403), (603), (703), and the video decoders (410), (510), (810) can be implemented using one or more processors that execute software instructions.

[0096] Various inter prediction modes are available in VVC. For an inter predicted CU, the motion parameters can include an MV, one or more reference picture indices, a reference picture list use index, and additional information for specific coding features used for inter prediction sample generation. The motion parameters can be signaled explicitly or implicitly. If a CU is coded in skip mode, the CU may be associated with a PU and may not have a significant residual coefficient, a coded motion vector delta or MV difference (e.g., MVD), or a reference picture index. A merge mode can be specified that includes spatial and / or temporal candidates and optionally additional information as introduced in VVC, where the motion parameters of the current CU are obtained from neighboring CUs. The merge mode is applicable not only to skip mode but also to inter predicted CUs. In one example, the selection for the merge mode is an explicit signaling of the motion parameters, where the MV, the corresponding reference picture indices of each reference picture list, and the reference picture list use flag, as well as other information, are signaled explicitly for each CU.

[0097] In embodiments such as VVC, the VVC Test Model (VTM) reference software includes one or more refined inter prediction coding tools including extended merge prediction, merge motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode with symmetric MVD signaling, affine motion compensation prediction, subblock-based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luma sample MV storage and 8x8 motion field compression), bi-prediction with CU-level weights (BCW), bi-directional optical flow (BDOF), prediction refinement using optical flow (PROF), decoder side motion vector refinement (DMVR), combined inter and intra prediction (CIIP), geometric partitioning mode (GPM), etc. In the following, inter prediction and related methods will be described in detail.

[0098] In some cases, extended merge prediction can be used. In an example like VTM4, the merge candidate list is constructed by including, in order, five types of candidates: a spatial motion vector predictor (MVP) from spatially adjacent CUs, a temporal MVP from the CU at the same position, a history-based MVP from a first-in-first-out (FIFO) table, a pairwise average MVP, and a zero MV.

[0099] The size of the merge candidate list can be signaled in the slice header. In one example, the maximum allowable size of the merge candidate list is 6 in VTM4. For each CU coded in merge mode, the index of the optimal merge candidate (e.g., the merge index) can be coded using truncated unary binarization (TU). The first bin of the merge index can be coded in context (e.g., using context-adaptive binary arithmetic coding (CABAC)), and bypass coding can be used for the other bins.

[0100] Some examples of the generation process for each category of merge candidates are provided below. In one embodiment, the spatial candidates are derived as follows. The derivation of spatial merge candidates in VVC can be the same as that in HEVC. In one example, up to four merge candidates are selected from the candidates at the positions shown in FIG. 9. FIG. 9 shows the positions of the spatial merge candidates according to one embodiment of the present disclosure. Referring to FIG. 9, the order of derivation is B1, A1, B0, A0, and B2. The position B2 is considered only if any of the CUs at positions A0, B0, B1, and A1 are not available (e.g., because the CU belongs to another slice or another tile) or if it is intra-coded. After the candidate at position A1 is added, the addition of the remaining candidates is subject to redundancy checking, which ensures that candidates with the same motion information are excluded from the candidate list, thereby improving coding efficiency.

[0101] To reduce the computational complexity, not all possible candidate pairs are considered in the above redundancy check. Instead, only the pairs linked by the arrows in FIG. 10 are considered, and a candidate is added to the candidate list only if the corresponding candidates used in the redundancy check do not have the same motion information. FIG. 10 shows candidate pairs considered for the redundancy check of spatial merge candidates according to an embodiment of the present disclosure. Referring to FIG. 10, the pairs linked by each arrow include A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Therefore, candidates at positions B1, A0, and / or B2 can be compared with candidates at position A1, and candidates at positions B0 and / or B2 can be compared with candidates at position B1.

[0102] In one embodiment, the temporal candidates are derived as follows. In one example, only one temporal merge candidate is added to the candidate list. FIG. 11 shows an exemplary motion vector scaling for temporal merge candidates. To derive the temporal merge candidate of the current CU (1111) in the current picture (1101), the scaled MV (1121) (e.g., shown by the dotted line in FIG. 11) can be derived based on the CU (1112) at the same position belonging to the reference picture (1104) at the same position. The reference picture list used to derive the CU (1112) at the same position can be explicitly signaled in the slice header. The scaled MV (1121) for the temporal merge candidate can be obtained as shown by the dotted line in FIG. 11. The scaled MV (1121) can be scaled from the MV of the CU (1112) at the same position using the picture order count (POC) distances tb and td. The POC distance tb can be defined as the POC difference between the current reference picture (1102) of the current picture (1101) and the current picture (1101). The POC distance td can be defined as the POC difference between the reference picture (1104) at the same position of the picture (1103) at the same position and the picture (1103) at the same position. The reference picture index of the temporal merge candidate can be set to 0.

[0103] FIG. 12 shows exemplary candidate positions (e.g., C0 and C1) of the current CU's temporal merge candidates. The position of the temporal merge candidate can be selected between candidate position C0 and candidate position C1. Candidate position C0 is located at the lower right corner of the CU (1210) at the same position as the current CU. Candidate position C1 is located at the center of the CU (1210) at the same position as the current CU. If the CU at candidate position C0 is not available, is intra-coded, or is outside the current row of the CTU, candidate position C1 is used to derive the temporal merge candidate. Otherwise, for example, if the CU at candidate position C0 is available, is intra-coded, and is in the current row of the CTU, candidate position C0 is used to derive the temporal merge candidate.

[0104] In some examples, a translational motion model is applied for motion compensation prediction (MCP). However, the translational motion model may not be suitable for modeling other types of motion such as zoom-in / out, rotation, perspective motion, and other irregular motions. In some embodiments, block-based affine transform motion compensation prediction is applied. In FIG. 13A, when a 4-parameter affine model is used, the affine motion field of the block is described by two control point motion vectors (CPMVs) CPMV0 and CPMV1 of two control points (CPs) CP0 and CP1. In FIG. 13B, when a 6-parameter affine model is used, the affine motion field of the block is described by three CPMVs CPMV0, CPMV1, and CPMV3 of three CPs CP0, CP1, and CP2.

[0105] For the 4-parameter affine motion model, the motion vector at the sample position (x, y) within the block is derived as follows.

Equation

[0106] In the case of a 6-parameter affine motion model, the motion vector at the sample position (x, y) within the block is derived as follows. [Number]

[0107] In Equations 1-2, (mv 0x , mv 0y ) is the motion vector of the top-left control point, (mv 1x , mv 1y ) is the motion vector of the top-right control point, and (mv 2x , mv 2y ) is the motion vector of the bottom-left control point. Further, the coordinates (x, y) are based on the top-left corner of each block, and W and H indicate the width and height of each block, respectively.

[0108] In some embodiments, to simplify motion compensation prediction, sub-block-based affine transformation prediction is applied. For example, in FIG. 14, a 4-parameter affine motion model is used and two CPMVs [Number] are determined. To derive the motion vector for each 4×4 (sample) luma sub-block (1402) divided from the current block (1410), the motion vector (1401) of the central sample of each sub-block (1402) is calculated according to Equation 1 and rounded to 1 / 16 fractional precision. Then, a motion compensation interpolation filter is applied to generate a prediction for each sub-block (1402) with the derived motion vector (1401). The sub-block size of the chroma component is set to 4×4. The MV of the 4×4 chroma sub-block is calculated as the average of the MVs of the corresponding 4×4 luma sub-block.

[0109] In some embodiments, similar to translational motion inter-prediction, two affine motion inter-prediction modes, namely the affine merge mode and the affine AMVP mode, are used.

[0110] In some embodiments, the affine merge mode can be applied to a CU where both the width and height are 8 or more. The current CU's affine merge candidates can be generated based on the motion information of spatially adjacent CUs. There can be up to five affine merge candidates, and an index is signaled to indicate the one used for the current CU. For example, the following three types of affine merge candidates are used to form the affine merge candidate list. (i) Inherited affine merge candidates extrapolated from the CPMV of adjacent CUs, (ii) Constructed affine merge candidates derived using the translational MV of adjacent CUs, and (iii) Zero MV

[0111] In some embodiments, there are at most two inherited affine candidates derived from the affine motion model of adjacent blocks, one derived from the left adjacent CU and one derived from the upper adjacent CU. For example, the candidate blocks can be arranged at the positions shown in FIG. 9. For the left predictor, the scan order is A0 > A1, and for the upper predictor, the scan order is B0 > B1 > B2. Only the candidates inherited first from each side are selected. No pruning check is performed between the two inherited candidates.

[0112] When an adjacent affine CU is identified, the CPMV of the identified adjacent affine CU is used to derive the CPMV candidates within the current CU's affine merge list. As shown in FIG. 15, the adjacent lower-left block A of the current CU (1510) is coded in the affine mode. The motion vectors at the upper-left corner, upper-right corner, and lower-left corner of the CU (1520) containing block A

Number

Number

Number

Number

[0113] The constructed affine candidates are constructed by combining the adjacent translational motion information of each control point. The motion information of the control point is derived from the specified spatial adjacency and temporal adjacency shown in FIG. 16. CPMV k (k = 1, 2, 3, 4) represents the k-th control point. For CPMV1, blocks with B2 > B3 > A2 are inspected in sequence, and the MV of the first available block is used. For CPMV2, blocks with B1 > B0 are inspected, and for CPMV3, blocks with A1 > A0 are inspected. The TMVP in block T is used as CPMV4 if available.

[0114] After the MVs of the four control points are obtained, affine merge candidates are constructed based on the motion information. The following combinations of control point MVs, namely, {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, {CPMV1, CPMV3} are used to construct in sequence.

[0115] The combination of three CPMVs constructs a six-parameter affine merge candidate, and the combination of two CPMVs constructs a four-parameter affine merge candidate. To avoid the motion scaling process, if the reference indices of the control points are different, the combination of the associated control point MVs is discarded.

[0116] After the inherited affine merge candidates and the constructed affine merge candidates are examined, if the list is still not full, zero MVs are inserted at the end of the merge candidate list.

[0117] In some embodiments, the affine AMVP mode can be applied to a CU where both the width and height are 16 or more. An affine flag at the CU level is signaled in the bitstream to indicate whether the affine AMVP mode is used, and then another flag is signaled to indicate whether a four-parameter affine or a six-parameter affine is used. The difference between the CPMV of the current CU and its predictor is signaled in the bitstream. The size of the affine AVMP candidate list is 2 and can be generated by using the following four types of CPVM candidates in order. (i) Inherited affine AMVP candidates extrapolated from the CPMV of the adjacent CU, (ii) Constructed affine AMVP candidates derived using the translational MVs of the adjacent CU, (iii) Translational MVs from the adjacent CU, and (iv) Zero MVs

[0118] In one example, the inspection order of the inherited affine AMVP candidates is the same as the inspection order of the inherited affine merge candidates. The difference is that for the AVMP candidates, an affine CU having the same reference picture as the current block is considered. When inserting an inherited affine motion predictor into the candidate list, the pruning process is not applied.

[0119] The constructed AMVP candidates are derived from the specified spatial adjacencies shown in Figure 16. The same inspection order as that performed in affine merge candidate construction is used. Further, the reference picture indices of the adjacent blocks are also inspected. The first block in the inspection order that has the same reference picture as the current CU and is inter-coded is used. If the current CU is coded with a 4-parameter affine model and both CPMV0 and CPMV1 are available, the available CPMV is added as one candidate in the affine AMVP list. If the current CU is coded with a 6-parameter affine mode and all three CPMVs (CPMV0, CPMV1, and CPMV2) are available, the available CPMV is added as one candidate in the affine AMVP list. Otherwise, the constructed AMVP candidates are set to unavailable.

[0120] After the inherited affine AMVP candidates and the constructed AMVP candidates are inspected, if the number of affine AMVP list candidates is still less than 2, if the translational motion vectors adjacent to the control points are available, they are added to predict all the control point MVs of the current CU. Finally, if the affine AMVP list is still not full, zero MVs are used to fill the affine AMVP list.

[0121] In video / image coding, template matching (TM) technology can be used. To further improve the compression efficiency of the VVC standard, for example, TM can be used to refine the motion vectors (MVs). In one example, TM is used on the decoder side. In the TM mode, the MV can be refined by constructing a template (e.g., the current template) of a block (e.g., the current block) within the current picture and determining the closest match between the template of the block in the current picture and multiple possible templates (e.g., multiple possible reference templates) in the reference picture. In one embodiment, the template of a block in the current picture can include the left adjacent reconstructed samples and the upper adjacent reconstructed samples of the block. TM can be used in video / image coding beyond VVC.

[0122] Figure 17 shows an example of template matching (1700). TM can be used to derive the motion information of the current coding unit (CU) (e.g., the current block) (1701) by determining the closest match between a template (e.g., the current template) (1721) of the current CU in the current picture (1710) and templates (e.g., reference templates) of multiple possible templates in the reference picture (1711) (e.g., one of the multiple possible templates is template (1725)). The template (1721) of the current CU (1701) can have any suitable shape and any suitable size.

[0123] In one embodiment, the template (1721) of the current CU (1701) includes an upper template (1722) and a left template (1723). Each of the upper template (1722) and the left template (1723) can have any suitable shape and any suitable size.

[0124] The upper template (1722) can include samples within one or more upper adjacent blocks of the current CU (1701). In one example, the upper template (1722) includes samples of four rows within one or more upper adjacent blocks of the current CU (1701). The left template (1723) can include samples within one or more left adjacent blocks of the current CU (1701). In one example, the left template (1723) includes samples of four columns within one or more left adjacent blocks of the current CU (1701).

[0125] Each of the plurality of possible templates in the reference picture (1711) (e.g., template (1725)) corresponds to the template (1721) in the current picture (1710). In one embodiment, the initial MV (1702) points to the reference block (1703) within the reference picture (1711) from the current CU (1701). Each of the plurality of possible templates in the reference picture (1711) and the template (1721) in the current picture (1710) (e.g., template (1725)) can have the same shape and the same size. For example, the template (1725) of the reference block (1703) includes the upper template (1726) in the reference picture (1711) and the left template (1727) in the reference picture (1711). The upper template (1726) can include samples within one or more upper adjacent blocks of the reference block (1703). The left template (1727) can include samples within one or more left adjacent blocks of the reference block (1703).

[0126] The TM cost can be determined based on a pair of templates such as a template (e.g., the current template) (1721) and a template (e.g., the reference template) (1725). The TM cost can indicate the match between template (1721) and template (1725). An optimized MV (or the final MV) can be determined based on a search around the initial MV (1702) of the current CU (1701) within the search range (1715). The search range (1715) can have any appropriate shape and any appropriate number of reference samples. In one example, the search range (1715) in the reference picture (1711) includes a [-L,L] pel range, where L is a positive integer such as 8 (e.g., 8 samples). For example, a difference (e.g., [0,1]) is determined based on the search range (1715), and an intermediate MV is determined by the sum of the initial MV (1702) and the difference (e.g., [0,1]). The intermediate reference block and the corresponding template in the reference picture (1711) can be determined based on the intermediate MV. The TM cost can be determined based on the template (1721) and the intermediate template in the reference picture (1711). The TM cost can be made to correspond to a difference (e.g., [0,0], [0,1], etc. corresponding to the initial MV (1702)) determined based on the search range (1715). In one example, the difference corresponding to the minimum TM cost is selected, and the optimized MV is the sum of the difference corresponding to the minimum TM cost and the initial MV (1702). As described above, TM can derive the final motion information (e.g., the optimized MV) from the initial motion information (e.g., the initial MV 1702).

[0127] TM can be appropriately modified. In one example, the search step size is determined by the AMVR mode. In one example, TM can be cascaded (e.g., used together) with other coding methods such as a bilateral matching process.

[0128] TM can be applied in affine modes such as the Affine AMVP mode and the Affine Merge mode, and can be called Affine TM. FIG. 18 shows an example of TM (1800) in the Affine Merge mode etc. The template (1821) of the current block (e.g., the current CU) (1801) can correspond to the template (e.g., the template (1721) in FIG. 17) in the TM applied to the translational motion model. The reference template (1825) of the reference block in the reference picture can include a plurality of sub-block templates (e.g., 4×4 sub-blocks) pointed to by the MVs derived from the control point MVs (CPMV, control point MV) of the adjacent sub-blocks (e.g., A0 to A3 and L0 to L3 as shown in FIG. 18) at the block boundary.

[0129] The search process of TM applied in the affine mode (e.g., the Affine Merge mode) can start from CPMV0 while keeping other CPMVs (e.g., (i) CPMV1 when a 4-parameter model is used, or (ii) CPMV1 and CPMV2 when a 6-parameter model is used) constant. The search can be performed in the horizontal and vertical directions. In one example, the diagonal search continues only if the zero vector is not the best difference vector found from the horizontal and vertical searches. Affine TM can repeat the same search process for CPMV1. Affine TM can repeat the same search process for CPMV2 when a 6-parameter model is used. Based on the refined CPMV, if the zero vector is not the best difference vector from the previous iteration and the search process has been repeated less than 3 times, the entire search process can be restarted from the refined CPMV0.

[0130] In one embodiment, the BCW technique is designed to predict a block by weighted-averaging two motion compensation prediction blocks. Weighting prediction (WP) can indicate weights at the slice level, but the weights used in BCW can be signaled at the CU level by using an index (e.g., the BCW index denoted as bcwIdx). The index in BCW can refer to a selected weight from a predefined list of candidate weights (e.g., a weight list). The list (e.g., the BCW list) can pre-define a plurality of (e.g., five) candidate weights such as {-2, 3, 4, 5, 10} / 8 that are selected for reference pictures in a reference list (e.g., reference list 1 or L1). The two weights -2 / 8 and 10 / 8 can be used to reduce negatively correlated noise between prediction blocks used in bidirectional prediction. When forward and backward reference pictures in both reference lists (e.g., L0 and L1) are used to achieve a better trade-off between performance and complexity, the list may be reduced to a list of {3, 4, 5} / 8. Generally, the list can include any appropriate number of candidate weights. Since the unit-gain constraint is applied, when the weight (referred to as the second weight denoted as w) pointed to by the index (e.g., bcwIdx) corresponding to reference list 1 is determined, the other weight (referred to as the first weight) corresponding to the other reference list (e.g., L0) is (1 - w). In one example, each luma prediction sample or each chroma prediction sample of BCW is determined as follows.

Number

[0131] In Equation 3, P0 and P1 are prediction samples pointed to by motion vectors from the first reference picture in reference list 0 (L0) and the second reference picture in reference list 1 (L1), respectively. P BCW is the final prediction of the sample in the current block, and P BCWis the weighted average of P0 and P1. In one embodiment, BCW becomes effective only for a bi-directionally predicted CU by at least 256 luma samples when WP is off for the bi-directionally predicted CU. The above BCW can be extended to a bi-directionally predicted CU coded in the affine AMVP mode.

[0132] The use of an index (e.g., bcwIdx) can be buffered for subsequent CUs within the same picture or the same frame to perform spatial motion merge, such as in the normal merge mode (e.g., the overall block-based merge mode) or the affine merge mode. When a spatial adjacent merge candidate is bi-directionally predicted and the current CU selects the spatial adjacent merge candidate, motion information including one or more of (i) one or more reference indices, (ii) a motion vector (or CPMV in the inherited affine merge mode), and (iii) the corresponding BCW index (e.g., bcwIdx) can be inherited by the current CU. In one example, all of (i) one or more reference indices, (ii) a motion vector (or CPMV in the inherited affine merge mode), and (iii) the corresponding BCW index (e.g., bcwIdx) can be inherited by the current CU. In one example, when the current CU has a valid CIIP flag, the weight index (e.g., bcwIdx) is not inherited. In the construction affine merge mode, the BCW index (e.g., bcwIdx) can be inherited from the weight index associated with the upper left CPMV(s) (or the upper right CPMV if the upper left CPMV is not used). In one example, when the inferred BCW index (e.g., bcwIdx) points to a weight other than 0.5, the DMVR mode and the BDOF mode are turned off.

[0133] In some embodiments, such as VVC and EE2, the BCW index is encoded in a fixed order. For example, the relationship between the BCW index (e.g., bcwIdx) and the corresponding BCW candidate weight in the BCW list is fixed. In one example, the fact that the BCW index (e.g., bcwIdx) is i corresponds to the i-th candidate weight in the BCW list (e.g., {-2, 3, 4, 5, 10} / 8), where i is an integer greater than or equal to 0. For example, when the BCW index is 0, 1, 2, 3, or 4, w is -2 / 8, 3 / 8, 4 / 8, 5 / 8, or 10 / 8, respectively. Encoding the BCW index in a fixed order may result in a high signaling cost for BCW (e.g., signaling of the BCW index), and thus, in some examples, the BCW mode may not be used efficiently.

[0134] The current CU or current block within the current picture can be coded with bi-directional prediction in the BCW mode. The current block can be coded based on a first reference block within a first reference picture in a first reference list (e.g., L0) and a second reference block within a second reference picture in a second reference list (e.g., L1). According to one embodiment of the present disclosure, in order to improve the efficiency of the BCW mode (e.g., reduce the cost of signaling the BCW index), the TM can be applied to BCW candidate weights such as -2 / 8, 3 / 8, 4 / 8, 5 / 8, or 10 / 8 in the BCW list (e.g., {-2, 3, 4, 5, 10} / 8). The TM cost corresponding to each BCW candidate weight can be determined, and the BCW candidate weight can be selected as the BCW weight used to code (e.g., encode or reconstruct) the current CU or current block based on the TM cost. In one example, the BCW candidate weights can be ranked or sorted based on the determined TM costs. The BCW weight can be selected from the ranked or sorted BCW candidate weights. For example, the BCW candidate weights are ranked or sorted based on the ascending order of the determined TM costs. The current block or current CU can be coded (e.g., encoded or reconstructed) based on the selected BCW weight (e.g., w) as shown in Equation 3.

[0135] According to one embodiment of the present disclosure, the TM cost in the TM cost corresponding to each BCW candidate weight in the BCW candidate weight is determined based on the reconstruction samples in the adjacent reconstruction blocks of the current block, the first reference block, and the second reference block, respectively. For example, the TM cost corresponding to each BCW candidate weight is determined based on the current template of the current block and each bidirectional predictor (e.g., bidirectional predictor template). The bidirectional predictor template can be determined based on each BCW candidate weight, the first reference template of the first reference block in the first reference picture, and the second reference template of the second reference block in the second reference picture. The first reference block and the second reference block correspond to the current block, and the first reference template and the second reference template correspond to the current template.

[0136] In one embodiment, when the current CU (e.g., current block) is encoded in bidirectional prediction and the BCW mode is valid for the current CU (e.g., when sps_bcw_enabled_flag is true), TM is applied to derive the BCW index. The TM search procedure can be performed for BCW weights (e.g., all possible BCW weights) to sort the BCW indexes in ascending order, for example, by using the TM cost between the current template in the current picture and the reference templates in the reference pictures such as the first reference template and the second reference template.

[0137] In one embodiment, the adjacent reconstruction regions of the current block, the first reference block, and the second reference block can be used as the current template, the first reference template, and the second reference template, respectively, to calculate each TM cost between the current picture (e.g., current reconstructed picture) and the reference pictures (e.g., the first reference picture and the second reference picture).

[0138] FIG. 19 shows an example of a template matching-based BCW index rearrangement process (1900). The current block (1902) being reconstructed in the current picture (1901) is coded with bidirectional prediction in BCW mode. The first MV (1916) can point from the current block (1902) to the first reference block (1912) in the first reference picture (1911) within the first reference list (e.g., L0). The second MV (1926) can point from the current block (1902) to the second reference block (1922) in the second reference picture (1921) within the second reference list (e.g., L1). The current block (1902) can be predicted based on a weighted average of the first reference block (1912) in the first reference picture (1911) and the second reference block (1922) in the second reference picture (1921) as explained in Equation 3. For example, a sample (e.g., P C ) in the current block (1902) is predicted based on a weighted average (e.g., P BCW ) of the first sample (e.g., P0) in the first reference block (1912) (e.g., with a weight of (1 - w)) and the second sample (e.g., P1) in the second reference block (1922) (e.g., with a weight of w). In one example, P C is equal to P BCW . In one example, P C is equal to the sum of P BCW and the corresponding residual.

[0139] TM can be applied to determine the BCW weight w used in calculating the weighted average of the first reference block (1912) and the second reference block (1922), such as the BCW weight w used in Equation 3. TM can be executed for BCW candidate weights such as -2 / 8, 3 / 8, 4 / 8, 5 / 8, or 10 / 8 in the BCW list (e.g., {-2,3,4,5,10} / 8). For example, TM is executed to determine each TM matching cost (also called TM cost) between the current template (1905) of the current block (1902) and each bidirectional predictor template corresponding to the BCW candidate weight in the BCW candidate weight. Each bidirectional predictor template can be determined based on (i) the first reference template (1915) of the first reference block (1912), (ii) the second reference template (1925) of the second reference block (1922), and the corresponding BCW candidate weight.

[0140] The current template (1905) can include samples (e.g., reconstructed samples) within the adjacent reconstructed block of the current block (1902). The current template (1905) can have any suitable shape and any suitable size. The shapes and sizes of the first reference template (1915) and the second reference template (1925) can be made to match the shape and size of the current template (1905), respectively.

[0141] In the example shown in FIG. 19, the current template (1905) includes the upper template (1904) and the left template (1903). Accordingly, the first reference template (1915) includes the first upper reference template (1914) and the first left reference template (1913), and the second reference template (1925) includes the second upper reference template (1924) and the second left reference template (1923).

[0142] The shape and / or size of the current template may change, and thus, for example, it can be adapted to the adjacent reconstruction data of the current block (1902) (e.g., the reconstructed samples in the adjacent reconstruction block), the decoding order of the current block (1902), the size of the current block (1902) (e.g., the number of samples in the current block (1902), the width of the current block (1902), the height of the current block (1902), etc.), the availability of the reconstructed samples in the adjacent reconstruction block, etc. In one example, when the width of the current block is greater than a threshold value, the current template includes only the upper template (e.g., (1904)) and does not include the left template (e.g., (1903)). In one example, the left template (e.g., (1903)) is not available, and the current template includes only the upper template (e.g., (1904)) and does not include the left template (e.g., (1903)). As described above, the shape and / or size of the reference template (e.g., the first reference template (1915) or the second reference template (1925)) may change, and thus it can be adapted according to the shape and / or size of the current template.

[0143] Figures 20A - 20D show examples of the current template of the current block that can be used in template - matching - based BCW index rearrangement. The current template (1905) in Figure 20A is the same as that shown in Figure 19, and the current template (1905) includes the upper template (1904) and the left template (1903). The current template (2005) of the current block (1902) in Figure 20B includes the upper template (1904), the left template (1903), and the upper - left template (2001). The current template (2015) of the current block (1902) in Figure 20C is the upper template (1904). The current template (2025) of the current block (1902) in Figure 20D is the left template (1903).

[0144] The current templates (1905) and (2005) in FIGS. 20A and 20B each have an L-shaped form. The current templates (2015) and (2025) in FIGS. 20C and 20D each have a rectangular form.

[0145] Returning to FIG. 19, the current template (1905), the first reference template (1915), and the second reference template (1925) have an L-shaped form and can be used to calculate the TM cost. The TM process can be performed for BCW candidate weights (e.g., all BCW candidate weights in the BCW list). Each TM cost can be performed between the current template (1905) in the current picture (1901) and each bidirectional predictor template predicted from the first reference template (1915) in the first reference picture (1911) in the first reference list (e.g., L0) and the second reference template (1925) in the second reference picture (1921) in the second reference list (e.g., L1). Each TM cost can be calculated according to the distortion between the current template (shown as TC) (1905) and each bidirectional predictor template (TP BCW shown as) of the first reference template (1915) and the second reference template (1925) each having a respective BCW candidate weight (e.g., a predefined BCW candidate weight). In the BCW mode, the bidirectional predictor template TP BCW of the first reference template (1915) and the second reference template (1925) can be derived as follows.

Equation

[0146] The parameter TP0 can represent the first reference template (1915). The parameter TP1 can represent the second reference template (1925). Based on Equation 4, the bidirectional predictor template TP BCWThe value of the predictor sample in [context] can be a weighted average of the first reference sample value (1915) in the first reference template and the second reference sample value (1925) in the second reference template based on the BCW candidate weight w.

[0147] In the above example shown in Equation 4, the weights are normalized by 8. Other normalization coefficients such as 16, 32, etc. may be used.

[0148] In one example, Equation 4 can be rewritten as follows.

Equation

[0149] The TM cost corresponding to the BCW candidate weight w can be calculated based on the current template TC (1905) and the bidirectional predictor template TP, for example, using the following Equation 6. BCW based on.

Equation

[0150] The sum of absolute differences (SAD) represents, for example, a function of the sum of the absolute differences between the sample value (1905) in the current template and the corresponding values of the predictor samples in the bidirectional predictor template TP. BCW and the corresponding values of the predictor samples in.

[0151] To determine the TM cost, other functions such as the sum of squared errors (SSE), variance, partial SAD, etc. may be used. In the example of partial SAD, a part of the current template (1905), the corresponding part of the first reference template (1915), and the corresponding part of the second reference template (1925) are used to determine the TM cost.

[0152] In an example of partial SAD, a part or all of the current template (1905), a part or all of the first reference template (1915), and a part or all of the second reference template (1925) are downsampled before being used to determine the TM cost.

[0153] In one example, the BCW list is {-2,3,4,5,10} / 8 including five BCW candidate weights -2 / 8, 3 / 8, 4 / 8, 5 / 8, and 10 / 8. As described above, when there is no TM, when the BCW index is 0, 1, 2, 3, or 4 respectively, w is -2 / 8, 3 / 8, 4 / 8, 5 / 8, or 10 / 8, and the relationship between the BCW index (e.g., bcwIdx) and the corresponding BCW candidate weight in the BCW list is fixed. According to one embodiment of the present disclosure, TM is executed for the five BCW candidate weights -2 / 8, 3 / 8, 4 / 8, 5 / 8, and 10 / 8, and the five corresponding TM costs (TM0 to TM4) are determined using Equations 4 and 6.

[0154] The five BCW candidate weights can be ranked (e.g., sorted) based on the corresponding TM cost. For example, the TM costs are TM3, TM4, TM0, TM2, and TM1 in ascending order, where TM1 is the largest among TM0 to TM4 and TM3 is the smallest among TM0 to TM4. The five BCW candidate weights are ranked (e.g., sorted) as 5 / 8, 10 / 8, -2 / 8, 4 / 8, and 3 / 8. Thus, when the BCW index is 0, 1, 2, 3, or 4 respectively, w is 5 / 8, 10 / 8, -2 / 8, 4 / 8, or 3 / 8. As described above, when TM is used, the relationship between the BCW index (e.g., bcwIdx) and the corresponding BCW candidate weight in the BCW list is not fixed. The relationship between the BCW index (e.g., bcwIdx) and the corresponding BCW candidate weight in the BCW list can be adapted to the values of the reconstructed samples in the current template (1905), the first reference template (1915), and / or the second reference template (1925). For example, when the transmitted BCW index (e.g., bcwIdx) is 0, 5 / 8 is selected to be the BCW weight (e.g., w in Equation 3) used to code the current block (1902) with TM. In contrast, the transmitted BCW index 3 indicates 5 / 8 without TM. Thus, compared to signaling the BCW index without TM, when TM is used to rank (e.g., sort) the BCW candidate weights, fewer bits can be used to signal the BCW index, and thus the signaling cost of the BCW mode can be reduced. For example, TM-based sorting is advantageous because the most useful BCW candidate weights (e.g., (i) 5 / 8 or (ii) BCW candidate weights with relatively small TM costs such as 5 / 8 and 10 / 8 in the above example) can have shorter codewords for entropy coding.

[0155] In one example, the BCW candidate weight (e.g., 5 / 8) corresponding to the minimum TM cost (e.g., TM3) is selected as the BCW weight used to code the current block (1902) as shown in Equation 7 below. In one example, the BCW index (e.g., bcwIdx) is not signaled, thus reducing the signaling cost of the BCW mode.

Number

[0156] The TM applied to the BCW candidate weights described in FIG. 19 can be adapted to sub-block based TMs such as those used in affine modes such as the affine AMVP mode or the affine merge mode.

[0157] In one embodiment, when the current block or current CU is coded in an affine mode (e.g., the affine AMVP mode), the TM is executed for all BCW candidate weights to reorder the BCW index by using the TM cost between the current template and the reference template. In the affine mode, the current template can be divided into a plurality of N×N sub-block templates. N is a positive integer. In one example, N is 4.

[0158] FIG. 21 shows an example of a sub-block based TM (2100) applied to BCW candidate weights in an affine mode (e.g., the affine AMVP mode). The TM (2100) can be applied to the BCW candidate weights in a BCW list (e.g., {-2, 3, 4, 5, 10} / 8).

[0159] The current block (2110) includes a plurality of sub-blocks (2101). The current block (2110) can be coded in a sub-block based bidirectional prediction mode. In one example, each sub-block (2101) within the current block (2110) is associated with a respective MV pair including a first MV pointing to a respective first reference sub-block (2103) within a first reference block (2111) and a second MV pointing to a respective second reference sub-block (2105) within a second reference block (2113).

[0160] The MV pair associated with each sub-block (2101) can be determined based on the affine parameters of the current block (2110) and the position of each sub-block (2101), as described with reference to FIG. 14. In the example, the affine parameters of the current block (2110) are determined based on the CPMV of the current block (2110) (e.g., CPMV0 to CPMV1 or CPMV0 to CPMV2).

[0161] The current template (2121) of the current block (2110) can include reconstructed samples within the adjacent reconstructed blocks of the current block (2110). The current template (2121) can have any suitable shape and / or any suitable size. The shape and / or size of the current template (2121) can vary, as described with reference to FIGS. 20A - 20D.

[0162] In the affine mode (e.g., the affine AMVP mode), the current template (2121) can include a plurality of sub-blocks (also referred to as sub-block templates). The current template (2121) can include any appropriate number of sub-blocks at any appropriate position. Each of the plurality of sub-block templates can have any appropriate size such as N×N. In one example, the plurality of sub-block templates includes an upper sub-block template (e.g., A0 to A3) and / or a left sub-block template (e.g., L0 to L3). For example, the current template (2121) can include (i) an upper template including an upper sub-block template, and / or (ii) a left template including a left sub-block template. In the example of FIG. 21, the current template (2121) includes upper sub-block templates A0 to A3 and left sub-block templates L0 to L3.

[0163] Each sub-block template (e.g., one of A0 to A3 or one of L0 to L3) within the current template (2121) can be associated with a respective MV pair including a first MV and a second MV. The MV pair associated with the sub-block template can be determined based on the affine parameters of the current block (2110) and the respective positions of the sub-block templates, as described in FIG. 14. Thus, the MV pairs associated with the respective sub-block templates can be different.

[0164] The first reference template (2123) associated with the first reference block (2111) can be determined based on the plurality of sub-block templates within the current template (2121) and the respective MV pairs (e.g., the associated first MVs) associated with the plurality of sub-block templates. Referring to FIG. 21, the first reference sub-block template (e.g., the first upper reference sub-block template A 00 ~A 03 and / or the first left reference sub-block template L00 ~L 03 ) can be determined based on a plurality of sub-block templates (e.g., A0 to A3 and / or L0 to L3) and respective MV pairs associated with each of the plurality of sub-block templates (e.g., A0 to A3 and / or L0 to L3). In one example, when the first MVs associated with the plurality of sub-block templates are different, the shape of the first reference template (2123) is different from the current template (2121).

[0165] Similarly, the second reference template (2125) associated with the second reference block (2113) can be determined based on a plurality of sub-block templates within the current template (2121) and respective MV pairs (e.g., the associated second MV) associated with each of the plurality of sub-block templates. Referring to FIG. 21, the second reference sub-block template (e.g., the second upper reference sub-block template A 10 ~A 13 and / or the second left reference sub-block template L 10 ~L 13 ) in the second reference template (2125) can be determined based on a plurality of sub-block templates (e.g., A0 to A3 and / or L0 to L3) within the current template (2121) and respective MV pairs associated with each of the plurality of sub-block templates (e.g., A0 to A3 and / or L0 to L3). In one example, when the second MVs associated with the plurality of sub-block templates are different, the shape of the second reference template (2125) is different from the current template (2121).

[0166] For example, the first MV pair of the sub-block template (e.g., A0) is the first MV pointing to the first reference sub-block template (e.g., A 00 ) in the first reference template (2123) and the second reference sub-block template (e.g., A 10It includes a second MV pointing to (e.g., A1). The second MV pair of the sub-block template (e.g., A1) is the first reference sub-block template (e.g., A 01 ) in the first reference template (2123), a third MV pointing to, and a second reference sub-block template (e.g., A 11 ) in the second reference template (2125), a fourth MV pointing to. In the example shown in FIG. 21, the first MV is different from the third MV, and the second MV is different from the fourth MV.

[0167] The TM embodiment described in FIG. 19 can be applied to the BCW candidate weights when the current block (2110) is coded in an affine mode such as the affine AMVP mode. For the BCW candidate weights in the BCW list, the TM cost can be determined based on the current template (2121) and the bidirectional predictors (e.g., bidirectional predictor templates) of the first reference template (2123) and the second reference template (2125) based on the BCW candidate weights. The bidirectional predictor template can be determined based on the weighted average of the first reference template (2123) and the second reference template (2125) by the BCW candidate weights as shown in Equations 4 - 5. The TM cost can be determined based on the current template (2121) and the bidirectional predictor template using, for example, Equation 6 as described in FIG. 19. In one embodiment, the TM cost corresponding to each BCW candidate weight in the BCW list is determined. Based on the determined corresponding TM costs, such as in ascending order of the determined TM costs, the BCW candidate weights can be ranked (e.g., sorted). From the ranked BCW candidate weights, the BCW candidate weight can be selected as the BCW weight used to code the current block (2110). The current block (2110) can be reconstructed based on the selected BCW weight as shown in Equation 3.

[0168] The differences between the embodiments in FIGS. 19 and 21 are described below.

[0169] In the example of FIG. 19, the current block (1902) is coded in a non-sub-block-based mode such as a translational motion mode. Therefore, the first reference template (1915) is determined based on a single MV (e.g., MV (1916)), and the shape of the first reference template (1915) is the same as the current template (1905). Similarly, the second reference template (1925) is determined based on a single MV (e.g., MV (1926)), and the shape of the second reference template (1925) is the same as the current template (1905).

[0170] In the example of FIG. 21, the current block (2110) is coded in an affine mode (e.g., affine AMVP mode). Therefore, the first reference template (2123) can be determined based on different MVs, and the shape of the first reference template (2123) (e.g., including A 00 ~A 03 and L 00 ~L 03 can be different from the current template (2121) (e.g., including A0~A3 and L0~L3). In the example in FIG. 21, two different MVs point to A0 and A1 to A 00 and A 01 respectively, and therefore, the relative displacement between A 00 and A 01 is different from the relative displacement between A0 and A1. In one example, the second reference template (2125) is determined based on different MVs, and the shape of the second reference template (2125) (e.g., including A 10 ~A 13 and L 10 ~L 13 is different from the current template (2121).

[0171] Since the current template (2121) includes a plurality of sub-block templates (e.g., including A0 to A3 and L0 to L3), the TM cost calculated using Equations 4 and 6 can be rewritten based on the sub-block-based TM cost, and each sub-block-based TM cost is based on the corresponding sub-block-based bidirectional predictor template and the corresponding sub-block template.

[0172] The sub-block-based bidirectional predictor template (e.g., the k-th sub-block-based bidirectional predictor template) TP BCW,k can be determined based on the k-th first reference sub-block template TP0 k (e.g., A 00 ) associated with the sub-block template (e.g., A0) using Equation 8 and the k-th second reference sub-block template TP1 k (e.g., A 10 ).

Number

[0173] The sub-block-based TM cost (e.g., the k-th sub-block-based TM cost TM k ) can be determined based on the k-th sub-block-based bidirectional predictor template TP BCW,k and the corresponding k-th sub-block template (e.g., A0) within the current template (2121). For example, TM k = SAD(TC k - TP BCW,k ), where the parameter TC k represents the k-th sub-block template (e.g., A0).

[0174] In one example, the TM cost corresponding to the BCW candidate weight is determined based on a part of the current template (2121), a part of the first reference template (2123), and a part of the second reference template (2125). In other examples, the TM cost corresponding to the BCW candidate weight is determined based on the entirety of the current template (2121), the entirety of the first reference template (2123), and the entirety of the second reference template (2125). Therefore, as shown in Equation 9, the TM cost can be accumulated based on a part (subset) or all of the sub-block template-based TM costs of the sub-block templates.

Number

[0175] Equation 9 can conform to the following Equation 10. In one example, the TM cost is rewritten as follows.

Number

[0176] Parameter TC Ap represents the p-th upper sub-block template within the current template (2121) (for example, A0 when p is 0), and parameter TP0 Ap represents the p-th first upper reference sub-block template (for example, A 00 ) when p is 0, and parameter TP1 Ap represents the p-th second upper reference sub-block template (for example, A 10 ) when p is 0. The first sum in Equation 10 is executed for the upper sub-block templates within the current template (2121) such as A0 to A3 when the parameter p in Equation 10 is from 0 to 3.

[0177] Parameter TC Lm represents the m-th left sub-block template within the current template (2121) (for example, L0 when m is 0), and parameter TP0 Lmrepresents the m-th left reference sub-block template (e.g., when m is 0, it is L 00 ), and the parameter TP1 Lm represents the m-th left reference sub-block template (e.g., when l is 0, it is L 10 ). The second sum in Equation 10 is executed for the left sub-block templates of the current template (2121), such as L0 to L3 when the parameter m in Equation 10 is 0 to 3.

[0178] As described above, the weighting is normalized by 8 as shown in Equations 8 and 10. Other normalization coefficients such as 16, 32, etc. may be used.

[0179] Other functions such as SSE, variance, partial SAD, etc. may be used to determine the TM cost in Equation 9 or 10.

[0180] In an example of partial SAD, a part of the current template (2121) (e.g., A0 to A3), the corresponding part of the first reference template (2123) (e.g., A 00 ~A 03 ) and the corresponding part of the second reference template (2125) (e.g., A 10 ~A 13 ) are used to determine the TM cost.

[0181] In an example of partial SAD, a part or all of the current template (2121), a part or all of the first reference template (2123), and a part or all of the second reference template (2125) are downsampled before being used to determine the TM cost.

[0182] In the example shown in Figure 21, the first number of the upper sub-block template (e.g., 4) is equal to the second number of the left sub-block template (e.g., 4).

[0183] In other examples, the first number of the upper sub-block template is different from the second number of the left sub-block template.

[0184] In one embodiment, the inherited affine parameters of the current block (2110) can be applied (e.g., directly applied) to reference templates such as the first reference template (2123) and / or the second reference template (2125) in sub-block based TM. For example, the first reference sub-block template (e.g., A 00 ~A 03 and L 00 ~L 03 ) in the first reference template (2123) is determined based on the affine parameters (e.g., inherited affine parameters) of the current block (2110), and samples in each first reference sub-block template (e.g., one of A 00 ~A 03 or one of L 00 ~L 03 ) can have the same motion information (e.g., the same MV).

[0185] FIG. 22 shows an example of the PROF method. In some embodiments, the PROF method is implemented to improve sub-block based affine motion compensation to have finer-grained motion compensation. According to the PROF method, after sub-block based affine motion compensation is performed (shown in FIG. 14), the prediction samples (e.g., luma prediction samples) can be refined by adding a set of adjustment values derived by an optical flow equation.

[0186] Referring to FIG. 22, the current block (2210) is divided into four sub-blocks (2212, 2214, 2216, and 2218). For example, each of the sub-blocks (2212, 2214, 2216, and 2218) has a size of 4×4 pixels. The sub-block MV SB of the sub-block (2212) can be derived according to affine prediction and points to the reference sub-block (2232). The initial sub-block prediction samples can be determined according to the reference sub-block (2232). The refinement value applied to the initial sub-block prediction samples is such that each prediction sample is the sub-block MV of the sub-block 2212 adjusted by the adjustment vector ΔMVSB It can be calculated as if it were at the position (e.g., the position (2232a) of the sample (2212a)) indicated by the refined MV (e.g., pixel MV) (2242) determined according to []. Referring to FIG. 22, the MV SB Based on the initial sub-block prediction sample (2252) is refined so as to be a refined sample at the position (2232a) based on the pixel MV (2242).

[0187] In some embodiments, the PROF method may start by performing sub-block-based affine motion compensation to generate an initial sub-block prediction sample I(i1,i2) (2252), where (i1,i2) corresponds to a specific sample within the current sub-block. Next, the spatial gradient g of the initial sub-block prediction sample I(i1,I2) (2252) x (i1,i2) and g y (i1,i2) can be calculated using a 3-tap filter [-1,0,1] according to the following.

Equation

[0188] Sub-block prediction is extended by one pixel on each side for gradient calculation. In some embodiments, to reduce memory bandwidth and complexity, the pixels on the extended boundary can be copied from the nearest integer pixel position in the reference picture. Thus, further interpolation for the padding area is avoided.

[0189] Prediction refinement can be calculated by an optical flow formula.

Equation

[0190] Δmv(i1,i2) (e.g., ΔMV) is the pixel MV (2242) at the sample position (i1,i2) and the sub-block MV of the sub-block to which the pixel position (i1,i2) belongs SBis the difference from. Since the affine model parameters and the pixel positions with respect to the sub-block centers do not change for each sub-block, Δmv(i1,i2) is calculated for the first sub-block (e.g., (2212)) and can be reused for other sub-blocks (e.g., (2214), (2216), and (2218)) within the same coded block or CU (e.g., (2210)). In some examples, if x and y are the horizontal and vertical positions of Δ(i1,i2) with respect to the center of sub-block (2212), then Δ(i1,i2) can be derived by the following formula. [Number] Here, Δmv x (x,y) is the x-component of Δmv(i1,i2), and Δmv y (x,y) is the y-component of Δmv(i1,i2).

[0191] For the 4-parameter affine model, it is as follows. [Number]

[0192] For the 6-parameter affine model, it is as follows. [Number]

[0193] (v 0x ,v 0y ), (v 1x ,v 1y ) and (v 2x ,v 2y ) are the motion vectors of the upper-left, upper-right, and lower-left control points, and w and h are the width and height of the coded block or CU.

[0194] Prediction refinement can be added to the initial sub-block prediction sample I(i1,i2). The final prediction sample I’ by the PROF method can be generated using Equation 17.

Number

[0195] In one embodiment, returning to FIG. 21, for example, each first reference sub-block template (e.g., one of A 00 ~A 03 or one of L 00 ~L 03 in the first reference template (2123)) or each second reference sub-block template (e.g., one of A 10 ~A 13 or one of L 10 ~L 13 in the second reference template (2125)) is determined by applying PROF to each sub-block template. For example, the first reference sub-block template (e.g., one of A 00 ~A 03 or one of L 00 ~L 03 ) is determined using the PROF mode, and two samples within the same first reference sub-block template can have different motion information (e.g., two different MVs).

[0196] In one embodiment, referring to FIG. 23, when the current CU (e.g., the current block) (2310) is encoded in the affine mode (e.g., the affine AMVP mode), the TM (2300) is executed for the BCW candidate weights in the BCW list (e.g., all BCW candidate weights), and the BCW index can be rearranged by using a pair of translational MVs (e.g., including the first MV and the second MV) for the entire current template (2321) of the current block (2310). The current block (2310) includes a plurality of sub-blocks (2301). The first reference block (2311) includes a first sub-block (2303) predicted based on the CPMV (e.g., CPMV0 to CPMV1 or CPMV0 to CPMV2) of the current block (2310), for example, based on a plurality of sub-blocks (2301) using the affine mode. The second reference block (2313) includes a second sub-block (2305) predicted based on the CPMV (e.g., CPMV0 to CPMV1 or CPMV0 to CPMV2) of the current block (2310), for example, based on a plurality of sub-blocks (2301) using the affine mode.

[0197] The current template (2321) of the current block (2310) includes the upper template (2341) and the left template (2342). According to one embodiment of the present disclosure, the first reference template (2323) can be predicted from the current template (2321) using a single MV (e.g., the first MV of a pair of translational MVs), and the second reference template (2325) can be predicted from the current template (2321) using another single MV (e.g., the second MV of a pair of translational MVs). The pair of MVs (e.g., the first MV and the second MV) can be determined based on CPMV0, CPMV1, or CPMV2. In one example, the first reference template (2323) includes the first upper reference template (2351) and the first left reference template (2352). In one example, the second reference template (2325) includes the second upper reference template (2353) and the second left reference template (2354). The first reference template (2323) and the second reference template (2325) can have the same shape as the shape of the current template (2321), and can have the same size as the size of the current template (2321).

[0198] TM (2300) can be executed on the BCW candidate weights using the current template (2321), the first reference template (2323), and the second reference template (2325) as described with reference to FIG. 19.

[0199] Regarding the TM in FIG. 21 or FIG. 23, the current template (2121) or the current template (2321) is taken as an example for explanation. When the current block is predicted using the affine mode (e.g., the affine AMVP mode), other shapes as described in FIGS. 20A-20D can be used as the current template, and FIGS. 21 and 23 can be appropriately adapted.

[0200] The various embodiments in FIGS. 19, 21, and 23 are described using all or the entire current template and the corresponding first and second reference templates. The embodiments in FIGS. 19, 21, and 23 can be appropriately adapted when a portion of the current template and the corresponding portions of the first and second reference templates are used in determining the TM cost as used in Equations 4-10.

[0201] The TM described in FIGS. 19, 21, and 23 can be applied to a subset or all of the BCW candidate weights within the BCW list.

[0202] FIG. 24 shows a flowchart illustrating an overview of an encoding process (2400) according to an embodiment of the present disclosure. In various embodiments, the process (2400) is executed by processing circuits such as processing circuits within terminal devices (310), (320), (330), and (340), and processing circuits that execute the functions of video encoders (e.g., (403), (603), (703)). In some embodiments, the process (2400) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (2400). The process starts from (S2401) and proceeds to (S2410).

[0203] (S2410), for the current block in the current picture encoded by bi-prediction with CU-level weights (BCW) at the coding unit (CU) level, as described in FIGS. 19, 21, and 23, for example, template matching (TM) can be performed for each BCW candidate weight within the BCW list. TM can be performed by (i) determining a respective TM cost corresponding to each BCW candidate weight and (ii) selecting, based on each determined TM cost, the BCW candidate weight that will be the BCW weight used to encode the current block.

[0204] In one embodiment, as described in FIGS. 21 and 23, each TM cost can be determined based at least on a part of the current template of the current block and each bidirectional predictor template. The bidirectional predictor template can be based on each BCW candidate weight, a part of the first reference template in the first reference picture, and a part of the first reference template in the second reference picture, and the first reference template and the second reference template correspond to the current template.

[0205] In one example, as described in FIGS. 19, 21, and 23, each TM cost can be determined based on the entire current template of the current block (i.e., the entire current template) and each bidirectional predictor template. The bidirectional predictor template can be based on each BCW candidate weight, the entire first reference template, and the entire second reference template.

[0206] In one embodiment, the BCW candidate weights are ranked or sorted based on the respectively determined TM costs, and the BCW candidate weight that is ranked or sorted is selected from the ranked or sorted BCW candidate weights to be the BCW weight used to encode the current block.

[0207] In one embodiment, all of the current template is used to determine each TM cost. For each BCW candidate weight, all of the first reference template determined based on the first motion vector (MV) of the current block is used to calculate the bidirectional predictor template, and all of the second reference template determined based on the second MV of the current block is used to calculate the bidirectional predictor template. In one example, for each BCW candidate weight, the bidirectional predictor template is a weighted average of all of the first reference template and all of the second reference template based on each BCW candidate weight.

[0208] In one example, the current block is predicted in an affine AMVP mode having a plurality of control points, and the first MV and the second MV are associated with a certain control point among the plurality of control points.

[0209] In one embodiment, the shape of the current template is based on one or more of (i) the reconstructed samples of the adjacent blocks of the current block, (ii) the decoding order of the current block, or (iii) the size of the current block.

[0210] In one example, the current template includes a reconstruction region that is an adjacent region of the current block. For example, the reconstruction region is one of (i) the left adjacent region and the upper adjacent region, (ii) the left adjacent region, the upper adjacent region, and the upper left adjacent region, (iii) the upper adjacent region, or (iv) the left adjacent region.

[0211] In one embodiment, the current block is predicted in an affine mode (e.g., affine AMVP mode). The current template includes the current sub-blocks, and a part of the current template used to determine each TM cost is one of the current sub-blocks. For each BCW candidate weight, the first reference template includes the first reference sub-blocks corresponding to the current sub-blocks respectively, and a part of the first reference template used to calculate the bidirectional predictor template is one of the first reference sub-blocks. The second reference template includes the second reference sub-blocks corresponding to the current sub-blocks respectively, and a part of the second reference template used to calculate the bidirectional predictor template is one of the second reference sub-blocks. The bidirectional predictor template can be based on each BCW candidate weight, one of the first reference sub-blocks, and one of the second reference sub-blocks. In one example, for each BCW candidate weight, the bidirectional predictor template is a weighted average of one of the first reference sub-blocks and one of the second reference sub-blocks based on each BCW candidate weight.

[0212] In one example, the BCW candidate weights are normalized to 8, 16, or 32.

[0213] (S2420), the current block can be encoded based on the selected BCW weight. Prediction information indicating that the current block is predicted by BCW can be encoded. In one example, the prediction information indicates a BCW index that points to a BCW candidate weight among the ranked BCW candidate weights.

[0214] In one example, the prediction information indicates that the current block is predicted in an affine mode (e.g., affine AMVP mode).

[0215] (S2430), the encoded prediction information and the encoded current block can be included in the video bitstream. Then, the process (2400) proceeds to (S2499) and ends.

[0216] The process (2400) can be appropriately adapted to various scenarios, and the steps in the process (2400) can be adjusted accordingly. One or more steps in the process (2400) can be adapted, omitted, repeated, and / or combined. Any appropriate order can be used to implement the process (2400). Additional steps can be added.

[0217] FIG. 25 shows a flowchart illustrating an overview of a decoding process (2500) according to an embodiment of the present disclosure. In various embodiments, the process (2500) is executed by a processing circuit such as a processing circuit in the terminal devices (310), (320), (330), and (340), a processing circuit that executes the functions of the video encoder (403), a processing circuit that executes the functions of the video decoder (410), a processing circuit that executes the functions of the video decoder (510), a processing circuit that executes the functions of the video encoder (603), etc. In some embodiments, the process (2500) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes the process (2500). The process starts from (S2501) and proceeds to (S2510).

[0218] In (S2510), the prediction information of the current block in the current picture can be decoded from the coded video bitstream. The prediction information can indicate that the current block is predicted by bi-prediction with CU-level weights (BCW).

[0219] In (S2520), as described in FIGS. 19, 21, and 23, for example, template matching (TM) can be performed for each BCW candidate weight in the BCW list. TM can be performed by (i) determining respective TM costs for each of the respective BCW candidate weights, and (ii) selecting, based on the respectively determined TM costs, the BCW candidate weight that becomes the BCW weight used to reconstruct the current block.

[0220] As described in FIGS. 21 and 23, each TM cost can be determined based at least in part on the current template of the current block and part or all of each of the bi-prediction sub-templates. The bi-prediction sub-template can be based on each BCW candidate weight, part or all of the first reference template in the first reference picture, and part or all of the first reference template in the second reference picture, and the first reference template and the second reference template correspond to the current template.

[0221] In one embodiment, the BCW candidate weights are ranked (e.g., sorted) based on the respectively determined TM costs, and the BCW candidate weight is selected from the ranked BCW candidate weights to be the BCW weight.

[0222] In one embodiment, all of the current template (i.e., the entire current template) is used to determine each TM cost. For each BCW candidate weight, all of the first reference template determined based on the first motion vector (MV) of the current block (i.e., the entire first reference template) is used to calculate the bidirectional predictor template, and all of the second reference template determined based on the second MV of the current block (i.e., the entire second reference template) is used to calculate the bidirectional predictor template. In one example, for each BCW candidate weight, the bidirectional predictor template is a weighted average of all of the first reference template and all of the second reference template based on their respective BCW candidate weights.

[0223] In one example, the prediction information decoded in (S2510) indicates that the current block is predicted in an affine AMVP mode having a plurality of control points. The first MV and the second MV are associated with a certain control point among the plurality of control points.

[0224] In one embodiment, the shape of the current template is based on one or more of (i) the reconstructed samples of the adjacent blocks of the current block, (ii) the decoding order of the current block, or (iii) the size of the current block.

[0225] In one embodiment, the current template includes a reconstructed region that is an adjacent region of the current block. For example, the reconstructed region is one of (i) the left adjacent region and the upper adjacent region, (ii) the left adjacent region, the upper adjacent region, and the upper left adjacent region, (iii) the upper adjacent region, or (iv) the left adjacent region.

[0226] In one embodiment, the prediction information decoded in (S2510) indicates that the current block is predicted in affine mode (e.g., affine AMVP mode). The current template includes the current sub-block, and a part of the current template used to determine each TM cost is one of the current sub-blocks. For each BCW candidate weight, the first reference template includes a first reference sub-block corresponding to the current sub-block respectively, and a part of the first reference template used to calculate the bidirectional predictor template is one of the first reference sub-blocks. The second reference template includes a second reference sub-block corresponding to the current sub-block respectively, and a part of the second reference template used to calculate the bidirectional predictor template is one of the second reference sub-blocks. The bidirectional predictor template can be based on each BCW candidate weight, one of the first reference sub-blocks, and one of the second reference sub-blocks. In one example, for each BCW candidate weight, the bidirectional predictor template is a weighted average of one of the first reference sub-blocks and one of the second reference sub-blocks based on each BCW candidate weight.

[0227] In one example, the BCW candidate weights are normalized to 8, 16, or 32.

[0228] (S2530), the current block can be reconstructed based on the selected BCW weight.

[0229] The process (2500) proceeds to (S2599) and ends.

[0230] The process (2500) can be appropriately adapted to various scenarios, and the steps in the process (2500) can be adjusted accordingly. One or more steps in the process (2500) can be adapted, omitted, repeated, and / or combined. Any appropriate order can be used to implement the process (2500). Additional steps can be added.

[0231] FIG. 26 shows a flowchart illustrating an overview of a decoding process (2600) according to an embodiment of the present disclosure. In various embodiments, process (2600) is executed by a processing circuit such as in terminal devices (310), (320), (330), and (340), a processing circuit that executes the functions of video encoder (403), a processing circuit that executes the functions of video decoder (410), a processing circuit that executes the functions of video decoder (510), a processing circuit that executes the functions of video encoder (603), and the like. In some embodiments, process (2600) is implemented by software instructions, and thus, when the processing circuit executes the software instructions, the processing circuit executes process (2600). The process starts at (S2601) and proceeds to (S2610).

[0232] (S2610), prediction information of the current block in the current picture can be decoded from the coded video bitstream.

[0233] (S2620), it is determined that the current block is predicted by bidirectional prediction and the prediction information indicates that bi-prediction with CU-level weights (BCW) is valid for the current block at the coding unit (CU) level.

[0234] (S2630), as described in FIGS. 19, 21, and 23, for example, template matching (TM) can be performed for each BCW candidate weight in the BCW list. TM can be performed by (i) determining respective TM costs corresponding to each of the respective BCW candidate weights and (ii) sorting the BCW candidate weights based on the respectively determined TM costs.

[0235] Each TM cost can be determined based at least in part on the current template of the current block and on part or all of each bidirectional predictor template, where each bidirectional predictor template is based on each BCW candidate weight, on part or all of the first reference template in the first reference picture, and on part or all of the first reference template in the second reference picture. The first reference template and the second reference template correspond to the current template.

[0236] (In S2640), the current block can be reconstructed based on the sorted BCW candidate weights.

[0237] Process (2600) proceeds to (S2699) and ends.

[0238] Process (2600) can be appropriately adapted to various scenarios, and the steps in process (2600) can be adjusted accordingly. One or more steps in process (2600) can be adapted, omitted, repeated, and / or combined. Any suitable order can be used to implement process (2600). Further steps can be added.

[0239] The embodiments in the present disclosure may be used individually or may be combined in any order. Further, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.

[0240] The above-described techniques can be implemented as computer software using computer-readable instructions and can be physically stored on one or more computer-readable media. For example, FIG. 27 shows a computer system (2700) suitable for implementing a particular embodiment of the disclosed subject matter.

[0241] Computer software can be coded using any suitable machine code or computer language and be the subject of assembly, compilation, linking, or similar mechanisms to create code containing instructions executable directly, or through interpretation, microcode execution, etc., by one or more computer central processing units (CPUs), graphics processing units (GPUs), and the like.

[0242] Instructions can be executed on various types of computers or their components, including, for example, personal computers, tablet computers, servers, smartphones, game devices, Internet of Things devices, and the like.

[0243] The components shown in FIG. 27 for the computer system (2700) are illustrative in nature and are not intended to suggest any limitation as to the use or functionality of the computer software implementing embodiments of the present disclosure. Nor should the configuration of the components be construed as having any dependency or requirement regarding any one or combination of the components shown in the exemplary embodiments of the computer system (2700).

[0244] The computer system (2700) can include specific human interface input devices. Such human interface input devices can respond to input by one or more human users through, for example, tactile input (e.g., keystrokes, swipes, movement of a data glove), voice input (e.g., voice, clapping), visual input (e.g., gesture), olfactory input (not shown). Also, the human interface device can be used to capture certain media not necessarily directly related to conscious human input, such as voice (e.g., speech, music, ambient sound), images (e.g., scanned images, photographic images obtained from a still image camera), video (e.g., two-dimensional video, three-dimensional video including stereoscopic video).

[0245] The input human interface device may include one or more of a keyboard (2701), a mouse (2702), a trackpad (2703), a touch screen (2710), a data glove (not shown), a joystick (2705), a microphone (2706), a scanner (2707), and a camera (2708) (only one of each is shown).

[0246] The computer system (2700) may also include certain human interface output devices. Such human interface output devices may, for example, stimulate the senses of one or more human users through tactile output, sound, light, and smell / taste. Such human interface output devices may include tactile output devices (e.g., tactile feedback by a touch screen (2710), a data glove (not shown), or a joystick (2705); however, there may also be tactile feedback devices that do not function as input devices), audio output devices (e.g., speakers (2709), headphones (not shown)), visual output devices (e.g., a screen (2710) including a CRT screen, an LCD screen, a plasma screen, an OLED screen; each may or may not have a touch screen input function, each may or may not have a tactile feedback function, and some of them can output higher than three-dimensional output through means such as two-dimensional visual output or stereoscopic output; virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and a printer (not shown).

[0247] The computer system (2700) can also include an optical medium including a CD / DVD ROM / RW (2720) together with a human-accessible memory device and associated media, such as a CD / DVD or similar media (2721), a thumb drive (2722), a removable hard drive or solid state drive (2723), legacy magnetic media such as tapes and floppy disks (not shown), specialized ROM / ASIC / PLD-based devices such as security dongles (not shown), etc.

[0248] One of ordinary skill in the art should also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include a transmission medium, a carrier wave, or other transient signals.

[0249] The computer system (2700) can also include an interface (2754) to one or more communication networks (2755). The network can be, for example, wireless, wired, or optical. The network can further be local, wide area, metropolitan area, in-vehicle and industrial, real-time, delay tolerant, etc. Examples of networks include Ethernet®, wireless LAN, GSM, 3G, 4G, 5G, LTE, etc. cellular networks, cable TV, satellite TV, TV wired or wireless wide area digital networks including terrestrial broadcast TV, in-vehicle and industrial including CANBus, etc. A particular network typically requires an external network interface adapter attached to a particular general-purpose data port or peripheral bus (2749) (e.g., a USB port of the computer system (2700), etc.). Others are typically integrated into the core of the computer system (2700) by attachment to a system bus as described later (e.g., an Ethernet interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (2700) can communicate with other entities. Such communication can be unidirectional, receive-only (e.g., broadcast TV), dedicated unidirectional transmission (e.g., CANbus to a particular CANbus device), or bidirectional to other computer systems using, for example, local or wide area digital networks. For each of the networks and network interfaces as described above, a particular protocol and protocol stack can be used.

[0250] The aforementioned human interface device, human-accessible memory device, and network interface can be attached to the core (2740) of the computer system (2700).

[0251] The core (2740) can include one or more central processing units (CPUs) (2741), a graphics processing unit (GPU) (2742), a specialized programmable processing device in the form of a field programmable gate array (FPGA) (2743), a hardware accelerator (2744) for a specific task, a graphics adapter (2750), etc. These devices can be connected through a system bus (2748) together with a read-only memory (ROM) (2745), a random access memory (2746), an internal mass storage device such as an internal hard drive or a solid state drive (SSD) that is not accessible to the user (2747). In some computer systems, the system bus (2748) may be accessible in the form of one or more physical plugs to enable expansion by additional CPUs, GPUs, etc. Peripheral devices can be attached directly to the system bus (2748) of the core or through a peripheral bus (2749). In one example, a screen (2710) can be connected to the graphics adapter (2750). Architectures for the peripheral bus include PCI, USB, etc.

[0252] The CPU (2741), GPU (2742), FPGA (2743), and accelerator (2744) can execute specific instructions that can together constitute the above-described computer code. That computer code can be stored in the ROM (2745) or the RAM (2746). Temporary data can be stored in the RAM (2746), while persistent data can be stored, for example, in the internal mass storage device (2747). By using a cache memory that can be closely associated with one or more CPUs (2741), GPUs (2742), the mass storage device (2747), the ROM (2745), the RAM (2746), etc., fast storage and retrieval to any of the memory devices can be enabled.

[0253] A computer-readable medium can have computer code thereon for performing various computer-implemented operations. The medium and the computer code may be specially designed and constructed for the purposes of this disclosure, or they may be of the kind well known and available to those having skill in the computer software arts.

[0254] By way of example and not limitation, a computer system having an architecture (2700), specifically a core (2740), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied on one or more tangible computer-readable media. Such computer-readable media can be media related to user-accessible mass storage as introduced above and specific storage of the core (2740) of a non-transitory nature such as a mass storage device (2747) inside the core or a ROM (2745). The software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (2740). The computer-readable media can include one or more memory devices or chips depending on specific needs. The software can include defining data structures stored in a RAM (2746) and modifying such data structures according to processes defined by the software, and causing the core (2740) and specifically the processors (including a CPU, GPU, FPGA, etc.) therein to execute specific processes or specific portions described herein. Additionally or alternatively, the computer system can provide functionality as a result of logic wired within a circuit (e.g., an accelerator (2744)) or otherwise embodied, which can operate instead of or in conjunction with software for executing specific processes or specific portions of specific processes described herein. References to software include logic and vice versa as appropriate. References to computer-readable media can, as appropriate, include circuits (e.g., integrated circuits (ICs)) storing software for execution, circuits embodying logic for execution, or both. The present disclosure encompasses any suitable combination of hardware and software.

[0255] Appendix A: Acronyms JEM: joint exploration model VVC: versatile video coding BMS: benchmark set MV: Motion Vector HEVC: High Efficiency Video Coding SEI: Supplementary Enhancement Information VUI: Video Usability Information GOP: Group of Pictures TU: Transform Unit PU: Prediction Unit CTU: Coding Tree Unit CTB: Coding Tree Block PB: Prediction Block HRD: Hypothetical Reference Decoder SNR: Signal Noise Ratio CPU: Central Processing Unit GPU: Graphics Processing Unit CRT: Cathode Ray Tube LCD: Liquid-Crystal Display OLED: Organic Light-Emitting Diode CD: Compact Disc DVD: Digital Video Disc ROM: Read-Only Memory RAM: Random Access Memory ASIC: Application-Specific Integrated Circuit PLD: Programmable Logic Device LAN: Local Area Network GSM: Global System for Mobile communications LTE: Long-Term Evolution CANBus: Controller Area Network Bus USB: Universal Serial Bus PCI: Peripheral Component Interconnect FPGA: Field Programmable Gate Areas SSD: solid-state drive IC: Integrated Circuit CU: Coding Unit R-D: Rate-Distortion

[0256] Although the present disclosure has described several exemplary embodiments, there are changes, substitutions, and various alternative equivalents that fall within the scope of the present disclosure. Thus, it will be understood by those skilled in the art that many systems and methods that embody the principles of the present disclosure but are not explicitly shown or described herein can be devised and are thus within the spirit and scope of the present disclosure.

Claims

Claim 1. A method for video encoding in a video encoder, comprising: determining a respective template matching (TM) cost corresponding to each of bidirectional prediction (BCW) candidate weights for each coding unit (CU) level weight, wherein each TM cost is determined based at least on a current template of a current block and some or all of a respective bidirectional predictor template, wherein the bidirectional predictor template is based on each BCW candidate weight, some or all of a first reference template in a first reference picture, and some or all of a second reference template in a second reference picture, and the first reference template and the second reference template correspond to the current template; sorting the BCW candidate weights based on their respective determined TM costs; performing TM for each BCW candidate weight of the current block in the current picture by encoding the current block based on the sorted BCW candidate weights; A method comprising:

2. The method described in claim 1, wherein the step of performing the TM further includes a step of selecting a BCW candidate weight from the sorted BCW candidate weights to be the BCW weight used to encode the current block. all of the current templates are used to determine each TM cost; For each BCW candidate weight, all of the first reference templates determined based on a first motion vector (MV) of the current block are used to calculate the bidirectional predictor template; The method of claim 2 , wherein all of the second reference templates determined based on a second MV of the current block are used to calculate the bidirectional predictor template.

4. The method described in claim 3, wherein for each BCW candidate weight, the bidirectional predictor template is a weighted average of all of the first reference templates and all of the second reference templates, and the weights of the weighted average are based on the respective BCW candidate weights.

5. The current block is predicted in an affine adaptive motion vector prediction (AMVP) mode with multiple control points; The method of claim 3 , wherein the first MV and the second MV are associated with a control point among the plurality of control points.

6. The method described in claim 1, wherein the shape of the current template is based on one or more of: (i) reconstructed samples of adjacent blocks of the current block, (ii) the decoding order of the current block, or (iii) the size of the current block.

7. The method described in claim 1, wherein the current template includes one or more reconstruction regions that are adjacent regions of the current block.

8. The method described in claim 7, wherein the one or more reconstruction regions that are adjacent regions of the current block are one of (i) a left adjacent region and an upper adjacent region, (ii) the left adjacent region, the upper adjacent region and the upper-left adjacent region, (iii) the upper adjacent region, or (iv) the left adjacent region.

9. The current block is predicted in affine mode; the current template includes a current sub-block, and the portion of the current template used to determine each TM cost is one of the current sub-blocks; For each BCW candidate weight, the first reference template includes first reference sub-blocks each corresponding to the current sub-block, and the portion of the first reference template used to calculate the bidirectional predictor template is one of the first reference sub-blocks; the second reference template includes second reference sub-blocks each corresponding to the current sub-block, and the portion of the second reference template used to calculate the bidirectional predictor template is one of the second reference sub-blocks; The method of claim 2 , wherein the bidirectional predictor template is based on respective BCW candidate weights, the one of the first reference sub-blocks, and the one of the second reference sub-blocks.

10. The method described in claim 9, wherein for each BCW candidate weight, the bidirectional predictor template is a weighted average of one of the first reference subblocks and one of the second reference subblocks, and the weights of the weighted average are based on the respective BCW candidate weights.

11. The method described in claim 9, wherein the BCW candidate weights are normalized by 8, 16, or 32.

12. A method for processing visual media data, comprising: generating a bitstream including the visual media data according to a format rule; transmitting the bitstream; Including, The bitstream includes prediction information for a current block in a current picture, the prediction information indicating (1) that the current block is predicted using bidirectional prediction, and (2) that coding unit (CU) level weighted bidirectional prediction (BCW) is enabled for the current block; The formatting rules are: determining a respective template matching (TM) cost corresponding to each of the respective BCW candidate weights, wherein each TM cost is determined based at least on a current template of the current block and some or all of a respective bidirectional predictor template, the bidirectional predictor template being based on each BCW candidate weight, some or all of a first reference template in a first reference picture, and some or all of a second reference template in a second reference picture, the first reference template and the second reference template corresponding to the current template; sorting the BCW candidate weights based on their respective determined TM costs; specifies that TM is performed for each BCW candidate weight by The current block is processed based on the reordered BCW candidate weights.

13. The method described in claim 12, wherein performing the TM further includes selecting a BCW candidate weight from the sorted BCW candidate weights to be the BCW weight used to encode the current block.

14. The method of claim 13, wherein all of the current templates are used to determine each TM cost; For each BCW candidate weight, all of the first reference templates determined based on a first motion vector (MV) of the current block are used to calculate the bidirectional predictor template; The method of claim 13 , wherein all of the second reference templates determined based on a second MV of the current block are used to calculate the bidirectional predictor template.

15. The method described in claim 14, wherein for each BCW candidate weight, the bidirectional predictor template is a weighted average of all of the first reference templates and all of the second reference templates, and the weights of the weighted average are based on the respective BCW candidate weights.

16. The prediction information indicates that the current block is predicted in an affine adaptive motion vector prediction (AMVP) mode with multiple control points; The method of claim 14 , wherein the first MV and the second MV are associated with a control point of the plurality of control points.

17. The method described in claim 12, wherein the shape of the current template is based on one or more of: (i) reconstructed samples of adjacent blocks of the current block, (ii) the decoding order of the current block, or (iii) the size of the current block.

18. The method described in claim 12, wherein the current template includes one or more reconstruction regions that are adjacent regions of the current block.

19. The method described in claim 18, wherein the one or more reconstruction regions that are adjacent regions of the current block are one of (i) a left adjacent region and an upper adjacent region, (ii) the left adjacent region, the upper adjacent region and the upper-left adjacent region, (iii) the upper adjacent region, or (iv) the left adjacent region.

20. The prediction information indicates that the current block is predicted in an affine mode, the current template includes a current sub-block, and the portion of the current template used to determine each TM cost is one of the current sub-blocks; For each BCW candidate weight, the first reference template includes first reference sub-blocks each corresponding to the current sub-block, and the portion of the first reference template used to calculate the bidirectional predictor template is one of the first reference sub-blocks; the second reference template includes second reference sub-blocks each corresponding to the current sub-block, and the portion of the second reference template used to calculate the bidirectional predictor template is one of the second reference sub-blocks; The method of claim 13 , wherein the bidirectional predictor template is based on respective BCW candidate weights, the one of the first reference sub-blocks, and the one of the second reference sub-blocks.

21. An apparatus for encoding video, including processing circuitry, 12. Apparatus, wherein the processing circuitry is configured to perform the method of any one of claims 1 to 11.

22. An apparatus for processing visual media data, including processing circuitry, 21. Apparatus, wherein the processing circuitry is configured to perform a method according to any one of claims 12 to 20.

23. A program causing at least one processor to execute a method according to any one of claims 1 to 11.

24. A program causing at least one processor to execute a method according to any one of claims 12 to 20.