Video decoding method, video encoder and computer-readable storage medium

By receiving and transmitting weight information at the encoding unit level and enabling or disabling weighted two-way prediction, the compatibility issues of different encoding tools in the video encoding system are solved, and encoding efficiency and time prediction accuracy are improved, especially when video content is attenuated.

CN116347097BActive Publication Date: 2025-08-12ALIBABA GROUP HOLDING LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310311779.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-12-20
Filing Date
2019-12-19
Publication Date
2025-08-12
Estimated Expiration
2039-12-19

AI Technical Summary

Technical Problem

Existing video encoding systems are difficult to compatible when using different encoding tools, especially when bidirectional prediction and weighted prediction, resulting in low encoding efficiency, especially when video content decays, the accuracy of bidirectional prediction is insufficient.

Method used

At the encoding unit level, weighted bidirectional prediction is enabled or disabled by receiving and transmitting weight information, enabling or disabling weighted bidirectional prediction, ensuring the exclusive use of weighted average bidirectional prediction and other tools, coordinating compatibility between different encoding tools, and optimizing the encoding process using merge candidate lists and history-based motion vector prediction.

Benefits of technology

Improve the encoding efficiency and compression performance of video encoding, especially when the video content is attenuated, the accuracy of time prediction and the compression capability of the encoder are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116347097B_ABST
    Figure CN116347097B_ABST
Patent Text Reader

Abstract

The present application provides a video decoding method, a video encoder, and a computer-readable storage medium. Disclosed are video encoding and decoding techniques for bidirectional prediction with weighted averaging. According to certain embodiments, a computer-implemented video decoding method includes the following steps: receiving a bitstream including weight information for predicting a coding unit (CU); and based on the weight information, performing weighted bidirectional prediction for the CU and disabling bidirectional prediction with weighted averaging (BWA) for the CU, or performing BWA for the CU and disabling weighted bidirectional prediction for the CU.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application with application number 201980085317.0 (International application number: PCT / US2019 / 067619, application date: December 19, 2019, invention name: Block-level bidirectional prediction with weighted averaging).

[0002] Related applications

[0003] This application claims priority to U.S. Patent Application No. 16 / 228,741, filed on December 20, 2018, which is hereby incorporated by reference in its entirety. Technical Field

[0004] The present disclosure relates generally to video processing, and more particularly, to video encoding and decoding using bi-prediction with weighted averaging (BWA) at the block (or coding unit) level. Background Art

[0005] Video coding systems are commonly used to compress digital video signals, for example to reduce storage space consumed or to reduce transmission bandwidth consumption associated with such signals.

[0006] Video coding systems can use a variety of tools or techniques to solve different problems. For example, temporal motion prediction is an effective method to improve coding efficiency and provide high compression. Temporal motion prediction can be unidirectional prediction using a single reference picture or bidirectional prediction using two reference pictures. In some cases, such as when fading occurs, bidirectional prediction may not produce the most accurate prediction. To compensate for this, weighted prediction can be used to weight the two prediction signals differently.

[0007] However, different coding tools are not always compatible. For example, it may be inappropriate to apply the above-mentioned temporal prediction, bidirectional prediction, or weighted prediction to the same coding block (e.g., coding unit), the same slice, or the same picture. Therefore, it is desirable to make different coding tools interact appropriately with each other. Summary of the Invention

[0008] Embodiments of the present disclosure relate to methods for encoding and signaling weights for bidirectional prediction based on weighted averaging at the coding unit (CU) level. In some embodiments, a computer-implemented video decoding method is provided, the video decoding method comprising the following steps: receiving a bitstream including weight information for predicting a coding unit (CU); and based on the weight information, performing weighted bidirectional prediction for the CU and disabling bidirectional prediction with weighted averaging (BWA) for the CU, or performing BWA for the CU and disabling weighted bidirectional prediction for the CU.

[0009] In some embodiments, a video encoder is provided, comprising: a memory storing instructions; and a processor configured to execute the instructions to cause the video encoder to: transmit a bitstream including weight information for predicting a coding unit CU to a video decoder, wherein the weight information indicates: performing weighted bidirectional prediction of the CU and disabling bidirectional prediction with weighted averaging (BWA) for the CU, or performing the BWA of the CU and disabling the weighted bidirectional prediction for the CU.

[0010] In some embodiments, a non-transitory computer-readable storage medium storing an instruction set is provided, wherein the non-transitory computer-readable storage medium stores a bit stream generated by encoding, and the instruction set can be executed by one or more processors of a video encoder to cause the video encoder to perform a method including the following steps: transmitting the bit stream including weight information for predicting a coding unit CU to a video decoder, wherein the weight information indicates: performing weighted bidirectional prediction of the CU and disabling bidirectional prediction BWA with weighted averaging for the CU, or performing the BWA of the CU and disabling the weighted bidirectional prediction for the CU.

[0011] In some embodiments, a computer-implemented video transmission method is provided. The video transmission method includes the following steps: a processor transmits a bitstream to a video encoder, the bitstream including weight information for predicting a coding unit (CU). The weight information indicates that if weighted prediction is enabled for a bidirectional prediction mode of the CU, weighted averaging for the bidirectional prediction mode is disabled.

[0012] In some embodiments, a computer-implemented video encoding method is provided. The video encoding method includes the following steps: a processor constructs a merge candidate list for a coding unit, the merge candidate list including motion information of a non-affine inter-frame coded block of the coding unit, the motion information including a bidirectional prediction weight associated with the non-affine inter-frame coded block. The video encoding method also includes the following steps: the processor performs encoding based on the motion information.

[0013] In some embodiments, a computer-implemented video transmission method is provided. The video transmission method includes the following steps: a processor determines the value of a bidirectional prediction weight for a coding unit (CU) of a video frame. The video transmission method also includes the following steps: the processor determines whether the bidirectional prediction weight is an equal weight. The video transmission method also includes the following steps: in response to the determination, the processor transmits the following to a video decoder: when the bidirectional prediction weight is an equal weight, transmitting a bitstream including a first syntax element, the first syntax element indicating the equal weight, or after determining that the bidirectional prediction weight is an unequal weight, transmitting a bitstream including a second syntax element, the second syntax element indicating the value of the bidirectional prediction weight corresponding to the unequal weight.

[0014] In some embodiments, a computer-implemented signal transmission method performed by a decoder is provided. The signal transmission method includes the following steps: the decoder receives a bitstream including weight information for predicting a coding unit (CU) from a video encoder. The signal transmission method also includes the following steps: if weighted prediction is enabled for a bidirectional prediction mode of the CU, determining to disable weighted averaging for the bidirectional prediction mode based on the weight information.

[0015] In some embodiments, a computer-implemented video encoding method performed by a decoder is provided. The video encoding method includes the following steps: the decoder receives a merge candidate list for a coding unit from an encoder, the merge candidate list including motion information of non-adjacent inter-frame coded blocks of the coding unit. The video encoding method also includes the following steps: determining bidirectional prediction weights associated with the non-adjacent inter-frame coded blocks based on the motion information.

[0016] In some embodiments, a computer-implemented signal transmission method performed by a decoder is provided. The signal transmission method includes the following steps: the decoder receives from a video encoder: a bitstream including a first syntax element corresponding to a bidirectional prediction weight for a coding unit (CU) of a video frame, or a bitstream including a second syntax element corresponding to the bidirectional prediction weight. The signal transmission method also includes the following steps: in response to receiving the first syntax element, the processor determines that the bidirectional prediction weight is an equal weight. The signal transmission method also includes the following steps: in response to receiving the first syntax element, the processor determines that the bidirectional prediction weight is an unequal weight, and the processor determines the value of the unequal weight based on the second syntax element.

[0017] Aspects of the disclosed embodiments may include a non-transitory tangible computer-readable medium storing software instructions that, when executed by one or more processors, are configured to and capable of performing and executing one or more of the methods, operations, etc., consistent with the disclosed embodiments. Furthermore, aspects of the disclosed embodiments may be performed by one or more processors configured as special-purpose processors based on software instructions programmed with logic and instructions that, when executed, perform one or more operations consistent with the disclosed embodiments.

[0018] Additional objects and advantages of the disclosed embodiments will be set forth in part in the following description and in part will be obvious from the description or may be learned by practice of the embodiments. The objects and advantages of the disclosed embodiments may be realized and obtained by the elements and combinations set forth in the claims.

[0019] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosed embodiments, as claimed. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a schematic diagram illustrating an exemplary video encoding and decoding system consistent with embodiments of the present disclosure.

[0021] Figure 2 It is an example of what can be consistent with the embodiment of the present disclosure. Figure 1 Schematic diagram of an exemplary video encoder that is part of an exemplary system.

[0022] Figure 3 It is an example of what can be consistent with the embodiment of the present disclosure. Figure 1 Schematic diagram of an exemplary video decoder that is part of an exemplary system.

[0023] Figure 4 This is a table of syntax elements for weighted prediction (WP) consistent with an embodiment of the present disclosure.

[0024] Figure 5 is a schematic diagram illustrating bidirectional prediction consistent with an embodiment of the present disclosure.

[0025] Figure 6 is a table of syntax elements for bidirectional prediction with weighted averaging (BWA) consistent with an embodiment of the present disclosure.

[0026] Figure 7 is a schematic diagram illustrating spatial neighbors used in merging candidate list construction consistent with an embodiment of the present disclosure.

[0027] Figure 8 is a table of syntax elements for transmitting enabling or disabling of WP at the picture level consistent with an embodiment of the present disclosure.

[0028] Figure 9 is a table of syntax elements for transmitting activation or deactivation of WP at a slice level consistent with an embodiment of the present disclosure.

[0029] Figure 10 is a table of syntax elements for maintaining exclusivity of WP and BWA at the CU level consistent with an embodiment of the present disclosure.

[0030] Figure 11 is a table of syntax elements for maintaining exclusivity of WP and BWA at the CU level consistent with an embodiment of the present disclosure.

[0031] Figure 12 Flowchart of a BWA weight transfer process for LD pictures consistent with an embodiment of the present disclosure.

[0032] Figure 13 is a flowchart of a BWA weight transfer process for non-LD pictures consistent with an embodiment of the present disclosure.

[0033] Figure 14 is a block diagram of a video processing device consistent with an embodiment of the present disclosure. DETAILED DESCRIPTION

[0034] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, in which, unless otherwise indicated, like reference numerals in different figures represent like or similar elements. The implementations set forth in the following description of the exemplary embodiments are not intended to represent all implementations consistent with the present invention. Rather, they are merely examples of apparatus and methods consistent with aspects related to the present invention as described in the appended claims.

[0035] Figure 1 is a block diagram illustrating an example video encoding and decoding system 100 that can utilize techniques consistent with various video coding standards such as HEVC / H.265 and WC / H.266. Figure 1As shown, system 100 includes a source device 120 that provides encoded video data for later decoding by a destination device 140. Consistent with the disclosed embodiments, each of source device 120 and destination device 140 may include any of a wide variety of devices, including a desktop computer, a notebook (e.g., laptop) computer, a tablet computer, a set-top box, a mobile phone, a television, a camera, a wearable device (e.g., a smartwatch or wearable camera), a display device, a digital media player, a video game console, a video streaming device, etc. Source device 120 and destination device 140 may be equipped for wireless or wired communication.

[0036] refer to Figure 1 , source device 120 may include a video source 122, a video encoder 124, and an output interface 126. Destination device 140 may include an input interface 142, a video decoder 144, and a display device 146. In other examples, the source device and the destination device may include other components or arrangements. For example, source device 120 may receive video data from an external video source (not shown), such as an external camera. Similarly, destination device 140 may interface with an external display device, rather than including an integrated display device.

[0037] Although the disclosed techniques are explained in the following description as being performed by a video encoding device, the techniques may also be performed by a video encoder / decoder, commonly referred to as a "CODEC." Furthermore, the techniques of the present disclosure may also be performed by a video preprocessor. Source device 120 and destination device 140 are merely examples of such encoding devices, where source device 120 generates encoded video data for transmission to destination device 140. In some examples, source device 120 and destination device 140 may operate in a generally symmetrical manner, such that each of source device 120 and destination device 140 includes video encoding and decoding components. Thus, system 100 may support one-way or two-way video transmission between source device 120 and destination device 140, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0038] The video source 122 of the source device 120 can include a video capture device (such as a camera), a video archive containing previously captured video, or a video feed interface for receiving video from a video content provider. As another alternative, the video source 122 can generate computer graphics-based data as the source video, or a combination of real-time video, archived video, and computer-generated video. The captured, pre-captured, or computer-generated video can be encoded by a video encoder 124. The encoded video information can then be output to a communication medium 160 by an output interface 126.

[0039] Output interface 126 may include any type of medium or device capable of transmitting encoded video data from source device 120 to destination device 140. For example, output interface 126 may include a transmitter or transceiver configured to transmit encoded video data directly from source device 120 to destination device 140 in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and transmitted to destination device 140.

[0040] The communication medium 160 may include a transient medium such as a wireless broadcast or a wired network transmission. For example, the communication medium 160 may include a radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). The communication medium 160 may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. In some embodiments, the communication medium 160 may include a router, a switch, a base station, or any other device that may be useful for facilitating communication from the source device 120 to the destination device 140. For example, a network server (not shown) may receive encoded video data from the source device 120 and provide the encoded video data to the destination device 140, for example, via network transmission.

[0041] Communication medium 160 may also be in the form of a storage medium (e.g., a non-transitory storage medium) such as a hard drive, a flash drive, an optical disc, a digital video disc, a Blu-ray disc, a volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In some embodiments, a computing device at a media production facility, such as a disc imprinting facility, may receive the encoded video data from source device 120 and produce a disc containing the encoded video data.

[0042] Input interface 142 of destination device 140 receives information from communication medium 160. The received information may include syntax information, which includes syntax elements that describe the characteristics or processing of blocks and other coding units. The syntax information is defined by video encoder 124 and used by video decoder 144. Display device 146 displays the decoded video data to a user and may include any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0043] In another example, the encoded video generated by source device 120 can be stored on a file server or storage device. Input interface 142 can access the stored video data from the file server or storage device via streaming or downloading. The file server or storage device can be any type of computing device capable of storing encoded video data and transmitting the encoded video data to destination device 140. Examples of file servers include a web server supporting a website, a file transfer protocol (FTP) server, a network attached storage (NAS) device, or a local disk drive. The transmission of the encoded video data from the storage device can be streaming, downloading, or a combination thereof.

[0044] The video encoder 124 and the video decoder 144 can each be implemented as any of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the techniques are partially implemented in software, the device can store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the techniques of the present disclosure. Each of the video encoder 124 and the video decoder 144 can be included in one or more encoders or decoders, either of which can be integrated as part of a combined encoder / decoder (CODEC) in the corresponding device.

[0045] The video encoder 124 and the video decoder 144 may operate according to any video coding standard, such as the Versatile Video Coding (VVC / H.266) standard, the High Efficiency Video Coding (HEVC / H.265) standard, the ITU-T H.264 (also known as MPEG-4) standard, etc. Figure 1 Not shown, but in some embodiments, the video encoder 124 and the video decoder 144 may each be integrated with an audio encoder and decoder, and may include appropriate MUX-DEMUX units or other hardware and software to handle the encoding of audio and video in a common data stream or in separate data streams.

[0046] Figure 2 is a diagram illustrating an exemplary video encoder 200 consistent with the disclosed embodiments. For example, the video encoder 200 can be used as the system 100 ( Figure 1) in the video encoder 124. The video encoder 200 may perform intra-frame or inter-frame coding of blocks within video frames (including video blocks or partitions or subpartitions of video blocks). Intra-frame coding may rely on spatial prediction to reduce or remove spatial redundancy in video within a given video frame. Inter-frame coding may rely on temporal prediction to reduce or remove temporal redundancy in video within adjacent frames of a video sequence. Intra-frame mode may refer to various spatial-based compression modes, and inter-frame mode (such as unidirectional prediction or bidirectional prediction) may refer to various temporal-based compression modes.

[0047] refer to Figure 2 , the input video signal 202 can be processed block by block. For example, a video block unit can be a 16×16 pixel block (e.g., a macroblock (MB)). In HEVC, extended block sizes (e.g., coding units (CUs)) can be used to compress video signals with resolutions of, for example, 1080p and higher. In HEVC, a CU can include up to 64×64 luma samples and corresponding chroma samples. In VVC, the size of the CU can be further increased to include 128×128 luma samples and corresponding chroma samples. The CU can be partitioned into prediction units (PUs), to which a separate prediction method can be applied. Individual input video blocks (e.g., MB, CU, PU, etc.) can be processed using a spatial prediction unit 260 or a temporal prediction unit 262.

[0048] The spatial prediction unit 260 performs spatial prediction (e.g., intra prediction) on the current CU using information about the same picture / slice containing the current CU. Spatial prediction can use pixels from already coded neighboring blocks in the same video picture / slice to predict the current video block. Spatial prediction can reduce spatial redundancy inherent in video signals. Temporal prediction (e.g., inter-frame prediction or motion-compensated prediction) can use samples from already coded video pictures to predict the current video block. Temporal prediction can reduce temporal redundancy inherent in video signals.

[0049] The temporal prediction unit 262 performs temporal prediction (e.g., inter-frame prediction) on the current CU using information from a different picture / slice than the one containing the current CU. The temporal prediction for a video block may be conveyed via one or more motion vectors. A motion vector may indicate the amount and direction of motion between the current block and one or more of its prediction blocks in a reference frame. If multiple reference pictures are supported, one or more reference picture indexes may be sent for the video block. One or more reference indexes may be used to identify which reference pictures in the reference picture store or decoded picture buffer (DPB) 264 the temporal prediction signal may come from. After spatial or temporal prediction, a mode decision in the encoder and the encoder control unit 280 may select a prediction mode, for example, based on a rate-distortion optimization approach. The prediction block may be subtracted from the current video block at the adder 216. The prediction residual may be transformed by the transform unit 204 and quantized by the quantization unit 206. The quantized residual coefficients may be inverse quantized at inverse quantization unit 210 and inverse transformed at inverse transform unit 212 to form a reconstructed residual. The reconstructed block may be added to the prediction block at adder 226 to form a reconstructed video block. In-loop filtering, such as a deblocking filter and an adaptive loop filter 266, may be applied to the reconstructed video block before it is placed in reference picture store 264 and used to encode future video blocks. To form the output video bitstream 220, the coding mode (e.g., inter or intra), prediction mode information, motion information, and quantized residual coefficients may be sent to entropy coding unit 208 for compression and packing to form the bitstream 220. The systems, methods, and apparatuses described herein may be implemented, at least in part, within temporal prediction unit 262.

[0050] Figure 3 3 is a schematic diagram illustrating a video decoder 300 consistent with the disclosed embodiments. For example, the video decoder 300 can be used as the system 100 ( Figure 1 ) in the video decoder 144. Figure 3 The video bitstream 302 may be unpacked or entropy decoded at the entropy decoding unit 308. The coding mode or prediction information may be sent to the spatial prediction unit 360 (e.g., if intra-coded) or the temporal prediction unit 362 (e.g., if inter-coded) to form a prediction block. If inter-coded, the prediction information may include a prediction block size, one or more motion vectors (e.g., which may indicate a direction and amount of motion), or one or more reference indexes (e.g., which may indicate from which reference picture the prediction signal is to be obtained).

[0051] Motion compensated prediction may be applied by a temporal prediction unit 362 to form a temporal prediction block. The residual transform coefficients may be sent to an inverse quantization unit 310 and an inverse transform unit 312 to reconstruct the residual block. The prediction block and the residual block may be added together at 326. The reconstructed block may be in-loop filtered (via a loop filter 366) before being stored in a reference picture store 364. The reconstructed video in the reference picture store 364 may be used to drive a display device or to predict future video blocks. The decoded video 320 may be displayed on a display.

[0052] Consistent with the disclosed embodiments, the video encoder and video decoder can use various video encoding / decoding tools to process, compress, and decompress video data. Three tools are described below: weighted prediction (WP), bidirectional prediction with weighted averaging (BWA), and history-based motion vector prediction (HMVP).

[0053] Weighted Prediction (WP) is used to provide significantly better temporal prediction when there is fade in the video sequence. Fading refers to the phenomenon when the average illumination level of a picture in the video content exhibits significant changes in the temporal domain, such as fading to white or fading to black. Content creators often use fade to create desired special effects and express their artistic vision. Fading causes the average illumination level of the reference picture and the average illumination level of the current picture to differ significantly, making it more difficult to obtain an accurate prediction signal from temporally neighboring pictures. As part of the effort to address this problem, WP can provide powerful tools to adjust the illumination level of the prediction signal obtained from the reference picture and match it to the illumination level of the current picture, thereby significantly improving the temporal prediction accuracy.

[0054] Consistent with the disclosed embodiments, parameters for WP (or "WP parameters") are transmitted for each of the reference pictures used to encode the current picture. For each reference picture, the WP parameter includes a pair of weights and offsets (w, o), which can be transmitted for each color component of the reference picture. Figure 4 Table 400 depicts syntax elements for WP according to the disclosed embodiment. Referring to Table 400, the pred_weight_table() syntax is transmitted as part of the slice header. For the i-th reference picture in the reference picture list Lx (x can be 0 or 1), the flags luma_weight_lx_flag[i] and chroma_weight_lx_flag[i] are transmitted to indicate whether weighted prediction is applied to the luma and chroma components, respectively, of the i-th reference picture.

[0055] Without loss of generality, the following description uses luma as an example to illustrate the transmission of WP parameters. Specifically, if the flag luma_weight_lx_flag[i] is 1 for the i-th reference picture in the reference picture list Lx, the WP parameters (w[i], o[i]) are transmitted for the luma component. Then, when temporal prediction is applied using a given reference picture, the following equation (1) applies:

[0056] WP(x, y)=w·P(x, y)+o Equation (1),

[0057] Where: WP(x, y) is the weighted prediction signal at the sample position (x, y); (w, o) is the WP parameter pair associated with the reference picture; and P(x, y) = ref(x-mvx, y-mvy) is the prediction before applying WP, (mvx, mvy) is the motion vector associated with the reference picture, and ref(x, y) is the reference signal at the position (x, y). If the motion vector (mvx, mvy) has fractional sample precision, interpolation can be applied, such as using an 8-tap luma interpolation filter in HEVC.

[0058] Bidirectional prediction with weighted averaging (BWA) is a tool that can be used to improve temporal prediction accuracy, thereby improving the compression performance of video encoders. It is used in various video coding standards such as H.264 / AVC, HEVC, and VVC. Figure 5 is a schematic diagram illustrating an exemplary bidirectional prediction. Figure 5 , CU 503 is part of the current picture. CU 501 is from reference picture 511, and CU 502 is from reference picture 512. In some embodiments, reference pictures 511 and 512 can be selected from two different reference picture lists, L0 and L1, respectively. Two motion vectors (mvx0, mvy0) and (mvx1, mvy1) can be generated with reference to CU 501 and CU 502, respectively. These two motion vectors form two prediction signals, which can be averaged to obtain a bidirectional prediction signal, i.e., a prediction corresponding to CU 503.

[0059] In the disclosed embodiment, reference pictures 511 and 512 may be from the same or different picture sources. Figure 5 The reference pictures 511 and 512 are depicted as two different physical reference pictures corresponding to different time points, but in some embodiments, the reference pictures 511 and 512 may be the same physical reference picture, because the same physical reference picture is allowed to appear one or more times in one or both of the reference picture lists L0 and L1. Figure 5The reference pictures 511 and 512 are depicted as being from the past and the future, respectively, in the time domain, but in some embodiments, the reference pictures 511 and 512 are allowed to be both from the past or both from the future relative to the current picture.

[0060] Specifically, refer to Figure 5 , bidirectional prediction can be performed based on the following equation:

[0061]

[0062] Wherein: (mvx0, mvy0) is the motion vector associated with a reference picture selected from reference picture list L0 (e.g., reference picture 511); (mvx1, mvy1) is the motion vector associated with a reference picture selected from reference picture list L1 (e.g., reference picture 512); and ref0(x, y) is the reference signal at position (x, y) in reference picture 511; and ref1(x, y) is the reference signal at position (x, y) in reference picture 512.

[0063] Still refer to Figure 5 , weighted prediction can be applied to bidirectional prediction. In some embodiments, an equal weight of 0.5 is assigned to each prediction signal so that the prediction signals are averaged based on the following equation:

[0064]

[0065] Here, (w0, o0) and (w1, o1) are WP parameters associated with reference pictures 511 and 512, respectively.

[0066] In some embodiments, BWA is used to apply unequal weights with weighted averaging to bidirectional prediction, which can improve coding efficiency. BWA can be applied adaptively at the block level. For each CU, if certain conditions are met, the weight index gbi_idx is transmitted. Figure 6 Table 600 depicts syntax elements for BWA according to the disclosed embodiment. Referring to Table 600, the syntax at 601 and 602 can be used to perform bidirectional prediction on a CU containing, for example, at least 256 luma samples. Based on the value of gbi_idx, a weight w is determined and applied to the reference signal according to the following equation:

[0067] P(x, y)=(1-w)·P0(x, y)+w·P1(x, y)

[0068] =(1-w)·ref0(x-mvx0,y-mvy0)+w·ref1(x-mvx1,y-mvy1)

[0069] Equation (4).

[0070] In some embodiments, the value of the BWA weight w may be selected from five possible values, for example, A low-delay (LD) picture is defined as one whose reference pictures precede it in display order. For LD pictures, all five values above can be used for BWA weights. That is, when transmitting BWA weights, the value of the weight index gbi_idx is in the range [0, 4], with the center value (gbi_idx=2) corresponding to equal weights. For non-low-latency (non-LD) pictures, only 3 BWA weights are used. In this case, the values of the weight index gbi_idx are in the range [0, 2], with the center value (gbi_idx=1) corresponding to equal weights. The value of .

[0071] If explicit transmission of the weight index gbi_idx is used, the value of the BWA weight of the current CU is selected by the encoder, for example, by rate-distortion optimization. One approach is to try all allowed weight values w and select the one with the lowest rate-distortion cost. However, an exhaustive search for the best combination of weights and motion vectors may significantly increase encoding time. Therefore, a fast encoding method can be applied to reduce encoding time without reducing encoding efficiency.

[0072] For each bidirectionally predicted CU, the BWA weight w may be determined and transmitted in one of the following two ways: 1) For a non-merged CU, the weight index is transmitted after the motion vector difference, as shown in Table 600 ( Figure 6 ) as shown; and 2) for the merged CU, the weight index gbi_idx is inferred from the neighboring blocks based on the merge candidate index. The merge mode will be described in detail below.

[0073] The merge candidates for a CU may come from neighboring blocks of the current CU, or from collocated blocks in a temporally collocated picture of the current CU. Figure 7 is a diagram illustrating spatial neighbors used in merging candidate list construction according to an exemplary embodiment. Figure 7 The following figure depicts the locations of five example spatial candidates for motion information. To construct a list of merge candidates, the five spatial candidates can be checked and added to the list, for example, in the order A1, B1, B0, A0, and A2. If the block at a spatial location is intra-coded or outside the boundaries of the current slice, the block can be considered unavailable. Redundant entries (e.g., locations where candidates have the same motion information) can be excluded from the merge candidate list.

[0074] Merge mode has been supported since the HEVC standard. Merge mode is an effective way to reduce motion transmission overhead. Instead of explicitly transmitting the motion information of the current CU (prediction mode, motion vector, reference index, etc.), the motion information of the neighboring blocks from the current CU is used to construct a merge candidate list. Both spatial and temporal neighboring blocks can be used to construct the merge candidate list. After constructing the merge candidate list, an index is transmitted to indicate which of the merge candidates is used to encode the current CU. The motion information from the merge candidate is then used to predict the current CU.

[0075] When BWA is enabled, if the current CU is in merge mode, the motion information it inherits from its merge candidate may include not only the motion vector and reference index, but also the weight index gbi_idx of the merge candidate. In other words, when performing motion compensated prediction, a weighted average is performed on the two prediction signals for the current CU according to the weight index gbi_idx of the neighboring block of the current CU. In some embodiments, if the merge candidate is a spatial neighbor, the weight index gbi_idx is only inherited from the merge candidate, and if the candidate is a temporal neighbor, the weight index gbi_idx is not inherited.

[0076] The merge mode in HEVC uses spatially adjacent blocks and temporally adjacent blocks to construct a list of merge candidates. Figure 7 In the example shown, all spatially adjacent blocks are adjacent to (i.e., connected to) the current CU. However, in some embodiments, non-adjacent neighbors can be used in merge mode to further improve the coding efficiency of merge mode. The merge mode using non-adjacent neighbors is called extended merge mode. In some embodiments, the history-based motion vector prediction (HMVP) method in VVC can be used for inter-frame coding in extended merge mode to improve compression performance with minimal implementation cost. In HMVP, a table of HMVP candidates is continuously maintained and updated during the video encoding / decoding process. The HMVP table can include up to six entries. The HMVP candidate is inserted in the middle of the merge candidate list of the spatial neighbor and can be selected as other merge candidates using the merge candidate index to encode the current CU.

[0077] Entries are removed from and added to the table using a first-in, first-out (FIFO) rule. After decoding a non-affine inter-coded block, the table is updated by adding the associated motion information as a new HMVP candidate to the last entry in the table and removing the oldest HMVP candidate in the table. The table is cleared when a new slice is encountered. In some embodiments, the table may be cleared more frequently, for example, when a new coding tree unit (CTU) is encountered or when a new CTU row is encountered.

[0078] The above description of WP, BWA, and HMVP shows that these tools need to be coordinated during video encoding and transmission. For example, both the BWA tool and the WP tool introduce weighting factors into the inter-frame prediction process to improve the prediction accuracy of motion compensation. However, the function of the BWA tool is different from that of the WP tool. According to equation (4), BWA applies weights in a normalized manner. That is, the weights applied to the L0 prediction and the L1 prediction are (1-w) and w, respectively. Because the weights add up to 1, BWA defines how the two prediction signals are combined together without changing the total energy of the bidirectional prediction signal. On the other hand, according to equation (3), WP does not have a normalization constraint. That is, w0 and w1 do not need to add up to 1. In addition, WP can add constant offsets o0 and o1 according to equation (3). In addition, BWA and WP are applicable to different types of video content. However, WP effectively attenuates video sequences (or other video content with global illumination changes in the time domain) and does not improve the coding efficiency of normal sequences when the illumination level does not change in the time domain. In contrast, BWA is a block-level adaptive tool that adaptively selects how to combine two prediction signals. While BWA is effective on normal sequences without illumination variations, it is far less effective than the WP method on decaying sequences. For these reasons, in some implementations, BWA and WP tools may be supported in video coding standards, but they operate in a mutually exclusive manner. Therefore, a mechanism is needed to disable one tool when the other is present.

[0079] In addition, as discussed above, if the selected merge candidate is a spatial neighbor of the current CU, the BWA tool can be combined with the merge mode by allowing the weight index gbi_idx from the selected merge candidate to be inherited. In order to take advantage of HMVP, it is necessary to combine BWA with HMVP to use a method of non-adjacent neighbors in extended merge mode.

[0080] At least some of the disclosed embodiments provide a solution to maintain the exclusivity of WP and BWA.A combination of syntax in the Picture Parameter Set (PPS) and the slice header is used to indicate whether WP is enabled for a picture. Figure 8 Table 800 is a syntax element for transmitting the activation or deactivation of WP at the picture level consistent with an embodiment of the present disclosure. As shown in 801 in Table 800, weighted_pred_flag and weighted_bipred_flag are sent in the PPS to indicate whether WP is enabled for unidirectional prediction and bidirectional prediction, respectively, according to the slice type of the slice referencing the PPS. Figure 9Table 900 is a syntax element for transmitting the activation or deactivation of WP at the slice level consistent with an embodiment of the present disclosure. As shown in 901 in Table 900, at the slice / picture level, if the PPS referenced by the slice (determined by matching the slice_pic_parameter_set_id of the slice header with the pps_pic_parameter_set_id of the PPS) has WP enabled, then Table 400 ( Figure 4 ) is sent to the decoder to indicate the WP parameters of each reference picture in the reference pictures of the current picture.

[0081] Based on such transmission, in some embodiments of the present disclosure, additional conditions may be added to the transmission of the CU-level weight index gbi_idx: Additional conditional transmission: If WP is enabled for the picture containing the current CU, weighted averaging is disabled for the bidirectional prediction mode of the current CU. Figure 10 Table 1000 is a syntax element for maintaining exclusivity of WP and BWA at the CU level consistent with embodiments of the present disclosure. Referring to Table 1000, a condition 1001 may be added to indicate that if the PPS referenced by the current slice allows WP for bidirectional prediction, then BWA is completely disabled for all CUs in the current slice. This ensures that WP and BWA are exclusive.

[0082] However, the above method can completely disable BWA for all CUs in the current slice, regardless of whether the current CU uses a reference picture with WP enabled. This may reduce coding efficiency. At the CU level, whether WP is enabled for its reference picture can be determined by the values of luma_weight_l0_flag[ref_idx_l0], / chroma_weight l0_flag[ref_idx_l0], luma_weight_l1_flag[ref_idx_l1], and luma / chroma_weight_l1_flag[ref_idx_l1], where ref_idx_l0 and ref_idx_l1 are the reference picture indices of the current CU in L0 and L1, respectively. For the L0 and L1 reference pictures of the current slice, luma / chroma_weight_l0_flag and luma / chroma_weight_l1_flag are transmitted in pred_weight_table(), as shown in Table 400( Figure 4 ) as shown. Figure 11Table 1100 is a syntax element for maintaining the exclusivity of WP and BWA at the CU level consistent with an embodiment of the present disclosure. Referring to Table 1100, a condition 1101 is added to control the exclusivity of WP and BWA at the CU level, regardless of whether the weight index gbi_idx is transmitted. When the weight index gbi_idx is not transmitted, it can be inferred that it is a default value indicating the equal weight case (i.e., 1 or 2, depending on whether 3 or 5 BWA weights are allowed).

[0083] The methods illustrated in Table 1000 and Table 1100 both add conditions to the transmission of the weight index gbi_idx at the CU level, which may complicate the parsing process of the decoder. Therefore, in the third embodiment, the transmission conditions of the weight index gbi_idx remain the same as those in Table 600 ( Figure 6 ) in the same transmission conditions. If WP is enabled for the luma or chroma components of the L0 or LI reference pictures, the default value of the weight index gbi_idx of the current CU is always sent, which becomes a bitstream consistency constraint for the encoder. That is, the weight index gbi_idx values corresponding to unequal weights can be sent only when WP is not enabled for the luma or chroma components of the L0 and LI reference pictures. Although this transmission is redundant, the actual bit cost of this redundant transmission is negligible because the context-adaptive binary arithmetic coding (CABAC) engine in the entropy coding stage can adapt to the statistical information of the weight index gbi_idx value. In addition, this simplifies the parsing process.

[0084] In the decoder (e.g. Figure 3 After the encoder 300 in FIG. 1 receives a bitstream including the above-described syntax for maintaining the exclusivity of WP and BWA, the decoder may parse the bitstream and determine whether to disable BWA based on the syntax.

[0085] At least some of the embodiments of the present disclosure may provide a solution for symmetrical transmission of BWA at the CU level. As discussed above, in some embodiments, the CU-level weight used in BWA is transmitted as a weight index gbi_idx, and for low-latency (LD) pictures, the value of gbi_idx is in the range of [0, 4], and for non-LD pictures, the value of gbi_idx is in the range of [0, 2]. However, this will produce inconsistencies between LD pictures and non-LD pictures, as shown below:

[0086]

[0087] Here, the same BWA weight value is represented by different gbi_idx values in LD and non-LD pictures.

[0088] To improve transmission consistency, according to some disclosed embodiments, the transmission of the weight index gbi_idx may be modified to first indicate whether the BWA weights are equal weighted, followed by an index or flag for unequal weights. Figure 12 and Figure 13 Flowcharts illustrating exemplary BWA weight transfer processes for LD and non-LD pictures, respectively. For LD pictures that allow 5 BWA weight values, use Figure 12 The transmission flow chart in , and for non-LD pictures that allow 3 BWA weight values, use Figure 13 The first flag, gbi_ew_flag, indicates whether equal weighting is applied in BWA. If gbi_ew_flag is 1, no additional transmission is required because equal weighting (w=1 / 2) is applied; otherwise, a flag (1 bit for 2 values) or an index (2 bits for 4 values) is transmitted to indicate which unequal weighting is applied. Figure 12 and Figure 13 This is only one example of a possible mapping relationship between BWA weight values and index / flag values. It is conceivable that other mappings between weight values and index / flag values can be used. Another benefit of splitting the weight index gbi_idx into two syntax elements (gbi_ew_flag and gbi_uew_val_idx (or gbi_uew_val_flag)) is that these values can be encoded using separate CABAC contexts. In addition, for LD pictures, when using a 2-bit value gbi_uew_val_idx, a separate CABAC context can be used to encode the first and second bits.

[0089] In the decoder (e.g. Figure 3 After the encoder 300 in FIG. 1 receives the above-mentioned transmission of BWA at the CU level, the decoder can parse the transmission and determine whether BWA uses equal weights based on the transmission. If BWA is determined to be unequal weights, the decoder can further determine the value of the unequal weights based on the transmission.

[0090] Some embodiments of the present disclosure provide a solution that combines BWA and HMVP. If the motion information stored in the HMVP table only includes the motion vector, reference index, and prediction mode (e.g., unidirectional prediction vs. bidirectional prediction) of the merge candidate, the merge candidate cannot be used with BWA because the BWA weights are not stored or updated in the HMVP table. Therefore, according to some disclosed embodiments, the BWA weights are included as part of the motion information stored in the HMVP table. When the HMVP table is updated, the BWA weights are also updated along with other motion information (such as motion vector, reference index, and prediction mode).

[0091] In addition, partial pruning can be applied to avoid having too many identical candidates in the merge candidate list. Identical candidates are defined as candidates that have the same motion information as at least one of the existing merge candidates in the merge candidate list. Identical candidates take up space in the merge candidate list but do not provide any additional motion information. Partial pruning can detect some of these situations and can prevent some of these identical candidates from being added to the merge candidate list. By including the BWA weights in the HMVP table, the pruning process also takes the BWA weights into account when determining whether two merge candidates are the same. Specifically, if a new candidate has the same motion vector, reference index, and prediction mode as another candidate in the merge candidate list, but has a different BWA weight than the other candidate, then the new candidate can be considered to be different and may not be pruned.

[0092] In the decoder (e.g. Figure 3 After the encoder 300 in FIG. 1 receives a bitstream including the above-mentioned HMVP table, the decoder may parse the bitstream and determine the BWA weights of the merge candidates included in the HMVP table.

[0093] Figure 14 is a block diagram of a video processing apparatus 1400 consistent with an embodiment of the present disclosure. For example, the apparatus 1400 may implement the above-described video encoder (e.g., Figure 2 ) or a video decoder (e.g., Figure 3 In the disclosed embodiment, the apparatus 1400 may be configured to perform the above-described method for encoding and transmitting BWA weights. Figure 14 , device 1400 may include a processing component 1402, a memory 1404, and an input / output (I / O) interface 1406. Device 1400 may also include one or more of a power supply component and a multimedia component (not shown), or any other suitable hardware or software components.

[0094] The processing component 1402 can control the overall operation of the device 1400. For example, the processing component 1402 can include one or more processors that execute instructions to perform the above-described method for encoding and transmitting BWA weights. In addition, the processing component 1402 can include one or more modules that facilitate interaction between the processing component 1402 and other components. For example, the processing component 1402 can include an I / O module to facilitate interaction between an I / O interface and the processing component 1402.

[0095] The memory 1404 is configured to store various types of data or instructions to support the operation of the device 1400. The memory 1404 may include a non-transitory computer-readable storage medium including instructions for an application or method operating on the device 1400, which instructions may be executed by one or more processors of the device 1400. Common forms of non-transitory media include, for example, a floppy disk, a foldable disk, a hard disk, a solid-state drive, a magnetic tape or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROM and EPROM, cloud storage, FLASH-EPROM or any other flash memory, NVRAM, a cache, a register, any other memory chip or tape cartridge, and network versions thereof.

[0096] The I / O interface 1406 provides an interface between the processing component 1402 and peripheral interface modules (such as a camera or display). The I / O interface 1406 can use communication protocols / methods such as audio, analog, digital, serial bus, universal serial bus (USB), infrared, PS / 2, BNC, coaxial, RF antenna, Bluetooth, etc. The I / O interface 1406 can also be configured to facilitate wired or wireless communication between the apparatus 1400 and other devices (such as devices connected to the Internet). The apparatus can access wireless networks based on one or more communication standards (such as WiFi, LTE, 2G, 3G, 4G, 5G, etc.).

[0097] The following terms may be used to further describe the implementation:

[0098] 1. A computer-implemented signal transmission method, comprising the following steps:

[0099] The processor transmits a bitstream to a video decoder, the bitstream including weight information for predicting a coding unit, the weight information indicating:

[0100] If weighted prediction is enabled for the bi-directional prediction mode of the coding unit, weighted averaging for the bi-directional prediction mode is disabled.

[0101] 2. The signal transmission method according to clause 1, wherein the weight information indicates:

[0102] If weighted prediction is enabled for bidirectional prediction of a picture including the coding unit, weighted averaging for the bidirectional prediction mode is disabled.

[0103] 3. The signal transmission method according to clause 2, wherein the bitstream includes a flag indicating whether weighted prediction is enabled for bidirectional prediction of the picture including the coding unit.

[0104] 4. The signal transmission method according to clause 1, wherein the weight information indicates:

[0105] If weighted prediction is enabled for at least one of the luma component and the chroma component of the reference picture of the coding unit, weighted averaging for the bi-directional prediction mode is disabled.

[0106] 5. The signal transmission method according to clause 4, wherein the bitstream includes a flag indicating whether weighted prediction is enabled for at least one of the luma component and the chroma component of the reference picture.

[0107] 6. The signal transmission method according to clause 1, wherein the weight information includes a value of a bidirectional prediction weight associated with the coding unit, the signal transmission method further comprising the following steps:

[0108] If weighted prediction is enabled for at least one of the luma component and the chroma component of the reference picture of the coding unit, the value of the bidirectional prediction weight is set to a default value.

[0109] 7. The signal transmission method according to clause 6, wherein the default value corresponds to equal weight.

[0110] 8. A computer-implemented video encoding method, the video encoding method comprising the following steps:

[0111] The processor constructs a merge candidate list for a coding unit, the merge candidate list including motion information of non-adjacent inter-frame coding blocks of the coding unit, the motion information including bidirectional prediction weights associated with the non-adjacent inter-frame coding blocks; and

[0112] The processor performs encoding based on the motion information.

[0113] 9. The video encoding method according to clause 8, further comprising the following steps:

[0114] The processor determines a new non-adjacent inter-frame coding block of the coding unit: and

[0115] The processor updates the merge candidate list by inserting the bidirectional prediction weight associated with the new non-adjacent inter-coded block into the merge candidate list.

[0116] 10. The video encoding method of clause 8, wherein:

[0117] The processor determines motion information of a new non-adjacent inter-coded block of the coding unit, the motion information comprising bidirectional prediction weights associated with the non-adjacent inter-coded block:

[0118] The processor compares the motion information of the new non-adjacent inter-coded block with motion information of each inter-coded block included in the merge candidate list;

[0119] In response to a comparison result that none of the inter-frame coding blocks included in the merge candidate list has the same motion information as the motion information of the new non-adjacent inter-frame coding block, the processor adds the new non-adjacent inter-frame coding block to the merge candidate list; or

[0120] In response to a comparison result that motion information of the new non-adjacent inter coding block is identical to motion information of at least one inter coding block included in the merge candidate list, determining that the new non-adjacent inter coding block is redundant with respect to the merge candidate list.

[0121] 11. The video coding method of clause 8, wherein the non-adjacent inter-frame coded blocks are from a history-based motion vector prediction (HMVP) table.

[0122] 12. A video encoding method according to clause 8, wherein the non-adjacent inter-frame coded blocks are located in a video frame including the coding unit and are spatial non-adjacent neighbors of the coding unit.

[0123] 13. The video encoding method of clause 8, wherein the motion information of the non-adjacent inter-coded blocks further comprises:

[0124] a reference index associated with said non-adjacent inter-coded block,

[0125] the motion vector associated with the non-affine inter-coded block, and

[0126] At least one of an indicator of a uni-directional prediction mode and an indicator of a bi-directional prediction mode.

[0127] 14. The video encoding method according to clause 8, wherein the step of encoding based on the motion information comprises:

[0128] Selecting the non-adjacent inter-frame coding blocks from the merge candidate list for encoding the coding unit; and

[0129] An index indicating that the non-adjacent inter-coded blocks are selected for encoding the coding unit is transmitted to a decoder.

[0130] 15. A computer-implemented signal transmission method, comprising the following steps:

[0131] The processor determines a value of a bidirectional prediction weight for a coding unit of a video frame;

[0132] The processor determines whether the bidirectional prediction weights are equal weights; and

[0133] In response to the determination, the processor transmits the following to the video decoder:

[0134] When the bidirectional prediction weights are equal weights, transmitting a bitstream including a first syntax element indicating the equal weights, or

[0135] After determining that the bidirectional prediction weights are unequal weights, a bitstream is transmitted including a second syntax element indicating values of the bidirectional prediction weights corresponding to the unequal weights.

[0136] 16. The signal transmission method according to clause 15, wherein the first syntax element is a flag having one bit.

[0137] 17. The signal transmission method according to clause 15, wherein the step of transmitting the bitstream including the second syntax element further comprises:

[0138] determining a number of unequal weights that can be used by the coding units; and

[0139] In response to determining that there are more than two unequal weights that can be used by the coding unit, the processor transmits the second syntax element to the video decoder as an index having at least two bits; or

[0140] In response to determining that the unequal weights usable by the coding unit are one or two, the processor transmits the second syntax element as a one-bit flag to the video decoder.

[0141] 18. The signal transmission method according to clause 17, further comprising the following steps:

[0142] The processor encodes each bit of the second syntax element using a different context-adaptive binary arithmetic coding (CABAC) context.

[0143] 19. The signal transmission method according to clause 15, further comprising the following steps:

[0144] determining whether the coding unit is part of a low-latency picture;

[0145] Determine the following:

[0146] In response to determining that the coding unit is part of the low-delay picture, determining that the coding unit uses more than two unequal weights, or

[0147] In response to determining that the coding unit is not part of the low-delay picture, determining that the coding unit uses one or two unequal weights.

[0148] 20. The signal transmission method according to clause 19, wherein:

[0149] When the coding unit is part of a low-delay picture and the value of the bidirectional prediction weight is 3 / 8, determining that the value of the second syntax element is 0;

[0150] When the coding unit is part of a low-delay picture and the value of the bidirectional prediction weight is 5 / 8, determining that the value of the second syntax element is 1;

[0151] When the coding unit is not part of a low-delay picture and the value of the bidirectional prediction weight is 3 / 8, determining that the value of the second syntax element is 0; and

[0152] When the coding unit is not part of a low-delay picture and the value of the bidirectional prediction weight is 5 / 8, the value of the second syntax element is determined to be 1.

[0153] 21. The signal transmission method according to clause 15, further comprising the following steps:

[0154] The processor allocates different numbers of bits to the second syntax element for low-delay pictures and non-low-delay pictures, respectively.

[0155] 22. A signal transmission method according to clause 21, wherein the value of the second syntax element corresponds to the same bidirectional prediction weight value for the low-delay picture and the non-low-delay picture.

[0156] 23. A signal transmission method according to clause 15, wherein each of the first syntax element and the second syntax element corresponds to one or more pre-allocated bits in a coding unit level weight index.

[0157] 24. The signal transmission method according to clause 15, further comprising the following steps:

[0158] The processor encodes a value of the first syntax element and a value of the second syntax element using different context-adaptive binary arithmetic coding (CABAC) contexts.

[0159] 25. A device comprising:

[0160] a memory storing instructions; and

[0161] a processor configured to execute the instructions to cause the device to:

[0162] Transmitting a bitstream to a video encoder, the bitstream including weight information for predicting a coding unit, the weight information indicating:

[0163] If weighted prediction is enabled for the bi-directional prediction mode of the coding unit, weighted averaging for the bi-directional prediction mode is disabled.

[0164] 26. The apparatus of clause 25, wherein the weight information indicates:

[0165] If weighted prediction is enabled for bidirectional prediction of a picture including the coding unit, weighted averaging for the bidirectional prediction mode is disabled.

[0166] 27. The apparatus of clause 26, wherein the bitstream comprises a flag indicating whether weighted prediction is enabled for bidirectional prediction of the picture comprising the coding unit.

[0167] 28. The apparatus of clause 25, wherein the weight information indicates:

[0168] If weighted prediction is enabled for at least one of the luma component and the chroma component of the reference picture of the coding unit, weighted averaging for the bi-directional prediction mode is disabled.

[0169] 29. The apparatus of clause 28, wherein the bitstream comprises a flag indicating whether weighted prediction is enabled for at least one of the luma component and the chroma component of the reference picture.

[0170] 30. The apparatus of clause 25, wherein the weight information comprises a value of a bidirectional prediction weight associated with the coding unit, and the processor is further configured to execute the instructions to:

[0171] If weighted prediction is enabled for at least one of the luma component and the chroma component of the reference picture of the coding unit, the value of the bidirectional prediction weight is set to a default value.

[0172] 31. The apparatus of clause 30, wherein the default value corresponds to equal weight.

[0173] 32. A device comprising:

[0174] a memory storing instructions; and

[0175] a processor configured to execute the instructions to cause the device to:

[0176] Constructing a merge candidate list for a coding unit, the merge candidate list including motion information of non-adjacent inter-frame coding blocks of the coding unit, the motion information including bidirectional prediction weights associated with the non-adjacent inter-frame coding blocks; and

[0177] Encoding is performed based on the motion information.

[0178] 33. The apparatus of clause 32, wherein the processor is further configured to execute the instructions to:

[0179] determining a new non-adjacent inter-frame coding block of the coding unit: and

[0180] The merge candidate list is updated by inserting the bidirectional prediction weights associated with the new non-adjacent inter-coded block into the merge candidate list.

[0181] 34. The apparatus of clause 32, wherein the processor is further configured to execute the instructions to:

[0182] Determining motion information of a new non-adjacent inter-frame coded block of the coding unit, the motion information including bidirectional prediction weights associated with the non-adjacent inter-frame coded block:

[0183] comparing the motion information of the new non-adjacent inter-frame coding block with the motion information of each inter-frame coding block included in the merge candidate list;

[0184] In response to a comparison result that none of the inter-frame coding blocks included in the merge candidate list has the same motion information as the motion information of the new non-adjacent inter-frame coding block, adding the new non-adjacent inter-frame coding block to the merge candidate list; or

[0185] In response to a comparison result that motion information of the new non-adjacent inter coding block is identical to motion information of at least one inter coding block included in the merge candidate list, determining that the new non-adjacent inter coding block is redundant with respect to the merge candidate list.

[0186] 35. The apparatus of clause 32, wherein the non-adjacent inter-coded blocks are from a history-based motion vector prediction (HMVP) table.

[0187] 36. The apparatus of clause 32, wherein the non-adjacent inter-coded blocks are located in a video frame that includes the coding unit and are spatially non-adjacent neighbors of the coding unit.

[0188] 37. The apparatus of clause 32, wherein the motion information of the non-adjacent inter-coded blocks further comprises:

[0189] a reference index associated with said non-adjacent inter-coded block,

[0190] the motion vector associated with the non-affine inter-coded block, and

[0191] At least one of an indicator of a uni-directional prediction mode and an indicator of a bi-directional prediction mode.

[0192] 38. The apparatus of clause 32, wherein the processor is further configured to execute the instructions to:

[0193] Selecting the non-adjacent inter-frame coding blocks from the merge candidate list for encoding the coding unit; and

[0194] An index indicating that the non-adjacent inter-coded blocks are selected for encoding the coding unit is transmitted to a decoder.

[0195] 39. A device comprising:

[0196] a memory storing instructions; and

[0197] a processor configured to execute the instructions to cause the device to:

[0198] determining values of bidirectional prediction weights for coding units of a video frame;

[0199] determining whether the bidirectional prediction weights are equal weights; and

[0200] In response to the determination, the following is transmitted to the video decoder:

[0201] When the bidirectional prediction weights are equal weights, transmitting a bitstream including a first syntax element indicating the equal weights, or

[0202] After determining that the bidirectional prediction weights are unequal weights, a bitstream is transmitted including a second syntax element indicating values of the bidirectional prediction weights corresponding to the unequal weights.

[0203] 40. The apparatus of clause 39, wherein the first syntax element is a flag having one bit.

[0204] 41. The apparatus of clause 39, wherein the processor is further configured to execute the instructions to:

[0205] determining a number of unequal weights that can be used by the coding units; and

[0206] In response to determining that there are more than two unequal weights that can be used by the coding unit, transmitting the second syntax element as an index having at least two bits to the video decoder; and

[0207] In response to determining that the unequal weights that can be used by the coding unit are one or two, the second syntax element is transmitted to the video decoder as a flag having one bit.

[0208] 42. The apparatus of clause 41, wherein the processor is further configured to execute the instructions to:

[0209] Each bit of the second syntax element is encoded using a different context-adaptive binary arithmetic coding (CABAC) context.

[0210] 43. The apparatus of clause 39, wherein the processor is further configured to execute the instructions to:

[0211] determining whether the coding unit is part of a low-latency picture;

[0212] Determine the following:

[0213] In response to determining that the coding unit is part of the low-delay picture, determining that the coding unit uses more than two unequal weights, and

[0214] In response to determining that the coding unit is not part of the low-delay picture, determining that the coding unit uses one or two unequal weights.

[0215] 44. The apparatus of clause 43, wherein the processor is further configured to execute the instructions to:

[0216] When the coding unit is part of a low-delay picture and the value of the bidirectional prediction weight is 3 / 8, determining that the value of the second syntax element is 0;

[0217] When the coding unit is part of a low-delay picture and the value of the bidirectional prediction weight is 5 / 8, determining that the value of the second syntax element is 1;

[0218] When the coding unit is not part of a low-delay picture and the value of the bidirectional prediction weight is 3 / 8, determining that the value of the second syntax element is 0; and

[0219] When the coding unit is not part of a low-delay picture and the value of the bidirectional prediction weight is 5 / 8, the value of the second syntax element is determined to be 1.

[0220] 45. The apparatus of clause 39, wherein the processor is further configured to execute the instructions to:

[0221] Different numbers of bits are allocated to the second syntax element for low-delay pictures and non-low-delay pictures, respectively.

[0222] 46. An apparatus according to clause 45, wherein the value of the second syntax element corresponds to the same bidirectional prediction weight value for the low-delay picture and the non-low-delay picture.

[0223] 47. The apparatus of clause 39, wherein each of the first syntax element and the second syntax element corresponds to one or more pre-allocated bits in a coding-unit level weight index.

[0224] 48. The apparatus of clause 39, wherein the processor is further configured to execute the instructions to:

[0225] The value of the first syntax element and the value of the second syntax element are encoded using different context-adaptive binary arithmetic coding (CABAC) contexts.

[0226] 49. A non-transitory computer-readable medium having stored thereon a set of instructions, the set of instructions being executable by one or more processors of a device to cause the device to perform a method comprising the steps of:

[0227] Transmitting a bitstream to a video decoder, the bitstream including weight information for predicting a coding unit, the weight information indicating:

[0228] If weighted prediction is enabled for the bi-directional prediction mode of the coding unit, weighted averaging for the bi-directional prediction mode is disabled.

[0229] 50. The medium of clause 49, wherein the weight information indicates:

[0230] If weighted prediction is enabled for bidirectional prediction of a picture including the coding unit, weighted averaging for the bidirectional prediction mode is disabled.

[0231] 51. The medium of clause 50, wherein the bitstream includes a flag indicating whether weighted prediction is enabled for bidirectional prediction of the picture including the coding unit.

[0232] 52. The medium of clause 49, wherein the weight information indicates:

[0233] If weighted prediction is enabled for at least one of the luma component and the chroma component of the reference picture of the coding unit, weighted averaging for the bi-directional prediction mode is disabled.

[0234] 53. The medium of clause 52, wherein the bitstream includes a flag indicating whether weighted prediction is enabled for at least one of the luma component and the chroma component of the reference picture.

[0235] 54. The medium of clause 49, wherein the weight information comprises a value of a bidirectional prediction weight associated with the coding unit, and wherein the set of instructions is executable by the one or more processors of the device to cause the device to further perform:

[0236] If weighted prediction is enabled for at least one of the luma component and the chroma component of the reference picture of the coding unit, the value of the bidirectional prediction weight is set to a default value.

[0237] 55. The medium of clause 54, wherein the default value corresponds to equal weight.

[0238] 56. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by one or more processors of a device to cause the device to perform a method comprising the steps of:

[0239] Constructing a merge candidate list for a coding unit, the merge candidate list including motion information of non-adjacent inter-frame coding blocks of the coding unit, the motion information including bidirectional prediction weights associated with the non-adjacent inter-frame coding blocks; and

[0240] Encoding is performed based on the motion information.

[0241] 57. The medium of clause 56, wherein the method further comprises the steps of:

[0242] determining a new non-adjacent inter-frame coding block of the coding unit: and

[0243] The merge candidate list is updated by inserting the bidirectional prediction weights associated with the new non-adjacent inter-coded block into the merge candidate list.

[0244] 58. The medium of clause 56, wherein the set of instructions is executable by the one or more processors of the device to cause the device to further perform:

[0245] Determining motion information of a new non-adjacent inter-frame coded block of the coding unit, the motion information including bidirectional prediction weights associated with the non-adjacent inter-frame coded block:

[0246] comparing the motion information of the new non-adjacent inter-frame coding block with the motion information of each inter-frame coding block included in the merge candidate list;

[0247] In response to a comparison result that none of the inter-frame coding blocks included in the merge candidate list has the same motion information as the motion information of the new non-adjacent inter-frame coding block, adding the new non-adjacent inter-frame coding block to the merge candidate list; and

[0248] In response to a comparison result that motion information of the new non-adjacent inter coding block is identical to motion information of at least one inter coding block included in the merge candidate list, determining that the new non-adjacent inter coding block is redundant with respect to the merge candidate list.

[0249] 59. The medium of clause 56, wherein the non-adjacent inter-coded blocks are from a history-based motion vector prediction table.

[0250] 60. The medium of clause 56, wherein the non-adjacent inter-coded blocks are located in a video frame that includes the coding unit and are spatially non-adjacent neighbors of the coding unit.

[0251] 61. The medium of clause 56, wherein the motion information for the non-adjacent inter-coded blocks further comprises:

[0252] a reference index associated with said non-adjacent inter-coded block,

[0253] the motion vector associated with the non-affine inter-coded block, and

[0254] At least one of an indicator of a uni-directional prediction mode and an indicator of a bi-directional prediction mode.

[0255] 62. The medium of clause 56, wherein encoding based on the motion information comprises:

[0256] Selecting the non-adjacent inter-frame coding blocks from the merge candidate list for encoding the coding unit; and

[0257] An index indicating that the non-adjacent inter-coded blocks are selected for encoding the coding unit is transmitted to a decoder.

[0258] 63. A non-transitory computer-readable medium storing a set of instructions, the set of instructions being executable by one or more processors of a device to cause the device to perform a method comprising the steps of:

[0259] determining values of bidirectional prediction weights for coding units of a video frame;

[0260] determining whether the bidirectional prediction weights are equal weights; and

[0261] In response to the determination, the following is transmitted to the video decoder:

[0262] When the bidirectional prediction weights are equal weights, transmitting a bitstream including a first syntax element indicating the equal weights, or

[0263] After determining that the bidirectional prediction weights are unequal weights, a bitstream is transmitted including a second syntax element indicating values of the bidirectional prediction weights corresponding to the unequal weights.

[0264] 64. The medium of clause 63, wherein the first syntax element is a one-bit flag.

[0265] 65. The medium of clause 63, wherein transmitting the bitstream including the second syntax element further comprises:

[0266] determining a number of unequal weights that can be used by the coding units; and

[0267] In response to determining that there are more than two unequal weights that can be used by the coding unit, transmitting the second syntax element as an index having at least two bits to the video decoder; and

[0268] In response to determining that the unequal weights that can be used by the coding unit are one or two, the second syntax element is transmitted to the video decoder as a flag having one bit.

[0269] 66. The medium of clause 65, wherein the set of instructions is executable by the one or more processors of the device to cause the device to further perform:

[0270] Each bit of the second syntax element is encoded using a different context-adaptive binary arithmetic coding (CABAC) context.

[0271] 67. The medium of clause 63, wherein the set of instructions is executable by the one or more processors of the device to cause the device to further perform:

[0272] determining whether the coding unit is part of a low-latency picture;

[0273] Determine the following:

[0274] In response to determining that the coding unit is part of the low-delay picture, determining that the coding unit uses more than two unequal weights, and

[0275] In response to determining that the coding unit is not part of the low-delay picture, determining that the coding unit uses one or two unequal weights.

[0276] 68. The medium of clause 67, wherein the set of instructions is executable by the one or more processors of the device to cause the device to further perform:

[0277] When the coding unit is part of a low-delay picture and the value of the bidirectional prediction weight is 3 / 8, determining that the value of the second syntax element is 0;

[0278] When the coding unit is part of a low-delay picture and the value of the bidirectional prediction weight is 5 / 8, determining that the value of the second syntax element is 1;

[0279] When the coding unit is not part of a low-delay picture and the value of the bidirectional prediction weight is 3 / 8, determining that the value of the second syntax element is 0; and

[0280] When the coding unit is not part of a low-delay picture and the value of the bidirectional prediction weight is 5 / 8, the value of the second syntax element is determined to be 1.

[0281] 69. The medium of clause 63, wherein the set of instructions is executable by the one or more processors of the device to cause the device to further perform:

[0282] Different numbers of bits are allocated to the second syntax element for low-delay pictures and non-low-delay pictures, respectively.

[0283] 70. The medium of clause 69, wherein the value of the second syntax element corresponds to the same bidirectional prediction weight value for the low-delay picture and the non-low-delay picture.

[0284] 71. The medium of clause 63, wherein each of the first syntax element and the second syntax element corresponds to one or more pre-allocated bits in a coding-unit level weight index.

[0285] 72. The medium of clause 63, wherein the set of instructions is executable by the one or more processors of the device to cause the device to further perform:

[0286] The value of the first syntax element and the value of the second syntax element are encoded using different context-adaptive binary arithmetic coding (CABAC) contexts.

[0287] 73. A computer-implemented signal transmission method performed by a decoder, the signal transmission method comprising the steps of:

[0288] The decoder receives a bitstream including weight information for predicting a coding unit from a video encoder;

[0289] If weighted prediction is enabled for the bidirectional prediction mode of the coding unit, disabling weighted averaging for the bidirectional prediction mode is determined based on the weight information.

[0290] 74. A computer-implemented video encoding method performed by a decoder, the method comprising the steps of:

[0291] The decoder receives a merge candidate list for a coding unit from the encoder, the merge candidate list including motion information of non-adjacent inter-coded blocks of the coding unit; and

[0292] Bi-directional prediction weights associated with the non-adjacent inter-coded blocks are determined based on the motion information.

[0293] 75. A computer-implemented signal transmission method performed by a decoder, the signal transmission method comprising the steps of:

[0294] The decoder receives from the video encoder:

[0295] a bitstream comprising a first syntax element corresponding to bidirectional prediction weights for a coding unit of a video frame, or

[0296] a bitstream comprising a second syntax element corresponding to the bidirectional prediction weight;

[0297] In response to receiving the first syntax element, the decoder determines that the bidirectional prediction weights are equal weights; and

[0298] In response to receiving the first syntax element, the decoder determines that the bidirectional prediction weights are unequal weights, and the decoder determines values of the unequal weights based on the second syntax element.

[0299] As used herein, unless expressly stated otherwise, the term "or" encompasses all possible combinations unless not feasible. For example, if it is stated that a database may include either A or B, then unless expressly stated otherwise or not feasible, the database may include either A, or B, or A and B. As a second example, if it is stated that a database may include either A, B, or C, then unless expressly stated otherwise or not feasible, the database may include either A, or B, or C, or A and B, or A and C, or B and C, or A, B, and C.

[0300] It will be understood that the invention is not limited to the exact construction that has been described above and illustrated in the accompanying drawings and that various modifications and changes may be made without departing from the scope of the invention, which is intended to be limited only by the appended claims.

Claims

1. A computer-implemented video decoding method, the video decoding method comprising the following steps: receiving a bitstream including weight information for predicting a coding unit CU; and Based on the weight information, Perform weighted bidirectional prediction for the CU and disable bidirectional prediction with weighted averaging (BWA) for the CU, or The BWA of the CU is performed and the weighted bi-directional prediction is disabled for the CU.

2. The video decoding method according to claim 1, further comprising: In response to the weight information indicating that weighted prediction is enabled for the CU, the weighted bidirectional prediction of the CU is performed and the BWA is disabled for the CU.

3. The video decoding method according to claim 1, further comprising: In response to the weight information indicating that weighted prediction is enabled for a picture including the CU, the weighted bidirectional prediction of the CU is performed and the BWA is disabled for the CU.

4. The video decoding method according to claim 3, wherein: The weight information includes: a picture parameter set indicating that weighted prediction is enabled for the picture, and A slice header indicating that the CU is associated with the picture.

5. The video decoding method according to claim 1, further comprising: In response to transmitting an index associated with the BWA in the bitstream, performing the BWA of the CU and disabling the weighted bi-directional prediction for the CU.

6. The video decoding method according to claim 1, further comprising: In response to determining that an index associated with the BWA is not transmitted in the bitstream, determining whether weighted prediction is enabled for reference pictures of the CU; as well as In response to determining that the weighted prediction is enabled for at least one of the luma component and the chroma component of the reference picture of the CU, performing the weighted bidirectional prediction of the CU and disabling the BWA for the CU, or, In response to determining that the weighted prediction is not enabled for the reference pictures of the CU, performing BWA with equal weights on the CU and disabling the weighted bi-directional prediction for the CU.

7. The video decoding method according to claim 1, further comprising: In response to transmitting in the bitstream an index associated with the unequal weights of the BWA, performing the BWA of the CU and disabling the weighted bi-directional prediction for the CU.

8. The video decoding method according to claim 1, further comprising: In response to transmitting in the bitstream an index associated with the equal weights of the BWAs, performing the weighted bi-directional prediction of the CU and disabling the BWA for the CU.

9. A video encoder, comprising: a memory storing instructions; as well as a processor configured to execute the instructions to cause the video encoder to: Transmitting a bitstream including weight information for predicting a coding unit CU to a video decoder, wherein the weight information indicates: Perform weighted bidirectional prediction for the CU and disable bidirectional prediction with weighted averaging (BWA) for the CU, or The BWA of the CU is performed and the weighted bi-directional prediction is disabled for the CU.

10. The video encoder according to claim 9, wherein The weight information indicates: If weighted prediction is enabled for the CU, performing the weighted bidirectional prediction of the CU and disabling the BWA for the CU. The video encoder according to claim 9 , wherein: The weight information indicates: If weighted prediction is enabled for a picture including the CU, performing the weighted bidirectional prediction of the CU and disabling the BWA for the CU.

12. The video encoder according to claim 11, wherein The weight information includes: a picture parameter set indicating that weighted prediction is enabled for the picture, and A slice header indicating that the CU is associated with the picture.

13. The video encoder according to claim 9, wherein The processor is configured to execute the instructions to cause the video encoder to: An index associated with the BWA is transmitted in the bitstream to a video decoder to cause the video decoder to perform the BWA for the CU and disable the weighted bi-directional prediction for the CU.

14. The video encoder according to claim 9, wherein The weight information indicates: If weighted prediction is enabled for at least one of the luma component and the chroma component of the reference picture of the CU, performing the weighted bidirectional prediction of the CU and disabling the BWA for the CU, or, If the weighted prediction is not enabled for the reference pictures of the CU, BWA with equal weight is performed on the CU and the weighted bi-directional prediction is disabled for the CU.

15. The video encoder according to claim 9, wherein The processor is configured to execute the instructions to cause the video encoder to: Indices associated with the unequal weights of the BWA are transmitted in the bitstream to a video decoder to cause the video decoder to perform the BWA for the CU and disable the weighted bi-directional prediction for the CU.

16. The video encoder according to claim 9, wherein The processor is configured to execute the instructions to cause the video encoder to: Indices associated with the equal weights of the BWAs are transmitted in the bitstream to a video decoder to cause the video decoder to perform the weighted bi-directional prediction for the CU and disable the BWA for the CU.

17. A non-transitory computer-readable storage medium storing an instruction set, the non-transitory computer-readable storage medium storing a bitstream generated by encoding, the instruction set being executable by one or more processors of a video encoder to cause the video encoder to perform a method comprising the following steps: The bitstream including weight information for predicting the coding unit CU is transmitted to a video decoder, wherein The weight information indicates: Perform weighted bidirectional prediction for the CU and disable bidirectional prediction with weighted averaging (BWA) for the CU, or The BWA of the CU is performed and the weighted bi-directional prediction is disabled for the CU.

18. The non-transitory computer-readable storage medium of claim 17, wherein: The weight information indicates: If weighted prediction is enabled for the CU, performing the weighted bidirectional prediction of the CU and disabling the BWA for the CU.

19. The non-transitory computer-readable storage medium of claim 17, wherein: The weight information indicates: If weighted prediction is enabled for a picture including the CU, performing the weighted bidirectional prediction of the CU and disabling the BWA for the CU.

20. The non-transitory computer-readable storage medium of claim 19, wherein: The weight information includes: a picture parameter set indicating that weighted prediction is enabled for the picture, and A slice header indicating that the CU is associated with the picture.