Video decoding method, video encoding method, and computer readable storage medium

By employing a weighted average bidirectional prediction method at the coding unit level, the incompatibility problem of different video coding tools in the same coding block or picture is solved, improving the efficiency and accuracy of video coding, especially when there is attenuation in the video sequence.

CN116347096BActive Publication Date: 2026-05-26ALIBABA GROUP HOLDING LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ALIBABA GROUP HOLDING LTD
Filing Date
2019-12-19
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Different video coding tools, such as temporal motion prediction, bidirectional prediction, and weighted prediction, are incompatible when applied to the same coding block or the same image, leading to reduced coding efficiency.

Method used

A weighted average bidirectional prediction method is adopted at the coding unit level. By constructing a merged candidate list and transmitting relevant motion information, the weighted prediction mode can be disabled or enabled to ensure the mutually exclusive use of different tools.

Benefits of technology

It improves the efficiency and accuracy of video coding, especially when there is decay in the video sequence, enhances the accuracy of time prediction, and reduces coding time and resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116347096B_ABST
    Figure CN116347096B_ABST
Patent Text Reader

Abstract

This application provides a video decoding method, a video encoding method, and a computer-readable storage medium. Video encoding and decoding techniques for bidirectional prediction with weighted averaging are disclosed. According to some embodiments, a computer-implemented video decoding method includes the following steps: decoding a bitstream associated with a target coding block; reconstructing the target coding block based on a merge candidate list, the merge candidate list including one or more merge candidates associated with one or more inter-frame coding blocks other than the target coding block, each of the one or more merge candidates being motion information associated with the corresponding inter-frame coding block, wherein at least one of the one or more merge candidates includes bidirectional prediction weights.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of the invention patent application No. 201980085317.0 (International Application No.: PCT / US2019 / 067619, Application Date: December 19, 2019, Invention Title: Block-level Bidirectional Prediction with Weighted Average).

[0002] Related applications

[0003] This application claims priority to U.S. Patent Application No. 16 / 228,741, filed December 20, 2018, the entire contents of which are incorporated herein by reference. Technical Field

[0004] This disclosure generally relates to video processing, and more specifically, to video encoding and decoding at the block (or coding unit) level using bi-prediction (BWA) with weighted average. Background Technology

[0005] Video coding systems are typically used to compress digital video signals, for example, to reduce the storage space consumed or the transmission bandwidth consumption associated with such signals.

[0006] Video coding systems can use various tools or techniques to solve different problems. For example, temporal motion prediction is an effective method to improve coding efficiency and provide high compression. Temporal motion prediction can be unidirectional prediction using one reference image or bidirectional prediction using two reference images. In some cases, such as when attenuation occurs, bidirectional prediction may not produce the most accurate prediction. To compensate for this, weighted prediction can be used to apply different weights to the two predicted signals.

[0007] However, different coding tools are not always compatible. For example, applying the aforementioned temporal prediction, bidirectional prediction, or weighted prediction to the same coding block (e.g., coding unit), the same slice, or the same picture may be inappropriate. Therefore, it is desirable to enable different coding tools to interact appropriately with each other. Summary of the Invention

[0008] Embodiments of this disclosure relate to methods for encoding and transmitting signals at the coding unit (CU) level using weighted average bidirectional prediction weights. In some embodiments, a computer-implemented video decoding method is provided, comprising the steps of: decoding a bitstream associated with a target coding block; reconstructing the target coding block based on a merge candidate list, the merge candidate list including one or more merge candidates associated with one or more inter-frame coding blocks other than the target coding block, each of the one or more merge candidates being motion information associated with a corresponding inter-frame coding block, wherein at least one of the one or more merge candidates includes bidirectional prediction weights.

[0009] In some embodiments, a computer-implemented video coding method is provided, the video coding method comprising the steps of: encoding a target coding block based on a merge candidate list, the merge candidate list comprising one or more merge candidates associated with one or more inter-frame coding blocks other than the target coding block, each of the one or more merge candidates being motion information associated with the corresponding inter-frame coding block, wherein at least one of the one or more merge candidates includes bidirectional prediction weights.

[0010] In some embodiments, a non-transitory computer-readable storage medium is provided that stores a bitstream of video, the bitstream being generated according to a method comprising the steps of: encoding a target coding block based on a merge candidate list, the merge candidate list comprising one or more merge candidates associated with one or more inter-coding blocks other than the target coding block, each of the one or more merge candidates being motion information associated with the corresponding inter-coding block, wherein at least one of the one or more merge candidates includes bidirectional prediction weights.

[0011] In some implementations, a computer-implemented video transmission method is provided. The video transmission method includes the steps of: a processor transmitting a bitstream to a video encoder, the bitstream including weight information for predictive coding units (CUs). The weight information indicates that if weighted prediction is enabled for a bidirectional prediction mode for the CU, then weighted averaging for the bidirectional prediction mode is disabled.

[0012] In some embodiments, a computer-implemented video coding method is provided. The video coding method includes the following steps: a processor constructs a merging candidate list for coding units, the merging candidate list including motion information of non-affine inter-coding blocks of the coding units, the motion information including bidirectional prediction weights associated with the non-affine inter-coding blocks. The video coding method further includes the step of the processor performing encoding based on the motion information.

[0013] In some embodiments, a computer-implemented video transmission method is provided. The video transmission method includes the step of: a processor determining values ​​of bidirectional prediction weights for coding units (CUs) of a video frame. The video transmission method further includes the step of: the processor determining whether the bidirectional prediction weights are equal weights. The video transmission method further includes the step of: in response to the determination, the processor transmitting the following to a video decoder: when the bidirectional prediction weights are equal weights, transmitting a bitstream including a first syntax element indicating the equal weights; or after determining that the bidirectional prediction weights are unequal weights, transmitting a bitstream including a second syntax element indicating a value of the bidirectional prediction weights corresponding to the unequal weights.

[0014] In some implementations, a computer-implemented signal transmission method is provided, executed by a decoder. The signal transmission method includes the step of the decoder receiving a bitstream from a video encoder, including weight information for prediction coding units (CUs). The signal transmission method further includes the step of determining, based on the weight information, to disable weighted averaging for the bidirectional prediction mode if weighted prediction is enabled for the CU.

[0015] In some implementations, a computer-implemented video coding method is provided, executed by a decoder. The video coding method includes the following steps: the decoder receives from an encoder a list of merging candidates for coding units, the list including motion information of non-adjacent inter-frame coding blocks of the coding unit. The video coding method further includes the step of determining bidirectional prediction weights associated with the non-adjacent inter-frame coding blocks based on the motion information.

[0016] In some embodiments, a computer-implemented signal transmission method executed by a decoder is provided. The signal transmission method includes the steps of: the decoder receiving from a video encoder either a bitstream comprising a first syntax element corresponding to bidirectional prediction weights of coding units (CUs) for a video frame, or a bitstream comprising a second syntax element corresponding to the bidirectional prediction weights. The signal transmission method further includes the steps of: in response to receiving the first syntax element, the processor determining that the bidirectional prediction weights are equal weights. The signal transmission method further includes the steps of: in response to receiving the first syntax element, the processor determining that the bidirectional prediction weights are unequal weights, and the processor determining the value of the unequal weights based on the second syntax element.

[0017] Aspects of the disclosed embodiments may include a non-transitory tangible computer-readable medium storing software instructions that, when executed by one or more processors, are configured to perform and execute one or more methods, operations, etc., consistent with the disclosed embodiments. Furthermore, aspects of the disclosed embodiments may be implemented by one or more processors configured as dedicated processors based on software instructions programmed with logic and instructions that, when executed, perform one or more operations consistent with the disclosed embodiments.

[0018] Additional objects and advantages of the disclosed embodiments will be set forth in part in the description which follows, and will be apparent in part from the description, or may be learned by practice of the embodiments. The objects and advantages of the disclosed embodiments may be realized and obtained by means of the elements and combinations set forth in the claims.

[0019] It should be understood that the foregoing general description and the following detailed description are merely exemplary and explanatory, and do not limit the claimed disclosed implementations. Attached Figure Description

[0020] Figure 1 This is a schematic diagram illustrating an exemplary video encoding and decoding system consistent with embodiments of the present disclosure.

[0021] Figure 2 This illustrates embodiments consistent with the present disclosure. Figure 1 A schematic diagram of an exemplary video encoder, which is part of an exemplary system.

[0022] Figure 3 This illustrates embodiments consistent with the present disclosure. Figure 1 A schematic diagram of an exemplary video decoder, which is part of an exemplary system.

[0023] Figure 4 It is a table of syntactic elements for weighted prediction (WP) consistent with the embodiments of this disclosure.

[0024] Figure 5 This is a schematic diagram illustrating bidirectional prediction consistent with embodiments of the present disclosure.

[0025] Figure 6 It is a table of grammatical elements for bidirectional prediction with weighted average (BWA) consistent with the embodiments of this disclosure.

[0026] Figure 7 This is a schematic diagram illustrating spatial neighbors used in the construction of a merge candidate list, consistent with embodiments of this disclosure.

[0027] Figure 8 It is a table of enabled or disabled syntax elements for WP in image-level transmission, consistent with the embodiments of this disclosure.

[0028] Figure 9 It is a table of enabled or disabled syntax elements for transmitting WP at the slice level, consistent with the embodiments of this disclosure.

[0029] Figure 10 It is a table of syntactic elements for maintaining the exclusivity of WP and BWA at the CU level, consistent with the embodiments of this disclosure.

[0030] Figure 11 It is a table of syntactic elements for maintaining the exclusivity of WP and BWA at the CU level, consistent with the embodiments of this disclosure.

[0031] Figure 12 This is a flowchart of BWA weight transfer processing for LD images, consistent with the embodiments of this disclosure.

[0032] Figure 13 This is a flowchart of BWA weight transfer processing for non-LD images, consistent with the embodiments of this disclosure.

[0033] Figure 14 This is a block diagram of a video processing apparatus consistent with embodiments of the present disclosure. Detailed Implementation

[0034] Reference will now be made in detail to exemplary embodiments, examples of which are illustrated in the accompanying drawings. The following description refers to the accompanying drawings, wherein, unless otherwise indicated, the same reference numerals in different drawings denote the same or similar elements. The implementations set forth in the following description of the exemplary embodiments do not represent all implementations consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with aspects of the invention as described in the appended claims.

[0035] Figure 1 This is a block diagram illustrating an example video encoding and decoding system 100, which can utilize technologies conforming to various video coding standards (such as HEVC / H.265 and WC / H.266). Figure 1 As shown, system 100 includes a source device 120 that provides encoded video data for later decoding by a destination device 140. Consistent with the disclosed embodiments, each of the source device 120 and the destination device 140 may include any of a wide variety of devices, including desktop computers, laptops (e.g., notebook computers), tablet computers, set-top boxes, mobile phones, televisions, cameras, wearable devices (e.g., smartwatches or wearable cameras), display devices, digital media players, video game consoles, video streaming devices, etc. The source device 120 and the destination device 140 may be equipped for wireless or wired communication.

[0036] refer to Figure 1 Source device 120 may include a video source 122, a video encoder 124, and an output interface 126. Destination device 140 may include an input interface 142, a video decoder 144, and a display device 146. In other examples, the source and destination devices may include other components or arrangements. For example, source device 120 may receive video data from an external video source (not shown), such as an external camera. Similarly, destination device 140 may interface with an external display device instead of including an integrated display device.

[0037] Although the disclosed techniques are interpreted in the following description as being performed by a video encoding device, the techniques can also be performed by a video encoder / decoder, commonly referred to as a "CODEC". Furthermore, the techniques disclosed herein can also be performed by a video preprocessor. Source device 120 and destination device 140 are merely examples of such encoding devices, wherein source device 120 generates encoded video data for transmission to destination device 140. In some examples, source device 120 and destination device 140 may operate in a substantially symmetrical manner, such that each of source device 120 and destination device 140 includes video encoding and decoding components. Therefore, system 100 can support one-way or two-way video transmission between source device 120 and destination device 140, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0038] The video source 122 of source device 120 may include a video capture device (such as a camera), a video archive containing previously captured video, or a video feed interface for receiving video from a video content provider. Alternatively, video source 122 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. The captured, pre-captured, or computer-generated video may be encoded by video encoder 124. The encoded video information may then be output to communication medium 160 via output interface 126.

[0039] Output interface 126 may include any type of medium or device capable of transmitting encoded video data from source device 120 to destination device 140. For example, output interface 126 may include a transmitter or transceiver configured to transmit encoded video data directly from source device 120 to destination device 140 in real time. The encoded video data may be modulated according to a communication standard such as a wireless communication protocol and transmitted to destination device 140.

[0040] Communication medium 160 may include transient media, such as wireless broadcasting or wired network transmissions. For example, communication medium 160 may include radio frequency (RF) spectrum or one or more physical transmission lines (e.g., cables). Communication medium 160 may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. In some embodiments, communication medium 160 may include a router, switch, base station, or any other device that may be useful in facilitating communication from source device 120 to destination device 140. For example, a network server (not shown) may receive encoded video data from source device 120 and provide the encoded video data to destination device 140, for example, via network transmission.

[0041] The communication medium 160 may also be in the form of a storage medium (e.g., a non-transitory storage medium), such as a hard disk, flash drive, optical disk, digital video disk, Blu-ray disc, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In some embodiments, a computing device of a media production facility, such as a disc imprinting facility, may receive encoded video data from source device 120 and produce a disc containing the encoded video data.

[0042] The input interface 142 of the destination device 140 receives information from the communication medium 160. The received information may include syntax information, which includes characteristics of description blocks and other coding units or grammatical elements of the processing. The syntax information is defined by the video encoder 124 and used by the video decoder 144. The display device 146 displays the decoded video data to the user and may include any of a variety of display devices, such as a cathode ray tube (CRT), liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or another type of display device.

[0043] In another example, the encoded video generated by source device 120 can be stored on a file server or storage device. Input interface 142 can access the stored video data from the file server or storage device via streaming or downloading. The file server or storage device can be any type of computing device capable of storing the encoded video data and transferring it to destination device 140. Examples of file servers include web servers supporting websites, file transfer protocol (FTP) servers, network attached storage (NAS) devices, or local disk drives. The transfer of encoded video data from the storage device can be streaming, downloading, or a combination thereof.

[0044] The video encoder 124 and video decoder 144 can each be implemented as any of a variety of suitable encoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is partially implemented in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the technology of this disclosure. Each of the video encoder 124 and video decoder 144 may be included in one or more encoders or decoders, and either encoder or decoder may be integrated as part of a combined encoder / decoder (CODEC) in the respective device.

[0045] The video encoder 124 and video decoder 144 can operate according to any video coding standard, such as the Universal Video Coding (VVC / H.266) standard, the High Efficiency Video Coding (HEVC / H.265) standard, the ITU-T H.264 (also known as MPEG-4) standard, etc. Although in Figure 1As not shown in the figure, but in some embodiments, the video encoder 124 and the video decoder 144 may each be integrated with the audio encoder and decoder, and may include appropriate MUX-DEMUX units or other hardware and software to process the encoding of audio and video in a common data stream or a separate data stream.

[0046] Figure 2 This is a schematic diagram illustrating an exemplary video encoder 200 consistent with the disclosed implementation. For example, the video encoder 200 can be used as system 100 ( Figure 1 The video encoder 124 is located in the video encoder 200. The video encoder 200 can perform intra-frame or inter-frame coding of blocks (including video blocks or partitions or sub-partitions of video blocks) within a video frame. Intra-frame coding can rely on spatial prediction to reduce or remove spatial redundancy in the video within a given video frame. Inter-frame coding can rely on temporal prediction to reduce or remove temporal redundancy in the video within adjacent frames of a video sequence. Intra-frame modes can refer to various spatial-based compression modes, and inter-frame modes (such as one-way prediction or two-way prediction) can refer to various temporal-based compression modes.

[0047] refer to Figure 2 The input video signal 202 can be processed block by block. For example, a video block unit can be a 16×16 pixel block (e.g., a macroblock (MB)). In HEVC, an expanded block size (e.g., a coding unit (CU)) can be used to compress video signals with resolutions of, for example, 1080p and higher. In HEVC, a CU can include up to 64×64 luma samples and corresponding chroma samples. In VVC, the size of the CU can be further increased to include 128×128 luma samples and corresponding chroma samples. CUs can be partitioned into prediction units (PUs), and individual prediction methods can be applied to the PUs. Individual input video blocks (e.g., MB, CU, PU, ​​etc.) can be processed using a spatial prediction unit 260 or a temporal prediction unit 262.

[0048] Spatial prediction unit 260 performs spatial prediction (e.g., intra-frame prediction) on the current CU using information about the same picture / slice containing the current CU. Spatial prediction can use pixels from already encoded neighboring blocks in the same video picture / slice to predict the current video block. Spatial prediction can reduce the spatial redundancy inherent in the video signal. Temporal prediction (e.g., inter-frame prediction or motion-compensated prediction) can use samples from already encoded video pictures to predict the current video block. Temporal prediction can reduce the temporal redundancy inherent in the video signal.

[0049] Timing prediction unit 262 performs timing prediction (e.g., inter-frame prediction) on the current CU using information from a picture / slice different from the picture / slice containing the current CU. Timing prediction of a video block can be transmitted via one or more motion vectors. Motion vectors can indicate the amount and direction of motion between the current block in a reference frame and one or more of its predicted blocks. If multiple reference pictures are supported, one or more reference picture indices can be sent for the video block. One or more reference indices can be used to identify which reference pictures the timing prediction signal can come from in the reference picture storage unit or the decoded picture buffer (DPB) 264. After spatial or temporal prediction, mode decision and encoder control unit 280 in the encoder can select a prediction mode, for example, based on a rate distortion optimization method. The predicted block can be subtracted from the current video block at adder 216. The prediction residual can be transformed by transform unit 204 and quantized by quantization unit 206. The quantized residual coefficients can be inversely quantized at inverse quantization unit 210 and inversely transformed at inverse transform unit 212 to form reconstructed residuals. The reconstructed block can be added to the prediction block at adder 226 to form a reconstructed video block. In-loop filtering, such as deblocking filters and adaptive loop filters 266, can be applied to the reconstructed video block before it is placed in reference image storage unit 264 and used for encoding future video blocks. To form the output video bitstream 220, the coding mode (e.g., inter-frame or intra-frame), prediction mode information, motion information, and quantized residual coefficients can be sent to entropy coding unit 208 for compression and packing to form the bitstream 220. The systems, methods, and means described herein can be implemented at least partially within time prediction unit 262.

[0050] Figure 3 This is a schematic diagram illustrating a video decoder 300 consistent with the disclosed implementation. For example, the video decoder 300 can be used as system 100 ( Figure 1 The video decoder 144 in ( ). Reference Figure 3 The video bitstream can be unpacked or entropy-decoded at entropy decoding unit 308. Encoding modes or prediction information can be sent to spatial prediction unit 360 (e.g., if intra-frame coding) or temporal prediction unit 362 (e.g., if inter-frame coding) to form prediction blocks. If inter-frame coding is used, the prediction information may include the prediction block size, one or more motion vectors (e.g., indicating the direction and amount of motion) or one or more reference indices (e.g., indicating from which reference image the prediction signal will be obtained).

[0051] Motion compensation prediction can be applied by time prediction unit 362 to form a time prediction block. Residual transform coefficients can be sent to inverse quantization unit 310 and inverse transform unit 312 to reconstruct the residual block. The prediction block and residual block can be added together at 326. Before storing the reconstructed block in reference image storage unit 364, the reconstructed block can be filtered in-loop (via loop filter 366). The reconstructed video in reference image storage unit 364 can be used to drive a display device or to predict future video blocks. The decoded video 320 can be displayed on a monitor.

[0052] Consistent with the disclosed implementations, the video encoder and video decoder described above can use various video encoding / decoding tools to process, compress, and decompress video data. Three tools are described below: Weighted Prediction (WP), Bidirectional Prediction with Weighted Average (BWA), and History-Based Motion Vector Prediction (HMVP).

[0053] Weighted Prediction (WP) is used to provide significantly better temporal predictions when attenuation is present in a video sequence. Attenuation refers to the phenomenon where the average illumination level of images in video content exhibits a significant change in the temporal domain, such as attenuating to white or black. Content creators often use attenuation to create desired special effects and express their artistic vision. Attenuation causes the average illumination level of the reference image to differ significantly from the average illumination level of the current image, making it more difficult to obtain accurate prediction signals from temporally adjacent images. As part of efforts to address this problem, WP provides powerful tools to adjust the illumination level of the prediction signal obtained from the reference image and match it to the illumination level of the current image, thereby significantly improving the accuracy of temporal predictions.

[0054] Consistent with the disclosed implementation, parameters (or "WP parameters") for the WP are transmitted for each reference image in the reference images used to encode the current image. For each reference image, the WP parameters include a pair of weights and offsets (w, o), which can be transmitted for each color component of the reference image. Figure 4 Table 400 depicts the syntax elements for WP according to the disclosed implementation. Referring to Table 400, the pred_weight_table0 syntax is transmitted as part of the slice header. For the i-th reference image in the reference image list Lx (x can be 0 or 1), the flags luma_weight_lx_flag[i] and chroma_weight_lx_flag[i] are transmitted to indicate whether weighted prediction is applied to the luminance and chrominance components of the i-th reference image, respectively.

[0055] Without loss of generality, the following description uses luminance as an example to illustrate the transmission of WP parameters. Specifically, if luma_weight_lx_flag[i] is set to 1 for the i-th reference image in the reference image list Lx, then the WP parameters (w[i], o[i]) are transmitted for the luminance component. Then, when applying temporal prediction using a given reference image, the following equation (1) applies:

[0056] Equation (1) is given by WP(x, y) = w·P(x, y) + o, where: WP(x, y) is the weighted prediction signal at sample location (x, y); (w, o) is the WP parameter pair associated with the reference image; and P(x, y) = ref(x-mvx, y-mvy) is the prediction before applying WP, (mvx, mvy) is the motion vector associated with the reference image, and ref(x, y) is the reference signal at location (x, y). If the motion vector (mvx, mvy) has fractional sample accuracy, interpolation can be applied, such as using an 8-tap luminance interpolation filter in HEVC.

[0057] For bidirectional prediction (BWA) tools with weighted averages, bidirectional prediction can be used to improve temporal prediction accuracy, thereby improving the compression performance of video encoders. It is used in various video coding standards such as H.264 / AVC, HEVC, and VVC. Figure 5 This is a schematic diagram illustrating an exemplary bidirectional prediction. (See reference) Figure 5 CU 503 is a portion of the current image. CU 501 is from reference image 511, and CU 502 is from reference image 512. In some embodiments, reference images 511 and 512 may be selected from two different lists of reference images L0 and L1, respectively. Two motion vectors (mvx0, mvy0) and (mvx1, mvy1) can be generated by referring to CU 501 and CU 502, respectively. These two motion vectors form two prediction signals, which can be averaged to obtain a bidirectional prediction signal, i.e., the prediction corresponding to CU 503.

[0058] In the disclosed embodiments, reference images 511 and 512 may be from the same or different image sources. Specifically, although Figure 5 Reference images 511 and 512 are depicted as two different physical reference images corresponding to different points in time. However, in some embodiments, reference images 511 and 512 may be the same physical reference image, as it is permissible for the same physical reference image to appear once or more in one or both of the reference image lists L0 and L1. Furthermore, although... Figure 5Reference images 511 and 512 are depicted as originating from the past and future in the time domain, respectively. However, in some implementations, reference images 511 and 512 may both originate from the past or both originate from the future, relative to the current image.

[0059] Specifically, refer to Figure 5 Bidirectional prediction can be performed based on the following equation:

[0060]

[0061] Wherein: (mvx0, mvy0) is a motion vector associated with a reference image (e.g., reference image 511) selected from the reference image list L0; (mvx1, mvy1) is a motion vector associated with a reference image (e.g., reference image 512) selected from the reference image list L1; and ref0(x, y) is a reference signal at position (x, y) in reference image 511; and ref1(x, y) is a reference signal at position (x, y) in reference image 512.

[0062] Still referencing Figure 5 Weighted prediction can be applied to bidirectional prediction. In some implementations, each prediction signal is assigned an equal weight of 0.5, such that the prediction signals are averaged based on the following equation:

[0063]

[0064] Where (w0,o0) and (w1,o1) are the WP parameters associated with reference images 511 and 512, respectively.

[0065] In some implementations, BWA is used to apply unequal weights with a weighted average to bidirectional prediction applications, which can improve coding efficiency. BWA can be applied adaptively at the block level. For each CU, the weight index gbi_idx is transmitted if certain conditions are met. Figure 6 Table 600 depicts the syntax elements for BWA according to the disclosed implementation. Referring to Table 600, bidirectional prediction of a CU containing, for example, at least 256 luminance samples can be performed using the syntax at 601 and 602. Based on the value of gbi_idx, a weight w is determined and applied to the reference signal according to the following equation:

[0066] P(x, y)=(1-w)·P0(x, y)+w·P1(x, y)

[0067] =(1-w)·ref0(x-mvx0,y-mvy0)+w·ref1(x-mvx1,y-mvy1) Equation (4).

[0068] In some implementations, the value of the BWA weight w can be selected from five possible values, for example, Low-latency (LD) images are defined as those whose reference images always precede their own in the display order. For LD images, all five values ​​mentioned above can be used for BWA weights. That is, when transmitting BWA weights, the weight index gbi_idx is in the range [0, 4], and the center value (gbi_idx = 2) corresponds to equal weights. The value of . For non-low latency (non-LD) images, only 3 BWA weights are used. It is used. In this case, the value of the weight index gbi_idx is in the range [0, 2], and the center value (gbi_idx = 1) corresponds to equal weights. The value of .

[0069] If an explicit transfer of the weight index gbi_idx is used, the values ​​of the BWA weights for the current CU are selected by the encoder, for example, through rate distortion optimization. One approach is to try all allowed weight values ​​w and select one with the lowest rate distortion cost. However, an exhaustive search for the optimal combination of weights and motion vectors can significantly increase encoding time. Therefore, fast encoding methods can be applied to reduce encoding time without sacrificing encoding efficiency.

[0070] For each bidirectionally predicted CU, the BWA weight w can be determined and transmitted in one of the following two ways: 1) For non-merged CUs, transmit the weight index after the motion vector difference, as shown in Table 600. Figure 6 As shown in Figure 1); and 2) for merging CUs, the weight index gbi_idx is inferred from the adjacent blocks based on the merge candidate index. The merge mode will be described in detail below.

[0071] Merge candidates for a CU can come from neighboring blocks of the current CU, or from juxtaposed blocks in the temporal juxtaposition image of the current CU. Figure 7 This is a schematic diagram illustrating the use of spatial neighbors in the construction of a merge candidate list according to an exemplary implementation. Figure 7 The locations of five spatial candidates for motion information are described. To construct a list of merged candidates, the five spatial candidates can be examined and added to the list, for example, in the order A1, B1, B0, A0, and A2. If a block located at a spatial location is internally encoded or outside the boundary of the current slice, that block can be considered unavailable. Redundant entries (e.g., candidates with the same motion information) can be excluded from the list of merged candidates.

[0072] Merging mode has been supported since the HEVC standard. Merging mode is an efficient method to reduce motion transmission overhead. Instead of explicitly transmitting motion information of the current CU (prediction mode, motion vectors, reference indices, etc.), motion information from neighboring blocks of the current CU is used to construct a merge candidate list. Both spatially and temporally neighboring blocks can be used to construct the merge candidate list. After constructing the merge candidate list, an index is transmitted to indicate which of the merge candidates to use for encoding the current CU. Motion information from that merge candidate is then used to predict the current CU.

[0073] When BWA is enabled, if the current CU is in merge mode, the motion information it inherits from its merge candidate can include not only motion vectors and reference indices, but also the weight index gbi_idx of that merge candidate. In other words, when performing motion compensation prediction, a weighted average of the two predicted signals is performed for the current CU based on the weight index gbi_idx of its neighboring blocks. In some implementations, if the merge candidate is a spatial neighbor, the weight index gbi_idx is inherited only from the merge candidate, and if the candidate is a temporal neighbor, the weight index gbi_idx is not inherited.

[0074] The merge pattern in HEVC constructs a list of merge candidates using spatially adjacent blocks and temporally adjacent blocks. Figure 7 In the example shown, all spatially adjacent blocks are adjacent to the current CU (i.e., connected). However, in some implementations, non-adjacent neighbors can be used in the merging mode to further improve the coding efficiency of the merging mode. The merging mode using non-adjacent neighbors is called the extended merging mode. In some implementations, the History-Based Motion Vector Prediction (HMVP) method from VVC can be used for inter-frame coding in the extended merging mode to improve compression performance with minimal implementation cost. In HMVP, an HMVP candidate table is continuously maintained and updated during video encoding / decoding processing. The HMVP table can include up to six entries. HMVP candidates are inserted in the middle of the list of merging candidates for spatial neighbors and can be selected as other merging candidates for encoding the current CU using the merging candidate index.

[0075] The First-In-First-Out (FIFO) rule is applied to remove and add entries to the table. After decoding a non-affine inter-coded block, the table is updated by adding the associated motion information as a new HMVP candidate to the last entry and removing the oldest HMVP candidate from the table. The table is cleared when a new slice is encountered. In some implementations, the table may be cleared more frequently, for example, when a new coding tree unit (CTU) or a new CTU row is encountered.

[0076] The above descriptions of WP, BWA, and HMVP indicate the need to coordinate these tools during video coding and transmission. For example, both BWA and WP introduce weighting factors into inter-frame prediction processing to improve the prediction accuracy of motion compensation. However, BWA differs from WP in its functionality. According to equation (4), BWA applies weights in a normalized manner. That is, the weights applied to L0 prediction and L1 prediction are (1-w) and w, respectively. Since the weights add up to 1, BWA limits how the two prediction signals are combined without changing the total energy of the bidirectional prediction signals. On the other hand, according to equation (3), WP does not have normalization constraints. That is, w0 and w1 do not need to add up to 1. Furthermore, WP can add constant offsets o0 and o1 according to equation (3). In addition, BWA and WP are suitable for different types of video content. However, WP effectively attenuates video sequences (or other video content with global illumination changes in the temporal domain) and does not improve the coding efficiency of normal sequences when the illumination level does not change in the temporal domain. In contrast, BWA is a block-level adaptive tool that adaptively selects how to combine two predicted signals. While BWA is effective on normal sequences without illumination changes, it is far less effective than the WP method on attenuated sequences. For these reasons, in some implementations, both BWA and WP tools can be supported in video coding standards, but they operate in a mutually exclusive manner. Therefore, a mechanism is needed to disable one tool when the other is present.

[0077] Furthermore, as discussed above, if the selected merge candidate is a spatial neighbor adjacent to the current CU, the BWA tool and merge pattern can be combined by allowing inheritance of the weight index gbi_idx from the selected merge candidate. To leverage the advantages of HMVP, BWA needs to be combined with HMVP to utilize non-adjacent neighbors in the extended merge pattern.

[0078] At least some of the disclosed implementations provide solutions for maintaining the exclusivity of WP and BWA. A combination of syntax in the Picture Parameter Set (PPS) and slice headers is used to indicate whether WP is enabled for a picture. Figure 8 Table 800 is a list of syntax elements for enabling or disabling WP at the picture level, consistent with embodiments of this disclosure. As shown in 801 of Table 800, weighted_pred_flag and weighted_bipred_flag are sent in the PPS to indicate whether WP is enabled for unidirectional and bidirectional prediction, respectively, depending on the slice type of the slice referencing the PPS. Figure 9Table 900 is a list of syntax elements for enabling or disabling WP at the slice level, consistent with embodiments of this disclosure. As shown at 901 in Table 900, at the slice / picture level, if the PPS referenced by the slice (determined by matching the slice_pic_parameter_set_id in the slice header with the pps_pic_parameter_set_id in the PPS) has WP enabled, then Table 400 ( Figure 4 The pred_weight_table() function in the current image is sent to the decoder to indicate the WP parameters of each reference image in the reference images of the current image.

[0079] Based on this transmission, in some embodiments of this disclosure, additional conditions can be added to the CU-level weighted index gbi_idx transmission. Additional condition transmission: If WP is enabled for an image containing the current CU, then weighted averaging is disabled for the bidirectional prediction mode of that current CU. Figure 10 Table 1000 is a set of syntactic elements consistent with the embodiments of this disclosure for maintaining the exclusivity of WP and BWA at the CU level. Referring to Table 1000, condition 1001 can be added to indicate that if the PPS referenced by the current slice allows WP for bidirectional prediction, then BWA is completely disabled for all CUs in that current slice. This ensures that WP and BWA are exclusive.

[0080] However, the above method can completely disable BWA for all CUs in the current slice, regardless of whether the current CU uses a reference image with WP enabled. This may reduce encoding efficiency. At the CU level, the values ​​of luma_weight_l0_flag[ref_idx_l0], / chroma_weight_l0_flag[ref_idx_l0], luma_weight_l1_flag[ref_idx_l1], and luma / chroma_weight_l1_flag[ref_idx_l1] can be used to determine whether WP is enabled for its reference image. Here, ref_idx_l0 and ref_idx_l1 are the reference image indices of the current CUs in L0 and L1, respectively. For the L0 and L1 reference images of the current slice, luma / chroma_weight_l0_flag and luma / chroma_weight_l1_flag are transmitted in pred_weight_table(), as shown in Table 400. Figure 4 As shown in the figure. Figure 11Table 1100 is a syntax element consistent with the implementation of this disclosure for maintaining the exclusivity of WP and BWA at the CU level. Referring to Table 1100, condition 1101 is added to control the exclusivity of WP and BWA at the CU level, regardless of whether the weight index gbi_idx is transmitted. When the weight index gbi_idx is not transmitted, it can be inferred that it is the default value representing the case of equal weights (i.e., 1 or 2, depending on whether 3 or 5 BWA weights are allowed).

[0081] The methods illustrated in Tables 1000 and 1100 both add conditions to the transmission of the weight index gbi_idx at the CU level, which may complicate the parsing process of the decoder. Therefore, in the third embodiment, the transmission conditions of the weight index gbi_idx remain the same as in Table 600 ( Figure 6 The transmission conditions are the same in both cases. If WP is enabled for the luma or chroma components of the L0 or LI reference picture, the default value of the weight index gbi_idx of the current CU is always sent as a bitstream consistency constraint for the encoder. That is, the weight index gbi_idx value corresponding to unequal weights can only be sent when WP is not enabled for the luma or chroma components of the L0 and LI reference pictures. Although this transmission is redundant, the actual bit cost of this redundant transmission is negligible because the context-adaptive binary arithmetic coding (CABAC) engine in the entropy coding stage can adapt to the statistics of the weight index gbi_idx value. In addition, this simplifies the parsing process.

[0082] In the decoder (e.g., Figure 3 After the encoder 300 receives the bitstream containing the above-mentioned syntax for maintaining the exclusivity of WP and BWA, the decoder can parse the bitstream and determine whether to disable BWA based on the syntax.

[0083] At least some embodiments of this disclosure provide a solution for symmetric transmission of BWA at the CU level. As discussed above, in some embodiments, the CU-level weights used in the BWA are transmitted as a weight index gbi_idx, with the value of gbi_idx in the range [0, 4] for low-latency (LD) images and in the range [0, 2] for non-LD images. However, this can create inconsistencies between LD and non-LD images, as follows:

[0084]

[0085] Here, the same BWA weight value is represented by different gbi_idx values ​​in LD and non-LD images.

[0086] To improve transmission consistency, according to some disclosed implementations, the transmission of the weight index gbi_idx can be modified to include a first flag indicating whether the BWA weights are equal weights, followed by an index or flag indicating unequal weights. Figure 12 and Figure 13 Flowcharts illustrating exemplary BWA weight transfer processing for LD images and non-LD images are shown respectively. For LD images allowing 5 BWA weight values, the following flowcharts are used... Figure 12 The transmission flowchart in the document, and for non-LD images that allow 3 BWA weight values, uses... Figure 13 The transmission flowchart is shown below. The first flag, gbi_ew_flag, indicates whether equal weights were applied in the BWA. If gbi_ew_flag is 1, no additional transmission is needed because equal weights (w = 1 / 2) were applied; otherwise, a flag (1 bit for 2 values) or an index (2 bits for 4 values) is transmitted to indicate which unequal weights were applied. Figure 12 and Figure 13 This is just one example illustrating a possible mapping between BWA weight values ​​and index / tag values. It is conceivable that other mappings between weight values ​​and index / tag values ​​could be used. Another advantage of splitting the weight index gbi_idx into two syntactic elements (gbi_ew_flag and gbi_uew_val_idx (or gbi_uew_val_flag)) is that these values ​​can be encoded using a separate CABAC context. Furthermore, for LD images, when using the 2-bit value gbi_uew_val_idx, a separate CABAC context can be used to encode the first and second bits.

[0087] In the decoder (e.g., Figure 3 After the encoder 300 receives the aforementioned transmission of the BWA at the CU level, the decoder can parse the transmission and determine whether the BWA uses equal weights based on the transmission. If the BWA is determined to have unequal weights, the decoder can further determine the value of the unequal weights based on the transmission.

[0088] Some embodiments of this disclosure provide a solution for combining BWA and HMVP. If the motion information stored in the HMVP table only includes the motion vectors, reference indices, and prediction patterns (e.g., one-way and two-way predictions) of the merge candidates, the merge candidates cannot be used with BWA because BWA weights are not stored or are not updated in the HMVP table. Therefore, according to some disclosed embodiments, BWA weights are included as part of the motion information stored in the HMVP table. When the HMVP table is updated, the BWA weights are also updated along with other motion information, such as motion vectors, reference indices, and prediction patterns.

[0089] Furthermore, partial pruning can be applied to avoid having too many identical candidates in the merge candidate list. Identical candidates are defined as those whose motion information is the same as at least one of the existing merge candidates in the merge candidate list. Identical candidates occupy space in the merge candidate list but do not provide any additional motion information. Partial pruning can detect some of these cases and can prevent some of these identical candidates from being added to the merge candidate list. By including BWA weights in the HMVP table, the pruning process also considers BWA weights when determining whether two merge candidates are identical. Specifically, if a new candidate has the same motion vector, reference index, and prediction pattern as another candidate in the merge candidate list, but has different BWA weights, then the new candidate can be considered different and can be left unpruned.

[0090] In the decoder (e.g., Figure 3 After the encoder 300 receives the bitstream including the HMVP table mentioned above, the decoder can parse the bitstream and determine the BWA weights of the merge candidates included in the HMVP table.

[0091] Figure 14 This is a block diagram of a video processing apparatus 1400 consistent with embodiments of the present disclosure. For example, the apparatus 1400 may implement the video encoder described above (e.g., Figure 2 The video encoder 200 or video decoder (e.g., in the video encoder 200) ... Figure 3 (Video decoder 300 in the image). In the disclosed embodiment, apparatus 1400 can be configured to perform the above-described method for encoding and transmitting BWA weights. (See also: [reference needed]) Figure 14 The device 1400 may include a processing unit 1402, a memory 1404, and an input / output (I / O) interface 1406. The device 1400 may also include one or more of a power supply unit and a multimedia unit (not shown) or any other suitable hardware or software unit.

[0092] Processing unit 1402 can control the overall operation of device 1400. For example, processing unit 1402 may include one or more processors that execute instructions to perform the methods described above for encoding and transmitting BWA weights. Furthermore, processing unit 1402 may include one or more modules that facilitate interaction between processing unit 1402 and other components. For example, processing unit 1402 may include I / O modules to facilitate interaction between I / O interfaces and processing unit 1402.

[0093] Memory 1404 is configured to store various types of data or instructions to support the operation of device 1400. Memory 1404 may include non-transitory computer-readable storage media including instructions for applications or methods operating on device 1400, which can be executed by one or more processors of device 1400. Common forms of non-transitory media include, for example, floppy disks, collapsible disks, hard disks, solid-state drives, magnetic tape or any other magnetic data storage media, CD-ROMs, any other optical data storage media, any physical media with a perforated pattern, RAM, PROMs and EPROMs, cloud storage units, FLASH-EPROMs or any other flash memory, NVRAM, cache units, registers, any other memory chips or cassette tapes and their network versions.

[0094] I / O interface 1406 provides an interface between processing unit 1402 and peripheral interface modules (such as cameras or displays). I / O interface 1406 can employ communication protocols / methods such as audio, analog, digital, serial bus, Universal Serial Bus (USB), infrared, PS / 2, BNC, coaxial, RF antenna, Bluetooth, etc. I / O interface 1406 can also be configured to facilitate wired or wireless communication between device 1400 and other devices (such as devices connected to the Internet). The device can access wireless networks based on one or more communication standards (such as WiFi, LTE, 2G, 3G, 4G, 5G, etc.).

[0095] The implementation method may be further described using the following terms:

[0096] 1. A computer-implemented signal transmission method, the signal transmission method comprising the following steps:

[0097] The processor transmits a bitstream to the video decoder, the bitstream including weight information for predictive coding units, the weight information indicating:

[0098] If weighted prediction is enabled for the bidirectional prediction mode of the coding unit, then weighted averaging for the bidirectional prediction mode is disabled.

[0099] 2. The signal transmission method according to Clause 1, wherein the weighting information indicates:

[0100] If weighted prediction is enabled for bidirectional prediction of an image including the coding unit, then weighted averaging for the bidirectional prediction mode is disabled.

[0101] 3. The signal transmission method according to Clause 2, wherein the bit stream includes a flag indicating whether weighted prediction is enabled for bidirectional prediction of the image including the coding unit.

[0102] 4. The signal transmission method according to Clause 1, wherein the weighting information indicates:

[0103] If weighted prediction is enabled for at least one of the luminance and chrominance components of the reference image for the coding unit, then weighted averaging for the bidirectional prediction mode is disabled.

[0104] 5. The signal transmission method according to Clause 4, wherein the bit stream includes a flag indicating whether weighted prediction is enabled for at least one of the luminance component and the chrominance component of the reference image.

[0105] 6. The signal transmission method according to Clause 1, wherein the weight information includes values ​​of bidirectional prediction weights associated with the coding unit, and the signal transmission method further includes the following steps:

[0106] If weighted prediction is enabled for at least one of the luminance and chrominance components of the reference image for the coding unit, the value of the bidirectional prediction weight is set to the default value.

[0107] 7. The signal transmission method according to Clause 6, wherein the default value corresponds to equal weights.

[0108] 8. A computer-implemented video encoding method, the video encoding method comprising the following steps:

[0109] The processor constructs a merging candidate list for each coding unit. The merging candidate list includes motion information of non-adjacent inter-coding blocks of the coding unit, the motion information including bidirectional prediction weights associated with the non-adjacent inter-coding blocks; and

[0110] The processor encodes based on the motion information.

[0111] 9. The video encoding method according to Clause 8, wherein the video encoding method further comprises the following steps:

[0112] The processor determines a new non-adjacent inter-frame coding block for the coding unit: and

[0113] The processor updates the merge candidate list by inserting bidirectional prediction weights associated with the new non-adjacent inter-frame coding block into the merge candidate list.

[0114] 10. The video coding method according to Clause 8, wherein:

[0115] The processor determines motion information for a new non-adjacent inter-coding block of the coding unit, the motion information including bidirectional prediction weights associated with the non-adjacent inter-coding block:

[0116] The processor compares the motion information of the new non-adjacent inter-coded blocks with the motion information of each inter-coded block included in the merge candidate list.

[0117] In response to a comparison result that none of the inter-coded blocks included in the merge candidate list have the same motion information as the new non-adjacent inter-coded block, the processor adds the new non-adjacent inter-coded block to the merge candidate list; or

[0118] In response to a comparison result showing that the motion information of the new non-adjacent inter-coded block is the same as the motion information of at least one inter-coded block included in the merge candidate list, it is determined that the new non-adjacent inter-coded block is redundant to the merge candidate list.

[0119] 11. The video coding method according to Clause 8, wherein the non-adjacent inter-frame coding blocks are derived from a history-based motion vector prediction HMVP table.

[0120] 12. The video coding method according to Clause 8, wherein the non-adjacent inter-frame coding block is located in a video frame including the coding unit and is a spatially non-adjacent neighbor of the coding unit.

[0121] 13. The video coding method according to Clause 8, wherein the motion information of the non-adjacent inter-frame coded block further includes:

[0122] The reference index associated with the non-adjacent inter-frame coded block,

[0123] Motion vectors associated with non-affine inter-coded blocks, and

[0124] At least one of the indicators for unidirectional prediction mode and bidirectional prediction mode.

[0125] 14. The video coding method according to Clause 8, wherein the step of encoding based on the motion information includes:

[0126] Select the non-adjacent inter-frame coding block from the list of merging candidates for encoding the coding unit; and

[0127] The decoder is transmitted an index indicating that the non-adjacent inter-frame coding block is selected for encoding the coding unit.

[0128] 15. A computer-implemented signal transmission method, the signal transmission method comprising the following steps:

[0129] The processor determines the values ​​of the bidirectional prediction weights for the coding units used in a video frame;

[0130] The processor determines whether the bidirectional prediction weights are equal weights; and

[0131] In response to the determination, the processor transmits the following to the video decoder:

[0132] When the bidirectional prediction weights are equal weights, a bitstream including a first syntax element is transmitted, the first syntax element indicating the equal weights, or

[0133] After determining that the bidirectional prediction weights are unequal weights, a bitstream including a second syntax element is transmitted, the second syntax element indicating the value of the bidirectional prediction weights corresponding to the unequal weights.

[0134] 16. The signal transmission method according to Clause 15, wherein the first syntactic element is a tag having one bit.

[0135] 17. The signal transmission method according to Clause 15, wherein the step of transmitting a bit stream including the second syntax element further comprises:

[0136] Determine the number of unequal weights that can be used by the encoding unit; and

[0137] In response to determining that more than two unequal weights are available for use by the encoding unit, the processor transmits the second syntax element as an index with at least two bits to the video decoder; or

[0138] In response to determining that the unequal weights that can be used by the encoding unit are one or two, the processor transmits the second syntax element as a tag with one bit to the video decoder.

[0139] 18. The signal transmission method according to Clause 17, further comprising the following steps:

[0140] The processor uses a different context-adaptive binary arithmetic encoding (CABAC) context to encode each bit of the second syntax element.

[0141] 19. The signal transmission method according to Clause 15, the signal transmission method further comprising the following steps:

[0142] Determine whether the encoded unit is part of a low-latency image;

[0143] Confirm the following:

[0144] In response to determining that the coding unit is part of the low-latency image, it is determined that the coding unit uses more than two unequal weights, or

[0145] In response to determining that the coding unit is not part of the low-latency image, it is determined that the coding unit uses one or two unequal weights.

[0146] 20. The signal transmission method according to Clause 19, wherein:

[0147] When the coding unit is part of a low-latency image and the value of the bidirectional prediction weight is 3 / 8, the value of the second syntax element is determined to be 0;

[0148] When the coding unit is part of a low-latency image and the value of the bidirectional prediction weight is 5 / 8, the value of the second syntax element is determined to be 1;

[0149] When the coding unit is not part of the low-latency image and the value of the bidirectional prediction weight is 3 / 8, the value of the second syntax element is determined to be 0; and

[0150] When the encoding unit is not part of a low-latency image and the bidirectional prediction weight is 5 / 8, the value of the second grammatical element is determined to be 1.

[0151] 21. The signal transmission method according to Clause 15, further comprising the following steps:

[0152] The processor allocates different numbers of bits to the second syntax element for low-latency images and non-low-latency images, respectively.

[0153] 22. The signal transmission method according to Clause 21, wherein, for the low-latency image and the non-low-latency image, the value of the second syntax element corresponds to the same bidirectional prediction weight value.

[0154] 23. The signal transmission method according to Clause 15, wherein each of the first syntax element and the second syntax element corresponds to one or more pre-allocated bits in the coding unit level weight index.

[0155] 24. The signal transmission method according to Clause 15, the signal transmission method further comprising the following steps:

[0156] The processor uses different context-adaptive binary arithmetic encoding (CABAC) contexts to encode the values ​​of the first syntax element and the second syntax element.

[0157] 25. An apparatus, the apparatus comprising:

[0158] A memory, wherein the memory stores instructions; and

[0159] Processor, the processor being configured to execute the instructions to cause the device to:

[0160] A bitstream is transmitted to the video encoder, the bitstream including weight information for predictive coding units, the weight information indicating:

[0161] If weighted prediction is enabled for the bidirectional prediction mode of the coding unit, then weighted averaging for the bidirectional prediction mode is disabled.

[0162] 26. The device according to clause 25, wherein the weighting information indicates:

[0163] If weighted prediction is enabled for bidirectional prediction of an image including the coding unit, then weighted averaging for the bidirectional prediction mode is disabled.

[0164] 27. The device according to Clause 26, wherein the bitstream includes a flag indicating whether weighted prediction is enabled for bidirectional prediction of the image including the coding unit.

[0165] 28. The device according to Clause 25, wherein the weighting information indicates:

[0166] If weighted prediction is enabled for at least one of the luminance and chrominance components of the reference image for the coding unit, then weighted averaging for the bidirectional prediction mode is disabled.

[0167] 29. The device according to Clause 28, wherein the bitstream includes a flag indicating whether weighted prediction is enabled for at least one of the luminance component and the chrominance component of the reference image.

[0168] 30. The apparatus according to claim 25, wherein the weight information includes values ​​of bidirectional prediction weights associated with the coding unit, and the processor is further configured to execute the instructions to:

[0169] If weighted prediction is enabled for at least one of the luminance and chrominance components of the reference image for the coding unit, the value of the bidirectional prediction weight is set to the default value.

[0170] 31. The device according to Clause 30, wherein the default value corresponds to equal weights.

[0171] 32. An apparatus, the apparatus comprising:

[0172] A memory, wherein the memory stores instructions; and

[0173] Processor, the processor being configured to execute the instructions to cause the device to:

[0174] A merging candidate list is constructed for each coding unit, the merging candidate list including motion information of non-adjacent inter-coding blocks of the coding unit, the motion information including bidirectional prediction weights associated with the non-adjacent inter-coding blocks; and

[0175] Encoding is performed based on the motion information.

[0176] 33. The device according to clause 32, wherein the processor is further configured to execute the instructions to:

[0177] Determine the new non-adjacent inter-frame coding block of the coding unit: and

[0178] The merge candidate list is updated by inserting the bidirectional prediction weights associated with the new non-adjacent inter-coding blocks into the merge candidate list.

[0179] 34. The device according to clause 32, wherein the processor is further configured to execute the instructions to:

[0180] Determine motion information for a new non-adjacent inter-coding block of the coding unit, the motion information including bidirectional prediction weights associated with the non-adjacent inter-coding block:

[0181] The motion information of the new non-adjacent inter-coded blocks is compared with the motion information of each inter-coded block included in the merge candidate list;

[0182] In response to a comparison result that none of the inter-coded blocks included in the merge candidate list have the same motion information as the new non-adjacent inter-coded block, the new non-adjacent inter-coded block is added to the merge candidate list; or

[0183] In response to a comparison result showing that the motion information of the new non-adjacent inter-coded block is the same as the motion information of at least one inter-coded block included in the merge candidate list, it is determined that the new non-adjacent inter-coded block is redundant to the merge candidate list.

[0184] 35. The apparatus according to clause 32, wherein the non-adjacent inter-frame coded blocks are derived from a history-based motion vector prediction HMVP table.

[0185] 36. The apparatus according to clause 32, wherein the non-adjacent inter-frame coding block is located in a video frame including the coding unit and is a spatially non-adjacent neighbor of the coding unit.

[0186] 37. The apparatus according to clause 32, wherein the motion information of the non-adjacent inter-frame coded block further includes:

[0187] The reference index associated with the non-adjacent inter-frame coded block,

[0188] Motion vectors associated with non-affine inter-coded blocks, and

[0189] At least one of the indicators for unidirectional prediction mode and bidirectional prediction mode.

[0190] 38. The device according to clause 32, wherein the processor is further configured to execute the instructions to:

[0191] Select the non-adjacent inter-frame coding block from the list of merging candidates for encoding the coding unit; and

[0192] The decoder is transmitted an index indicating that the non-adjacent inter-frame coding block is selected for encoding the coding unit.

[0193] 39. An apparatus, the apparatus comprising:

[0194] A memory, wherein the memory stores instructions; and

[0195] Processor, the processor being configured to execute the instructions to cause the device to:

[0196] Determine the values ​​of the bidirectional prediction weights for the coding units used in video frames;

[0197] Determine whether the bidirectional prediction weights are equal; and

[0198] In response to the determination, the following is transmitted to the video decoder:

[0199] When the bidirectional prediction weights are equal weights, a bitstream including a first syntax element is transmitted, the first syntax element indicating the equal weights, or

[0200] After determining that the bidirectional prediction weights are unequal weights, a bitstream including a second syntax element is transmitted, the second syntax element indicating the value of the bidirectional prediction weights corresponding to the unequal weights.

[0201] 40. The device according to clause 39, wherein the first syntactic element is a tag having one bit.

[0202] 41. The device according to clause 39, wherein the processor is further configured to execute the instructions to:

[0203] Determine the number of unequal weights that can be used by the encoding unit; and

[0204] In response to determining that more than two unequal weights are available for use by the encoding unit, the second syntax element, as an index with at least two bits, is transmitted to the video decoder; and

[0205] In response to determining that the unequal weights that can be used by the encoding unit are one or two, the second syntax element, as a tag with one bit, is transmitted to the video decoder.

[0206] 42. The apparatus according to clause 41, wherein the processor is further configured to execute the instructions to:

[0207] The individual bits of the second syntax element are encoded using a different context-adaptive binary arithmetic encoding CABAC context.

[0208] 43. The device according to clause 39, wherein the processor is further configured to execute the instructions to:

[0209] Determine whether the encoded unit is part of a low-latency image;

[0210] Confirm the following:

[0211] In response to determining that the coding unit is part of the low-latency image, it is determined that the coding unit uses more than two unequal weights, and

[0212] In response to determining that the coding unit is not part of the low-latency image, it is determined that the coding unit uses one or two unequal weights.

[0213] 44. The device according to clause 43, wherein the processor is further configured to execute the instructions to:

[0214] When the coding unit is part of a low-latency image and the value of the bidirectional prediction weight is 3 / 8, the value of the second syntax element is determined to be 0;

[0215] When the coding unit is part of a low-latency image and the value of the bidirectional prediction weight is 5 / 8, the value of the second syntax element is determined to be 1;

[0216] When the coding unit is not part of the low-latency image and the value of the bidirectional prediction weight is 3 / 8, the value of the second syntax element is determined to be 0; and

[0217] When the encoding unit is not part of a low-latency image and the bidirectional prediction weight is 5 / 8, the value of the second grammatical element is determined to be 1.

[0218] 45. The apparatus according to clause 39, wherein the processor is further configured to execute the instructions to:

[0219] Different numbers of bits are allocated to the second syntax element for low-latency images and non-low-latency images respectively.

[0220] 46. ​​The device according to Clause 45, wherein, for the low-latency image and the non-low-latency image, the value of the second grammatical element corresponds to the same bidirectional prediction weight value.

[0221] 47. The apparatus according to Clause 39, wherein each of the first syntactic element and the second syntactic element corresponds to one or more pre-allocated bits in the coding unit level weight index.

[0222] 48. The device according to clause 39, wherein the processor is further configured to execute the instructions to:

[0223] The values ​​of the first syntax element and the second syntax element are encoded using different context-adaptive binary arithmetic encoding CABAC contexts.

[0224] 49. A non-transitory computer-readable medium storing an instruction set executable by one or more processors of a device to cause the device to perform a method comprising the steps of:

[0225] A bitstream is transmitted to the video decoder, the bitstream including weight information for predictive coding units, the weight information indicating:

[0226] If weighted prediction is enabled for the bidirectional prediction mode of the coding unit, then weighted averaging for the bidirectional prediction mode is disabled.

[0227] 50. The medium according to clause 49, wherein the weighting information indicates:

[0228] If weighted prediction is enabled for bidirectional prediction of an image including the coding unit, then weighted averaging for the bidirectional prediction mode is disabled.

[0229] 51. The medium according to Clause 50, wherein the bitstream includes a flag indicating whether weighted prediction is enabled for bidirectional prediction of the image including the coding unit.

[0230] 52. The medium according to clause 49, wherein the weighting information indicates:

[0231] If weighted prediction is enabled for at least one of the luminance and chrominance components of the reference image for the coding unit, then weighted averaging for the bidirectional prediction mode is disabled.

[0232] 53. The medium according to Clause 52, wherein the bitstream includes a flag indicating whether weighted prediction is enabled for at least one of the luminance component and the chrominance component of the reference image.

[0233] 54. The medium according to clause 49, wherein the weight information includes values ​​of bidirectional prediction weights associated with the coding unit, and the instruction set is executable by the one or more processors of the device to enable the device to further perform:

[0234] If weighted prediction is enabled for at least one of the luminance and chrominance components of the reference image for the coding unit, the value of the bidirectional prediction weight is set to the default value.

[0235] 55. The medium as described in Clause 54, wherein the default value corresponds to equal weights.

[0236] 56. A non-transitory computer-readable medium storing an instruction set executable by one or more processors of a device to cause the device to perform a method, the method comprising the steps of:

[0237] A merging candidate list is constructed for each coding unit, the merging candidate list including motion information of non-adjacent inter-coding blocks of the coding unit, the motion information including bidirectional prediction weights associated with the non-adjacent inter-coding blocks; and

[0238] Encoding is performed based on the motion information.

[0239] 57. The method further comprises the following steps: (Based on the medium described in Clause 56)

[0240] Determine the new non-adjacent inter-frame coding block of the coding unit: and

[0241] The merge candidate list is updated by inserting the bidirectional prediction weights associated with the new non-adjacent inter-coding blocks into the merge candidate list.

[0242] 58. The medium according to clause 56, wherein the instruction set is executable by the one or more processors of the device to cause the device to further perform:

[0243] Determine motion information for a new non-adjacent inter-coding block of the coding unit, the motion information including bidirectional prediction weights associated with the non-adjacent inter-coding block:

[0244] The motion information of the new non-adjacent inter-coded blocks is compared with the motion information of each inter-coded block included in the merge candidate list;

[0245] In response to a comparison result that none of the inter-coded blocks included in the merge candidate list have the same motion information as the new non-adjacent inter-coded block, the new non-adjacent inter-coded block is added to the merge candidate list; and

[0246] In response to a comparison result showing that the motion information of the new non-adjacent inter-coded block is the same as the motion information of at least one inter-coded block included in the merge candidate list, it is determined that the new non-adjacent inter-coded block is redundant to the merge candidate list.

[0247] 59. The medium according to Clause 56, wherein the non-adjacent inter-frame coded blocks are derived from a history-based motion vector prediction table.

[0248] 60. The medium according to Clause 56, wherein the non-adjacent inter-frame coding block is located in a video frame including the coding unit and is a spatially non-adjacent neighbor of the coding unit.

[0249] 61. The medium according to clause 56, wherein the motion information of the non-adjacent inter-frame coded blocks further includes:

[0250] The reference index associated with the non-adjacent inter-frame coded block,

[0251] Motion vectors associated with non-affine inter-coded blocks, and

[0252] At least one of the indicators for unidirectional prediction mode and bidirectional prediction mode.

[0253] 62. The medium according to clause 56, wherein the step of encoding based on the motion information includes:

[0254] Select the non-adjacent inter-frame coding block from the list of merging candidates for encoding the coding unit; and

[0255] The decoder is transmitted an index indicating that the non-adjacent inter-frame coding block is selected for encoding the coding unit.

[0256] 63. A non-transitory computer-readable medium storing an instruction set executable by one or more processors of a device to cause the device to perform a method, the method comprising the steps of:

[0257] Determine the values ​​of the bidirectional prediction weights for the coding units used in video frames;

[0258] Determine whether the bidirectional prediction weights are equal; and

[0259] In response to the determination, the following is transmitted to the video decoder:

[0260] When the bidirectional prediction weights are equal weights, a bitstream including a first syntax element is transmitted, the first syntax element indicating the equal weights, or

[0261] After determining that the bidirectional prediction weights are unequal weights, a bitstream including a second syntax element is transmitted, the second syntax element indicating the value of the bidirectional prediction weights corresponding to the unequal weights.

[0262] 64. The medium according to clause 63, wherein the first syntactic element is a token having one bit.

[0263] 65. The medium according to clause 63, wherein the step of transmitting the bit stream including the second syntax element further comprises:

[0264] Determine the number of unequal weights that can be used by the encoding unit; and

[0265] In response to determining that more than two unequal weights are available for use by the encoding unit, the second syntax element, as an index with at least two bits, is transmitted to the video decoder; and

[0266] In response to determining that the unequal weights that can be used by the encoding unit are one or two, the second syntax element, as a tag with one bit, is transmitted to the video decoder.

[0267] 66. The medium according to clause 65, wherein the instruction set is executable by the one or more processors of the device to cause the device to further perform:

[0268] The individual bits of the second syntax element are encoded using a different context-adaptive binary arithmetic encoding CABAC context.

[0269] 67. The medium according to clause 63, wherein the instruction set is executable by the one or more processors of the device to cause the device to further perform:

[0270] Determine whether the encoded unit is part of a low-latency image;

[0271] Confirm the following:

[0272] In response to determining that the coding unit is part of the low-latency image, it is determined that the coding unit uses more than two unequal weights, and

[0273] In response to determining that the coding unit is not part of the low-latency image, it is determined that the coding unit uses one or two unequal weights.

[0274] 68. The medium according to clause 67, wherein the instruction set is executable by the one or more processors of the device to cause the device to further perform:

[0275] When the coding unit is part of a low-latency image and the value of the bidirectional prediction weight is 3 / 8, the value of the second syntax element is determined to be 0;

[0276] When the coding unit is part of a low-latency image and the value of the bidirectional prediction weight is 5 / 8, the value of the second syntax element is determined to be 1;

[0277] When the coding unit is not part of the low-latency image and the value of the bidirectional prediction weight is 3 / 8, the value of the second syntax element is determined to be 0; and

[0278] When the encoding unit is not part of a low-latency image and the bidirectional prediction weight is 5 / 8, the value of the second grammatical element is determined to be 1.

[0279] 69. The medium according to clause 63, wherein the instruction set is executable by the one or more processors of the device to cause the device to further perform:

[0280] Different numbers of bits are allocated to the second syntax element for low-latency images and non-low-latency images respectively.

[0281] 70. The medium according to Clause 69, wherein, for the low-latency image and the non-low-latency image, the value of the second syntax element corresponds to the same bidirectional prediction weight value.

[0282] 71. The medium according to Clause 63, wherein each of the first syntactic element and the second syntactic element corresponds to one or more pre-allocated bits in the coding unit-level weight index.

[0283] 72. The medium according to clause 63, wherein the instruction set is executable by the one or more processors of the device to cause the device to further perform:

[0284] The values ​​of the first syntax element and the second syntax element are encoded using different context-adaptive binary arithmetic encoding CABAC contexts.

[0285] 73. A computer-implemented signal transmission method executed by a decoder, the signal transmission method comprising the following steps:

[0286] The decoder receives a bitstream from the video encoder that includes weight information for predicting coding units;

[0287] If weighted prediction is enabled for the bidirectional prediction mode of the coding unit, then the weighted average used for the bidirectional prediction mode is disabled based on the weight information.

[0288] 74. A computer-implemented video encoding method executed by a decoder, the method comprising the following steps:

[0289] The decoder receives from the encoder a list of merging candidates for a coding unit, the list including motion information of non-adjacent inter-frame coding blocks of the coding unit; and

[0290] Based on the motion information, bidirectional prediction weights associated with the non-adjacent inter-frame coding blocks are determined.

[0291] 75. A computer-implemented signal transmission method executed by a decoder, the signal transmission method comprising the following steps:

[0292] The decoder receives from the video encoder:

[0293] A bitstream comprising first syntactic elements corresponding to the bidirectional prediction weights of the coding units used for video frames, or

[0294] This includes a bitstream of second syntactic elements corresponding to the bidirectional prediction weights;

[0295] In response to receiving the first grammatical element, the decoder determines that the bidirectional prediction weights are equal weights; and

[0296] In response to receiving the first grammatical element, the decoder determines that the bidirectional prediction weights are unequal weights, and the decoder determines the value of the unequal weights based on the second grammatical element.

[0297] As used herein, unless otherwise expressly stated, the term "or" covers all possible combinations unless impractical. For example, if a statement declares that a database may include A or B, then unless otherwise expressly stated or impractical, the database may include A, or B, or A and B. As a second example, if a statement declares that a database may include A, B, or C, then unless otherwise stated or impractical, the database may include A, or B, or C, or A and B, or A and C, or B and C, or A and B and C.

[0298] It will be understood that the invention is not limited to the exact construction described above and illustrated in the accompanying drawings, and various modifications and changes can be made without departing from the scope of the invention. The scope of the invention is intended to be limited only by the appended claims.

Claims

1. A computer-implemented video decoding method, the video decoding method comprising the following steps: Decoding a bitstream associated with a target coded block, wherein the decoding generates decoded motion information associated with the target coded block; The target coding block is reconstructed based on a merge candidate list, which includes one or more merge candidates associated with one or more inter-coding blocks other than the target coding block, and each of the merge candidates represents motion information associated with the corresponding inter-coding block. Wherein, at least one of the one or more merging candidates includes bidirectional prediction weights, which are used to perform weighted prediction on the target coding block based on two inter-frame coding blocks; and The merge candidate list is updated based on the decoded motion information, which includes bidirectional prediction weights associated with the target coding block.

2. The video decoding method according to claim 1, further comprising: The list of merge candidates is generated before reconstructing the target coding block.

3. The video decoding method of claim 1, wherein, Updating the merge candidate list based on the decoded motion information further includes: The merge candidate list is updated by inserting the decoded motion information into the merge candidate list.

4. The video decoding method of claim 1, wherein, Updating the merge candidate list based on the decoded motion information further includes: Determine whether the list of merge candidates includes merge candidates that are identical to the decoded motion information; and If it is determined that the merge candidate list does not include a merge candidate that is identical to the decoded motion information, then the decoded motion information is inserted into the merge candidate list, or... If it is determined that the merge candidate list includes merge candidates that are identical to the decoded motion information, then the decoded motion information is not inserted into the merge candidate list.

5. The video decoding method of claim 1, wherein, The list of merged candidates is based on the historical motion vector prediction HMVP table.

6. The video decoding method of claim 1, wherein, The one or more inter-frame coded blocks are located in the video frame that includes the target coded block, and are spatially non-adjacent neighbors of the target coded block.

7. The video decoding method of claim 1, wherein, The motion information associated with each inter-coding block in the one or more inter-coding blocks further includes one or more of the following: Reference Index Motion vector, Indicators of unidirectional prediction patterns, or Indicator of bidirectional prediction mode.

8. A computer-implemented video encoding method, the video encoding method comprising the following steps: The target coding block is encoded based on a merge candidate list, which includes one or more merge candidates associated with one or more inter-coding blocks other than the target coding block, and each of the merge candidates represents motion information associated with the corresponding inter-coding block. Wherein, at least one of the one or more merging candidates includes bidirectional prediction weights, which are used to perform weighted prediction on the target coding block based on two inter-frame coding blocks; and Encoding the target coded block based on the merged candidate list includes: Motion information of the target coding block is determined based on the merging candidate list, and the merging candidate list is updated based on the motion information of the target coding block. The motion information of the target coding block includes bidirectional prediction weights associated with the target coding block.

9. The video encoding method according to claim 8, further comprising: The list of merge candidates is generated before the target coded block is encoded.

10. The video coding method of claim 8, wherein, The steps for updating the list of merge candidates include: Determine whether the list of merging candidates includes merging candidates that have the same motion information as the target coded block; and If it is determined that the merge candidate list does not include a merge candidate with the same motion information as the target coding block, then the motion information of the target coding block is inserted into the merge candidate list, or If it is determined that the merge candidate list includes merge candidates that have the same motion information as the target coding block, then the motion information of the target coding block is not inserted into the merge candidate list.

11. The video coding method of claim 8, wherein, The list of merged candidates is based on the historical motion vector prediction HMVP table.

12. The video coding method of claim 8, wherein, The one or more inter-frame coded blocks are located in the video frame that includes the target coded block, and are spatially non-adjacent neighbors of the target coded block.

13. The video coding method of claim 8, wherein, The motion information associated with each inter-coding block in the one or more inter-coding blocks further includes one or more of the following: Reference Index Motion vector, Indicators of unidirectional prediction patterns, or Indicator of bidirectional prediction mode.

14. A non-transitory computer-readable storage medium storing an instruction set, the non-transitory computer-readable storage medium storing a bitstream of video, the instruction set being executable by one or more processors of a video encoder to cause the video encoder to perform a method of generating the bitstream by: The target coding block is encoded based on a merge candidate list, which includes one or more merge candidates associated with one or more inter-coding blocks other than the target coding block, and each of the merge candidates represents motion information associated with the corresponding inter-coding block. wherein At least one of the one or more merging candidates includes bidirectional prediction weights, which are used to perform weighted prediction on the target coding block based on two inter-frame coding blocks; as well as Encoding the target coded block based on the merged candidate list includes: Motion information of the target coding block is determined based on the merging candidate list, and the merging candidate list is updated based on the motion information of the target coding block. The motion information of the target coding block includes bidirectional prediction weights associated with the target coding block.

15. The non-transitory computer-readable storage medium according to claim 14, wherein, The method further includes: The list of merge candidates is generated before the target coded block is encoded.

16. The non-transitory computer-readable storage medium according to claim 14, wherein, The steps for updating the list of merge candidates include: Determine whether the list of merging candidates includes merging candidates that have the same motion information as the target coded block; and If it is determined that the merge candidate list does not include a merge candidate with the same motion information as the target coding block, then the motion information of the target coding block is inserted into the merge candidate list, or If it is determined that the merge candidate list includes merge candidates that have the same motion information as the target coding block, then the motion information of the target coding block is not inserted into the merge candidate list.

17. The non-transitory computer-readable storage medium according to claim 14, wherein, The list of merged candidates is based on the historical motion vector prediction HMVP table.

18. The non-transitory computer-readable storage medium according to claim 14, wherein, The one or more inter-frame coded blocks are located in the video frame that includes the target coded block, and are spatially non-adjacent neighbors of the target coded block.