Video decoding method, apparatus, and program, and video encoding method
The method addresses inefficiencies in JMVD and compound weighted prediction by determining weight factors and motion information, enhancing video compression efficiency and decoding quality.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-06-29
- Publication Date
- 2026-03-11
AI Technical Summary
Existing video coding technologies face challenges in efficiently encoding and decoding videos using joint motion vector difference (JMVD) and compound weighted prediction modes, particularly in handling scaling factors and motion information for reference frames, leading to suboptimal compression efficiency and quality.
A method and apparatus for video encoding/decoding that utilize a processing circuit to determine weight factors and motion information based on predefined or subset scaling factors, enabling efficient reconstruction of video blocks using JMVD and combined weighted prediction modes.
Improves video compression efficiency by accurately determining weight factors and motion information, resulting in enhanced decoding quality and reduced data volume without quality degradation.
Smart Images

Figure 2026508486000001_ABST
Abstract
Description
[Technical Field]
[0001] This disclosure describes embodiments generally related to video coding. [Background technology]
[0002] The background description provided herein is intended to generally present the context for the present disclosure. The work of the currently named inventors is not expressly or implicitly admitted as prior art to the present disclosure to the extent that work is described in the background section, and aspects of the description that may not otherwise qualify as prior art at the time of filing.
[0003] Image / video compression can help transmit image / video files between different devices, storage, and networks with minimal quality degradation. In some examples, video codec technology can compress video based on spatial and temporal redundancy. In examples, video codecs can use a technique called intra-prediction, which can compress images based on spatial redundancy. For example, intra-prediction can use reference data from the current picture being reconstructed for sample prediction. In other examples, video codecs can use a technique called inter-prediction, which can compress images based on temporal redundancy. For example, inter-prediction can predict samples in a current picture from a previously reconstructed picture through motion compensation. Motion compensation can generally be indicated by motion vectors (MVs). Summary of the Invention
[0004] Aspects of the present disclosure provide a method and apparatus for encoding / decoding videos and / or pictures. The apparatus includes a processing circuit configured to receive a bitstream including a frame. Coding information for a block in the frame indicates that the block is coded in a Joint Motion Vector Difference (JMVD) coding mode and a compound weighted prediction mode. The coding information further includes scaling factor information for the JMVD coding mode. In response to the scaling factor information indicating that each of the scaling factors for at least one MVD component associated with at least one respective reference frame of the block is a predefined scaling factor, the processing circuit determines weight factors for the compound weighted prediction mode based on a list of weight factors. In response to the scaling factor information indicating that at least one of the scaling factors is different from the predefined scaling factor, the processing circuit determines weight factors for the compound weighted prediction mode based on a subset of the list of weight factors. The processing circuit may further determine motion information associated with each reference frame of the block based on a scaling factor of at least one MVD component associated with at least one respective reference frame of the block using a JMVD coding mode. The reference frames of the block include at least one respective reference frame. The processing circuit reconstructs the block based on the motion information associated with each reference frame of the block and the determined weighting factor using a joint weighted prediction mode. In an example, the list is signaled in the bitstream. In an example, only a subset of the list is signaled in the bitstream.
[0005] In an embodiment, a processing circuit is configured to receive a bitstream including a frame. Coding information for a block in the frame indicates that the block is coded in a joint MVD coding mode and a combined weighted prediction mode. The coding information further indicates weighting factor information for the combined weighted prediction mode. In response to the weighting factor information indicating that equal weighting factors are applied to each of the block's reference frames, the processing circuit determines a scaling factor for at least one MVD component associated with at least one of the reference frames based on a list of scaling factors. In response to the weighting factor information indicating that unequal weighting factors are applied to the block's reference frames, the processing circuit determines a scaling factor for at least one MVD component based on a subset of the list of scaling factors. The processing circuit further determines motion information associated with each of the block's reference frames based on the determined scaling factor using the joint MVD coding mode, and reconstructs the block based on the motion information and weighting factor information associated with each of the block's reference frames using the combined weighted prediction mode.
[0006] Aspects of the present disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video encoding / decoding, cause the computer to perform a method of video encoding / decoding.
[0007] Further features, nature and various advantages of the disclosed subject matter will become more apparent from the following detailed description and accompanying drawings. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a schematic diagram of an exemplary block diagram of a communication system (100). [Figure 2] FIG. 2 is a schematic diagram of an exemplary block diagram of a decoder. [Figure 3] FIG. 2 is a schematic diagram of an exemplary block diagram of an encoder. [Figure 4] 1 illustrates the relationship between the class of non-zero components of a motion vector differential (MVD) and the magnitude range of the non-zero components of the MVD, according to an embodiment of the present disclosure. [Figure 5] 1 illustrates the relationship between the class of non-zero components of the MVD and the magnitude of the non-zero components of the MVD, according to an embodiment of the present disclosure. [Figure 6] 1 illustrates supported motion vector (MV) precisions within two precision sets according to an embodiment of the present disclosure. [Figure 7] 10 illustrates exemplary scaling factors for JOINT_AMVDNEWMV mode according to an embodiment of the present disclosure. [Figure 8] 10 illustrates an example scaling factor for JOINT_NEWMV according to an embodiment of the present disclosure. [Figure 9] 10 illustrates an exemplary lookup table illustrating the relationship between scaling coefficient pairs and corresponding weighting factors according to an embodiment of the present disclosure. [Figure 10] 1 shows a flowchart illustrating a decoding process according to an embodiment of the present disclosure. [Figure 11] 1 shows a flowchart illustrating an encoding process according to an embodiment of the present disclosure. [Figure 12] 1 shows a flowchart illustrating a decoding process according to an embodiment of the present disclosure. [Figure 13] 1 shows a flowchart illustrating an encoding process according to an embodiment of the present disclosure. [Figure 14] FIG. 1 is a schematic diagram of a computer system according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] 1 shows a block diagram of a video processing system 100 in some examples. The video processing system 100 is an example of an application of the disclosed subject matter, namely, a video encoder and video decoder in a streaming environment. The disclosed subject matter can be similarly applicable to other video-enabled applications including, for example, video conferencing, digital TV, streaming services, storage of compressed video on digital media including CDs, DVDs, memory sticks, etc.
[0010] The video processing system (100) includes a capture subsystem (113) that may include, for example, a video source (101), such as a digital camera, that generates an uncompressed stream of video pictures (102). By way of example, the stream of video pictures (102) includes samples captured by the digital camera. The stream of video pictures (102) is represented by a bold line to emphasize its high data volume compared to the encoded video data (104) (or coded video bitstream) and may be processed by an electronic device (120) that includes a video encoder (103) coupled to the video source (101). The video encoder (103) may include hardware, software, or a combination thereof to enable or implement aspects of the disclosed subject matter, as described in more detail below. The encoded video data (104) (or coded video bitstream) is represented by a thin line to emphasize its lower data volume compared to the stream of video pictures (102) and may be stored on a streaming server (105) for future use. One or more streaming client subsystems, such as the client subsystems 106 and 108 of FIG. 1, can access the streaming server 105 to retrieve copies 107 and 109 of the encoded video data 104. The client subsystem 106 may include a video decoder 110, for example, in an electronic device 130. The video decoder 110 decodes the incoming copy 107 of the encoded video data and generates an outgoing stream 111 of video pictures that can be rendered on a display 112 (e.g., a display screen) or other rendering device (not shown). In some streaming systems, the encoded video data 104, 107, and 109 (e.g., a video bitstream) may be encoded according to a particular video coding / compression standard. An example of such a standard is ITU-T Recommendation H.265.By way of example, a video coding standard under development is commonly known as Versatile Video Coding (VVC), and the disclosed subject matter may be used in connection with VVC.
[0011] It should be noted that electronic devices 120 and 130 may include other components (not shown). For example, electronic device 120 may include a video decoder (not shown), and electronic device 130 may similarly include a video encoder (not shown).
[0012] 2 shows an example block diagram of a video decoder (210). The video decoder (210) may be included in an electronic device (230). The electronic device (230) may include a receiver (231) (e.g., a receiving circuit). The video decoder (210) may be used in place of the video decoder (110) in the example of FIG. 1.
[0013] The receiver (231) may receive one or more coded video sequences to be decoded by the video decoder (210). In an embodiment, one coded video sequence is received at a time, and the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequences may be received from a channel (201), which may be a hardware / software link to a storage device storing the coded video data. The receiver (231) may receive the coded video data along with other data, such as coded audio data and / or auxiliary data streams, which may be forwarded to their respective using entities (not shown). The receiver (231) may separate the coded video sequences from other data. To combat network jitter, a buffer memory (215) may be coupled between the receiver (231) and the entropy decoder / parser (220) (hereinafter "parser (220)"). In certain applications, the buffer memory 215 is part of the video decoder 210. In others, it can be external to the video decoder 210 (not shown). In still other applications, there can be a buffer memory (not shown) external to the video decoder 210, e.g., to combat network jitter, plus another buffer memory 215 within the video decoder 210, e.g., to manipulate playback timing. When the receiver 231 is receiving data from a storage / forwarding device with sufficient bandwidth and controllability, or from an isosynchronous network, the buffer memory 215 may not be required or may be small. For use with best-effort packet networks such as the Internet, the buffer memory 215 may be required, but it can be relatively large and advantageously adaptively sized, and may be implemented at least in part in an operating system or similar element (not shown) external to the video decoder 210.
[0014] The video decoder (210) may include a parser (220) for reconstructing symbols (221) from the coded video sequence. These symbol categories include information used to manage the operation of the video decoder (210) and, potentially, information for controlling a rendering device, such as a render device (212) (e.g., a display screen) that is not an essential part of the electronic device (230) but may be coupled to the electronic device (230) as shown in FIG. 2. Control information for the rendering device may take the form of a Supplemental Enhancement Information (SEI) message or a Video Usability Information (VUI) parameter set fragment (not shown). The parser (220) may parse / entropy decode the received coded video sequence. The coding of the coded video sequence may follow a video coding technique or standard and may follow various principles, including variable length coding, Huffman coding, context-dependent or non-context-dependent arithmetic coding, etc. The parser (220) may extract from the coded video sequence a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder based on at least one parameter corresponding to the group. The subgroups may include groups of pictures (GOPs), pictures, tiles, slices, macroblocks, coding units (CUs), blocks, transform units (TUs), prediction units (PUs), etc. The parser (220) may also extract information from the coded video sequence, such as transform coefficients, quantization parameter values, motion vectors, etc.
[0015] The parser (220) may perform entropy decoding / parsing operations on the video sequence received from the buffer memory (215) to generate symbols (221).
[0016] The reconstruction of the symbols (221) can have many different units depending on the type of coded video picture or portion thereof (e.g., inter and intra pictures, inter and intra blocks) and other factors. Which units are included and how may be controlled by subgroup control information parsed by the parser (220) from the coded video sequence. The flow of such subgroup control information between the parser (220) and the following units is not shown for clarity.
[0017] Beyond the functional blocks already described, the video decoder (210) may be conceptually subdivided into a number of functional units, which are described below. In an actual implementation operating under commercial constraints, many of these units may interact closely with each other and may be at least partially integrated with each other. However, for purposes of describing the disclosed subject matter, the following conceptual subdivision into functional units is appropriate.
[0018] The first unit is a scalar / inverse transform unit (251), which receives quantized transform coefficients as symbols (221) from the parser (220) along with control information including which transform to use, block size, quantization coefficients, quantization scaling matrices, etc. The scalar / inverse transform unit (251) can output blocks containing sample values that can be input to an aggregator (255).
[0019] In some cases, the output samples of the scaler / inverse transform unit (251) may relate to intra-coded blocks. Intra-coded blocks are blocks that do not use prediction information from a previously reconstructed picture, but can use prediction information from a previously reconstructed portion of the current picture. Such prediction information may be provided by the intra-picture prediction unit (252). In some cases, the intra-picture prediction unit (252) generates a block of the same size and shape as the block being reconstructed using surrounding, already reconstructed information fetched from the current picture buffer (258). The current picture buffer (258), for example, buffers a partially reconstructed and / or fully reconstructed current picture. The aggregator (255), in some cases, adds, on a sample-by-sample basis, the prediction information generated by the intra-prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251).
[0020] In other cases, the output samples of the scalar / inverse transform unit (251) may relate to an inter-coded, and potentially motion-compensated, block. In such cases, the motion-compensated prediction unit (253) may access the reference picture memory (257) to fetch samples used for prediction. After motion-compensating the fetched samples according to the symbols (221) related to the block, the samples may be added by the aggregator (255) to the output of the scalar / inverse transform unit (251) (in this case, referred to as residual samples or residual signals) to generate output sample information. The addresses in the reference picture memory (257) from which the motion-compensated prediction unit (253) fetches prediction samples may be controlled by motion vectors available to the motion-compensated prediction unit (253), for example, in the form of symbols (221), which may have X, Y, and reference picture components. Motion compensation can also include interpolation of sample values fetched from the reference picture memory (257) when sub-sample accurate motion vectors are used, as well as motion vector prediction mechanisms.
[0021] The output samples of the aggregator (255) may undergo various loop filtering techniques in a loop filter unit (256). Video compression techniques may include in-loop filtering techniques controlled by parameters contained in the coded video sequence (also called the coded video bitstream) and made available to the loop filter unit (256) as symbols (221) from the parser (220). Video compression may also respond to meta-information obtained during decoding of previous portions (in decoding order) of the coded picture or coded video sequence, and may also respond to previously reconstructed loop-filtered sample values.
[0022] The output of the loop filter unit (256) can be a sample stream that can be output to a render device (212) and further stored in a reference picture memory (257) for use in future inter-picture prediction.
[0023] Once a particular coded picture is fully reconstructed, it can be used as a reference picture for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified as a reference picture (e.g., by the parser (220)), the current picture buffer (258) can become part of the reference picture memory (257), and any unused current picture buffer can be reallocated before beginning reconstruction of a subsequent coded picture.
[0024] The video decoder (210) may perform decoding operations in accordance with a given video compression technology or standard, such as ITU-T Recommendation H.265. A coded video sequence may conform to the syntax prescribed by the video compression technology or standard in use, in the sense that the coded video sequence conforms to both the syntax of the video compression technology or standard and a profile documented in the video compression technology or standard. Specifically, a profile may select specific tools from all tools available in the video compression technology or standard as the only tools available for use under that profile. Compliance also requires that the complexity of the coded video sequence be within the boundaries defined by the level of the video compression technology or standard. In some cases, the level limits the maximum picture size, maximum frame rate, maximum reconstruction sample rate (e.g., measured in megasamples per second), maximum reference picture size, etc. The limits set by the level may, in some cases, be further constrained through a Hypothetical Reference Decoder (HRD) specification and metadata for HRD buffer management signaled in the coded video sequence.
[0025] In embodiments, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may also be included as part of the coded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or to more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant pictures, forward error correction codes, etc.
[0026] 3 shows an example block diagram of a video encoder (303). The video encoder (303) is included in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., a transmitting circuit). The video encoder (303) can be used in place of the video encoder (303) in the example of FIG. 1.
[0027] The video encoder (303) may receive video samples from a video source (301) (which is not part of the electronic device (320) in the example of FIG. 3) that may capture video images to be coded by the video encoder (303). In other examples, the video source (301) is part of the electronic device (320).
[0028] The video source (301) may provide a source video sequence to be coded by the video encoder (303) in the form of a digital video sample stream, which may be of any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601 YCrCB, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media serving system, the video source (301) may be a storage device storing prepared video. In a video conferencing system, the video source (301) may be a camera capturing local image information as a video sequence. The video data may be provided as multiple individual pictures that, when viewed in sequence, impart motion. The pictures themselves may be organized as a spatial array of pixels, each of which may have one or more samples depending on the sampling structure, color space, etc., in use. Those skilled in the art will readily understand the relationship between pixels and samples. This specification will focus on samples hereafter.
[0029] According to an embodiment, the video encoder (303) may code and compress pictures of a source video sequence into a coded video sequence (343) in real time or under any other time constraints as needed. Imposing an appropriate coding rate is one function of the controller (350). In some embodiments, the controller (350) controls and is operatively coupled to other functional units, as described below. The coupling is not shown for clarity. Parameters set by the controller (350) may include parameters related to rate control (picture skip, quantizer, lambda value for rate-distortion optimization techniques, etc.), picture size, group-of-picture (GOP) layout, maximum motion vector search range, etc. The controller (350) may be configured with other appropriate functions related to the video encoder (303) optimized for a particular system design.
[0030] In some embodiments, the video encoder (303) is configured to operate in a coding loop. As an overly simplified description, in an example, the coding loop can include a source coder (330) (e.g., responsible for generating symbols, such as a symbol stream, based on an input picture to be coded and a reference picture) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to generate sample data in a manner similar to what a (remote) decoder would also generate. The reconstructed sample stream (sample data) is input to a reference picture memory (334). Because decoding of the symbol stream produces bit-exact results independent of the location (local or remote) of the decoder, the contents of the reference picture memory (334) are also bit-perfect between the local and remote encoders. In other words, the predictive portion of the encoder "sees" exactly the same sample values as the decoder would "see" when using prediction during decoding. This basic principle of reference picture synchronicity (and the resulting drift when synchronicity cannot be maintained, for example due to channel errors) is also used in several related techniques.
[0031] The operation of the "local" decoder (333) can be the same as a "remote" decoder, such as the video decoder (210), already described in detail above in conjunction with Figure 2. Referring also momentarily to Figure 2, however, the entropy decoding portion of the video decoder (210), including the buffer memory (215) and parser (220), may not be fully implemented in the local decoder (233), given the availability of symbols and the fact that the encoding / decoding of symbols into a coded video sequence by the entropy coder (345) and parser (220) can be lossless.
[0032] In embodiments, decoder techniques, with the exception of parsing / entropy decoding, present in a decoder are present in the corresponding encoder in the same or substantially the same functional form. Therefore, the disclosed subject matter focuses on the operation of the decoder. Descriptions of encoder techniques may be omitted, as they are the inverse of the decoder techniques described generically. To the extent specified, more detailed descriptions are provided below.
[0033] In operation, in some examples, the source coder (330) may perform motion-compensated predictive coding, which predictively codes an input picture with reference to one or more previously coded pictures from a video sequence designated as "reference pictures." In this manner, the coding engine (332) codes differences between pixel blocks of the input picture and pixel blocks of the reference pictures that may be selected as predictive references for the input picture.
[0034] The local video decoder (333) may decode coded video data of pictures that may be designated as reference pictures based on symbols generated by the source coder (330). The operation of the coding engine (332) may advantageously be a lossy process. When the coded video data is decoded by a video decoder (not shown in FIG. 3), the reconstructed video sequence is typically a copy of the source video sequence, with some errors. The local video decoder (333) may replicate the decoding process that may be performed by the video decoder on the reference pictures, causing the reconstructed reference pictures to be stored in the reference picture memory (334). In this way, the video encoder (303) may locally store copies of reconstructed reference pictures that have content in common with reconstructed reference pictures that would be obtained by a far-end video decoder (without transmission errors).
[0035] The predictor (335) may perform a prediction search for the coding engine (332). That is, for a new picture to be coded, the predictor (335) may search the reference picture memory (334) for specific metadata, such as reference picture motion vectors, block shapes, or sample data (as candidate reference pixel blocks) that can serve as suitable prediction references for the new picture. The predictor (335) may operate on a sample block-by-pixel block basis to find suitable prediction references. In some cases, as determined by the search results obtained by the predictor (335), the input picture may have prediction references drawn from multiple reference pictures stored in the reference picture memory (334).
[0036] The controller (350) may manage the coding operations of the source coder (330), including, for example, setting the parameters and subgroup parameters used to encode the video data.
[0037] The output of all of the above functional units may undergo entropy coding in an entropy coder (345), which converts the symbols produced by the various functional units into a coded video sequence by losslessly compressing the symbols according to techniques such as Huffman coding, variable length coding, or arithmetic coding.
[0038] The transmitter (340) may buffer the coded video sequence produced by the entropy coder (345) to prepare it for transmission over a communication channel (360), which can be a hardware / software link to a storage device that stores the coded video data. The transmitter (340) may also merge the coded video data from the video encoder (303) with other data to be transmitted, such as coded audio data and / or auxiliary data streams (sources not shown).
[0039] The controller (350) may manage the operation of the video encoder (303). During coding, the controller (350) may assign a particular coding picture type to each coded picture, which may affect the coding technique that may be applied to each picture. For example, pictures may often be assigned as one of the following picture types:
[0040] An Intra Picture (I-picture) can be one that can be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow various types of Intra pictures, including, for example, Independent Decoder Refresh (IDR) pictures. Those skilled in the art are aware of these variations of I-pictures and their respective applications and characteristics.
[0041] A Predictive Picture (P-picture) can be coded and decoded by intra-prediction or inter-prediction, which uses at most one motion vector and reference index to predict the sample values of each block.
[0042] Bi-directionally Predictive Pictures (B-pictures) can be those that can be coded and decoded using intra- or inter-prediction, which uses at most two motion vectors and reference indices to predict the sample values of each block. Similarly, multiple-predictive pictures can use more than two reference pictures and associated metadata for the reconstruction of a single block.
[0043] A source picture is generally spatially subdivided into multiple sample blocks (e.g., blocks of 4x4, 8x8, 4x8, or 16x16 samples, respectively) and may be coded block by block. Blocks may be predictively coded with reference to other (already coded) blocks as determined by the coding assignment applied to each picture of the blocks. For example, blocks of an I-picture may be non-predictively coded, or they may be predictively coded with reference to already coded blocks of the same picture (spatial prediction or intra prediction). Pixel blocks of a P-picture may be predictively coded by spatial prediction or temporal prediction with reference to one previously coded reference picture. Blocks of a B-picture may be predictively coded by spatial prediction or temporal prediction with reference to one or two previously coded reference pictures.
[0044] The video encoder (303) may perform coding operations according to a predetermined video coding technique or standard, such as ITU-T Recommendation H.265. During its operation, the video encoder (303) may perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancies in the input video sequence. Thus, the coded video data may conform to a syntax defined by the video coding technique or standard being used.
[0045] In embodiments, the transmitter (340) may transmit additional data along with the encoded video. The source coder (330) may include such data as part of the coded video sequence. The additional data may include temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, etc.
[0046] Video may be captured as multiple source pictures (video pictures) in a time sequence. Intra-picture prediction (often abbreviated as intra-prediction) exploits spatial correlation within a given picture, while inter-picture prediction exploits correlation (temporal or otherwise) between pictures. As an example, a particular picture being encoded / decoded, called the current picture, is partitioned into blocks. If a block in the current picture is similar to a reference block in a previously coded reference picture in the video that is still buffered, the block in the current picture may be coded by a vector called a motion vector. A motion vector points to a reference block within a reference picture and may have a third dimension that identifies the reference picture if multiple reference pictures are used.
[0047] In some embodiments, bi-prediction techniques may be used in inter-picture prediction. According to bi-prediction techniques, two reference pictures are used, e.g., a first reference picture and a second reference picture, both of which precede the current picture in decoding order (but may be past and future, respectively, in display order) in the video. A block in the current picture may be coded with a first motion vector that points to a first reference block in the first reference picture and a second motion vector that points to a second reference block in the second reference picture. The block is predictable by a combination of the first and second reference blocks.
[0048] Furthermore, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.
[0049] According to some embodiments of the present disclosure, prediction, such as inter-picture prediction and intra-picture prediction, is performed in units of blocks. For example, according to the HEVC standard, pictures in a sequence of video pictures are partitioned into coding tree units (CTUs) for compression, and the CTUs within a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU includes three coding tree blocks (CTBs), one luma CTB and two chroma CTBs. Each CTU may be recursively quadtree partitioned into one or more coding units (CUs). For example, a 64x64 pixel CTU can be divided into one CU of 64x64 pixels, four CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. By way of example, each CU is analyzed to determine a prediction type for the CU, such as an inter-prediction type or an intra-prediction type. A CU is divided into one or more prediction units (PUs) according to temporal and / or spatial predictability. Generally, each PU includes one luma prediction block (PB) and two chroma PBs. In an embodiment, prediction operations in coding (encoding / decoding) are performed in units of prediction blocks. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, etc.
[0050] It should be noted that the video encoders (103) and (303) and the video decoders (110) and (210) may be implemented using any suitable technology. In some embodiments, the video encoders (103) and (203) and the video decoders (110) and (210) may be implemented using one or more integrated circuits. In other embodiments, the video encoders (103) and (203) and the video decoders (110) and (210) may be implemented using one or more processors executing software instructions.
[0051] In an embodiment, inter-mode coding, such as that used in Codec Working Group (CWG)-18, is described below. In an example, as seen in, for example, AOMedia Video 1 (AV1), for each coded block in an interframe (or interpicture), if the mode of the current block is an inter-coding mode rather than a skip mode, a flag is signaled to indicate whether a single reference mode or a mixed reference mode is used to code the current block. In a single reference mode, a predictive block may be generated by a motion vector (MV) of a block (e.g., a current block). In a mixed reference mode, a predictive block may be generated by an average (e.g., weighted average) of multiple predictive blocks (or multiple reference blocks), such as two predictive blocks derived from two MVs.
[0052] In the case of a single reference using a single reference mode, the following modes may be signaled, including but not limited to NEARMV mode, NEWMV mode, GLOBALMV mode, and / or the like.
[0053] In NEARMV mode, an MVP among the Motion Vector Predictors (MVPs) in a list may be used. The MVP may be indicated by an index, such as a Dynamic Reference List (DRL) index. In an example, the MVP is used as the MV of the block.
[0054] In NEWMV mode, an MVP among the MVPs in the list may be used as a reference. The MVP may be signaled by an index, such as a DRL index. A delta (e.g., a delta MV) may be applied to the reference (e.g., the MVP indicated by the DRL index). In an example, the MV of a block is determined based on the delta and the MVP. For example, the delta may indicate the MV difference (MVD) between the MVP and the MV of the block.
[0055] In GLOBALMV mode, MVs based on frame-level global motion parameters may be used, for example, to predict blocks.
[0056] In the case of a mixed reference mode, the following modes may be signaled, including but not limited to NEAR_NEARMV mode, NEAR_NEWMV mode, NEW_NEARMV mode, NEW_NEWMV mode, GLOBAL_GLOBALMV mode, and / or the like.
[0057] In NEAR_NERAMV mode, an MVP among the MVPs in the list may be used. The MVP may be signaled by an index, such as a DRL index. In an example, an MVP among the MVPs in the list may be used to predict a block. For example, the MVP may indicate a first MV and a second MV of a block, and the first MV and the second MV of the block may be used to generate a first reference block and a second reference block, respectively, to predict the block.
[0058] In NEAR_NEWMV mode, an MVP among the MVPs in the list may be used as a reference. The MVP may be signaled by an index, such as a DRL index. A delta (e.g., a delta MV) may be applied to the reference (e.g., the MVP indicated by the DRL index) to determine the second MV of the block. In an example, the reference may be used directly as the first MV of the block, for example.
[0059] In NEW_NEARMV mode, an MVP among the MVPs in the list may be used as a reference. The MVP may be signaled by an index, such as a DRL index. A delta (e.g., a delta MV) may be applied to the reference (e.g., the MVP indicated by the DRL index) to determine the first MV of the block. In an example, the reference may be used directly as, for example, the second MV of the block.
[0060] In NEW_NEWMV mode, an MVP among the MPVs in the list may be used as a reference. The MVP may be signaled by an index, such as a DRL index. Two deltas (e.g., two delta MVs) may be applied to the reference (e.g., the MVP indicated by the DRL index) to determine the first MV of the block and the second MV of the block. In the example, two deltas are signaled.
[0061] In the GLOBAL_GLOBALMV mode, the MVs from each reference based on the global motion parameters of each frame level may be used to predict a block, for example.
[0062] In the present disclosure, the terms first and second components of a 2D vector (e.g., MV, MVD, delta MV, etc.) may refer to values of the 2D vector along a first and second axis, respectively. The first and second axes may be two predefined directions that are orthogonal to each other. In an example, the first and second axes are the x-axis and y-axis, and the first and second components are horizontal and vertical components of the 2D vector. In an example, the first and second axes are the 45-degree and 135-degree axes, and the first and second components are components of the 2D vector along the 45-degree and 135-degree axes, respectively.
[0063] The embodiments of the present disclosure may be applied to the first and second components of a 2D vector (e.g., MV, MVD, delta MV, etc.). In an example, the x-axis and y-axis may be replaced by two other axes along two predefined directions that are orthogonal to each other, and the disclosed embodiments for the horizontal and vertical components of a 2D parameter may also be applied to two other axes. For example, the x-axis and y-axis may be replaced by a 45-degree axis and a 135-degree axis.
[0064] Various methods may be used to code the MVD. An exemplary method, such as that used in AV1, is described below. In an embodiment, as in AV1, a motion vector precision (or accuracy) of 1 / 8 pixel may be used, and the following syntax (e.g., including but not limited to one or more of mv_joint, mv_sign, mv_class, mv_bit, mv_fr, mv_hp, and / or the like) may be used to signal or indicate the MVD in reference frame list 0 or reference frame list 1.
[0065] The syntax mv_joint can specify which components of the MVD (e.g., the first and second components) are non-zero. For example, a first value (e.g., 0) for the syntax mv_joint indicates that there is no non-zero MVD along either the first axis (e.g., the horizontal direction) or the second axis (e.g., the vertical direction), a second value (e.g., 1) for the syntax mv_joint indicates that there is a non-zero MVD only along the first axis (e.g., the horizontal direction), a third value (e.g., 2) for the syntax mv_joint indicates that there is a non-zero MVD only along the second axis (e.g., the vertical direction), and a fourth value (e.g., 3) for the syntax mv_joint indicates that there is a non-zero MVD along both the first and second axes (e.g., the horizontal and vertical directions).
[0066] The syntax mv_sign can specify whether the MVD is positive or negative. For example, the syntax mv_sign can be associated with each non-zero component of the MVD and can indicate the sign of each non-zero component of the MVD.
[0067] In an embodiment, the syntax elements mv_class, mv_bit, mv_fr, and mv_hp are used to indicate the magnitude of the non-zero components of the MVD. The integer part of the magnitude of the non-zero components of the MVD can be indicated by the syntax elements mv_class and mv_bit. The fractional part of the magnitude of the non-zero components of the MVD can be indicated by the syntax elements mv_fr and mv_hp.
[0068] The syntax mv_class (MV class) can specify the class (e.g., magnitude class) of the MVD. For example, the syntax mv_class can be associated with each non-zero component of the MVD and can indicate the class of each non-zero component of the MVD, such as that shown in FIG.
[0069] FIG. 4 illustrates the relationship between the class of non-zero components of the MVD (or MV class) and the magnitude range of the non-zero components of the MVD. In an embodiment, the higher the class, the larger the magnitude of the MVD (e.g., the non-zero components of the MVD). The first column of FIG. 4 illustrates various MV classes, e.g., MV class 0 (MV_CLASS_0) through MV class 10 (MV_CLASS_10). MV class 10 is higher than MV classes 0-9. The second column in Figure 4 shows the corresponding magnitude ranges, denoted by (m1, m2]. The value m1 is called the starting magnitude of the MV class corresponding to the magnitude range (m1, m2]. In the example shown in Figure 4, the starting magnitudes of the MV classes (MV class 0 (MV_CLASS_0) to MV class 10 (MV_CLASS_10)) are 0, 2, 4, 8, 16, 32, 64, 128, 256, 512, and 1024, respectively. For example, MV_CLASS_0 corresponds to a magnitude range (0,2) that is greater than 0 and less than or equal to 2 (e.g., m1=0, m2=2). The value m1 (e.g., 0) is the starting magnitude of MV_CLASS_0. If the syntax mv_class of the non-zero components of MVD is MV_CLASS_0, the magnitude of the non-zero components of MVD is greater than 0 and less than or equal to 2.
[0070] The syntax mv_bit can specify the integer portion of the offset between the MVD and the starting magnitude of each MV class. For example, a syntax mv_bit can be associated with each non-zero component of the MVD and can indicate the integer portion of the difference (or offset) between each non-zero component of the MVD and the starting magnitude of the corresponding MV class. For example, the non-zero components of the MVD have an MV class of MV_CLASS_10, and the syntax mv_bit associated with the non-zero components of the MVD indicates the integer portion of the difference between the non-zero components of the MVD and 1024.
[0071] The syntax mv_fr can specify the first two fractional bits of the MVD. For example, the syntax mv_fr can be associated with each non-zero component of the MVD and can indicate the first two fractional bits of each non-zero component of the MVD.
[0072] The syntax mv_hp may specify the third fractional bit of the MVD. For example, the syntax mv_hp may be associated with each non-zero component of the MVD and may indicate the third fractional bit of each non-zero component of the MVD.
[0073] Adaptive MVD resolution (AMVD) may be used, for example, as found in CWG-B092. In embodiments, for certain inter-prediction modes, such as NEW_NEARMV and NEAR_NEWMV modes, the accuracy of the MVD may depend on the associated class (e.g., MV class) and / or magnitude of the MVD. For example, the accuracy of a non-zero component of the MVD may depend on the associated class (e.g., MV class) and / or magnitude of the non-zero component of the MVD. In examples, fractional MVD is allowed only if the magnitude of the MVD is less than or equal to one pixel. For example, fractional parts of non-zero components of the MVD are allowed only if the magnitude of the non-zero component of the MVD is less than or equal to one pixel.
[0074] In the AMVD example, only one MVD value is allowed if the value of the associated MV class is greater than or equal to MV_CLASS_1. The MVD value for each MV class may be derived as 4, 8, 16, 32, or 64 for MV class 1 (MV_CLASS_1), MV class 2 (MV_CLASS_2), MV class 3 (MV_CLASS_3), MV class 4 (MV_CLASS_4), or MV class 5 (MV_CLASS_5). For example, the non-zero components of the MVD for each MV class may be derived as 4, 8, 16, 32, or 64 for MV class 1 (MV_CLASS_1), MV class 2 (MV_CLASS_2), MV class 3 (MV_CLASS_3), MV class 4 (MV_CLASS_4), or MV class 5 (MV_CLASS_5).
[0075] 5 illustrates adaptive MVD for each MV magnitude class according to an embodiment of the present disclosure. For example, FIG. 5 illustrates example MVD values allowed for each MV class. For example, if the MV class is MV_CLASS_0, the magnitude of the non-zero components of the MVD can be between 0 and 1 (excluding 0), 1, or 2. If the MV class is one of the MV classes including MV_CLASS_1 to MV_CLASS_10, the magnitude of the non-zero components of the MVD can be a single value such as 4, 8, 16, 32, 64, 128, 256, 512, 1024, or 2048.
[0076] In an example, if the current block is coded in NEW_NEARMV mode or NEAR_NEWMV mode, one context (e.g., for Context-Adaptive Binary Arithmetic Coding (CABAC)) is used to signal syntax such as syntax mv_joint or syntax mv_class, otherwise another context (e.g., for CABAC) may be used to signal syntax such as syntax mv_joint or syntax mv_class.
[0077] Joint MVD coding (JMVD) mode may be applied in some embodiments, as seen for example in CWG-B092.
[0078] A new inter-coding mode called JOINT_NEWMV mode may be applied to indicate whether MVDs of multiple reference lists (e.g., two reference lists) are jointly signaled. When the inter-prediction mode is equal to JOINT_NEWMV mode, MVDs of multiple reference lists (e.g., reference list 0 and reference list 1) may be jointly signaled. Thus, in an example of JOINT_NEWMV mode, only one MVD (also referred to as joint MVD or joint_mvd) is signaled and transmitted to the decoder, and delta MVs (e.g., MVDs) of each of the multiple reference lists (e.g., reference list 0 and reference list 1) are derived from the single MVD (e.g., joint_mvd). In the example, joint MVD information indicating the joint MVD (e.g., joint_mvd) is signaled, and delta MVs of the multiple reference lists are derived from the joint MVD information.
[0079] In an embodiment, the JOINT_NEWMV mode is one of the composite reference modes. For example, the JOINT_NEWMV mode is signaled together with other composite reference modes, such as the NEAR_NEARMV mode, the NEAR_NEWMV mode, the NEW_NEARMV mode, the NEW_NEWMV mode, the GLOBAL_GLOBALMV mode, and / or the like. In an embodiment, no additional context is added. In an example, the JOINT_NEWMV mode is signaled as one of the composite reference modes.
[0080] When JOINT_NEWMV mode is signaled and the Picture Order Count (POC) distances between multiple reference frames (e.g., two reference frames) and the current frame are different, the MVD may be scaled for reference frame list 0 or reference frame list 1 based on the POC distances. In an embodiment, the POC distance between the first reference frame in reference frame list 0 and the current frame is represented by parameter td0, and the POC distance between the second reference frame in reference frame list 1 and the current frame is represented by parameter td1. If td0 is greater than or equal to td1, joint_mvd may be used directly as the first MVD (represented by MVP0) for the first reference frame in reference frame list 0. The second MVD (represented by MVP1 or driven_mvd1) for the second reference frame in reference frame list 1 may be derived from joint_mvd, for example, based on equation (1):
number
[0081] Otherwise, if td1 is greater than or equal to td0, then joint_mvd may be used directly as the second MVD (e.g., MVP1) for the second reference frame in reference list 1. The first MVD (e.g., represented by MVP0 or derived_MVD0) for the first reference frame in reference frame list 0 may be derived from joint_mvd based on equation (2):
number
[0082] In an embodiment, the adaptive MVD decomposition (AMVD) may be modified or improved, for example as seen in CWG-C011.
[0083] A new inter-coding mode called AMVDMV mode may be added to the single reference mode. When the AMVDMV mode is selected, it indicates that AMVD is applied to signal MVD.
[0084] In an embodiment, a flag called amvd_flag may be added to the JOINT_NEWMV mode to indicate whether AMVD is applied to the joint MVD coding mode. When adaptive MVD decomposition (AMVD) is applied to the joint MVD coding mode, the joint MVD coding mode may be referred to as the joint AMVD coding mode. In the joint AMVD coding mode, joint MVDs for multiple reference frames (e.g., two reference frames) may be jointly signaled, and the accuracy of the joint MVD may be implicitly determined by the magnitude of the MVD. Otherwise, for example, when AMVD is not applied to the joint MVD coding mode, joint MVDs for multiple reference frames (e.g., two reference frames) may be jointly signaled, and MVD coding different from AMVD (e.g., as described with reference to FIG. 4) is applied.
[0085] Alternatively, a new inter-prediction mode within the mixed reference mode, called JOINT_AMVDNEWMV mode, is added to indicate that AMVD is applied to the joint MVD coding mode. In an embodiment, in JOINT_AMVDNEWMV mode, joint MVDs for multiple reference frames (e.g., two reference frames) may be signaled together, and the accuracy of the joint MVD may be implicitly determined by the magnitude of the MVDs.
[0086] In an embodiment, adaptive motion vector decomposition (AMVR) may be applied, for example as found in CWG-C012 and CWG-C020.
[0087] In an example, as seen in, e.g., CWG-C012, AMVR is used, where a total of seven MV precisions are supported (e.g., 8 pixels or 8 pels, 4 pels, 2 pels, 1 pel, 1 / 2 pel, 1 / 4 pel, and 1 / 8 pel). For each precision block, an encoder (e.g., an AOM Video Model (AVM) encoder) can look up the supported MV precisions, e.g., all supported precision values, and signal the selected precision (e.g., the best precision) to the decoder.
[0088] To reduce the runtime of the encoder, two precision sets may be supported, as shown in FIG. 6. FIG. 6 shows the supported MV precisions within the two precision sets. The number of MV precisions within each of the two precision sets (e.g., four) is less than the number of MV precisions among the supported precision values (e.g., seven). In an example, each precision set includes four predefined precisions. The precision set may be adaptively selected at the frame level based on the maximum precision value of the frame. In an example, the maximum precision may be signaled in the frame header, as used in AV1, for example. FIG. 6 shows supported precision values based on the maximum precision at the frame level, according to an embodiment of the present disclosure. If the maximum precision at the frame level is 1 / 8 pel, the precision set includes 1 / 8 pel, 1 / 2 pel, 1 pel, and 4 pel. If the maximum precision at the frame level is 1 / 4 pel, the precision set includes 1 / 4 pel, 1 pel, 1 pel, and 8 pel.
[0089] In an example, as found in current AVM software similar to AV1, a flag such as a frame-level flag (e.g., cur_frame_force_integer_mv flag) is used to indicate whether the MV of a frame can include sub-pel precision, such as ½ pel, ¼ pel, and ⅛ pel. In an example, AMVR is valid only if the frame-level flag (e.g., cur_frame_force_integer_mv flag) has a value of 0. In an example, in AMVR, if the precision of a block is less than maximum precision, the motion model and interpolation filter are not signaled. If the precision of a block is less than maximum precision, the motion mode may be inferred as translation motion, and the interpolation filter may be inferred as a REGULAR interpolation filter. Similarly, in an example, if the precision of a block is either 4 pels or 8 pels, the inter-intra mode may not be signaled and may be inferred to be 0.
[0090] In embodiments, the joint MVD coding mode may be changed or improved, for example, as seen in CWG-C053. The joint MVD coding mode may refer to the JOINT_NEWMV mode or the JOINT_AMVDNEWMV mode.
[0091] If the block is coded in a joint MVD coding mode such as JOINT_NEWMV mode or JOINT_AMVDNEWMV mode, syntax such as mvd_scaling_factor_idx (e.g., a scaling factor index) may be signaled in the bitstream to explicitly indicate, for example, the scaling factor between the MVDs associated with reference frame 0 and reference frame 1 (e.g., represented by jmvd_scale).
[0092] As shown in Figures 7-8, two lookup tables (e.g., two predefined lookup tables) may be used to store supported or allowed scaling factors for the JOINT_AMVDNEWMV mode and the JOINT_NEWMV mode, respectively. The associated entry index of a selected scaling factor in each lookup table (e.g., the lookup table of Figure 7 or Figure 8) may be signaled in the bitstream.
[0093] 7 shows exemplary scaling factors for JOINT_AMVDNEWMV mode. For JOINT_AMVDNEWMV mode, the same scaling factor may be applied to the first and second components (e.g., vertical and horizontal components) of the MVD of reference frame list 0 and / or reference frame list 1.
[0094] 8 shows exemplary scaling factors for the JOINT_NEWMV mode. In the JOINT_NEWMV mode, the scaling factor of one component of the MVD (e.g., the vertical or horizontal component) may be limited to a first value (e.g., 1), and the scaling factor of the other component of the MVD may be another value (e.g., 2, 1 / 2, etc.) different from the first value (e.g., 1).
[0095] The MVD of reference frame list 0 or reference frame list 1 may be calculated based on equations (3) and (4), respectively:
number
[0096] In an example, in JOINT_AMVDNEWMV mode, the scaling factor (e.g., jmvd_scale) is the same for the first and second components (e.g., vertical and horizontal components) of the MVD (e.g., mvd_ref1), and equations (3) and (4) may be applied to the MVD or each component of the MVD. For example, the first MVD (e.g., MVD0) of reference frame list 0 is indicated by mvd_ref0 and obtained directly from joint_mvd (e.g., signaled), and the second MVD of reference frame list 1 is indicated by mvd_ref1 and obtained by modifying joint_mvd with the ratio of td1 / td0 and the scaling factor (e.g., jmvd_scale) as described in equation (1). Equations (3) and (4) can be modified to determine the first MVD based on the scaling factor and the ratio of td0 / td1 as described in equation (2), and to determine the second MVD directly from joint_mvd.
[0097] In an example, in JOINT_NEWMV mode, the scaling factors (e.g., jmvd_scale) of the first and second components (e.g., vertical and horizontal components) of the MVD are different. Equations (3) and (4) may be applied to each component of the MVD (e.g., MVD0 or MVD1).
[0098] In an embodiment, a bi-prediction mode with CU-level weight (BCW) mode (also called combined weighted prediction) may be applied to a block. In an example, for example in HEVC, a bi-prediction signal is generated by averaging two prediction signals obtained from two different reference pictures and / or using two different MVs. In an example, for example in VVC, the bi-prediction mode may be extended beyond simple averaging to allow a weighted averaging of two prediction signals, for example as shown in equation (5):
number
[0099] In the example, the bi-predictive signal P bi-pred are bi-predictive samples, and the two predicted signals P0 and P1 are two predicted samples.
[0100] Referring to equation (5), the bi-predictive signal P bi-pred can be a weighted average of two prediction signals P0 and P1 with weights (8-w) and w, respectively. In the example, the sum of weights (8-w) and w is a constant (e.g., 8).
[0101] In an embodiment, a list of weighting factors (e.g., including five weights) is allowed in weighted average bi-prediction or BCW mode. In an example, the list of weighting factors includes five weights -2, 3, 4, 5, and 10, e.g., w∈{-2, 3, 4, 5, 10}. When w is equal to 4, an equal weighting factor (e.g., 4) can be used in the weighted average of two predicted samples. For example, the weighting factors of the two predicted signals P0 and P1 are equal (e.g., 4), and Equation (5) becomes P bi-pred =(4×P0+4×P1+4)>>3.
[0102] For each bi-predicted CU, the weight w can be determined in one of two ways: 1) for non-merged CUs, the weight index can be signaled after MVD; 2) for merged CUs, the weight index can be inferred from neighboring blocks based on merge candidate indexes.
[0103] In an example, the BCW mode is applied only to CUs with more than 256 luma samples (e.g., CU width × CU height ≧ 256). In an example, for low latency pictures, a list of weighting factors (including all five weights) is used. For non-low latency pictures, a subset of the list of weighting factors is used. In an example, the subset of the list of weighting factors includes only three weights in the list of weighting factors (e.g., w∈{3,4,5}). In an example, a low latency picture may refer to a picture whose coding order is the same as its display order, e.g., the pictures are coded sequentially.
[0104] Equation (5) is, for example, a bi-predictive signal P bi-pred can be adapted so that it can be the weighted average of two prediction signals P0 and P1 with weights (16-w) and w, respectively, as shown in equation (6). In the example, the sum of weights (w-16) and w is a constant (e.g., 16):
number
[0105] In an embodiment, for a joint MVD coding mode (JMVD), an inter-MVD scaling factor (e.g., jmvd_scale) associated with reference frame list 0 and reference frame list 1 is used, e.g., as described above in Figures 7-8 and equation (4). In an example, when applying BCW mode and JMVD mode to a block, e.g., when applying BCW mode to JMVD mode, five scaling factors, such as w∈{-2, 3, 4, 5, 10}, are used for each supported scaling factor, so a large number of bits need to be signaled and the signaling cost is relatively high.
[0106] This disclosure describes a set of advanced image and video coding techniques, including embodiments related to block-based weighting factors for joint MVD coding modes.
[0107] In this disclosure, the direction of the reference frame may be determined based on whether the reference frame is before the current frame in display order or after the current frame in display order.
[0108] In this disclosure, JMVD mode may refer to JMVD with normal full MV resolution (e.g., the highest resolution such as 1 / 4 pel or 1 / 8 pel) or JMVD with AMVR.
[0109] In some embodiments, when a joint MVD coding mode is applied to code a block in a frame (or the current frame), the scaling factors used in the joint MVD coding mode may be determined based on a list of allowed or supported scaling factors L0s. When a BCM mode is applied to code a block, the weighting factors used in the BCM mode may be determined based on a list of allowed or supported weighting factors L0w. In accordance with embodiments of the present disclosure, when both a joint MVD coding mode and a BCM mode are applied to code a block, the scaling factors used in the joint MVD coding mode may be determined based on a subset of the list of scaling factors L0s and / or the weighting factors used in the BCM mode may be determined based on a subset of the list of weighting factors L0w, e.g., to reduce signaling costs and improve coding efficiency.
[0110] According to an embodiment of the present disclosure, when a joint MVD coding mode is selected to code a block in a frame and a scaling factor of at least one MVD (e.g., mvd_ref1 in Equation (4)) component (e.g., the first and second components) associated with at least one reference frame (e.g., the first reference frame) of the block is a given value (e.g., also referred to as a predefined scaling factor), a list of weighting factors L0w may be supported to perform a weighted average of prediction samples from the block's reference frame (e.g., multiple reference frames). The block's reference frame (e.g., multiple reference frames) may include at least one reference frame of the block. Otherwise, when another scaling factor is selected for the joint MVD coding mode, a subset of the list of weighting factors L0w may be permitted. In an example, weighting factor information (e.g., a weighting factor index) may be signaled to indicate (i) a subset of the list of weighting factors L0w or (ii) which weighting factor in the list of weighting factors L0w may be selected to perform a weighted average of prediction samples. The weighted averaging of the prediction samples may be performed by the BCW mode, for example, as described in equation (5) or equation (6). The list of weighting factors L0w may be predefined. A subset of the list of weighting factors L0w may be predefined.
[0111] In an embodiment, a block in a frame is coded in a joint MVD coding mode and a BCW mode. The joint MVD coding mode may refer to a JOINT_NEWMV mode (e.g., with or without AMVR) or a JOINT_AMVDNEWMV mode (e.g., with or without AMVR). In the example, the reference frames include a first reference frame from reference frame list 0 and a second reference frame from reference frame list 1, and the MVD of the block includes a first MVD associated with the first reference frame (e.g., MVD0 or mvd_ref0) and a second MVD associated with the second reference frame (e.g., MVD1 or mvd_ref1).
[0112] In an example, the joint MVD coding mode is JOINT_AMVDNEWMV mode, and Equations (3) and (4) are used to determine a first MVD (e.g., mvd_ref0) and a second MVD (e.g., mvd_ref1), respectively. At least one reference frame may refer to a second reference frame, and at least one MVD may refer to the second MVD. Referring to FIG. 7, the scaling factors of the first and second components of the second MVD may be the same, such as 1, 2, or 1 / 2. In an example, the predefined scaling factor is 1. If the scaling factor of the component of the second MVD is equal to the predefined scaling factor 1 (corresponding to index 0 in FIG. 7), the weighting factor for the BCW mode may be determined based on a list of weighting factors L0w. Otherwise, the weighting factor for the BCW mode may be determined based on a subset of the list of weighting factors L0w. In an example, L0w includes five weights, such as -2, 3, 4, 5, and 10.
[0113] In the example, the joint MVD coding mode is JOINT_NEWMV mode, and Equation (3) and Equation (4) are used to determine the first MVD (e.g., mvd_ref0) and the second MVD (e.g., mvd_ref1), respectively. For example, the first MVD (e.g., mvd_ref0) is determined directly from joint_mvd, and each component of the second MVD (e.g., mvd_ref1) is determined using Equation (4) with its respective scaling coefficient (e.g., jmvd_scale). In the example, the predefined scaling coefficient is 1, and the scaling coefficients of the first and second components (e.g., horizontal and vertical components) are 1 and 2 (index of 1 in FIG. 8 ). Therefore, the scaling coefficients of the components of the second MVD are different from the predefined scaling coefficient (e.g., 1) (e.g., 2 for the vertical component), and the weighting coefficients of the BCW may be determined based on a subset of the list of weighting coefficients L0w. If the scaling coefficients of the first and second components (e.g., vertical and horizontal components) are equal to 1 (corresponding to index 0 in FIG. 8), the weighting coefficients for the BCW mode may be determined based on the list of weighting coefficients L0w.
[0114] Motion information associated with each reference frame of the block may be determined using a joint MVD coding mode based on scaling coefficients of components of at least one MVD associated with at least one of the reference frames of the block. For example, the motion information includes a first MV associated with a first reference frame and a second MV associated with a second reference frame. The first MV may be determined based on the first MVD (e.g., using Equation (3)) and the first MVP, and the second MV may be determined based on the second MVD (e.g., using Equation (4)) and the second MVP. The first MVP and the second MVP may be determined based on an index (e.g., a DRL index) and a list of MVPs. In an example, an entry in the list of MVPs indicates the first MVP and the second MVP.
[0115] The blocks may be reconstructed using the BCW mode based on the motion information associated with each reference frame of the block and the determined weighting factor w. For example, a first reference block (or a first predicted block) (e.g., including sample values indicated by P0) is determined based on a first MV, and a second reference block (or a second predicted block) (e.g., including sample values indicated by P1) is determined based on a second MV. bi-pred The sample values denoted by ( ) can be determined by the BCW mode (eg, using equations (5) or (6)).
[0116] In an embodiment, if the scaling coefficient of either the first or second component of one of at least one MVD associated with at least one reference frame of the block is not equal to a given value (e.g., 1) for the joint MVD coding mode, only one weighting factor is allowed in the BCW mode. The one weighting factor may be predefined. In an example, signaling of the one weighting factor is not required. If the given value is 1, only one weighting factor is allowed if the index is 1 or 2 in FIG. 7 (e.g., JOINT_AMVDNEWMV mode), and only one weighting factor is allowed if the index is 1, 2, 3, or 4 in FIG. 8 (e.g., JOINT_NEWMV mode). In an example, if the scaling factor of either the first or second component of one of the at least one MVD associated with at least one reference frame of the block is not equal to a given value (e.g., 1) for the joint MVD coding mode, only equal weighting factors (e.g., equal weighting factors are 4 in equation (5) or 8 in equation (6)) are allowed. The equal weighting factors may be predefined. In an example, signaling of equal weighting factors is not required.
[0117] In an embodiment, if the scaling coefficient of either the first or second component of one of at least one MVD associated with at least one reference frame of a block is not equal to a given value (e.g., 1) for a joint MVD coding mode, a predefined lookup table may be used to determine an associated weighting coefficient for weighted averaging of multiple predicted samples. FIG. 9 illustrates an example of a predefined lookup table according to an embodiment of the present disclosure. The first column of the lookup table indicates a pair of scaling coefficients (two scaling coefficients) for the first and second components of the MVD. The second column of the lookup table indicates a weighting coefficient corresponding to each scaling coefficient included in the first column. The weighting coefficients in the second column may be 3 and / or 5 if the sum of the weights for averaging the two predicted samples is 8 (e.g., Equation (5)). The weighting coefficients in the second column may be 6 and / or 10 if the sum of the weights for averaging the two predicted samples is 16 (e.g., Equation (6)).
[0118] In an embodiment, if the scaling coefficient of either the first or second component of one of at least one MVD associated with at least one reference frame of the block is not equal to a given value (e.g., 1) for the joint MVD coding mode, unequal weighting factors may be used in the weighted averaging of multiple predicted samples. Equal weighting factors are not allowed. In an example, a subset of the list of weighting factors includes only unequal weighting factors. For example, only one or more sets of unequal weighting factors may be used in the weighted averaging of multiple predicted samples (e.g., two predicted samples). Each set of unequal weighting factors may be predefined and may be used for each scaling coefficient pair of the first and second components of at least one MVD. In an embodiment, only two unequal weighting factors may be used to perform the weighted averaging of two predicted samples. The two unequal weighting factors may be, for example, 3 and 5 when the sum of the weights for averaging two predicted samples is 8 (e.g., Equation (5)). The two unequal weighting factors can be, for example, 6 and 10, where the sum of the weights for averaging the two predicted samples is 16 (eg, equation (6)).
[0119] A joint MVD coding mode is selected for coding a block. In accordance with an embodiment of the present disclosure, if equal weighting factors are applied to each reference block of the block (e.g., equal weighting factors are used to average prediction samples from multiple reference frames), then a list of scaling coefficients L0s may be permitted for scaling at least one MVD component associated with at least one reference frame (e.g., reference frame 0 and / or reference frame 1). Otherwise, if unequal weighting factors are used to average prediction samples from multiple reference frames, then only a subset of the list of scaling coefficients L0s may be permitted. In an example, scaling coefficient information (e.g., a scaling coefficient index) may be signaled to indicate (i) a subset of the list of scaling coefficients L0s or (ii) which scaling coefficients in the list of scaling coefficients L0s may be used in the joint MVD coding mode.
[0120] In an embodiment, if unequal weighting factors are used to average prediction samples from multiple reference frames, the scaling factors used to scale the first and second components (e.g., vertical and horizontal components) of at least one MVD may be restricted to a given predefined value (e.g., 1).
[0121] In an embodiment, if unequal weighting factors are used to average prediction samples from multiple reference frames, the scaling factors used to scale the first and second components (e.g., vertical and horizontal components) of at least one MVD are not allowed to be a given predefined value (e.g., 1).
[0122] In an embodiment, if a joint MVD coding mode is selected for coding a block, for example, to blend (e.g., weighted average) multiple prediction samples, the context (e.g., for CABAC) for signaling a weighting factor index may depend on scaling factor information, such as a scaling factor of at least one MVD component associated with at least one reference frame. The weighting factor index may indicate (i) a list of weighting factors L0w or (ii) a weight within a subset of the list of weighting factors L0w.
[0123] In an embodiment, the context for signaling the weighting factor index is a first context if the scaling factor information indicates that each of the scaling factors of the at least one component of the MVD is a given value (or a predefined scaling factor such as 1). Otherwise, if the scaling factor information indicates that at least one of the scaling factors is different from the given value, the context for signaling the weighting factor index is another set of contexts different from the first context (e.g., a second context).
[0124] In an embodiment, the context for signaling the weighting factor index may depend on whether the block is coded in a joint MVD coding mode. In an example, the weighting factor index may indicate a weighting factor used to blend two prediction samples, e.g., by a weighted average of the two prediction samples.
[0125] In an embodiment, a high-level syntax may be signaled at a high level, such as a sequence level, a frame level, a slice level, a superblock level, a CTU level, etc., to indicate whether all supported weighting factors (e.g., a list of weighting factors L0w) or a subset of weighting factors (e.g., a subset of the list of weighting factors L0w) are allowed for a joint MVD coding mode when one of the scaling factors of at least one MVD component is not equal to a predefined value, such as 1.
[0126] In an embodiment, the context (e.g., for CABAC) for signaling scaling coefficient information (e.g., scaling coefficient index) indicating the scaling coefficient of at least one MVD component in JMVD mode depends on weighting coefficient information indicating the weighting coefficient index.
[0127] 10 shows a flowchart illustrating a process (1000) according to an embodiment of the present disclosure. The process (1000) may be used in a video / image decoder. In various embodiments, the process (1000) is performed by a processing circuit, such as a processing circuit performing the functions of the video decoder (110), a processing circuit performing the functions of the video decoder (210), etc. In some embodiments, the process (1000) is implemented by software instructions, such that the processing circuit performs the process (1000) when the processing circuit executes the software instructions. The process begins at (S1001) and proceeds to (S1010).
[0128] At (S1010), a bitstream including a frame (also called a picture) may be received. Coding information for a block in the frame may indicate that the block is coded in a joint MVD coding mode and a composite weighted prediction mode (or BCW mode). The coding information may further include scaling factor information for the joint MVD coding mode.
[0129] At (S1020), it may be determined whether the scaling factor information indicates that each of the scaling factors of the components of at least one MVD associated with at least one respective reference frame of the block is a predefined scaling factor (e.g., 1). If the scaling factor information indicates that each of the scaling factors of the components of at least one MVD associated with at least one respective reference frame of the block is a predefined scaling factor, the process (1000) proceeds to (S1030). Otherwise, if the scaling factor information indicates that at least one of the scaling factors is different from the predefined scaling factor, the process (1000) proceeds to (S1040).
[0130] At (S1030), a weighting factor for the BCW mode may be determined based on a list of weighting factors, such as L0w. In an example, the list of weighting factors is predefined and not signaled in the bitstream. In an example, the list is signaled in the bitstream. Process (1000) proceeds to (S1050).
[0131] At (S1040), weighting factors for the BCW mode may be determined based on a subset of the list of weighting factors. In an example, the subset of the list is predefined and not signaled in the bitstream. In an example, the subset of the list is signaled in the bitstream. Process (1000) proceeds to (S1050).
[0132] At (S1050), motion information associated with each reference frame of the block may be determined based on a scaling factor of at least one MVD component associated with at least one respective reference frame of the block using a joint MVD coding mode. The reference frames of the block may include at least one respective reference frame. The process then proceeds to (S1099) and ends.
[0133] Process 1000 may be adapted as appropriate. Steps of process 1000 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.
[0134] In an example, the block may be reconstructed using the BCW mode based on the determined weighting factors and motion information associated with each reference frame of the block.
[0135] In an example, steps (S1020), (S1030), and (S1040) may be adapted as follows: If the scaling coefficient information indicates that each of the scaling coefficients of the at least one component of the MVD associated with at least one respective reference frame of the block is a predefined scaling coefficient, the weighting coefficients for the BCW mode may be determined based on the list of weighting coefficients; otherwise, if the scaling coefficient information indicates that at least one of the scaling coefficients is different from the predefined scaling coefficient, the weighting coefficients for the BCW mode may be determined based on a subset of the list of weighting coefficients.
[0136] In an embodiment, the scaling factor information indicates that at least one of the scaling factors is different from a predefined scaling factor, the subset of the list of weighting factors includes only one weighting factor, and the determined weighting factor is that one weighting factor and is not signaled.
[0137] In an example, one weighting factor is an equal weighting factor. As described with reference to Equation (4) when the equal weighting factor is 4, a prediction block of a block is obtained by averaging the reference blocks in each reference frame with an equal weighting factor for each reference block. The block can be reconstructed based on the prediction block.
[0138] In an example, the weighting factor is determined based on a scaling factor of a component of at least one MVD, as described with reference to FIG. 9 . For example, the weighting factor is determined based on a predefined relationship (e.g., the lookup table of FIG. 9 ) between a scaling factor (e.g., the scaling factor pair of FIG. 9 ) and a weighting factor (e.g., the weighting factor of FIG. 9 ) of one component of the at least one MVD. In an example, the reference frame includes a first reference frame having a first weighting factor and a second reference frame having a second weighting factor, and the sum of the first weighting factor and the second weighting factor is a constant. If the constant is 8, the weighting factor is 3 or 5, and the first weighting factor and the second weighting factor include 3 and 5. If the constant is 16, the weighting factor is 6 or 10, and the first weighting factor and the second weighting factor include 6 and 10. A prediction block of a block can be obtained by averaging the first reference frame having the first weighting factor and the second reference frame having the second weighting factor, and the block can be reconstructed based on the prediction block.
[0139] In an embodiment, the subset of the list of weighting factors includes only unequal weighting factors. The weighting factors may be determined as one of the unequal weighting factors. A prediction block of a block may be obtained by averaging reference blocks in each reference frame with each unequal weighting factor determined based on one of the unequal weighting factors as shown in Equation (5) if the weighting factor is not 4. In an example, the unequal weighting factors are two unequal weighting factors including the weighting factor, the reference blocks include a first reference block and a second reference block, and the prediction block is obtained by averaging the first reference block and the second reference block with the two unequal weighting factors, respectively. In an example, the two unequal weighting factors include (i) 3 and 5, or (ii) 6 and 10.
[0140] In an embodiment, the context for signaling the weighting factor index (e.g., for CABAC) depends on the scaling factor information: the weighting factor may be determined based on (i) the weighting factor index, and (ii) the list of weighting factors or a subset of the list of weighting factors.
[0141] In an example, the context for signaling the weighting factor index is a first context based on scaling factor information indicating that each of the scaling factors of the at least one MVD component is a predefined scaling factor, and the context for signaling the weighting factor index is a second context different from the first context based on scaling factor information indicating that at least one of the scaling factors is different from the predefined scaling factor.
[0142] 11 shows a flowchart illustrating a process (1100) according to an embodiment of the present disclosure. The process (1100) may be used in a video encoder. In various embodiments, the process (1100) is performed by a processing circuit, such as a processing circuit performing the functions of the video encoder (103), a processing circuit performing the functions of the video encoder (303), etc. In some embodiments, the process (1100) is implemented by software instructions, such that the processing circuit performs the process (1100) when the processing circuit executes the software instructions. The process begins at (S1101) and proceeds to (S1110).
[0143] At (S1110), it may be determined whether each of the scaling coefficients of at least one MVD component associated with each reference frame of at least one of the blocks is a predefined scaling coefficient. The scaling coefficients may be used in a joint MVD coding mode to encode the block within the frame. Furthermore, a BCW mode may be applied to encode the block. If it is determined that each of the scaling coefficients of at least one MVD component associated with each reference frame of at least one of the blocks is a predefined scaling coefficient (e.g., 1), the process (1100) proceeds to (S1120). Otherwise, at least one of the scaling coefficients is different from the predefined scaling coefficient, and the process (1100) proceeds to (S1130).
[0144] At (S1120), a weighting factor for the BCW mode may be determined based on a list of weighting factors such as L0w. The process (1100) proceeds to (S1140).
[0145] At step S1130, weighting factors for the BCW mode may be determined based on a subset of the list of weighting factors. The process 1100 then proceeds to step S1140.
[0146] At (S1140), the block may be coded based on the BCW mode. For example, the block may be predicted using equation (5) or equation (6).
[0147] The scaling factor information for the joint MVD coding mode may be coded to indicate scaling factors. The scaling factor information may indicate scaling factors for at least one MVD component associated with at least one respective reference frame of the block. In an example, the coded scaling factor information is included in a video bitstream.
[0148] Then, the process (1100) proceeds to (S1199) and ends.
[0149] Process 1100 may be adapted as appropriate. Steps of process 1100 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.
[0150] In an example, prior to (S1110), a scaling factor for a component of at least one MVD associated with at least one respective reference frame may be determined. Motion information associated with each reference frame of the block may be determined. The reference frames of the block include at least one respective reference frame. At least one MVD associated with at least one respective reference frame of the block may be determined based on the motion information. A scaling factor for a component of at least one MVD associated with at least one respective reference frame may be determined, for example, based on the at least one MVD. In an example, the scaling factor is further determined based on a POC difference between each reference frame and a current frame including the block.
[0151] The various embodiments described for the process (1000) can be applied to the process (1100) used in the encoding process.
[0152] 12 shows a flowchart illustrating a process (1200) according to an embodiment of the present disclosure. The process (1200) may be used in a video / image decoder. In various embodiments, the process (1200) is performed by a processing circuit, such as a processing circuit performing the functions of the video decoder (110), a processing circuit performing the functions of the video decoder (210), etc. In some embodiments, the process (1200) is implemented by software instructions, such that the processing circuit performs the process (1200) when the processing circuit executes the software instructions. The process begins at (S1201) and proceeds to (S1210).
[0153] At (S1210), a bitstream including a frame may be received. Coding information for a block in the frame may indicate that the block is coded in a joint MVD coding mode and a BCW mode. The coding information may further include weighting factor information for the BCW mode.
[0154] At (S1220), it may be determined whether the weighting factor information indicates that equal weighting factors are applied to each of the block's reference frames. If it is determined that the weighting factor information indicates that equal weighting factors are applied to each of the block's reference frames, the process (1200) proceeds to (S1230). Otherwise, if it is determined that the weighting factor information indicates that unequal weighting factors are applied to the block's reference frames, the process (1200) proceeds to (S1240).
[0155] At (S1230), a scaling factor for at least one component of the MVD associated with each of the at least one reference frame may be determined based on a list of scaling factors, such as L0s. The process (1200) proceeds to (S1250).
[0156] At step S1240, a scaling factor for at least one component of the MVD associated with each of the at least one reference frame may be determined based on a subset of the list of scaling factors. The process 1200 then proceeds to step S1250.
[0157] At (S1250), motion information associated with each reference frame of the block may be determined based on the determined scaling coefficients using a joint MVD coding mode.
[0158] At (S1260), the blocks may be reconstructed using BCW mode based on the motion information and weighting factor information associated with each of the block's reference frames. The process (1200) then proceeds to (S1299) and ends.
[0159] Process 1200 may be adapted as appropriate. Steps of process 1200 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.
[0160] In an example, steps (S1220), (S1230), and (S1240) may be adapted as follows: if the weighting factor information indicates that equal weighting factors are applied to each of the reference frames of the block, scaling factors for the components of at least one MVD associated with each of the at least one reference frame within the reference frame may be determined based on the list of scaling factors; if the weighting factor information indicates that unequal weighting factors are applied to each of the reference frames of the block, scaling factors for the components of at least one MVD associated with each of the at least one reference frame within the reference frame may be determined based on a subset of the list of scaling factors.
[0161] In an embodiment, the weighting factor information indicates that unequal weighting factors are applied to the reference frames of the block, the subset of the list of scaling factors includes only predefined scaling factors, and the determined scaling factor is equal to the predefined scaling factor and is not signaled. In an example, the predefined scaling factor is 1.
[0162] In an embodiment, the weighting factors indicate that unequal weighting factors are applied to the reference frames of the block, and the subset of the list of scaling factors does not include a predefined scaling factor (e.g., 1).
[0163] In an embodiment, the context for signaling the scaling factor index (e.g., for CABAC) depends on the weighting factor information: the scaling factor may be determined based on (i) the scaling factor index, and (ii) a list of scaling factors or a subset of the list of scaling factors.
[0164] 13 shows a flowchart illustrating a process (1300) according to an embodiment of the present disclosure. The process (1300) may be used in a video encoder. In various embodiments, the process (1300) is performed by a processing circuit, such as a processing circuit performing the functions of the video encoder (103), a processing circuit performing the functions of the video encoder (303), etc. In some embodiments, the process (1300) is implemented by software instructions, such that the processing circuit performs the process (1300) when the processing circuit executes the software instructions. The process begins at (S1301) and proceeds to (S1310).
[0165] At (S1310), it may be determined whether equal weighting factors are applied to each of the block's reference frames within a frame (or current frame). The block is to be coded in BCW mode. If it is determined that equal weighting factors are applied to each of the block's reference frames to code the block, the process (1300) proceeds to (S1320). Otherwise, unequal weighting factors are applied to the block's reference frames to code the block, and the process (1300) proceeds to (S1330).
[0166] At (S1320), a scaling factor for at least one MVD component associated with at least one each of the reference frames may be determined based on a list of scaling factors (e.g., L0s). The scaling factor is used in the joint MVD coding mode. The process (1300) proceeds to (S1340).
[0167] At step S1330, a scaling factor for at least one component of the MVD associated with each of the at least one reference frame may be determined based on a subset of the list of scaling factors. The process 1300 then proceeds to step S1340.
[0168] At (S1340), the block may be coded based on the BCW mode. For example, the block may be predicted using equation (5) or equation (6).
[0169] The weighting factor information for the BCW mode may be coded to indicate the weighting factors (e.g., equal or unequal weighting factors). In an example, the coded weighting factor information is included in the video bitstream.
[0170] The process (1300) then proceeds to (S1399) and ends.
[0171] Process 1300 may be adapted as appropriate. Steps of process 1300 may be modified and / or omitted. Additional steps may be added. Any suitable order of performance may be used.
[0172] In an example, prior to (S1310), the weighting factors used in BCW mode to encode the block may be determined based on a predefined list of weighting factors, such as L0w or a subset of L0w.
[0173] The various embodiments described for the process (1200) may be applied to the process (1300) used in the encoding process.
[0174] The embodiments of the present disclosure may be used separately or combined in any order. Furthermore, each of the method (or embodiment), encoder, and decoder may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, the one or more processors execute a program stored on a non-transitory computer-readable medium. In the present disclosure, the term block may be interpreted as a prediction block, a coding block, a coding unit (CU), etc.
[0175] The techniques described above can be implemented as computer software using computer-readable instructions and physically stored on one or more computer-readable media. For example, Figure 14 illustrates a computer system (1400) suitable for implementing certain embodiments of the disclosed subject matter.
[0176] The computer software may be coded in any suitable machine code or computer language that may be subject to mechanisms such as assembly, compilation, linking, etc. to produce code containing instructions that may be executed by one or more central processing units (CPUs), graphics processing units (GPUs), etc., directly or through interpretation, microcode execution, etc.
[0177] The instructions may be executable by various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming consoles, Internet of Things devices, and the like.
[0178] 14 for computer system 1400 are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure. The arrangement of components should not be interpreted as having any dependency or requirement regarding any one or combination of components described in the exemplary embodiment of computer system 1400.
[0179] The computer system 1400 may include certain human interface input devices. Such human interface input devices may respond to input by one or more users through, for example, tactile input (e.g., keyboard, swipe, dataglove motion), audio input (e.g., voice, claps), visual input (e.g., gestures), or olfactory input (not shown). The human interface devices may also be used to capture certain media that are not necessarily directly related to conscious human input, such as audio (e.g., speech, music, ambient sounds), images (e.g., scanned images, photographic images obtained from a still camera), and video (e.g., two-dimensional video, three-dimensional video, including stereoscopic video).
[0180] The input human interface devices may include one or more of a keyboard (1401), a mouse (1402), a trackpad (1403), a touchscreen (1410), a data glove (not shown), a joystick (1405), a microphone (1406), a scanner (1407), and a camera (1408) (only one of each is shown).
[0181] The computer system 1400 may also include certain human interface output devices that may stimulate one or more of the user's senses through, for example, tactile output, sound, light, and smell / taste. Such human interface output devices may include haptic output devices (e.g., haptic feedback via a touchscreen (1410), data gloves (not shown), or joystick (1405), although haptic feedback devices that do not function as input devices may also exist), audio output devices (e.g., speakers (1409), headphones (not shown)), visual output devices (e.g., screens (1410) including CRT screens, LCD screens, plasma screens, OLED screens, each with or without touchscreen input capability, each with or without haptic feedback capability, some of which are capable of outputting two-dimensional visual output or output in more than three dimensions by means of stereoscopic output, virtual reality glasses (not shown), holographic displays, and smoke tanks (not shown)), and printers (not shown).
[0182] The computer system (1400) may also include human-accessible storage devices and their associated media, such as CD / DVD or similar media (1421), CD / DVD ROM / RW (1420), including thumb drives (1422), removable hard disks or solid state drives (1423), legacy magnetic media, such as tape and floppy disks (not shown), dedicated ROM / ASIC / PLD-based devices, such as security dongles (not shown), and the like.
[0183] Those skilled in the art will also understand that the term "computer-readable medium" as used in connection with the presently disclosed subject matter does not include transmission media, carrier waves, or other transitory signals.
[0184] The computer system 1400 may also include interfaces 1454 to one or more communications networks 1455. Networks may be, for example, wireless, wireline, or optical. Networks may also be local, wide-area, metropolitan, vehicular, and industrial, real-time, delay-tolerant, and the like. Examples of networks include local area networks such as Ethernet; wireless LANs; cellular networks including GSM, 3G, 4G, 5G, LTE, and the like; TV wireline or wireless wide-area digital networks including cable TV, satellite TV, and terrestrial broadcast TV; and vehicle and factory networks including CAN bus. Certain networks generally require an external network interface adapter attached to a particular general-purpose digital port or peripheral bus 1449 (e.g., a USB port on the computer system 1400). Others are generally integrated into the core of the computer system 1400 by attachment to a system bus as described below (e.g., an Ethernet network interface to a PC computer system or a cellular network interface to a smartphone computer system). Using any of these networks, the computer system (1400) can communicate with other entities. Such communication can be one-way receive-only (e.g., broadcast TV) or one-way transmit-only (e.g., a CAN bus to a specific CAN bus device), or it can be two-way to other computer systems, for example, using a local or wide-area digital network. Specific protocols or protocol stacks can be used with each of the networks and network interfaces described above.
[0185] The above-mentioned human interface devices, human-accessible storage devices, and network interfaces may be attached to the core 1440 of the computer system 1400 .
[0186] The core (1440) may include one or more central processing units (CPUs) (1441), graphics processing units (GPUs) (1442), dedicated programmable processing units in the form of field programmable gate arrays (FPGAs) (1443), task-specific hardware accelerators (1444), graphics adapters (1450), etc. These devices may be connected through a system bus (1448), along with read-only memory (ROM) (1445), random access memory (RAM) (1446), internal mass storage devices such as internal non-user-accessible hard drives, SSDs, etc. (1447). In some computer systems, the system bus (1448) may be accessible in the form of one or more physical plugs to allow expansion with additional CPUs, GPUs, etc. Peripheral devices may be attached to the core's system bus (1448) directly or through a peripheral bus (1449). In an example, a display 1410 may be connected to a graphics adapter 1450. Architectures for peripheral buses include PCI, USB, and the like.
[0187] The CPU (1441), GPU (1442), FPGA (1443), and accelerator (1444) can execute specific instructions that, in combination, can constitute the above-mentioned computer code. The computer code can be stored in ROM (1445) or RAM (1446). Temporary data can also be stored in RAM (1446), while persistent data can be stored, for example, in an internal mass storage device (1447). Rapid storage and retrieval from any of the memory devices can be enabled through the use of cache memory. Cache memory can be closely associated with one or more of the CPU (1441), GPU (1442), mass storage device (1447), ROM (1445), RAM (1446), etc.
[0188] The computer-readable medium can bear computer code for performing various computer-implemented operations. The medium and computer code can be those specially designed and constructed for the purposes of the present disclosure, or they can be of the kind well known and available to those skilled in the computer software arts.
[0189] By way of example, and not limitation, a computer system having the architecture (1400), and in particular the core (1440), can provide functionality as a result of a processor (including a CPU, GPU, FPGA, accelerator, etc.) executing software embodied in one or more tangible computer-readable media. Such computer-readable media can be media associated with the user-accessible mass storage devices previously introduced, in addition to specific storage of the core (1440) that is non-transitory in nature, such as the core's internal mass storage device (1447) or ROM (1445). Software implementing various embodiments of the present disclosure can be stored on such devices and executable by the core (1440). The computer-readable media can include one or more memory devices or chips, depending on particular needs. Software can cause the cores (1440), and specifically the processors therein (including CPUs, GPUs, FPGAs, etc.), to perform particular processes or portions of particular processes described herein, including defining data structures stored in RAM (1446) and modifying such data structures according to software-defined processes. Additionally or alternatively, the computer system can provide functionality as a result of hardwired or otherwise embodied logic in circuitry (e.g., accelerators (1444)) that can operate in place of or in conjunction with software to perform particular processes or portions of particular processes described herein. References to software can encompass logic, where appropriate, and vice versa. References to computer-readable media can encompass circuitry (e.g., integrated circuits (ICs)) storing software for execution, circuitry embodying logic for execution, or both, where appropriate. The present disclosure encompasses any appropriate combination of hardware and software.
[0190] The use of "at least one of" or "one of" in this disclosure is intended to include any one or combination of the listed elements. For example, reference to at least one of A, B, or C; at least one of A, B, and C; at least one of A, B, and / or C; and at least one of A through C is intended to include A only, B only, C only, or any combination thereof. Reference to one of A or B, and one of A and B is intended to include A or B or (A and B). The use of "one of" does not exclude any combination of the listed elements, where applicable, such as when the elements are not mutually exclusive.
[0191] While this disclosure has described several exemplary embodiments, alternatives, permutations, and various substitute equivalents exist and are included within the scope of this disclosure. Thus, it will be apparent to those skilled in the art that numerous systems and methods, although not explicitly shown or described herein, embody the principles of the present disclosure and are therefore within its spirit and scope.
[0192] [Incorporated by reference] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 451,807, filed March 13, 2023, and entitled "BLOCK BASED WEIGHTING FACTOR FOR JOINT MOTION VECTOR DIFFERENCE CODING MODE," which in turn claims the benefit of priority to U.S. Provisional Patent Application No. 18 / 215,704, filed June 28, 2023, and entitled "BLOCK BASED WEIGHTING FACTOR FOR JOINT MOTION VECTOR DIFFERENCE CODING MODE," the disclosures of which are incorporated herein by reference in their entirety.
Claims
1. 1. A method of video decoding performed by a video coder, comprising: receiving a bitstream including a frame, wherein coding information of a block in the frame indicates that the block is coded in a joint motion vector differential (JMVD) coding mode and a combined weighted prediction mode, the coding information further including scaling factor information for the JMVD coding mode; determining weighting factors for the composite weighted prediction mode based on a list of weighting factors signaled in the bitstream in response to the scaling factor information indicating that each of the scaling factors for at least one MVD component associated with at least one respective reference frame of the block is a predefined scaling factor; determining weighting factors for the composite weighted prediction mode based on a subset of the list of weighting factors in response to the scaling factor information indicating that at least one of the scaling factors is different from the predefined scaling factor, where only the subset of the list was signaled in the bitstream; using the JMVD coding mode to determine motion information associated with each reference frame of the blocks based on scaling factors of the at least one MVD component associated with each reference frame of the at least one blocks, the reference frames of the blocks including each of the at least one reference frame; reconstructing the block using the combined weighted prediction mode based on the motion information associated with the respective reference frames of the block and the determined weighting factors; A method having the following.
2. the scaling factor information indicates that the at least one of the scaling factors is different from the predefined scaling factor; the subset of the list of weighting factors includes only one weighting factor; the determined weighting factor is the one weighting factor and is not signaled; The method of claim 1.
3. the one weighting factor is an equal weighting factor; The reconstructing step includes: obtaining a prediction block of the block by averaging reference blocks in each of the reference frames with the equal weighting factors for each reference block; and reconstructing the block based on the predicted block. The method of claim 2.
4. determining the weighting factors includes determining the weighting factors based on scaling factors of the components of the at least one MVD; The method of claim 2.
5. determining the weighting factor includes determining the weighting factor based on a predefined relationship between a scaling factor of one component of the at least one MVD and the weighting factor; The method of claim 4.
6. the reference frames include a first reference frame having a first weighting factor and a second reference frame having a second weighting factor, the sum of the first weighting factor and the second weighting factor being a constant; In response to the constant being 8, the weighting factor is 3 or 5; the first weighting factor and the second weighting factor include 3 and 5; In response to the constant being 16, the weighting factor is 6 or 10; the first weighting factor and the second weighting factor include 6 and 10; The reconstructing step includes: obtaining a prediction block of the block by averaging the first reference frame with the first weighting factor and the second reference frame with the second weighting factor; and reconstructing the block based on the predicted block. The method of claim 4.
7. the subset of the list of weighting factors includes only unequal weighting factors; determining the weighting factor includes determining the weighting factor as one of the unequal weighting factors; The reconstructing step includes: obtaining a prediction block of the block by averaging reference blocks in the respective reference frames with respective unequal weighting factors determined based on the one of the unequal weighting factors; and reconstructing the block based on the predicted block. The method of claim 1.
8. the unequal weighting factors are two unequal weighting factors comprising the weighting factor; the reference blocks include a first reference block and a second reference block; obtaining the predicted block includes averaging the first reference block and the second reference block having the two unequal weighting factors, respectively; The method of claim 7.
9. The two unequal weighting factors include: (i) 3 and 5, or (ii) 6 and 10. The method of claim 8.
10. a context for signaling a weighting factor index depends on the scaling factor information; determining the weighting coefficients of the composite weighted prediction mode based on the list of weighting coefficients includes determining the weighting coefficients based on the weighting coefficient indexes and the list of weighting coefficients; determining the weighting factors for the composite weighted prediction mode based on the subset of the list of weighting factors includes determining the weighting factors based on the weighting factor index and the subset of the list of weighting factors; The method of claim 1.
11. the context for signaling the weighting factor index is a first context based on the scaling factor information indicating that each of the scaling factors of the components of the at least one MVD is the predefined scaling factor; the context for signaling the weighting factor index is a second context different from the first context based on the scaling factor information indicating that the at least one of the scaling factors is different from the predefined scaling factor. The method of claim 10.
12. 1. A method of video decoding performed by a video coder, comprising: receiving a bitstream including a frame, wherein coding information of a block in the frame indicates that the block is coded in a joint motion vector differential (MVD) coding mode and a combined weighted prediction mode, and the coding information further indicates weighting factor information of the combined weighted prediction mode; determining, in response to the weighting factor information indicating that equal weighting factors are applied to each of the block's reference frames, a scaling factor for at least one component of MVD associated with at least one respective reference frame based on a list of scaling factors; determining the scaling factor for the component of the at least one MVD based on a subset of the list of scaling factors in response to the weighting factor information indicating that unequal weighting factors have been applied to the reference frame of the block; determining motion information associated with each of the reference frames of the blocks based on the determined scaling factors using the joint MVD coding mode; reconstructing the blocks based on the motion information and the weighting factor information associated with the respective reference frames of the blocks using the combined weighted prediction mode; A method having the following.
13. the weighting factor information indicating that the unequal weighting factors are applied to the reference frame of the block; the subset of the list of scaling factors includes only predefined scaling factors; the determined scaling factor is equal to the predefined scaling factor and is not signaled; The method of claim 12.
14. the predefined scaling factor is 1; The method of claim 13.
15. the weighting factor information indicating that the unequal weighting factors are applied to the reference frame of the block; the subset of the list of scaling factors does not include predefined scaling factors; The method of claim 12.
16. the predefined scaling factor is 1; 16. The method of claim 15.
17. a context for signaling a scaling factor index depends on the weighting factor information; determining the scaling factor based on the list of scaling factors includes determining the scaling factor based on the scaling factor index and the list of scaling factors; determining the scaling factor based on the subset of the list of scaling factors comprises determining the scaling factor based on the scaling factor index and the subset of the list of scaling factors; The method of claim 12.
18. 1. An apparatus for video decoding, comprising: A method comprising: a memory storing a program; and a processing circuit configured to execute the program stored in the memory to perform the method of any one of claims 1 to 17. Device.
19. A program which, when executed by a processor, causes the processor to carry out a method according to any one of claims 1 to 17.
20. 1. A method of video encoding performed by an encoder, comprising: determining whether each of scaling factors of at least one motion vector differential (MVD) component associated with at least one respective reference frame of a block is a predefined scaling factor, the scaling factor of the at least one MVD component being used in a joint MVD coding mode to encode the block within a frame; determining a weighting factor of a composite weighted prediction mode to be applied to encode the block based on a list of weighting factors when each of the scaling factors of the at least one component of the MVD associated with the at least one respective reference frame of the block is the predefined scaling factor; determining the weighting factors for the composite weighted prediction mode based on a subset of the list of weighting factors if at least one of the scaling factors of the at least one MVD component associated with the at least one respective reference frame of the block is different from the predefined scaling factor; using the joint MVD coding mode to determine motion information associated with each reference frame of the blocks based on the scaling factors of the at least one MVD component associated with each reference frame of the at least one of the blocks, the reference frames of the blocks including each of the at least one reference frame; encoding the block using the composite weighted prediction mode based on the motion information associated with the respective reference frames of the block and the determined weighting factors; A method having the following.
21. 1. A method of video encoding performed by an encoder, comprising: determining whether equal weighting factors are applied to each of the reference frames of a block in a composite weighted prediction mode applied to encode the block within the frame; if it is determined that the equal weighting factors are to be applied to each of the reference frames of the block, determining a scaling factor for at least one motion vector differential (MVD) component associated with at least one respective one of the reference frames based on a list of scaling factors; if it is determined that unequal weighting factors are to be applied to the reference frames of the block, determining the scaling factor for the at least one component of the MVD associated with each of the at least one reference frame based on a subset of the list of scaling factors; determining motion information associated with each of the reference frames of the blocks based on the determined scaling factors using a joint MVD coding mode; encoding the block using the composite weighted prediction mode based on the motion information associated with the respective reference frames of the block and the equal weighting factors or the unequal weighting factors; A method having the following.