Video decoding method and apparatus, video encoding method and apparatus, and storage medium
By determining the motion vector of the motion compensation padding region outside the image during video encoding, the problem of inaccurate motion vectors in the motion compensation padding region is solved, thus improving the accuracy and efficiency of video decoding.
Patent Information
- Application Number
- CN202380011057.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-04-17
- Filing Date
- 2023-04-18
- Publication Date
- 2026-02-10
- Estimated Expiration
- 2043-04-18
AI Technical Summary
Existing video coding technologies often fail to accurately determine the motion vectors of motion-compensated filling regions when processing video boundaries, resulting in poor video decoding performance.
The motion vectors in the motion compensation filling region located outside the image are determined by the processing circuit, and samples are reconstructed based on the candidate motion vectors at the image boundary to achieve motion compensation boundary filling.
It improves the accuracy and efficiency of video decoding and optimizes the effect of video boundary processing.
Smart Images

Figure CN117256143B_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This application claims priority to U.S. Patent Application No. 18 / 135,398, filed April 17, 2023, entitled "Picture Boundary Filling with Motion Compensation," which claims priority to U.S. Provisional Application No. 63 / 331,941, filed April 18, 2022, also entitled "Picture Boundary Filling with Motion Compensation," and U.S. Provisional Application No. 63 / 452,651, filed March 16, 2023, also entitled "Picture Boundary Filling with Motion Compensation." The entire contents of these earlier applications are incorporated herein by reference. Technical Field
[0003] This disclosure describes embodiments that are generally related to video encoding and decoding. Background Technology
[0004] The background description provided in this disclosure is intended to present the overall context of this disclosure. The work of the currently named inventors described in the background section, as well as various aspects of the specification that may not have been prior art at the time of application, are neither intended nor implied to be recognized as prior art to this disclosure.
[0005] Image / video compression facilitates the transfer of image / video files between different devices, storage devices, and networks with minimal quality degradation. In some examples, video codec techniques can compress video based on spatial and temporal redundancy. In one example, a video codec can use a technique called intra-frame prediction, which compresses images based on spatial redundancy. For example, intra-frame prediction can use reference data from the reconstructed current image to predict samples. In another example, a video codec can use a technique called inter-frame prediction, which compresses images based on temporal redundancy. For example, inter-frame prediction can predict samples in the current image from a previously reconstructed image using motion compensation. Motion compensation is typically represented by a motion vector (MV). Summary of the Invention
[0006] Aspects of the disclosure provide methods and apparatuses for video coding / decoding. In some examples, an apparatus for video decoding includes receiving circuitry and processing circuitry. The processing circuitry receives a bitstream carrying a plurality of pictures, and determines at least one motion compensation padding (MCP) block in a MCP region located outside of a picture and close to a picture boundary of the picture. The processing circuitry determines a motion vector for the MCP block for motion compensation boundary padding according to a plurality of candidates within the picture having positions located at the picture boundary, and reconstructs at least one sample in the MCP block according to the determined motion vector for the motion compensation boundary padding.
[0007] Aspects of the disclosure also provide a non-transitory computer-readable medium storing instructions that, when executed by a computer for video decoding, cause the computer to perform a method for video decoding. BRIEF DESCRIPTION OF DRAWINGS
[0008] Further features, properties, and various advantages of the disclosed subject matter will become more apparent from the following detailed description and from the accompanying drawings, in which:
[0009] Figure 1 is a schematic diagram of an exemplary block diagram of a communication system.
[0010] Figure 2 is a schematic diagram of an exemplary block diagram of a decoder.
[0011] Figure 3 is a schematic diagram of an exemplary block diagram of an encoder.
[0012] Figure 4 shows positions of spatial merge candidates according to embodiments of the disclosure.
[0013] Figure 5 shows pairs of candidates considered for a redundancy check of a spatial merge candidate according to embodiments of the disclosure.
[0014] Figure 6 shows exemplary motion vector scaling for temporal merge candidates.
[0015] Figure 7 shows exemplary candidate positions (e.g., C0 and C1) of temporal merge candidates of a current CU.
[0016] Figures 8A-8B shows affine motion fields in some examples.
[0017] Figure 9 shows block-based affine transform prediction in some examples.
[0018] Figure 10A candidate block is shown in some examples.
[0019] Figure 11 A control point motion vector is shown in some examples.
[0020] Figure 12 Spatial and temporal neighbors are shown in some examples.
[0021] Figure 13 An example illustration of a difference between a sample motion vector and a subblock motion vector is shown.
[0022] Figures 14-15 An example SbTMVP process used in a subblock-based temporal motion vector prediction (SbTMVP) mode is shown.
[0023] Figure 16 An extended coding unit region is shown in some examples.
[0024] Figure 17 An example illustration of decoder-side motion vector refinement is shown in some examples.
[0025] Figure 18 An example of a split line for a geometric partition mode is shown in some examples.
[0026] Figure 19 A diagram of uni-prediction motion vector selection for a geometric partition mode is shown in some examples.
[0027] Figure 20 An illustration of generating a blending weight in a geometric partition mode is shown in some examples.
[0028] Figure 21 A table to determine a number of shift bits is shown in some examples.
[0029] Figure 22 A portion of a video coding standard is shown in some examples.
[0030] Figure 23 An illustration of a subblock-level bi-prediction constraint is shown in some examples.
[0031] Figures 24A-24B A diagram of an extended region of a picture for padding is shown in some examples.
[0032] Figure 25 An illustration of motion-compensated boundary padding is shown in some examples.
[0033] Figure 26A diagram illustrates an example of deriving motion information for motion- compensated padding blocks in some examples.
[0034] Figure 27 A diagram illustrates an example of encoding order in some examples.
[0035] Figure 28 A diagram illustrates collocated blocks in time in some examples.
[0036] Figure 29 A flowchart illustrating a process overview in accordance with some embodiments of the disclosure is shown.
[0037] Figure 30 A flowchart illustrating another process overview in accordance with some embodiments of the disclosure is shown.
[0038] Figure 31 A diagram of a computer system in accordance with embodiments is shown. DETAILED DESCRIPTION
[0039] Figure 1 A block diagram of a video processing system (100) in some examples is shown. The video processing system (100) is an example of an application of the disclosed subject matter, video encoders and video decoders in a streaming environment. The disclosed subject matter can be equally applicable to other video enabled applications, including, for example, video conferencing, digital TV, streaming video, storing compressed video on digital media including CD, DVD, memory stick and the like, and so on.
[0040] The video processing system (100) includes a capture subsystem (113) that can include a video source (101), for example a digital camera, creating a stream of video pictures (102) that are uncompressed. In one example, the stream of video pictures (102) includes samples taken by the digital camera. The stream of video pictures (102), depicted as a bold line to emphasize the high data volume when compared to the encoded video data (104) (or coded video bitstream), can be processed by an electronic device (120) that includes a video encoder (103) coupled to the video source (101). The video encoder (103) can include hardware, software, or a combination of hardware and software to enable or implement aspects of the disclosed subject matter as described in greater detail below. The encoded video data (104) (or encoded video bitstream (104)), depicted as a thin line to emphasize the lower data volume when compared to the stream of video pictures (102), can be stored on a streaming server (105) for future use. One or more streaming client subsystems, e.g., video decoding device (110), can access the streaming server (105) to retrieve individual coded pictures or video segments for decoding and playback. Figure 1Client subsystems (106) and (108) in the network environment (100) can access the streaming server (105) to retrieve copies (107) and (109) of the encoded video data (104). The client subsystem (106) can include, for example, a video decoder (110) in an electronic device (130). The video decoder (110) decodes the incoming copy (107) of encoded video data and generates an outgoing stream of video pictures (111) that can be presented on a display (112) (e.g., a display screen) or another presentation device (not depicted). In some streaming systems, the encoded video data (104), video data (107), and video data (109) (e.g., video bitstreams) can be encoded according to certain video coding / compression standards. Examples of those standards include ITU-T H.265. In one example, a video coding standard under development is informally known as Versatile Video Coding (VVC). The subject matter of the present disclosure can be used within the context of the VVC standard.
[0041] It is noted that the electronic devices (120) and (130) can include other components (not shown). For example, the electronic device (120) can include a video decoder (not shown) and the electronic device (130) can include a video encoder (not shown) as well.
[0042] Figure 2 An example block diagram of a video decoder (210) is shown. The video decoder (210) can be provided in an electronic device (230). The electronic device (230) can include a receiver (231) (e.g., receiving circuitry). The video decoder (210) can be used in place of Figure 1 the video decoder (110) in the embodiments.
[0043] The receiver (231) can receive one or more coded video sequences to be decoded by the video decoder (210). In an embodiment, the receiver (231) receives one coded video sequence at a time, where the decoding of each coded video sequence is independent of the decoding of other coded video sequences. The coded video sequences can be received from a channel (201), which can be a hardware / software link into a storage device having the encoded video data stored therein. The receiver (231) can receive the encoded video data along with other data, e.g., coded audio data and / or ancillary data streams that can be forwarded to their respective consuming entities (not shown). The receiver (231) can separate the coded video sequence from the other data. To protect against network jitter, a buffer memory (215) can be coupled in between the receiver (231) and the entropy decoder / parsing device (220), hereinafter “parser (220).” In some applications, the buffer memory (215) is part of the video decoder (210). In other cases, the buffer memory (215) can be located (not shown) outside the video decoder (210). In yet other cases, a buffer memory (not shown) is located outside the video decoder (210), e.g., to protect against network jitter, and another buffer memory (215) is configured within the video decoder (210), e.g., to handle playout timing. When the receiver (231) receives data from a store / forward device that has sufficient bandwidth and controllability, or from an isosychronous network, it can not be necessary to configure a buffer memory (215), or the buffer memory can be smaller. Of course, to operate over best effort packet networks, such as the Internet, the buffer memory (215) can be relatively large and can have an adaptive size, and can be implemented at least partly in operating system or similar elements (not shown) outside the video decoder (210).
[0044] The video decoder (210) can include a parser (220) to reconstruct symbols (221) from the coded video sequence. Categories of those symbols include information used to manage operation of the video decoder (210), and potential information to control a rendering device(s) (212) (e.g., a display screen) that can not be a part of the electronic device (230) but can be coupled to the electronic device (230), as Figure 2The control information for the display device can be Supplemental Enhancement Information (SEI messages) or Parameter Sets fragments (not depicted) of Video Usability Information (VUI). The parser (220) can parse / entropy-decode the received coded video sequence. The coding of the coded video sequence can be in accordance with a video coding technology or standard, and can follow various principles, including variable length coding, Huffman coding, arithmetic coding with or without context sensitivity, and so forth. The parser (220) can extract from the coded video sequence, a set of subgroup parameters for at least one of the subgroups of pixels in the video decoder, based on at least one parameter corresponding to the group. The subgroups can include Groups of Pictures (GOPs), pictures, tiles, slices, macroblocks, Coding Units (CUs), blocks, Transform Units (TUs), Prediction Units (PUs) and so forth. The parser (220) can also extract from the coded video sequence, information such as transform coefficients, quantizer parameter values, motion vectors and so forth.
[0045] The parser (220) can perform entropy-decoding / parsing operation on the video sequence received from the buffer memory (215), thereby creating symbols (221).
[0046] The reconstruction of the symbols (221) can involve a number of different units, depending on the type of coded video picture or portion thereof (e.g., inter and intra pictures, inter and intra blocks), and other factors. Which units are involved, and how, can be controlled by subgroup control information that the parser (220) parses from the coded video sequence. For the sake of brevity, such subgroup control information flow between the parser (220) and the following units is not depicted.
[0047] Beyond the functional blocks already mentioned, the video decoder (210) can be conceptually subdivided into a number of functional units as described below. In practical implementations, many of these units interact closely with each other, and may
[0048] The first unit is a scaler / inverse transform unit (251). The scaler / inverse transform unit (251) receives quantized transform coefficients as the symbol(s) (221) and control information, including which transform to use, block size, quantization factor, quantization scaling matrices, etc., from the parser (220). The scaler / inverse transform unit (251) can output a block comprising sample values, which can be input into the aggregator (255).
[0049] In some cases, the output samples of the scaler / inverse transform unit (251) can belong to an intra coded block. The intra coded block is a block that is coded without using predictive information from previously reconstructed pictures, but can use predictive information from previously reconstructed parts of the current picture. Such predictive information can be provided by an intra picture prediction unit (252). In some cases, the intra picture prediction unit (252) generates a block of the same size and shape as the block being reconstructed using surrounding already reconstructed information extracted from the current picture buffer (258). For example, the current picture buffer (258) buffers partially reconstructed current pictures and / or fully reconstructed current pictures. In some cases, the aggregator (255) adds, on a per sample basis, the predictive information generated by the intra prediction unit (252) to the output sample information provided by the scaler / inverse transform unit (251).
[0050] In other cases, the output samples of the scaler / inverse transform unit (251) can belong to an inter coded and potentially motion compensated block. In this case, the motion compensation prediction unit (253) can access the reference picture memory (257) to fetch samples for prediction. After motion compensation of the fetched samples according to the symbol (221) belonging to the block, these samples can be added by the aggregator (255) to the output of the scaler / inverse transform unit (251), referred to as residual samples or residual signal in this case, generating the output sample information. The fetching of the prediction samples by the motion compensation prediction unit (253) from addresses within the reference picture memory (257) can be controlled by motion vectors, and the motion vectors are available to the motion compensation prediction unit (253) in the form of the symbol (221), e.g., comprising X, Y, and reference picture component. Motion compensation can also include interpolation of sample values fetched from the reference picture memory (257), motion vector prediction mechanisms, etc., when sub-sample precision motion vectors are used.
[0051] Output samples of the aggregator (255) can be subject to various loop filtering techniques in the loop filter unit (256). Video compression technologies can include in-loop filter technologies that are controlled by parameters included in the coded video sequence (also referred to as coded video bitstream) and made available to the loop filter unit (256) from the symbols (221) as output by the parser (220). Video compression can also be responsive to meta-information obtained during the decoding of previous (in decoding order) parts of the coded picture or coded video sequence, as well as responsive to previously reconstructed and loop-filtered sample values.
[0052] The output of the loop filter unit (256) can be a stream of samples that can be output to the display device (212) and that can be stored into the reference picture memory (257) for future inter-picture prediction.
[0053] Once fully reconstructed, certain coded pictures can be used as reference pictures for future prediction. For example, once a coded picture corresponding to a current picture is fully reconstructed and the coded picture is identified (by, for example, the parser (220)) as a reference picture, the current picture buffer (258) can become part of the reference picture memory (257), and a new current picture buffer can be reallocated before commencing the reconstruction of the next coded picture.
[0054] The video decoder (210) can perform decoding operations according to a predetermined video compression technology or standard, such as ITU-T H.265. The coded video sequence can conform to a syntax specified by the video compression technology or standard being used, in the sense that the coded video sequence adheres to the syntax of the video compression technology or standard, and the configuration files recorded in the coded video sequence are within the range of levels and profiles / tiers that the video compression technology or standard supports. Specifically, a profile can select certain tools available from all the tools available in the video compression technology or standard, as the only tools to be used under that profile. Also required for conforming to the standard is that the complexity of coded video sequence is within the bounds as defined by the level as specified by the video compression technology or standard. In some cases, the limits set by the level can be further restricted through Hypothetical Reference Decoder (HRD) specifications and metadata for HRD buffer management signaled in the coded video sequence.
[0055] In this embodiment, the receiver (231) may receive additional (redundant) data along with the encoded video. The additional data may be a portion of the encoded video sequence. The additional data may be used by the video decoder (210) to properly decode the data and / or more accurately reconstruct the original video data. The additional data may take the form of, for example, temporal, spatial, or signal-to-noise ratio (SNR) enhancement layers, redundant slices, redundant images, forward error correction codes, etc.
[0056] Figure 3 An exemplary block diagram of a video encoder (303) is shown. The video encoder (303) is disposed in an electronic device (320). The electronic device (320) includes a transmitter (340) (e.g., transmission circuitry). The video encoder (303) can be used in place of... Figure 1 The video encoder (103) in the example.
[0057] The video encoder (303) can obtain data from the video source (301) (not) Figure 3 In one embodiment, a portion of the electronic device (320) receives video samples, the video source being capable of capturing video images to be encoded by a video encoder (303). In another embodiment, the video source (301) is a portion of the electronic device (320).
[0058] A video source (301) can provide a sequence of source videos encoded by a video encoder (303) in the form of a digital video sample stream, which can have any suitable bit depth (e.g., 8-bit, 10-bit, 12-bit, etc.), any color space (e.g., BT.601YCrCb, RGB, etc.), and any suitable sampling structure (e.g., YCrCb 4:2:0, YCrCb 4:4:4). In a media service system, the video source (301) can be a storage device storing previously prepared video. In a video conferencing system, the video source (301) can be a camera that captures local image information as a video sequence. Video data can be provided as multiple individual pictures, which are given motion when viewed in sequence. The pictures themselves can be constructed as spatial pixel arrays, where each pixel can include one or more samples depending on the sampling structure, color space, etc., used. Those skilled in the art can readily understand the relationship between pixels and samples. The following focuses on describing samples.
[0059] According to an embodiment, the video encoder (303) can code and compress the pictures of the source video sequence into a coded video sequence (343) in real time or under any other time constraints as required. Enforcing appropriate coding speed is one function of a controller (350). In some embodiments, the controller (350) controls other functional units as described below and is functionally coupled to these units. For brevity, no couplings are shown in the drawing. Parameters set by the controller (350) can include rate control related parameters (picture skip, quantizer, lambda value of rate-distortion optimization techniques, etc.), picture size, group of pictures (GOP) layout, maximum motion vector search range, and so on. The controller (350) can be used for other suitable functions related to the video encoder (303) optimized for a certain system design.
[0060] In some embodiments, the video encoder (303) operates in a coding loop. As a simple description, in one example, the coding loop can include a source coder (330) (e.g., responsible for creating symbols, such as a symbol stream, based on input pictures to be coded, and reference pictures) and a (local) decoder (333) embedded in the video encoder (303). The decoder (333) reconstructs the symbols to create the sample data in a similar manner as a (remote) decoder would create. The reconstructed sample stream (sample data) is input to the reference picture buffer (334). As the decoding of the symbol stream results in a bit-exact outcome independent of the decoder location (local or remote), the content in the reference picture buffer (334) is also bit exactly corresponding between the local encoder and remote encoder. In other words, the prediction part of an encoder "sees" the same sample values that a decoder would "see" when using the predictions during decoding. This fundamental principle of reference picture synchronization (and resulting drift if synchronization cannot be maintained, e.g., due to channel errors) is used by some related technologies as well.
[0061] The operation of the "local" decoder (333) can be the same as the "remote" decoder detailed above. Figure 2 The operation of the "local" decoder (333) can be the same as the "remote" decoder detailed above. Figure 2 When symbols are available and the entropy encoder (345) and parser (220) are able to losslessly encode / decode the symbols into the coded video sequence, the entropy decoding part of the video decoder (210), including the buffer (215) and the parser (220), can not be fully implemented in the local decoder (333).
[0062] In one embodiment, the decoder technology, other than parsing / entropy decoding present in the decoder, is present in the corresponding encoder in substantially the same functional form. Thus, the presently disclosed subject matter focuses on the decoder operations. The description of the encoder technology can be simplified because the encoder technology is reciprocal to the decoder technology that is described fully. A more detailed description in certain areas is provided below.
[0063] During operation, in some examples, the source coder (330) can perform motion compensated predictive coding of the input picture data to predict future pictures based on coded data of previously coded pictures from the video sequence. In this manner, the coding engine (332) codes differences between pixel blocks of an input picture and pixel blocks of reference pictures that can be selected as prediction references for the input picture.
[0064] The local video decoder (333) can decode coded video data of pictures that can be designated as reference pictures based on symbols created by the source coder (330). Operations of the coding engine (332) can be lossy processes. When the coded video data can be decoded at a video decoder (not shown), the reconstructed video sequence can typically be a replica of the source video sequence with some errors. Figure 3 The local video decoder (333) replicates decoding processes that can be performed by a video decoder on reference pictures and can cause reconstructed reference pictures to be stored in the reference picture memory (334). In this manner, the video encoder (303) can store copies of reconstructed reference pictures locally that have common content as the reconstructed reference pictures that will be obtained by a far-end video decoder (absent transmission errors).
[0065] The predictor (335) can perform a prediction search for the coding engine (332). That is, for each new picture to be coded, the predictor (335) can search the reference picture memory (334) for sample data (as candidate reference pixel blocks) or certain metadata, such as reference picture motion vectors, block shapes, and so on, that can serve as appropriate prediction references for the new pictures. The predictor (335) can operate on a sample block-by-pixel block basis to find appropriate prediction references. In some cases, depending on the search results obtained by the predictor (335), it can be determined that the input picture can have prediction references drawn from multiple reference pictures stored in the reference picture memory (334).
[0066] The controller (350) can manage coding operations of the source coder (330), including, for example, setting of parameters and subgroup parameters used for encoding the video data.
[0067] The outputs of all the above-described functional units can be entropy encoded in an entropy encoder (345). The entropy encoder (345) losslessly compresses the symbols generated by the various functional units according to, for example, Huffman coding, variable length coding, arithmetic coding, and so on, thereby converting the symbols into an encoded video sequence.
[0068] The transmitter (340) can buffer the encoded video sequence(s) as they are created by the entropy coder (345) in preparation for transmission via a communication channel (360), which can be a hardware / software link to a storage device that will store the encoded video data. The transmitter (340) can merge encoded video data from the video coder (303) with other data to be transmitted, for example, encoded audio data and / or ancillary data streams (sources not shown).
[0069] The controller (350) can manage operation of the video encoder (303). During coding, the controller (350) can assign to each coded picture a certain coded picture type, which can affect the coding techniques that can be applied to the corresponding picture. For example, pictures often can be assigned as one of the following picture types:
[0070] An Intra Picture (I picture) can be a picture that can be coded and decoded without using any other picture in the sequence as a source of prediction. Some video codecs allow different types of Intra pictures, including, for example Independent Decoder Refresh (“IDR”) pictures. A person of skill in the art is aware of the variants of I pictures and their respective applications and features.
[0071] A predictive picture (P picture) can be a picture that can be coded and decoded using either intra prediction or inter prediction that uses at most one motion vector and reference index to predict sample values of each block.
[0072] A bi-directionally predictive picture (B picture) can be a picture that can be coded and decoded using either intra prediction or inter prediction that uses at most two motion vectors and reference indices to predict sample values of each block. Similarly, multiple predictive pictures can use more than two reference pictures and associated metadata for reconstructing a single block.
[0073] A source picture can generally be spatially subdivided into a plurality of blocks of samples (e.g., 4x4, 8x8, 4x8, or 16x16 samples), and encoded block by block. These blocks can be predictively encoded with reference to other (already encoded) blocks, determined according to the coding assignment of the respective picture to which the block applies. For example, blocks of I pictures can be non-predictively encoded, or they can be predictively encoded with reference to already encoded blocks of the same picture (spatial, or intra, prediction). Blocks of P pictures can be predictively encoded with reference to one previously encoded reference picture, either by spatial prediction or by temporal prediction. Blocks of B pictures can be predictively encoded with reference to one or two previously encoded reference pictures, either by spatial prediction or by temporal prediction.
[0074] The video encoder (303) can perform encoding operations according to a predetermined video coding technology or standard, such as ITU-T H.265. In its operation, the video encoder (303) can perform various compression operations, including predictive coding operations that exploit temporal and spatial redundancies in the input video sequence. The coded video data, therefore, can conform to a syntax specified by the video coding technology or standard being used.
[0075] In one embodiment, the transmitter (340) can transmit additional data with the encoded video. The source coder (330) can include such data as part of the coded video sequence. Additional data can comprise temporal / spatial / SNR enhancement layers, other forms of redundant data such as redundant pictures and slices, SEI messages, VUI parameter set fragments, and so on.
[0076] Collected video can be in the form of a plurality of source pictures (video pictures) in temporal sequence. Intra-picture prediction (often simply referred to as intra prediction) exploits spatial correlation in a given picture, while inter-picture prediction exploits correlation between pictures (either temporal or other). In one example, a particular picture being encoded / decoded, called the current picture, is partitioned into blocks. A block in the current picture can be coded with a reference to a reference block in a reference picture that has been coded and remains buffered, using a vector referred to as a motion vector. The motion vector points to a reference block in the reference picture, and can have a third dimension identifying the reference picture, in cases where multiple reference pictures are in use.
[0077] In some embodiments, bi-prediction techniques can be used in inter-picture prediction. According to bi-prediction techniques, two reference pictures are used, e.g., a first reference picture and a second reference picture, both preceding the current picture in decoding order (but possibly past and future, respectively, in display order) in the video. A block in the current picture can be coded by a first motion vector pointing to a first reference block in the first reference picture and a second motion vector pointing to a second reference block in the second reference picture. Specifically, the block can be predicted by a combination of the first reference block and the second reference block.
[0078] Furthermore, merge mode techniques can be used in inter-picture prediction to improve coding efficiency.
[0079] According to some embodiments of the disclosure, the execution of the prediction, such as inter-picture prediction and intra-picture prediction, is in the unit of blocks. For example, according to the HEVC standard, a picture in a sequence of video pictures is partitioned into coding tree units (CTU) for compression, the CTUs in a picture have the same size, such as 64x64 pixels, 32x32 pixels, or 16x16 pixels. Generally, a CTU includes three coding tree blocks (CTB), one luma CTB and two chroma CTBs. Each CTU can be further split into one or more coding units (CU) in a quadtree manner. For example, a CTU of 64x64 pixels can be split into one CU of 64x64 pixels, or 4 CUs of 32x32 pixels, or 16 CUs of 16x16 pixels. In one example, each CU is analyzed to determine a prediction type for the CU, such as an inter prediction type or an intra prediction type. The CU is split into one or more prediction units (PU) depending on the temporal and / or spatial predictability. Generally, each PU includes a luma prediction block (PB), and two chroma PBs. In one embodiment, the prediction process in encoding (encoding / decoding) is performed in the unit of a prediction block. Using a luma prediction block as an example of a prediction block, the prediction block includes a matrix of pixel values (e.g., luma values), such as 8x8 pixels, 16x16 pixels, 8x16 pixels, 16x8 pixels, and / or the like.
[0080] It should be noted that the video encoder (103), the video encoder (303), and the video decoder (110) and the video decoder (210) can be implemented using any suitable technique. In one embodiment, the video encoder (103), the video encoder (303), and the video decoder (110) and the video decoder (210) can be implemented using one or more integrated circuits. In another embodiment, the video encoder (103), the video encoder (303), and the video decoder (110) and the video decoder (210) can be implemented using one or more processors that execute software instructions.
[0081] Aspects of the disclosure provide techniques for motion-compensated picture padding (also referred to as motion-compensated boundary padding). The techniques for motion- compensated picture padding can include deriving motion information for motion-compensated picture padding and / or signaling motion information for motion-compensated picture padding.
[0082] Various inter prediction modes can be used in VVC. For an inter-predicted CU, the motion parameters can include one or more MVs, one or more reference picture indices, a reference picture list usage index, and additional information for certain coding features used in inter-predicted sample generation. The motion parameters can be signaled explicitly or implicitly. When a CU is coded with a skip mode, the CU can be associated with a PU and the CU can have no valid residual coefficients, no coded motion vector delta or MV difference (MVD), or reference picture indices. A merge mode can be specified, in which the motion parameters of the current CU are obtained from one or more neighboring CUs, including spatial and / or temporal candidates, and optionally additional information introduced in VVC, for example. The merge mode can be applied not only to the skip mode, but also to the inter-predicted CUs. In one example, an alternative to the merge mode is explicit signaling of the motion parameters, in which, for each CU, one or more MVs, the corresponding reference picture indices for each reference picture list and the reference picture list usage flags, and other information are explicitly signaled.
[0083] In one embodiment, such as in VVC, the VVC Test model (VTM) reference software includes one or more refined inter prediction coding tools including extended merge prediction, merge motion vector difference (MMVD) mode, adaptive motion vector prediction (AMVP) mode with symmetric MVD signaling, affine motion compensation prediction, subblock-based temporal motion vector prediction (SbTMVP), adaptive motion vector resolution (AMVR), motion field storage (1 / 16 luma sample MV storage and 8x8 motion field compression), bi-prediction with CU-level weights (BCW), bi-directional optical flow (BDOF), prediction refinement using optical flow (PROF), decoder-side motion vector refinement (DMVR), combined inter and intra prediction (CIIP), geometric partitioning mode (GPM), etc. Inter prediction and related methods will be described in detail below.
[0084] Extended merge prediction can be used in some examples. In an example such as in VTM4, a merge candidate list is constructed by including the following five types of candidates in order: one or more spatial motion vector predictors (MVPs) from one or more spatial neighboring CUs, one or more temporal MVPs from one or more collocated CUs, one or more history-based MVPs (HMVPs) from a first-in-first-out (FIFO) table, one or more pair- wise average MVPs, and one or more zero MVs.
[0085] The size of the merge candidate list can be signaled in the slice header. In one example, in VTM4, the maximum allowed size of the merge candidate list is 6. The index of the best merge candidate (e.g., merge index) can be coded using truncated unary binarization TUs for each CU coded in merge mode. The first bin of the merge index can be context coded (e.g., context-adaptive binary arithmetic coding (CABAC)), and bypass coding can be used for the other bins.
[0086] Some examples of the generation process of each type of merge candidate are provided below. In one embodiment, the derivation process of spatial candidates is as follows. The derivation of spatial merge candidates in VVC can be the same as in HEVC. In one example, the four spatial merge candidates are selected from the candidates located at the positions shown in Figure 4 The most four merge candidates are selected from the candidates located at the positions shown in
[0087] Figure 4 The positions of spatial merge candidates are shown according to one embodiment of the disclosure. Referring to Figure 4 , the order of derivation is B1, A1, B0, A0, and B2. Position B2 is considered only when any of the CUs at positions A0, B0, B1, and A1 are not available (e.g., because the CU belongs to another slice or another tile) or are intra coded. After adding the candidate at position A1, a redundancy check is performed for the addition of the remaining candidates to ensure that candidates with the same motion information are excluded from the candidate list, thereby improving coding efficiency.
[0088] To reduce the computational complexity, not all possible pairs of candidates are considered in the above redundancy check. Instead, only the pairs linked with arrows in Figure 5 are considered, and a candidate object is only added to the candidate list if the corresponding candidate object used for the redundancy check does not have the same motion information.
[0089] Figure 5 The pairs of candidates considered for the redundancy check of spatial merge candidates according to an embodiment of the disclosure are shown. Referring to Figure 5 , the pairs linked with corresponding arrows include A1 and B1, A1 and A0, A1 and B2, B1 and B0, and B1 and B2. Thus, the candidate at position A1 can be compared with the candidates at positions B1, A0, and / or B2, and the candidate at position B1 can be compared with the candidates at positions B0 and / or B2.
[0090] In one embodiment, a temporal candidate is derived in the following way. In one example, only one temporal merge candidate is added to the candidate list. Figure 6 An exemplary motion vector scaling for temporal merge candidate is shown. To derive a temporal merge candidate for a current CU (611) in a current picture (601), a scaled MV (621) can be derived based on a collocated CU (612) belonging to a collocated reference picture (604) (e.g., by Figure 6 A reference picture list for deriving the collocated CU (612) can be explicitly signaled in the slice header. As Figure 6 A scaled MV (621) for a temporal merge candidate can be obtained as shown by the dashed line in FIG. 6B. The scaled MV (621) can be scaled from the MV of the collocated CU (612) using picture order count (POC) distances tb and td. The POC distance tb can be defined as the POC difference between the current reference picture (602) of the current picture (601) and the current picture (601). The POC distance td can be defined as the POC difference between the collocated picture (603) and the collocated reference picture (604). The reference picture index of the temporal merge candidate can be set to zero.
[0091] Figure 7 Exemplary candidate positions (e.g., C0 and C1) of a temporal merge candidate of a current CU are shown. The position of the temporal merge candidate can be selected in the candidate positions C0 and C1. The candidate position C0 is at the bottom-right corner of the collocated CU (710) of the current CU. The candidate position C1 is at the center of the collocated CU (710) of the current CU. If the CU at the candidate position C0 is unavailable, intra coded, or outside the current row of the CTU, the candidate position C1 is used to derive the temporal merge candidate. Otherwise, e.g., the CU at the candidate position C0 is available, intra coded, and in the current row of the CTU, the candidate position C0 is used to derive the temporal merge candidate.
[0092] In HEVC, a translational motion model is applied to motion compensated prediction (MCP). In the real world, however, there can be multiple kinds of motion, such as zooming in / out, rotation, perspective motion, and other irregular motion. Block-based affine transform motion compensated prediction can be applied, e.g., in VTM. Figure 8A An affine motion field of a block (802) described by motion information of two control points (4 parameters) is shown. Figure 8B An affine motion field of a block (804) described by three control point motion vectors (6 parameters) is shown.
[0093] As Figure 8AAs shown, in the 4-parameter affine motion model, the motion vector at the sample position (x, y) in block (802) can be derived from the following equation (1):
[0094]
[0095] Among them, mv x It can be the motion vector in the first direction (or the X direction), and mv y It can be a motion vector in the second direction (or the Y direction). The motion vector can also be described in equation (2):
[0096]
[0097] like Figure 8B As shown, in the 6-parameter affine motion model, the motion vector at the sample position (x, y) in block (1304) can be derived from the following equation (3):
[0098]
[0099] The 6-parameter affine motion model can also be described in the following equation (4):
[0100]
[0101] As shown in equations (1) and (3), (mv 0x ,mv 0y (mv) can be the motion vector of the top-left control point. 1x ,mv 1y (mv) can be the motion vector of the upper right control point. 2x ,mv 2y () can be the motion vector of the lower left control point.
[0102] like Figure 9 As shown, to simplify motion-compensated prediction, block-based affine transformation prediction can be applied. To derive the motion vector for each 4×4 luma sub-block, the motion vector of the center sample (e.g., (902)) of each sub-block (e.g., (904)) in the current block (900) can be calculated according to equations (1)-(4) and rounded to 1 / 16 fractional precision. A motion-compensated interpolation filter can then be applied to generate a prediction for each sub-block using the derived motion vector. The sub-block size for the chroma component can also be set to 4×4. The MV of the 4×4 chroma sub-blocks can be calculated as the average of the MVs of the four corresponding 4×4 luma sub-blocks.
[0103] In affine merge prediction, an affine merge (AF_MERGE) mode can be applied to a CU whose width and height are both greater than or equal to 8. CPMVs of the current CU can be generated based on motion information of spatial neighboring CUs. Up to five control point motion vector prediction (CPMVP) candidates can be applied to affine merge prediction, and an index can be signaled to indicate which of the five CPMVP candidates can be used for the current CU. In affine merge prediction, three types of CPMV candidates can be used to form an affine merge candidate list: (1) an inherited affine merge candidate inferred from CPMVs of neighboring CUs, (2) a constructed affine merge candidate with CPMVs derived using translational MVs of neighboring CUs, and (3) a zero MV.
[0104] In VTM3, up to two inherited affine candidates can be applied. The two inherited affine candidates can be derived from affine motion models of neighboring blocks. For example, one inherited affine candidate can be derived from a left neighboring CU, and another inherited affine candidate can be derived from an above neighboring CU. An example candidate block can be shown in Figure 10 Figure 10 As shown in, for a left predictor (or inherited affine candidate from the left), the scan order can be A0 -> A1, and for an above predictor (or inherited affine candidate from the above), the scan order can be B0 -> B1 -> B2. Thus, only the first available inherited candidate from each side can be selected. No pruning check can be performed among the two inherited candidates. When a neighboring affine CU is identified, the control point motion vectors of the neighboring affine CU can be used to derive CPMVP candidates in the affine merge list of the current CU. As shown in Figure 11 When a neighboring lower-left block A of the current block (1104) is encoded in affine mode, the motion vectors v2, v3, and v4 of the top-left, top-right, and bottom-left corners of the CU (1102) containing block A can be obtained. When block A is encoded with a 4-parameter affine model, two CPMVs of the current CU (1104) can be calculated from v2 and v3 of the CU (1102). In the case of encoding the block with a 6-parameter affine model, three CPMVs of the current CU (1104) can be calculated from v2, v3, and v4 of the CU (1102).
[0105] A constructed affine candidate of the current block can be a candidate constructed by combining neighboring translational motion information of each control point of the current block. The motion information of the control points can be derived from the designated spatial neighboring blocks and temporal neighboring blocks as shown in Figure 12 Figure 12 As shown in, the CPMVs k(k = 1, 2, 3, 4) denotes the k-th control point of the current block (1202). For CPMV1, the B2->B3->A2 block can be checked, and the MV of the first available block can be used. For CPMV2, the B1->B0 block can be checked. For CPMV3, the A1->A0 block can be checked. If CPMV4 is not available, TMVP can be used as CPMV4.
[0106] After obtaining the MVs of the four control points, an affine merge candidate can be constructed for the current block (1702) based on the motion information of the four control points. For example, the affine merge candidate can be constructed based on the combination of the MVs of the four control points in the following order: {CPMV1, CPMV2, CPMV3}, {CPMV1, CPMV2, CPMV4}, {CPMV1, CPMV3, CPMV4}, {CPMV2, CPMV3, CPMV4}, {CPMV1, CPMV2}, and {CPMV1, CPMV3}.
[0107] The combination of 3 CPMVs can construct a 6-parameter affine merge candidate, and the combination of 2 CPMVs can construct a 4-parameter affine merge candidate. To avoid the motion scaling process, if the reference indices of the control points are different, the related combination of the control point MVs can be discarded.
[0108] After checking the inherited affine merge candidates and the constructed affine merge candidate, if the list is still not full, a zero MV can be inserted at the end of the list.
[0109] In affine AMVP prediction, the affine AMVP mode can be applied to a CU whose width and height are both greater than or equal to 16. An affine flag at the CU level can be written in the bitstream to indicate whether the affine AMVP mode is used, and then another flag can be signaled to indicate whether a 4-parameter affine or a 6-parameter affine is applied. In affine AMVP prediction, the difference values of the CPMVs of the current CU and the predictors of the CPMVs of the current CU can be written in the bitstream. The size of the affine AMVP candidate list can be 2, and the affine AMVP candidate list can be generated by using four types of CPMV candidates in the following order:
[0110] (1) an inherited affine AMVP candidate inferred from the CPMVs of neighboring CUs,
[0111] (2) an affine AMVP candidate constructed with CPMVPs derived using the translational MVs of neighboring CUs,
[0112] (3) translational MVs from neighboring CUs, and
[0113] (4) zero MVs.
[0114] The checking order of inherited affine AMVP candidates can be the same as that of inherited affine merge candidates. To determine an AMVP candidate, only affine CUs with the same reference picture as the current block can be considered. When an inherited affine motion predictor is inserted into the candidate list, the pruning process can not be applied.
[0115] A constructed AMVP candidate can be derived from a specified spatial neighbor. As shown in Figure 12 the same checking order as in the construction of affine merge candidates can be applied. In addition, the reference picture index of the neighboring block can also be checked. The first block in the checking order can be inter coded and have the same reference picture as the current CU (1202). When the current CU (1202) is coded with the 4-parameter affine mode, one constructed AMVP candidate can be determined and both mv0 and mv1 are available. The constructed AMVP candidate can be further added to the affine AMVP list. When the current CU (1202) is coded with the 6-parameter affine mode and all three CPMVs are available, the constructed AMVP candidate can be added to the affine AMVP list as one candidate. Otherwise, the constructed AMVP candidate can be set as unavailable.
[0116] If the number of candidates in the affine AMVP list is still less than 2 after checking the inherited affine AMVP candidates and the constructed AMVP candidate, mv0, mv1 and mv2 can be added in turn. When mv0, mv1 and mv2 are available, mv0, mv1 and mv2 can be used as translational MVs to predict all control point MVs of the current CU (e.g., (1202)). Finally, if the affine AMVP list is still not full, zero MVs can be used to fill the affine AMVP list.
[0117] Compared with pixel-based motion compensation, subblock-based affine motion compensation can save memory access bandwidth and reduce computational complexity, but at the cost of loss of prediction accuracy. To achieve finer motion compensation granularity, prediction refinement with optical flow (PROF) can be used to refine the subblock-based affine motion compensation prediction without increasing the memory access bandwidth for motion compensation. In VVC, after performing subblock-based affine motion compensation, the luma prediction samples can be refined by adding the difference derived by the optical flow equation. PROF can be described in the following four steps.
[0118] Step (1): Subblock-based affine motion compensation can be performed to generate a subblock prediction I(i,j).
[0119] Step (2): The spatial gradient g of the sub-block prediction can be calculated at each sample position using a 3-tap filter [-1, 0, 1] x (i,j) and g y (i,j). The gradient calculation can be the same as in BDOF. For example, the spatial gradient g x (i,j) and g y (i,j) can be calculated based on equations (5) and (6), respectively.
[0120] g x (i,j) = (I(i+1,j) » shiftl) - (I(i-1,j) » shiftl) equation (5)
[0121] g y (i,j) = (I(i,j+1) » shiftl) - (I(i,j-1) » shiftl) equation (6)
[0122] As shown in equations (5) and (6), shiftl can be used to control the precision of the gradient. The sub-block (e.g., 4x4) prediction can be extended by one sample on each side for gradient calculation. To avoid extra memory bandwidth and extra interpolation calculation, the extended samples on the extended boundaries can be copied from the nearest integer pixel positions in the reference picture.
[0123] Step (3): The luma prediction refinement can be calculated by the optical flow equation as shown in equation (7).
[0124] ΔI(i,j) = g x (i,j) * Δv x (i,j) + g y (i,j) * Δv y (i,j) equation (7)
[0125] where Δv(i,j) can be the difference between the sample MV and the sub-block MV of the sub-block to which the sample position (i,j) belongs, the sample MV is denoted by v(i,j), and the sub-block MV is denoted by v SB . Figure 13 An exemplary diagram showing the difference between the sample MV and the sub-block MV is shown. As Figure 13 shown, a sub-block (1302) can be included in a current block (1300), and a sample (1304) can be included in the sub-block (1302). The sample (1304) can include a sample motion vector v(i,j) corresponding to a reference pixel (1306). The sub-block (1302) can include a sub-block motion vector v SB . Based on the sub-block motion vector v SB, the sample (1304) can correspond to the reference pixel (1308). The difference (denoted by Δv(i, j)) between the sample MV and the sub-block MV can be represented by the difference between the reference pixel (1306) and the reference pixel (1308). Δv(i, j) can be quantized in units of 1 / 32 luma sample precision.
[0126] Because the affine model parameters and the sample position relative to the sub-block center do not change from sub-block to another sub-block, Δv(i, j) can be computed for a first sub-block (e.g., (1302)) and used for other sub-blocks (e.g., (1310)) in the same CU (e.g., (1300)). Let dx(i, j) be the horizontal offset from the sample position (i, j) to the center of the sub-block (x SB , y SB ), and let dy(i, j) be the vertical offset from the sample position (i, j) to the center of the sub-block (x SB , y SB ), Δv(x, y) can be derived by the following equations (8) and (9):
[0127]
[0128]
[0129] To maintain accuracy, the center of the sub-block (x SB , y SB ) can be computed as ((WSB-1) / 2, (HSB-1) / 2), where WSB and HSB are the width and height of the sub-block, respectively.
[0130] Once Δv(x, y) is obtained, the parameters of the affine model can be obtained. For example, for a 4-parameter affine model, the parameters of the affine model can be shown in equation (10).
[0131]
[0132] For a 6-parameter affine model, the parameters of the affine model can be shown in equation (11).
[0133]
[0134] where (v 0x , v 0y ), (v 1x , v 1y ), and (v 2x , v 2y ) can be the top-left control point motion vector, the top-right control point motion vector, and the bottom-left control point motion vector, respectively, and w and h can be the width and height of the CU, respectively.
[0135] Step (4): Finally, the luminance prediction refinement ΔΙ(ί, j) can be added to the subblock prediction I(i, j). The final prediction I' can be generated as shown in Equation (12).
[0136] I'(i, j) = I(i, j) + ΔΙ(ί, j) Equation (12)
[0137] For affine coded CUs, PROF can not be applied in the following two cases: (1) all control point MVs are the same, which indicates that the CU has only translational motion, and (2) the affine motion parameters are larger than a specified limit, because subblock-based affine MC is downgraded to CU-based MC to avoid large memory access bandwidth requirements.
[0138] To improve coding efficiency and reduce the transmission overhead of MVs, subblock-level MV refinement can be applied to extend CU-level temporal motion vector prediction (TMVP). In one example, a subblock-based TMVP (SbTMVP) mode allows inheriting subblock-level motion information from a collocated reference picture. Each subblock of a current CU (e.g., a current CU with a large size) in a current picture can have corresponding motion information without explicitly sending the block partition structure or the corresponding motion information. In the SbTMVP mode, the motion information of each subblock can be obtained in the following manner, e.g., in the following three steps. In the first step, a displacement vector (DV) of the current CU can be derived. In the second step, the availability of SbTMVP candidates can be checked, and a center motion (e.g., a center motion of the current CU) can be derived. In the third step, the subblock motion information can be derived from the corresponding subblock in the collocated block using the DV. The above three steps can be combined into one or two steps, and / or the order of the three steps can be adjusted.
[0139] Unlike TMVP candidate derivation from a collocated block in a reference frame or picture, in the SbTMVP mode, a DV (e.g., a DV derived from a MV of a left neighboring CU of the current CU) can be applied to locate a corresponding subblock in a collocated picture for each subblock in the current CU in the current picture. If the corresponding subblock is not inter coded, the motion information of the current subblock can be set to the center motion information of the collocated block.
[0140] SbTMVP mode can be supported by various video coding standards including, for example, VVC. Similar to TMVP mode, in SbTMVP mode, for example, in HEVC, motion fields (also referred to as motion information fields or MV fields) in a collocated picture can be used to improve MV prediction and merge mode for CUs in the current picture. In one example, the same collocated picture used by TMVP mode is used in SbTMVP mode. In one example, SbTMVP mode differs from TMVP mode in that (i) TMVP mode predicts motion information at CU level, while SbTMVP mode predicts motion information at sub-CU level; (ii) TMVP mode obtains multiple temporal MVs from a collocated block in a collocated picture (e.g., the collocated block is the bottom-right block or the center block with respect to the current CU), while SbTMVP mode can apply motion shift before obtaining temporal motion information from the collocated picture. In one example, the motion shift used in SbTMVP mode is obtained from a MV of one of the spatial neighboring blocks of the current CU.
[0141] Figures 14-15 An example SbTMVP process used in SbTMVP mode is shown. The SbTMVP process can, for example, predict multiple MVs for sub-CUs (e.g., sub-blocks) within a current CU (e.g., current block) (1401) in a current picture (1511) in two steps. In the first step, spatial neighboring blocks (e.g., Al) of the current block (1401) in Figure 14 and Figure 15 are examined. If a spatial neighboring block (e.g., Al) has a MV (1521) that uses a collocated picture (1512) as a reference picture for the spatial neighboring block (e.g., Al), the MV (1521) can be selected as a motion shift (or DV) to be applied to the current block (1401). If no such MV (e.g., a MV that uses the collocated picture (1512) as a reference picture) is identified, the motion shift or DV can be set to a zero MV (e.g., (0, 0)). In some examples, if no such MV is identified for spatial neighboring block Al, MVs of additional spatial neighboring blocks such as A0, B0, Bl, etc. are examined.
[0142] In the second step, the motion shift or DV (1521) identified in the first step can be applied to the current block (1401) (e.g., the DV (1521) is added to the coordinates of the current block) to obtain sub-CU level motion information (e.g., including multiple MVs and reference indices) from the collocated picture (1512). In Figure 15In the illustrated example, the motion displacement or DV (1521) is set as the MV of a spatial neighboring block Al (e.g., block Al) of the current block (1401). For each sub-CU or sub-block (1531) in the current block (1401), the motion information of the sub-CU or sub-block (1531) can be derived using the motion information of a corresponding collocated block (1501) in the collocated picture (1512) (e.g., the motion information of the smallest motion grid covering the center sample of the collocated block (1501)). After identifying the motion information of a collocated sub-CU (1532) in the collocated block (1501), the motion information of the collocated sub-CU (1532) can be converted to the motion information of the current sub-CU (1531) (e.g., one or more MVs and one or more reference indices) using a scaling method, for example, in a manner similar to the TMVP process used in HEVC, where temporal motion scaling is applied to align the reference pictures of the temporal MVs with the reference pictures of the current CU.
[0143] The DV (1521) derived motion field of the current block (1401) can include the motion information of each sub-block (1531) in the current block (1401), such as one or more MVs and one or more associated reference indices. The motion field of the current block (1401) can also be referred to as an SbTMVP candidate, and corresponds to the DV (1521).
[0144] Figure 15 An example of a motion field or SbTMVP candidate for the current block (1401) is shown. The motion information of the bi-predicted sub-block (1531(1)) includes a first MV, a first index indicating a first reference picture in reference picture list 0 (L0), a second MV, and a second index indicating a second reference picture in reference picture list 1 (L1). In one example, the motion information of the uni-predicted sub-block (1531(2)) includes a MV and an index indicating a reference picture in L0 or L1.
[0145] In one example, DV (1521) is applied to the center position of the current block (1401) to locate the shifted center position in the collocated picture (1512). If the block including the shifted center position is not inter coded, the SbTMVP candidate is considered unavailable. Otherwise, if the block including the shifted center position (e.g., the collocated block (1501)) is inter coded, the motion information of the center position of the current block (1401) (referred to as the center motion of the current block (1401)) can be derived from the motion information of the block including the shifted center position in the collocated picture (1512). In one example, scaling process can be used to derive the center motion of the current block (1401) from the motion information of the block including the shifted center position in the collocated picture (1512). When the SbTMVP candidate is available, DV (1521) can be applied to find the corresponding sub-block (1532) in the collocated picture (1512) for each sub-block (1531) of the current block (1401). The motion information of the corresponding sub-block (1532) can be used to derive the motion information of the sub-block (1531) in the current block (1401), e.g., in the same manner as used to derive the center motion of the current block (1901). In one example, if the corresponding sub-block (1532) is not inter coded, the motion information of the current sub-block (1531) is set to the center motion of the current block (1401).
[0146] In some examples, bi-prediction with CU-level weight (BCW) can be used to differentially weight predictions from different reference pictures. In one example (e.g., HEVC), a bi-predicted signal is generated by averaging two prediction signals obtained from two different reference pictures and / or two prediction signals obtained using two different motion vectors. In some examples (e.g., VVC), the bi-prediction mode is extended beyond simple averaging to allow a weighted average of the two prediction signals, e.g., shown by equation (13):
[0147] P bi-pred = ((8 - w) x P0 + w x P1 + 4) » 3 Equation (13)
[0148] where P bi-pred represents the bi-predicted signal, P0 represents a first prediction signal from a first reference picture,
[0149] P1 represents a second prediction signal from a second reference picture, and w represents a weight parameter.
[0150] In some examples, the BCW weights in weighted bi-prediction allow for five values, e.g., w0e {-2, 3, 4, 5, 10}. In some examples, the BCW weights can be represented by a BCW weight index. For each bi-predicted CU, the value of the weight parameter w is determined in one of the following two ways: 1) for non-merge CUs, the weight index is written after the motion vector difference values, the weight index indicates the selected value from the list; 2) for merge CUs, the weight index is inferred from the merge candidate index based on the multiple neighboring blocks. In some examples, BCW is only applied to CUs with 256 or more luma samples (i.e., CU width times CU height is greater than or equal to 256). In one example, for low-delay pictures, all 5 weight values are used. For non-low-delay pictures, only 3 weight values are used (e.g., w e {3, 4, 5}).
[0151] In some examples, at the encoder side, a search algorithm such as a fast search algorithm is applied to find the weight index without significantly increasing the complexity of the encoder. When BCW is combined with AMVR, the unequal weights of 1-pel and 4-pel motion vector precision are conditionally checked only if the current picture is a low-delay picture.
[0152] In some examples, when BCW is combined with affine, affine ME is performed for unequal weights only when the affine mode is selected as the current best mode.
[0153] In some examples, when the two reference pictures in bi-prediction are the same, then the unequal weights are conditionally checked.
[0154] In some examples, the unequal weights are not searched when certain conditions are met, e.g., depending on the POC distance between the current picture and its reference pictures, the coding QP, and the temporal level.
[0155] In some examples, the weight index for BCW is coded using a first context-coded bin followed by a bypass-coded bin. The first context-coded bin indicates whether equal weights are used; and if unequal weights are used, the bypass-coded is used to signal an additional bin to indicate which unequal weight from the list of values of the weight parameter is used.
[0156] Some video codecs, such as the H.264 / AVC and HEVC standards, support an encoding tool called weighted prediction (WP) for efficient encoding of decaying video content. Support for WP is also added in the VVC standard. WP allows signaling of a weighting parameter (weight and offset) for each reference picture in each of the reference picture lists L0 and L1. Then, during motion compensation, the weight and offset of the corresponding reference picture are applied. WP and BCW are designed for different types of video content. To avoid the interaction between WP and BCW, which would complicate the VVC decoder design, if a CU uses WP, the BCW weight index is not signaled, and w is inferred to be 4 (i.e., equal weights are applied). For merge CUs, the weight index is inferred from neighboring blocks based on the merge candidate index. This can be applied to both the normal merge mode and the inherited affine merge mode. For the constructed affine merge mode, the affine motion information is constructed based on motion information of up to 3 blocks. The BCW index for a CU using the constructed affine merge mode is simply set to be equal to the BCW index of the first control point MV.
[0157] In VVC, CIIP and BCW cannot be jointly applied to a CU. When a CU is coded with CIIP mode, the BCW index of the current CU is set to 2, e.g., equal weights.
[0158] Bi-directional optical flow (BDOF) in VVC was previously known as BIO in JEM. Compared to the JEM version, the BDOF in VVC can be a simpler version that requires less computation, especially in terms of the number of multiplications and the multiplier size.
[0159] BDOF can be used to refine the bi-directional prediction signal of a CU at the 4x4 subblock level. BDOF can be applied to a CU if the following conditions are met:
[0160] (1) the CU is coded using the “true” bi-directional prediction mode, i.e., one of the two reference pictures is before the current picture in display order and the other is after the current picture in display order,
[0161] (2) the distance (e.g., POC difference) of the two reference pictures to the current picture is the same,
[0162] (3) both reference pictures are short-term reference pictures,
[0163] (4) the CU is not coded using affine mode or SbTMVP merge mode,
[0164] (5) the CU has more than 64 luma samples,
[0165] (6) CU height and CU width are both greater than or equal to 8 luma samples,
[0166] (7) BCW weight index indicates equal weights,
[0167] (8) Weighted prediction (WP) is not enabled for the current CU, and
[0168] (9) The current CU does not use CIIP mode.
[0169] BDOF can be applied only to the luma component. As the name BDOF indicates, the BDOF mode can be based on the concept of optical flow, which assumes that the motion of objects is smooth. For each 4x4 subblock, a motion refinement (v x ,v y ) can be calculated by minimizing the difference between the L0 and L1 prediction samples. The motion refinement can then be used to adjust the bi-prediction sample values in the 4x4 subblock. BDOF can include the following steps.
[0170] First, the horizontal and vertical gradients of the two prediction signals from reference list L0 and reference list L1 can be calculated by directly calculating the difference between two neighboring samples know k = 0, 1. The horizontal and vertical gradients can be provided in Equations (14) and (15) as follows:
[0171]
[0172]
[0173] where I (k) (i,j) can be the sample value at coordinates (i,j) of the prediction signal in list k, where k = 0, 1, and shift1 can be calculated based on the luma bit depth (bitDepth), shift1 = max(6, bitDepth - 6).
[0174] Then, the auto-correlation and cross-correlation of gradients S1, S2, S3, S5, and S6 can be calculated according to Equations (16)-(20) as follows:
[0175] S1 =∑ (i,j)∈Ω Abs(ψ x (i,j)), Equation (16)
[0176] S2 =∑ (i,j)∈Ω ψ x (i,j) · Sign(ψ y(i, j) Equation (17)
[0177] S3 = ∑ (i,j)∈Ω θ(i, j) · Sign(ψ x (i, j) Equation (18)
[0178] S5 = ∑ (i,j)∈Ω Abs(ψ y (i, j) Equation (19)
[0179] S6 = ∑ (i,j)∈Ω θ(i, j) · Sign(ψ y (i, j) Equation (20)
[0180] where ψ x (i, j), ψ y (i, j), and θ(i, j) can be provided in Equations (21)-(23), respectively:
[0181]
[0182]
[0183] θ(i, j) = (I (1) (i, j) » n b ) - (I (0) (i, j) » n b ) Equation (23)
[0184] where Ω can be a 6x6 window around the 4x4 subblock, and n a and n b values can be set equal to min(l, bitDepth - 11) and min(4, bitDepth - 8), respectively.
[0185] The motion refinement (v x , v y ) can then be derived using the cross-correlation terms and the auto-correlation terms using Equations (24) and (25) as follows:
[0186]
[0187]
[0188] where and is a floor function, and Based on the motion refinement and the gradients, an adjustment for each sample in the 4x4 subblock can be calculated based on Equation (26):
[0189]
[0190] Finally, the BDOF samples for a CU can be calculated from the adjusted bi-predicted samples according to equation (27) as follows:
[0191] pred BDOF (x, y) = (I (0) (x, y) + I (1) (x, y) + b(x, y) + o offset ) » shift equation (27)
[0192] These values can be chosen such that the multipliers in the BDOF process do not exceed 15 bits, and the maximum bit-width of intermediate parameters in the BDOF process can be kept within 32 bits.
[0193] To derive the gradient values, some prediction samples I (k) (i, j) outside the current CU boundary need to be generated in the list k (k = 0, 1). As shown in FIG. 16, the BDOF in VVC can use an extended row / column (1602) around the boundary (1606) of a CU (1604). To control the computational complexity of generating the out-of-boundary prediction samples, the prediction samples in the extended area (e.g., the un-shaded area in FIG. 16) can be generated by directly using the reference samples at the nearby integer positions (e.g., using floor() operation on the coordinates) without interpolation, and the prediction samples inside the CU (e.g., the shaded area in FIG. 16) can be generated using the normal 8-tap motion compensation interpolation filter. These extended sample values can only be used for gradient calculation. For the remaining steps in the BDOF process, if any sample and gradient values outside the CU boundary are needed, the sample and gradient values can be padded (e.g., repeated) from the nearest neighborhood of the sample and gradient values. Figure 16 Figure 16 Figure 16
[0194] When the width and / or height of a CU is greater than 16 luma samples, the CU can be partitioned into sub-blocks with width and / or height equal to 16 luma samples, and the sub-block boundaries can be treated as CU boundaries in the BDOF process. The maximum unit size of the BDOF process can be limited to 16x16. For each sub-block, the BDOF process can be skipped. The BDOF process can not be applied to a sub-block when the sum of absolute difference (SAD) between the initial L0 and L1 prediction samples is less than a threshold. The threshold can be set to equal to (8xWx(H»1), where W can represent the width of the sub-block and H can represent the height of the sub-block. To avoid the extra complexity of SAD calculation, the SAD between the initial L0 and L1 prediction samples calculated in the DMVR process can be reused in the BDOF process.
[0195] If BCW is enabled for the current block, i.e., the BCW weight index indicates unequal weights, bi-directional optical flow can be disabled. Similarly, if WP is enabled for the current block, i.e., the luma weight flag (e.g., luma_weight_lx_flag) is one for either of the two reference pictures, BDOF can also be disabled. BDOF can also be disabled when the CU is coded with symmetric MVD mode or CIIP mode.
[0196] To improve the precision of the MVs in merge mode, decoder-side motion vector refinement based on bilateral-matching (BM) can be applied, for example, in VVC. In bi-prediction operation, a refined MV can be searched around the initial MV in the reference picture list L0 and the reference picture list L1. The BM method can calculate the distortion between two candidate blocks in the reference picture list L0 and list L1.
[0197] Figure 17 An exemplary schematic diagram of BM-based decoder-side motion vector refinement is shown. As shown in FIG. 6, the BM-based decoder-side motion vector refinement can include the following steps: Figure 17As shown, a current picture (1702) can include a current block (1708). The current picture can include a reference picture list L0 (1704) and a reference picture list L1 (1706). The current block (1708) can include an initial reference block (1712) in the reference picture list L0 (1704) according to an initial motion vector MV0, and include an initial reference block (1714) in the reference picture list L1 (1706) according to an initial motion vector MV1. A search process can be performed around the initial MV0 in the reference picture list L0 (1704) and around the initial MV1 in the reference picture list L1 (1706). For example, a first candidate reference block (1710) can be identified in the reference picture list L0 (1704), and a first candidate reference block (1716) can be identified in the reference picture list L1 (1706). Based on each MV candidate (e.g., MV0’ and MV1’) around the initial MV (e.g., MV0 and MV1), a SAD between the candidate reference blocks (e.g., (1710) and (1716)) can be calculated. The MV candidate with the lowest SAD can become a refined MV and be used to generate a bi-prediction signal to predict the current block (1708).
[0198] The application of DMVR can be limited, and for example, in VVC, it can only be applied to CUs that are coded based on mode and features, as follows:
[0199] (1) CU-level merge mode with bi-predictive MVs,
[0200] (2) one reference picture is past and the other reference picture is future with respect to the current picture,
[0201] (3) the distance (e.g., POC difference) of the two reference pictures to the current picture is the same,
[0202] (4) both reference pictures are short-term reference pictures,
[0203] (5) the CU has more than 64 luma samples,
[0204] (6) both CU height and CU width are greater than or equal to 8 luma samples,
[0205] (7) the BCW weight index indicates equal weights,
[0206] (8) weighted prediction (WP) is not enabled for the current block, and
[0207] (9) CIIP mode is not used for the current block.
[0208] The refined MV derived from the DMVR process can be used to generate inter prediction samples and for temporal motion vector prediction for future picture coding. While the original MV can be used for deblocking process and for spatial motion vector prediction for future CU coding.
[0209] In DMVR, the search points can be centered around the initial MV, and the MV offset can follow the MV difference mirroring rule. In other words, any point examined by DMVR represented by the candidate MV pair (MV0, MV1) can follow the MV difference mirroring rule shown in equations (28) and (29):
[0210] MV0' = MV0 + MV_offset Equation (28)
[0211] MV1' = MV1 - MV_offset Equation (29)
[0212] where MV_offset can represent the refinement offset between the initial MV and the refined MV in one of the reference pictures. The refinement search range can be two integer luma samples from the initial MV. The search can include an integer sample offset search stage and a fractional sample refinement stage.
[0213] For example, a 25-point full search can be applied for the integer sample offset search. The SAD of the initial MV pair can be first calculated. If the SAD of the initial MV pair is less than a threshold, the integer sample stage of DMVR can terminate. Otherwise, the SAD of the remaining 24 points can be calculated and examined in a scan order, such as a raster scan order. The point with the smallest SAD can be selected as the output of the integer sample offset search stage. To reduce the loss due to the uncertainty of DMVR refinement, the original MV can be prioritized in the DMVR process. The SAD between the reference blocks referenced by the initial MV candidate can reduce the SAD value by 1 / 4.
[0214] The fractional sample refinement can be performed after the integer sample search. To save computational complexity, the fractional sample refinement can be derived by using a parametric error surface equation instead of additional search by SAD comparison. The fractional sample refinement can be conditionally invoked based on the output of the integer sample search stage. The fractional sample refinement can be further applied when the integer sample search stage ends with the center having the smallest SAD in the first iteration search or the second iteration search.
[0215] In the parametric error surface based sub-pixel offset estimation, the center position cost and the costs at four neighboring positions from the center can be used to fit a 2-D parabolic error surface equation based on equation (30):
[0216] E(x, y) = A(x - xmin ) 2 + B(y - y min ) 2 + C Equation (30)
[0217] where (x min , y min ) can correspond to the fractional position with the minimum cost, and C can correspond to the minimum cost value. By solving Equation (30) using the cost values of the five search points, (x min , y min ) can be calculated in Equations (31) and (32).
[0218] x min = (E(-1, 0) - E(1, 0)) / (2(E(-1, 0) + E(1, 0) - 2E(0, 0))) Equation (31)
[0219] y min = (E(0, -1) - E(0, 1)) / (2((E(0, -1) + E(0, 1) - 2E(0, 0))) Equation (32)
[0220] Because all cost values are positive, the minimum value is E(0, 0), so the values of x min and y min can be automatically limited between -8 and 8. The constraint of the values of x min and y min can correspond to the half-pel (or pixel) offset with 1 / 16-6 pel MV precision in VVC. The calculated fraction (x min , y min ) can be added to the integer distance refined MV to obtain the sub-pixel accurate refined delta MV.
[0221] For example, in VVC, bilinear interpolation and sample padding can be applied. For example, the resolution of the MV can be 1 / 16 luma sample. The samples at the fractional positions can be interpolated using an 8-tap interpolation filter. In DMVR, the search points can surround the initial fractional pel MV with integer sample offsets, so the samples at the fractional positions need to be interpolated for the DMVR search process. To reduce the computational complexity, a bilinear interpolation filter can be used to generate the fractional samples for the search process of DMVR. In another important effect, by using a bilinear filter with a 2-sample search range, DMVR does not access more reference samples compared to the normal motion compensation process. After obtaining the refined MV through the DMVR search process, a normal 8-tap interpolation filter can be applied to generate the final prediction. To not access more reference samples compared to the normal MC process, samples can be padded from the available samples that can not be needed for the interpolation process based on the original MV, but can be needed for the interpolation process based on the refined MV.
[0222] When the width and / or height of a CU is larger than 16 luma samples, the CU can be further partitioned into sub-blocks with width and / or height equal to 16 luma samples. The maximum unit size of the DMVR search process can be limited to 16x16.
[0223] In some examples (e.g., VVC), a geometric partitioning mode (GPM) is supported for inter prediction. The geometric partitioning mode is signaled using a CU-level flag as one of the merge modes with other merge modes (e.g., regular merge mode, MMVD mode, CIIP mode, sub-block merge mode, etc.). In some examples, for each possible CU size (w x h = 2 m ×2 n where n, n e {3…6}, except for 8x64 and 64x8, the geometric partitioning mode supports 64 partitions in total.
[0224] In some examples, when the geometric partitioning mode is used, the CU is partitioned into two parts by a straight line that is also referred to as a split line, which is geometrically positioned.
[0225] Figure 18 Examples of split lines for the geometric partitioning mode in some examples are shown. The split lines are grouped by the same angle. Specifically, Figure 18 Each rectangle (1801) in (1800) represents a CU. A plurality of parallel lines are shown in each rectangle. The plurality of parallel lines correspond to split lines of the same angle. The plurality of split lines have different offsets.
[0226] The position of the split line can be derived mathematically based on the specific split's angle and offset parameters. Each part of the two geometric splits that are split by the split line in the CU use its own motion for inter prediction; each split is only allowed single directional prediction. Thus, each part has one motion vector and one reference index. The single directional prediction motion constraint is applied to ensure that the CU in GPM mode can be coded as traditional bi-directional prediction, e.g., two motion compensated predictions are performed for each CU.
[0227] In some examples, when the geometric partition mode is used for the current CU, then further signaling the geometric partition index (e.g., indicating the angle and offset) of the partition mode that indicates the geometric partition and two merge indices (one for each split). In some examples, the number of maximum GPM candidate sizes is explicitly signaled in the SPS and the syntax binarization is specified for the GPM merge indices. After predicting each part of the geometric split, a hybrid process with adaptive weights is used to adjust the sample values along the geometric split edge to obtain the prediction signal for the whole CU. Then, like other prediction modes, the transform and quantization processes are applied to the whole CU. Finally, the motion field of the CU predicted using the geometric partition mode is stored.
[0228] In the geometric partition mode, in some examples, a candidate list referred to as the geometric uni-prediction candidate list is derived directly from the merge candidate list constructed according to the extended merge prediction process (also referred to as the extended merge candidate list). In one example, n denotes the index of the uni-prediction motion in the geometric uni-prediction candidate list (merge index). The LX motion vector of the nth candidate in the extended merge candidate list (X equals the parity of n) is used as the nth candidate of the geometric uni-prediction candidate list for the geometric partition mode.
[0229] Figure 19 A diagram showing uni-prediction MV selection for the geometric partition mode in some examples is shown. The motion vector is in the Figure 19For example, when the merge index n is 0 and the parity is 0, then the L0 motion vector of the n-th candidate in the extended merge candidate list is used as the n-th motion vector in the geometric uni-prediction candidate list for the geometric partition mode; when the merge index n is 1 and the parity is 1, then the L1 motion vector of the n-th candidate in the extended merge candidate list is used as the n-th motion vector in the geometric uni-prediction candidate list for the geometric partition mode; when the merge index n is 2 and the parity is 0, then the L0 motion vector of the n-th candidate in the extended merge candidate list is used as the n-th motion vector in the geometric uni-prediction candidate list for the geometric partition mode; when the merge index n is 3 and the parity is 1, then the L1 motion vector of the n-th candidate in the extended merge candidate list is used as the n-th motion vector in the geometric uni-prediction candidate list for the geometric partition mode; when the merge index n is 4 and the parity is 0, then the L0 motion vector of the n-th candidate in the extended merge candidate list is used as the n-th motion vector in the geometric uni-prediction candidate list for the geometric partition mode.
[0230] In some examples, when the LX motion vector of the n-th candidate in the extended merge candidate list does not exist, then the L(1-X) motion vector of the same candidate is used as the uni-prediction motion vector for the geometric partition mode.
[0231] In some examples, after each part of the geometric partition is predicted using its own motion, a blending is applied to the two prediction signals to derive the samples around the edges of the geometric partition. In one example, the blending weight for each position of the CU is derived based on the distance between the single position and the partition edge.
[0232] For example, the distance of a position (x, y) to the partition edge is derived according to the following equations (33)-(36):
[0233]
[0234]
[0235]
[0236]
[0237] where i, j are the indices of the angle and offset of the geometric partition, which depend on the signaled geometric partition index. x,j and the sign of the angle index i. y,j
[0238] The weight of each part of the geometric partition is derived according to the following equations (37)-(39):
[0239] wldxL(x, y) = partldx? 32 + d(x, y) : 32 - d(x, y) Equation (37)
[0240]
[0241] w1(x, y) = 1 - w0(x, y) Equation (39)
[0242] partldx depends on the angle index i.
[0243] Figure 20 A schematic diagram showing the blending weights w0 in the generated geometric partition mode in some examples is shown.
[0244] The motion field information of the geometric partition mode is stored appropriately. In some examples, a first motion vector Mv1 from a first portion of the geometric partition, a second motion vector Mv2 from a second portion of the geometric partition, and a combination My of Mv1 and Mv2 are stored in the motion field of the geometric partition mode that is coded as a CU.
[0245] The stored motion vector type at each individual position in the motion field is determined according to Equation (40):
[0246] sType = abs(motionldx) < 32? 2 : (motionldx < 0? (1 - partldx) : partldx) Equation (40)
[0247] where motionldx is equal to d(4x+2,4y+2), which is recomputed according to Equation (33). partldx depends on the angle index i.
[0248] When sType is equal to 0, Mv1 is stored in the respective motion field; when sType is equal to 1, Mv2 is stored in the respective motion field; when sType is equal to 2, a combination Mv of Mv1 and Mv2 is stored. The combined Mv is generated using the following procedure: when Mv1 and Mv2 are from different reference picture lists (one from L0 and the other from L1), then Mv1 and Mv2 are simply combined to form a bi-predictive motion vector. Otherwise, in one example, when Mv1 and Mv2 are from the same list, only the uni-predictive motion Mv2 is stored.
[0249] Some video coding standards, such as HEVC and AVC / H.264, use a fixed motion vector resolution of quarter luma samples. According to one aspect of the disclosure, an optimal trade-off between displacement vector rate and prediction error rate must be chosen to achieve overall rate-distortion optimality. Some video coding standards, such as VVC, allow the selection of motion vector resolution at the coding block level, thus, for the signaling of motion parameters, a trade-off between bit rate and fidelity. Specifically, VVC enables adaptive motion vector resolution (AMVR) in the AMVR mode. The AMVR mode is signaled at the coding block level if at least one component of the MVD is not equal to zero. The motion vector predictors are rounded to a given resolution to guarantee that the resulting motion vector falls on the grid of the given resolution. For each given resolution, a number of shift bits (denoted by AmvrShift) is defined for the motion vector difference to specify the resolution of the motion vector difference value by a left shift operation of AmvrShift-bits. When the AMVR mode is enabled, a given motion vector difference value (denoted as MvdL0 and MvdL1 in the AMVP mode, and MvdCpL0 and MvdCpL1 in the affine AMVP mode) is modified.
[0250] Figure 21 A table is shown that determines the number of shift bits AmvrShift according to a flag and / or syntax.
[0251] Figure 22 A portion of the VVC standard is shown that modifies the motion vector difference in the AMVP mode and the affine AMVP mode when the AMVR mode is enabled.
[0252] According to one aspect of the disclosure, a motion vector can point to a region outside the boundary of a reference picture. In some examples, when certain reference samples used in motion compensation go beyond the picture boundary, a constraint is applied to bi-prediction.
[0253] In one example, when a current pixel with bi-predicted motion vector has a motion vector pointing to a location outside the picture boundary by a distance threshold on one of the two reference lists, the motion vector of that reference list is considered to be out-of-bound and the inter prediction is changed to uni-prediction. In this example, only the motion vector of the other reference list that does not go beyond the boundary will be used for uni-prediction. In one example, when the MVs of both reference lists go beyond the boundary, the bi-prediction is not constrained.
[0254] In another example, the constraint is applied to bi-prediction at the sub-block level.
[0255] Figure 23A diagram showing a sub-block level bi-prediction constraint in some examples is shown.
[0256] As Figure 23 shown, for each of the NxN sub-blocks (e.g., 16 sub-blocks in the current block) within a coding block (e.g., the current block) with inter bi-predicted MVs, when the motion vector on one of the reference lists (e.g., L0) points outside the boundary of the reference picture by more than M pixel threshold, the sub-block can change to uni-prediction mode using only the MVs on the reference list that do not point outside the boundary threshold of the corresponding reference picture. As Figure 23 shown, sub-blocks 2310 at the edges of the current block switch to uni-prediction mode due to the constraint. Figure 23
[0257] In some examples, when bi-prediction changes to uni-prediction due to the boundary condition being exceeded, bi-prediction related tools can be disabled or modified. In one example, BDOF can be disabled when the bi-prediction restriction is applied and uni-prediction is used.
[0258] In some examples, techniques for picture boundary padding can be used. The padded portion outside the picture boundary can be used to predict other pictures.
[0259] For example, in VVC and ECM-4.0, the extended picture region is a region around the picture with a size of (maxCUwidth + 16) in each direction of the picture boundary. Pixels in the extended region are derived by repeated boundary padding. When a reference block is partially or fully outside the picture boundary (OOB), the repeated padded pixels in the extended region can be similar to the pixels in the reference picture for motion compensation.
[0260] Figure 24A A diagram showing a reference picture (2401) with an extended region (also referred to as a repeated padded region or a padded region) (2410) is shown, which is padded according to repeated boundary padding in some examples (e.g., ECM-4.0). In one example, in a first round of repeated boundary padding, a first pixel in the extended region that is immediately adjacent to a pixel within the reference picture boundary is padded based on the pixel within the reference picture boundary; in a second round of repeated boundary padding, a second pixel in the extended region that is immediately adjacent to the first pixel is padded based on the first pixel; in one example, multiple rounds of repeated boundary padding can continue until all pixels in the extended region are padded.
[0261] In related examples, techniques for motion compensated boundary padding can be used. For example, samples outside the picture boundary are obtained by motion compensation rather than just using repeated padding.
[0262] Figure 24B A schematic diagram of a reference picture with an extended area (also referred to as a padding area) is shown, which in some examples is padded according to motion-compensated padding and repeated padding. In Figure 24B In examples, the total padding area size is increased by 64 compared to Figure 24A the padding area (2410) in . The extended area includes a first portion (2470) that can be padded according to motion-compensated padding. The first portion (2470) is also referred to as a motion-compensation (MC) padding area. The extended area also includes a repeated padding area (2480) that is padded by non-normative repeated padding.
[0263] Figure 25 A schematic diagram of motion-compensation (MC) padding in some related examples is shown. In MC padding, an Lx4 or 4x4 padding block (shown as MCP Blk) in the MC padding area of the current picture is derived using the MV of a 4x4 boundary block (shown as BBlk) in the current picture. As shown in Figure 25 The value L is derived as the distance of the reference block (shown as Ref BBlk) of the boundary block to the picture boundary in the reference picture. In examples, when the boundary block is intra coded, then the MV is not available and L is set to equal 0. If L is less than 64, then the remaining portion of the MC padding area is padded with repeated padded samples using repeated padding in examples.
[0264] In the case of bi-directional inter prediction, only one prediction direction is used in MC (boundary) padding, which has a motion vector pointing to a pixel position further away from the picture boundary in the reference picture in the padding direction.
[0265] In some examples, the pixels in the MC padding block are corrected with an offset equal to the difference between the DC value of the reconstructed boundary block (e.g., BBlk in Figure 25 ) and the DC value of the reference block (e.g., Ref BBlk in Figure 25 ) in the reference picture corresponding to the reconstructed boundary block.
[0266] According to one aspect of the disclosure, the motion-compensated boundary padding techniques in the above related examples can not be efficient. For example, in the related examples, only the MV of the corresponding boundary block is used. There is a certain chance that repeated padding is used instead of motion-compensated boundary padding (also referred to as motion-compensation padding in some examples).
[0267] Some aspects of the disclosure provide additional techniques to derive the MVs for motion-compensated boundary padding.
[0268] According to some aspects of the disclosure, a simplified skip mode can be used for motion compensated padding blocks. In the simplified skip mode, a block in the MC padded region (also referred to as a MC padded block) is predicted according to a motion vector pointing to a reference block in a reference picture.
[0269] According to one aspect of the disclosure, the motion information of a MC padded block is derived from one of a plurality of neighboring candidates at the picture boundary. Note that since multiple neighboring candidates are used, the probability of motion compensation boundary padding without motion information is reduced.
[0270] Figure 26 A diagram showing an example of deriving motion information for a motion compensation padding (MCP) block in some examples is shown.
[0271] In Figure 26 In an example, 3 neighboring candidate positions A, B and C can be used to derive the motion information of a MCP block.
[0272] In Figure 26 In an example, when the MCP block (2610) is located to the left of a picture vertical boundary, position A is the right neighboring position, position B is the top right neighboring position, and position C is the bottom right neighboring position. When the MCP block (2620) is located to the right of a picture vertical boundary, position A is the left neighboring position, position B is the top left neighboring position, and position C is the bottom left neighboring position.
[0273] In Figure 26 In an example, when the MCP block (2630) is located at the top of a picture horizontal boundary, position A is the bottom neighboring position, position B is the left bottom neighboring position, and position C is the right bottom neighboring position. When the MCP block (2640) is located at the bottom of a picture horizontal boundary, position A is the top neighboring position, position B is the left top neighboring position, and position C is the right top neighboring position.
[0274] In some embodiments, each candidate position corresponds to a block size of NxN. In one example, N is equal to 4. In another example, N is equal to 8.
[0275] In some embodiments, a MCP region is defined by four rectangles extending from a picture boundary. In Figure 26 In an example, the MCP region includes four rectangles (2601)-(2604). The range parameter R of rectangles (2601)-(2604) defines the luma samples extending outward from the picture boundary. In one example, R is equal to 64 luma samples. In another example, R is equal to the padding range.
[0276] In some embodiments, as Figure 26As shown, for top / bottom boundaries, the MCP block size is P x L, and for left / right boundaries, the MCP block size is L x P. In examples where L is equal to the MCP range R or equal to the distance of the reference block to the reference picture boundary (e.g., similar to Figure 25 ) the smaller one is used. In some examples, P is a predefined size in luma samples. In one example, P is equal to 4. In another example, P is equal to 8.
[0277] In some embodiments, 3 neighboring positions (e.g., A, B, and C in Figure 26 ) can be examined in order, e.g., A→B→C. The first available candidate in the examination order can be used as the motion information predictor for the MCP block. The motion information can then be used for the prediction of the MCP block to generate the padded samples in the MCP block.
[0278] In some embodiments, the neighboring position selected for predicting the MCP block can be signaled. In some examples, one bin is used as a candidate index to signal one of the first 2 available neighboring positions (also referred to as neighboring candidates or neighboring candidate positions). In some examples, variable length coding can be used to signal one of the 3 available neighboring candidates. In some examples, signaling can not be necessary when only one candidate block has available motion information, and the available candidate can be used implicitly. In some examples, the signaling for all MCP blocks can be done once for the entire frame in a predefined order.
[0279] In one example, the predefined order starts from the leftmost MCP block in the top MCP region, and follows a clockwise direction for all MCP blocks until the topmost MCP block in the leftmost MCP region.
[0280] Figure 27 A picture (2700) with picture boundaries is shown in some examples. MCP regions (2701)-(2704) extend from the picture boundaries. Figure 27 A predefined order is also shown, which starts from the leftmost MCP block (2710) in the top MCP region (2701), and follows a clockwise direction for all MCP blocks until the topmost MCP block (2799) in the leftmost MCP region (2704).
[0281] In another example, the predefined order can be from left to right for the top MCP region (2701), from left to right for the bottom MCP region (2703), from top to bottom for the left MCP region (2704), and from top to bottom for the right MCP region (2702).
[0282] In some embodiments, the pixels in the MCP block are corrected with an offset. The offset can be determined as the difference between the DC value of the reconstructed candidate block (e.g., the candidate block at one of the A, B, C positions) and the DC value of the reference block in the reference picture that corresponds to the reconstructed candidate block.
[0283] According to another aspect of the disclosure, for the MCP blocks near a picture boundary (e.g., left boundary, right boundary, top boundary, bottom boundary) of a picture, the motion information is derived from all inter-miniblocks (e.g., each 4x4 block in VVC) at the picture boundary in the picture. For example, the most commonly used motion information from the inter-miniblocks at the picture boundary is determined and used for the motion compensation of all the MCP blocks near the picture boundary.
[0284] In some examples, the use of motion information is collected from all inter-miniblocks at the respective picture boundary. In some examples, the use of motion information is collected only from the inter-miniblocks of reference index 0 of L0 and / or L1. In one example, the use of motion information is collected only from the inter-miniblocks of reference index 0 of L0. In another example, the use of motion information is collected only from the inter-miniblocks of reference index 1 of L1.
[0285] According to another aspect of the disclosure, for the MCP blocks next to a picture boundary (e.g., left boundary, right boundary, top boundary, bottom boundary), the motion information is derived from all inter-miniblocks at the respective picture boundary. For example, the average motion vector is determined and used for the motion compensation of all the MCP blocks near the picture boundary.
[0286] In some examples, the average of the motion vectors is calculated based on all inter-miniblocks at the respective picture boundary. In some examples, the average of the motion vectors is calculated based only on the inter-miniblocks of reference index 0 of L0 and / or L1. In one example, the average of the motion vectors is calculated based only on the inter-miniblocks of reference index 0 of L0. In another example, the average of the motion vectors is calculated based only on the inter-miniblocks of reference index 1 of L1.
[0287] According to one aspect of the disclosure, when a boundary block is intra-coded and no motion information is available for the related MCP block, a repetition padding can be used for the related MCP block.
[0288] According to one aspect of the disclosure, the motion information of a MCP block is derived from spatial candidates and temporal candidates (e.g., one of the multiple neighboring candidates at the picture boundary and N temporal collocated blocks (temporal candidates) of the current block (e.g., the current MCP block)). Example values of N include, but are not limited to, 1, 2, 3, 4,....
[0289] In some examples, the N temporally collocated blocks refer to N sub-blocks located at N predefined relative positions of a block having the same coordinates as the current block in a reference picture.
[0290] Figure 28 A schematic diagram of 2 temporally collocated blocks in some examples is shown. In Figure 28 the block (2801) is located in a reference picture and has the same coordinates as a current MCP block in a current picture. Two temporally collocated blocks at predefined relative positions are shown as block C0 at the middle position and block C1 at the lower right position.
[0291] In some examples, a plurality of spatial neighboring blocks of the MCP block (e.g., at neighboring positions) and N temporally collocated blocks of the MCP block are scanned, and the plurality of spatial neighboring blocks are checked before the N temporally collocated blocks, and a first available candidate block can be used as a motion information predictor for the MCP block.
[0292] According to some aspects of the disclosure, for a sample located at a picture boundary in a padding region of a picture, a further adjustment to the sample value can be performed. In some examples, the adjustment to the sample value is based on a deblocking process on a padding sample located at the picture boundary. In some examples, the adjustment to the sample value is based on a smoothing filtering process on a padding sample located at the picture boundary.
[0293] Figure 29 A flowchart of a process (2900) summarizing an embodiment according to the disclosure is shown. The process (2900) can be used in a video encoder. In various embodiments, the process (2900) is performed by processing circuitry, such as processing circuitry performing the functions of a video encoder (103), processing circuitry performing the functions of a video encoder (303), and the like. In some embodiments, the process (2900) is implemented in software instructions, so when the processing circuitry executes the software instructions, the processing circuitry performs the process (2900). The process starts from (S2901) and proceeds to (S2910).
[0294] At (S2910), samples in the reconstructed picture are determined.
[0295] At (S2920), for a motion compensated padding (MCP) block located in a padding region of a picture outside the picture and close to a picture boundary of the picture, a particular candidate is selected from a plurality of candidates located at the picture boundary within the picture.
[0296] At (S2930), a motion vector of the MCP block is determined according to motion information of the particular candidate.
[0297] At (S2940), at least one sample in the MCP block is reconstructed according to the motion vector.
[0298] At (S2950), a syntax element is encoded in the bitstream carrying the picture, the syntax element indicating the particular candidate selected from the plurality of candidates.
[0299] In some examples, a binary file is included in the bitstream, and the binary file indicates one of the first two available candidates from the plurality of candidates.
[0300] In some examples, the syntax element indicating the particular candidate from the three available candidates is encoded by variable length coding.
[0301] In some examples, it is determined that only one candidate from the plurality of candidates has available motion information. Then, in response to the only one candidate, it is determined that the syntax element in the bitstream does not need to be encoded.
[0302] In some examples, a bit sequence is encoded in the bitstream, the bit sequence indicating respective candidates for a plurality of MCP blocks located outside a picture boundary of a picture according to a predefined order of the plurality of MCP blocks.
[0303] The process then proceeds to (S2999) and terminates.
[0304] The process (2900) can be adjusted as appropriate. One or more steps in the process (2900) can be modified and / or omitted. Additional steps can be added. Any suitable implementation order can be employed.
[0305] Figure 30 A flowchart showing an overview of a process (3000) in accordance with embodiments of the present disclosure is shown. The process (3000) can be used in a video decoder. In various embodiments, the process (3000) is performed by processing circuitry, such as processing circuitry performing the functions of the video decoder (110), processing circuitry performing the functions of the video decoder (210), etc. In some embodiments, the process (3000) is implemented in software instructions, so when the processing circuitry executes the software instructions, the processing circuitry performs the process (3000). The process starts from (S3001), and proceeds to (S3010).
[0306] At (S3010), a bitstream carrying a plurality of pictures is received.
[0307] At (S3020), samples in a picture from the plurality of pictures are reconstructed.
[0308] At (S3030), a motion vector for a motion compensated padding MCP block located outside the picture and close to a picture boundary of the picture is determined according to a plurality of candidates located at the picture boundary within the picture.
[0309] At (S3040), at least one sample in the MCP block is reconstructed according to the motion vector.
[0310] In some examples, the candidates of the plurality of candidates correspond to neighboring blocks of the MCP block, the neighboring blocks having a predetermined block size.
[0311] In some examples, the MCP region corresponds to a rectangle extending outward from a picture boundary by an MCP range for padding according to motion compensation.
[0312] To determine the motion vector of the MCP block, in some examples, the plurality of candidates is examined according to a predetermined order to determine a first available candidate having motion information. The motion vector of the MCP block is determined according to the motion information of the first available candidate.
[0313] To determine the motion vector of the MCP block, in some examples, a signal indicating a particular candidate of the plurality of candidates is decoded from a bitstream. The motion vector of the MCP block is then determined according to the motion information of the particular candidate. In one example, a bit sequence is decoded from the bitstream. The bit sequence indicates respective candidates of a plurality of MCP blocks according to a predetermined order of the plurality of MCP blocks located outside a picture boundary of a picture.
[0314] To reconstruct at least one sample in the MCP block according to the motion vector, in some examples, a first reference block in a reference picture corresponding to the MCP block is determined according to the motion vector, and the MCP block is reconstructed according to the first reference block. In some examples, an offset is determined based on a first DC value of a candidate block used to determine the motion vector and a second DC value of a second reference block in the reference picture corresponding to the candidate block. The offset is then applied to the at least one sample in the MCP block.
[0315] In some examples, to determine the motion vector of the MCP block, a most commonly used motion information in a set of inter-minimal blocks at the picture boundary is determined, and the motion vector is determined according to the most commonly used motion information.
[0316] In some examples, to determine the motion vector of the MCP block, an average motion vector is determined according to motion information of a set of inter-minimal blocks at the picture boundary, and the motion vector is determined according to the average motion vector.
[0317] In some examples, the motion vector of the MCP block is determined according to a plurality of candidates within the picture boundary, and one or more collocated blocks of the MCP block in a reference picture. For example, the plurality of candidates is examined before examining the one or more collocated blocks of the MCP block in the reference picture to determine first available motion information. The motion vector of the MCP block is then determined according to the first available motion information.
[0318] In some examples, at least one of a deblocking filter and a smoothing filter is applied to at least one sample in the MCP block.
[0319] The process then proceeds to (S3099) and terminates.
[0320] The process (3000) can be adjusted as appropriate. One or more steps in the process (3000) can be modified and / or omitted. Additional steps can be added. Any suitable implementation sequence can be used.
[0321] The techniques described above, can be implemented as computer software using computer readable instructions and physically stored in one or more computer-readable media. For example, Figure 31 A computer system (3100) suitable for implementing certain embodiments of the disclosed subject matter is shown.
[0322] The computer software can be coded using any suitable machine code or computer language that can be subject to assembly, compilation, linking, or similar
[0323] The instructions can be executed on various types of computers or components thereof, including, for example, personal computers, tablet computers, servers, smartphones, gaming devices, internet of things devices, and the like.
[0324] Figure 31 The components of the computer system (3100) shown are exemplary in nature and are not intended to suggest any limitation as to the scope of use or functionality of the computer software implementing embodiments of the present disclosure. Neither should the configuration of components be interpreted as having any dependency or requirement relating to any one or combination of components illustrated in the exemplary embodiment of a computer system (3100).
[0325] The computer system (3100) can include certain human interface input devices. Such a human interface input device can be responsive to
[0326] The input human interface devices can include one or more of (only one of each is shown): keyboard (3101), mouse (3102), touchpad (3103), touchscreen (3110), data glove (not shown), joystick (3105), microphone (3106), scanner (3107), camera (3108)
[0327] The computer system (3100) can also include certain human interface output devices. Such human interface output devices can be stimulating one or more of the human senses of sight, touch, taste, smell, and hearing. For example, such human interface output devices can include a display (3110) (which can be virtual or virtualized), a projector, speakers (3109), haptic feedback
[0328] The computer system (3100) can also include human accessible storage devices and their associated media and interfaces, for example, an optical disk drive (3120) with
[0329] Those skilled in the art should also understand that the term "computer readable medium" used in connection with the presently disclosed subject matter does not encompass transitory
[0330] The computer system (3100) can also include an interface (3154) to one or more communication networks (3155). Networks can for example be wireless networks, wire-line networks, optical networks. Networks can further be local networks, wide-area networks, metropolitan area networks, vehicular and industrial networks, real-time networks, delay-tolerant networks, etc. Examples of networks include local area networks such as Ethernet networks, wireless LANs, cellular networks to include GSM, 3G, 4G, 5G, LTE and the like, TV wireline or wireless wide area digital networks to include cable networks, satellite TV, and terrestrial TV, vehicular and industrial to include CANBus, etc. Certain networks commonly require various permissions, certificates, keys, and / or other credentials to prove access rights, and these permissions, certificates, keys, and / or other credentials can be stored in database (3120) and / or in a smartcard (3122) read by the computer system (3100). Certain networks also can use various encryption and / or decryption keys and algorithms, and these can be stored in database (3120) and / or in a smartcard (3122) read by the computer system (3100). These permissions, keys, and algorithms can be used by the network interface (3154) and / or by software applications running on the computer system (3100) to prove their right to use the network (3155) and / or to access certain features or data of the network (3155). The computer system (3100) can use any one of these networks (3155) to communicate with outside entities, including the user of the computer system (3100). Such communication can be done in the form of requests for and responses to retrieve, transmit, receive, and process data, control commands, and other communications. As described above, certain protocols and protocol stacks can be used on each of those networks and network interfaces.
[0331] The above-described human interface devices, human-accessible storage devices, and network interfaces can be attached to the core (3140) of the computer system (3100).
[0332] The core (3140) can include one or more Central Processing Units (CPUs) (3141), Graphics Processing Units (GPUs) (3142), specialized programmable processing units in the form of Field Programmable Gate Areas (FPGAs) (3143), hardware accelerators for certain tasks (3144), a graphics adapter (3150), and so forth. These devices, along with Read-only memory (ROM) (3145), Random-access memory (3146), internal mass storage such as internal non-user accessible hard drives, SSDs, and the like (3147), can be connected to the peripheral devices via an interconnect such as a system bus (3148). In some computer systems, the system bus (3148) can be accessible via one or more physical plugs to enable extensions via additional CPUs, GPUs, and the like. The peripheral devices can be connected to the system bus (3148) directly, or by way of a peripheral bus (3149). In one example, the touchscreen (3110) can be connected to the graphics adapter (3150). The architecture of the peripheral bus (3149) can include PCI, USB, and the like.
[0333] CPUs (3141), GPUs (3142), FPGAs (3143), and accelerators (3144) can execute certain instructions that can be combined to form the aforementioned computer code. That computer code can also be stored in ROM (3145) or RAM (3146). Transitional data can be also stored in RAM (3146), whereas permanent data can be stored for example, in the internal mass storage (3147). Fast storage and retrieval speeds of any storage devices can be facilitated by the use of caches, which can be closely associated with one or more CPU(s) (3141), GPU(s) (3142), mass storage (3147), ROM (3145), RAM (3146), and the like.
[0334] The computer software can be implemented as computer code that is stored on a computer readable medium, which can be any medium that is capable of storing information. The computer software can be implemented as computer code that is stored on a computer readable medium, which can be any medium that is capable of storing information.
[0335] As non-limiting examples, a computer system with architecture (3100), and specifically the core (3140), can provide functionality as a result of the execution of software contained in the one or more tangible computer-readable media by one or more processors including CPUs, GPUs, FPGAs, accelerators, and the like. Such computer-readable media can be media associated with user-accessible mass storage as introduced above, as well as certain non-transitory memory of the core (3140), such as the internal mass storage (3147) or ROM (3145). Software implementing various embodiments of the present disclosure can be stored in such devices and executed by the core (3140). A computer-readable medium containing the software, or physical media storing such software, may
[0336] While the present disclosure has described a number of exemplary embodiments, there are alterations, modifications, various replacements and various equivalent equivalents that fall within the scope of the present disclosure. Therefore, it should be understood that those skilled in the art will be able to design many systems and methods, although not explicitly shown or described herein, embodying the principles of the present disclosure, thus falling within the spirit and scope of the present disclosure.
Claims
1. A method for video decoding, characterized in that, include: Receive a bitstream carrying multiple images; Identify at least one motion compensation filled MCP block located outside the image and close to the image boundary within the motion compensation filled MCP region. Based on multiple candidates located at the image boundary within the image, a motion vector for the motion compensation filling MCP block is determined. These multiple candidates include at least a first candidate, a second candidate, and a third candidate. The first candidate is located at a first position at the image boundary within the image; the second candidate is located at a second position at the image boundary within the image; and the third candidate is located at a third position at the image boundary within the image. The first candidate is adjacent to a first corner of the MCP block; the second candidate is adjacent to a second corner of the MCP block; and the third candidate is adjacent to (i) the second corner of the MCP block and (ii) the second candidate. At least one sample in the motion-compensated filled MCP block is reconstructed based on the motion vector used for the motion-compensated boundary filling.
2. The method according to claim 1, characterized in that, The candidate among the plurality of candidates corresponds to an adjacent block used for the motion compensation filling MCP block, the adjacent block having a predetermined block size.
3. The method according to claim 1, characterized in that, The motion compensation fill (MCP) region corresponds to a rectangle extending outward from the image boundary to the MCP range, for filling according to motion compensation.
4. The method according to any one of claims 1 to 3, characterized in that, Determining the motion vector of the motion-compensated MCP block includes: A first available candidate with motion information is selected from the plurality of candidates according to a predetermined order; and The motion vector of the motion-compensated filling MCP block is determined based on the motion information of the first available candidate.
5. The method according to any one of claims 1 to 3, characterized in that, Determining the motion vector of the motion-compensated MCP block includes: Decode from the bitstream a signal indicating a specific candidate among the plurality of candidates; and The motion vector of the motion-compensated MCP block is determined based on the motion information of the specific candidate.
6. The method according to claim 5, characterized in that, The decoded signal indicating a specific candidate includes: A bit sequence is decoded from the bit stream, the bit sequence indicating corresponding candidates of the plurality of motion-compensated fill MCP blocks according to a predetermined order of the plurality of motion-compensated fill MCP blocks located outside the image boundary of the image.
7. The method according to any one of claims 1 to 3, characterized in that, The step of reconstructing at least one sample in the motion-compensated filled MCP block based on motion vectors includes: Based on the motion vector, a first reference block in the reference image is determined for use in the motion compensation filling MCP block; and The motion-compensated filled MCP block is reconstructed based on the first reference block.
8. The method according to claim 7, characterized in that, The reconstruction of the motion-compensated filled MCP block includes: The offset is determined based on a first DC value of a candidate block used to determine the motion vector and a second DC value of a second reference block in the reference image used for the candidate block; and The offset is applied to at least one sample in the motion-compensated filled MCP block.
9. The method according to any one of claims 1 to 3, characterized in that, Determining the motion vector of the motion-compensated MCP block includes: Determine the most frequently used motion information in a set of inter-frame minimum blocks at the image boundary; and The motion vector is determined based on the most commonly used motion information.
10. The method according to any one of claims 1 to 3, characterized in that, Determining the motion vector of the motion-compensated MCP block includes: Based on the motion information of a set of inter-frame minimum blocks at the image boundary, the average motion vector is determined; and The motion vector is determined based on the average motion vector.
11. The method according to any one of claims 1 to 3, characterized in that, Determining the motion vector of the motion-compensated MCP block includes: The motion vector of the motion-compensated MCP block is determined based on the plurality of candidates within the image boundary and one or more co-position blocks in the reference image used for the motion-compensated MCP block.
12. The method according to claim 11, characterized in that, Determining the motion vector of the motion-compensated filled MCP block further includes: The plurality of candidates are examined before examining one or more co-position blocks in the reference image used for the motion-compensated filling MCP block to determine first available motion information; and The motion vector of the motion compensation filler block is determined based on the first available motion information.
13. The method according to any one of claims 1 to 3, characterized in that, The step of reconstructing at least one sample in the motion-compensated filled MCP block based on motion vectors includes: At least one of a deblocking filter and a smoothing filter is applied to at least one sample in the motion-compensated filled MCP block.
14. A video decoding apparatus, characterized in that, It includes a processing circuit for performing the method according to any one of claims 1 to 13.
15. A video decoding apparatus, characterized in that, The device includes: The receiving module is used to receive a bitstream carrying multiple images; The first determining module is used to determine at least one motion compensation filled MCP block located in the motion compensation filled MCP region outside the image and close to the image boundary. The second determining module is configured to determine the motion vector of the motion compensation filling MCP block for motion compensation boundary filling based on multiple candidates located at the image boundary within the image. The multiple candidates include at least: a first candidate, a second candidate, and a third candidate. The first candidate is located at a first position at the image boundary within the image; the second candidate is located at a second position at the image boundary within the image; and the third candidate is located at a third position at the image boundary within the image. The first candidate is adjacent to a first corner of the MCP block; the second candidate is adjacent to a second corner of the MCP block; and the third candidate is adjacent to (i) the second corner of the MCP block and (ii) the second candidate. A reconstruction module is used to reconstruct at least one sample in the motion-compensated filled MCP block based on the motion vector used for the motion-compensated boundary filling.
16. A video encoding method, characterized in that, include: Identify at least one motion compensation filled MCP block, wherein the at least one MCP block is located outside the motion compensation filled MCP region of one of the multiple images and close to the image boundary of the image; Based on multiple candidates located at multiple positions on the image boundary within the image, a motion vector for motion-assisted boundary filling of the MCP block is determined. The multiple candidates include at least: a first candidate, a second candidate, and a third candidate. The first candidate is located at a first position on the image boundary within the image, the second candidate is located at a second position on the image boundary within the image, and the third candidate is located at a third position on the image boundary within the image. The first candidate is adjacent to a first corner of the MCP block, the second candidate is adjacent to a second corner of the MCP block, and the third candidate is adjacent to (i) the second corner of the MCP block and (ii) the second candidate. Reconstruct at least one sample in the MCP block based on the motion vector used for motion supplement boundary filling; and The plurality of images are encoded into a bitstream, at least in part based on the MCP block.
17. A video encoding apparatus, characterized in that, It includes a processing circuit for performing the method of claim 16.
18. A video encoding apparatus, characterized in that, include: The first determining module is used to determine at least one motion compensation filled MCP block, wherein the at least one MCP block is located in the motion compensation filled MCP region outside one of the multiple images and is close to the image boundary of the image; The second determining module is used to determine the motion vector of the MCP block for motion-assisted boundary filling based on multiple candidates located at multiple positions on the image boundary within the image. The multiple candidates include at least: a first candidate, a second candidate, and a third candidate. The first candidate is located at a first position on the image boundary within the image, the second candidate is located at a second position on the image boundary within the image, and the third candidate is located at a third position on the image boundary within the image. The first candidate is adjacent to a first corner of the MCP block, the second candidate is adjacent to a second corner of the MCP block, and the third candidate is adjacent to (i) the second corner of the MCP block and (ii) the second candidate. The reconstruction module is used to reconstruct at least one sample in the MCP block based on the motion vectors used for motion supplementation boundary filling, and An encoding module is used to encode the plurality of images into a bitstream, at least in part based on the MCP block.
19. A computer-readable storage medium, characterized in that, The device stores instructions that, when executed by at least one processor, cause the at least one processor to perform the method of any one of claims 1 to 13 or to perform the method of claim 16.
Citation Information
Patent Citations
Methods and apparatuses of video processing with overlapped block motion compensation in video coding systems
CN111989927A
Motion compensated boundary pixel padding
US20190082193A1