HMVC for Affine and SBTMVP Motion Vector Prediction Modes

By storing and utilizing the sub-block motion information data set in the video encoding process, the problem that inter-block-based sub-block inter-coding blocks in the prior art cannot benefit, and the video compression efficiency and the utilization efficiency of motion information are improved.

CN114097235BActive Publication Date: 2025-07-11INTERDIGITAL CE PATENT HOLDINGS SAS
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202080046519.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-06-25
Filing Date
2020-06-18
Publication Date
2025-07-11
Estimated Expiration
2040-06-18

AI Technical Summary

Technical Problem

In the prior art, inter-coded blocks that are not based on sub-blocks alone cannot benefit from the motion vectors of previously encoded/decoded blocks, resulting in a reduced video compression efficiency.

Method used

During the video encoding process, a set of sub-block motion information data associated with the current block is determined and stored and added to the motion information data list for motion information prediction of subsequent blocks.

Benefits of technology

The video compression efficiency is improved, especially the encoding and decoding performance of inter-frame decoding blocks based on sub-blocks, and the utilization efficiency of motion information is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114097235B_ABST
    Figure CN114097235B_ABST
Patent Text Reader

Abstract

An apparatus for encoding or decoding a block of a current picture encodes a sub-block of a first block of the current picture. The sub-block is encoded or decoded based on a motion vector determined according to motion information data associated with the first block. In a second step, a second set of motion information data is determined as a function of the first motion information data. These second motion information data are added to a list of motion information data which is used to determine motion information data of other blocks of the current picture to be encoded or decoded.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] At least one embodiment of the present invention generally relates to a method or apparatus for video encoding or decoding, and more particularly, to a method or apparatus for decoding and encoding blocks of a video picture based on motion vectors determined from previously encoded or decoded blocks. Background Art

[0002] To achieve high compression efficiency, image and video coding schemes typically employ prediction (including motion vector prediction) and transformation to exploit spatial and temporal redundancy in video content. Generally, intra-frame or inter-frame prediction is used to exploit intra-frame or inter-frame correlation, and then the difference between the original image and the predicted image, which is typically represented as a prediction error or prediction residue, is transformed, quantized, and entropy-coded. To reconstruct the video, the compressed data is decoded through inverse processes corresponding to entropy coding, quantization, transformation, and prediction.

[0003] Recent additions to high compression techniques include the use of motion models based on affine modeling and / or sub-block based temporal motion vector predictors (SbTMVP). In particular, these models are used for motion compensation in the encoding and decoding of video pictures. Generally, affine modeling is a model using at least two parameters, such as two control point motion vectors (CPMV) representing the motion at the respective corners of a picture block, which allows deriving a motion field for the entire block of the picture to simulate, for example, rotation and similarity (scaling). The motion field is typically discretized in a set of motion vectors associated with sub-blocks of the block.

[0004] Recent developments in the field also include the use of history-based motion vector prediction (HMVP) methods, where HMVP candidates are defined as the motion information of previously decoded blocks. A table with multiple HMVP candidates is maintained during the encoding / decoding process. Whenever there is an inter-frame decoded block that is not sub-block based, the associated motion information is added as a new HMVP candidate to the last entry of the table. These candidates can be used for encoding / decoding additional blocks, particularly adjacent blocks.

[0005] However, since only non-sub-block based inter-frame decoded blocks contribute to the HMVP list, when the block to be encoded or decoded is surrounded by sub-block based inter-frame decoded blocks, this block cannot benefit from the motion vectors of previously encoded / decoded blocks. There is a lack of a solution to this problem. Summary of the Invention

[0006] Disadvantages and drawbacks of the prior art are addressed by the general aspects described herein, which relate to storing motion information associated with sub-block based inter-frame decoded blocks for encoding and decoding additional blocks. According to a first aspect, a method is provided. The method includes the step of decoding sub-blocks of a first block of a current picture. The sub-blocks are decoded based on motion vectors determined according to first motion information data associated with the first block. The method further includes: determining a second set of motion information data as a function of the first motion information data; and adding the second set of motion information data to a list of motion information data. Items of the list can be used to determine motion information data for a second block of the current picture.

[0007] According to another aspect, a second method is provided. The method includes the step of encoding sub-blocks of a first block of a current picture. The sub-blocks are encoded based on motion vectors determined according to first motion information data associated with the first block. The method further includes: determining a second set of motion information data as a function of the first motion information data; and adding the second set of motion information data to a list of motion information data. Items of the list can be used to determine motion information data for a second block of the current picture.

[0008] According to another aspect, an apparatus is provided. The apparatus includes a processor. The processor can be configured to encode a block of a current picture of a video or decode a bitstream by executing any one of the above methods.

[0009] According to another general aspect of at least one embodiment, a device is provided that includes means according to any of the decoding embodiments and at least one of the following: (i) an antenna configured to receive a signal that includes a decoded block of a picture; (ii) a band limiter configured to limit the received signal to a band that includes the decoded block of the picture; and (iii) a display configured to display an output representing the decoded block of the picture.

[0010] According to another general aspect of at least one embodiment, a non-transitory computer-readable medium is provided that contains data content generated according to any of the described encoding embodiments or variations.

[0011] According to another general aspect of at least one embodiment, a signal is provided that includes video data generated according to any of the described encoding embodiments or variations.

[0012] According to another general aspect of at least one embodiment, a bitstream is formatted to include data content generated according to any of the described encoding embodiments or variations.

[0013] According to another general aspect of at least one embodiment, there is provided a computer program product including instructions that, when executed by a computer, cause the computer to perform any of the described decoding embodiments or variations.

[0014] These and other aspects, features, and advantages of the general aspect will become apparent from the following detailed description of exemplary embodiments read in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 An encoder is shown;

[0016] Figure 2 A block diagram of a video decoder is shown;

[0017] Figure 3 A block diagram of an example of a system in which various aspects and embodiments are implemented is shown;

[0018] Figure 4 A coding tree and coding tree unit (CTU) structure for representing a compressed picture according to, for example, the HEVC video compression standard is shown;

[0019] Figure 5A and Figure 5B An affine motion vector field based on 4×4 sub-blocks for two and three control points, respectively, is shown;

[0020] Figure 6 An example of updating a list in the HMVP method is shown;

[0021] Figure 7 A method 70 for encoding / decoding a block of a current picture is illustrated diagrammatically; DETAILED DESCRIPTION

[0022] The general aspects described herein are in the field of video compression. The purpose of these aspects is to improve compression efficiency compared to existing video compression systems.

[0023] This application describes a number of aspects, including tools, features, embodiments, models, methods, etc. Many of these aspects are described as being specific and are typically described in a way that may sound restrictive, at least for the purpose of showing individual characteristics. However, this is for the purpose of clarity of description and does not limit the application or scope of those aspects. In fact, all the different aspects can be combined and interchanged to provide additional aspects. In addition, these aspects can also be combined and interchanged with aspects described in earlier filed documents.

[0024] The aspects described and contemplated in this application can be implemented in many different forms. The following Figure 1 , Figure 2 and Figure 3Some embodiments are provided, but other embodiments are conceivable, and the discussion of Figure 1 、 Figure 2 and Figure 3 does not limit the breadth of implementation. At least one aspect generally relates to video encoding and decoding, and at least one other aspect generally relates to transmitting the generated or encoded bitstream. These and other aspects can be implemented as methods, apparatuses, computer-readable storage media storing instructions for encoding or decoding video data according to any of the described methods, and / or computer-readable storage media storing bitstreams generated according to any of the described methods.

[0025] In this application, the terms "reconstruction" and "decoding" can be used interchangeably, the terms "pixel" and "sample" can be used interchangeably, and the terms "image", "picture", and "frame" can be used interchangeably. Generally, but not necessarily, the term "reconstruction" is used on the encoder side, while "decoding" is used on the decoder side.

[0026] Various methods are described herein, and each method includes one or more steps or actions for implementing the described method. Unless the correct operation of the method requires steps or actions in a specific order, the order and / or use of specific steps and / or actions can be modified or combined.

[0027] The various methods and other aspects described in this application can be used to modify modules, such as Figure 1 and Figure 2 the motion compensation modules 170 and 275 of the video encoder 100 and decoder 200 shown. Additionally, aspects of the present invention are not limited to VVC or HEVC, and can be applied to (for example) other standards and recommendations (whether pre-existing or future-developed), as well as extensions of any such standards and recommendations (including VVC and HEVC). Unless otherwise indicated or technically excluded, the aspects described in this application can be used alone or in combination.

[0028] Figure 1 The encoder 100 is shown. Variants of the encoder 100 are conceivable, but for clarity, the encoder 100 is described below without describing all the expected variants.

[0029] Before being encoded, the video sequence can undergo pre-encoding processing 101, for example, applying a color transformation to the input color picture (e.g., conversion from RGB 4:4:4 to YCbCr 4:2:0), or performing remapping of the input picture components to obtain a signal distribution more resilient to compression (e.g., using histogram equalization of one of the color components). Metadata can be associated with the preprocessing and appended to the bitstream.

[0030] In encoder 100, as described below, pictures are encoded by encoder elements. The picture to be encoded is segmented (102) and processed in units such as CUs. Each unit is encoded using, for example, an intra or inter mode. When the unit is encoded in the intra mode, intra prediction (160) is performed. In the inter mode, motion estimation (175) and motion compensation (170) are performed. The encoder determines (105) to encode the unit using one of the intra mode or the inter mode, and indicates the intra / inter decision by, for example, a prediction mode flag. The prediction residual is calculated, for example, by subtracting (110) the prediction block from the original image block.

[0031] Then, the prediction residual is transformed (125) and quantized (130). The quantized transform coefficients, motion vectors, and other syntax elements are entropy decoded (145) to output a bitstream. The encoder may also skip the transformation and directly apply quantization to the untransformed residual signal. The encoder may bypass the transformation and quantization, i.e., directly decode the residual without applying the transformation or quantization process.

[0032] The encoder decodes the encoded block to provide a reference for further prediction. The quantized transform coefficients are inverse quantized (140) and inverse transformed (150) to decode the prediction residual. The decoded prediction residual and the prediction block are combined (155) to reconstruct the image block. An in-loop filter (165) is applied to the reconstructed picture to perform, for example, deblocking / SAO (sample adaptive offset) filtering to reduce encoding artifacts. The filtered image is stored in the reference picture buffer (180).

[0033] Figure 2 A block diagram of video decoder 200 is shown. In decoder 200, the bitstream is decoded by decoder elements as described below. Video decoder 200 generally performs a decoding pass that is reciprocal to the encoding pass described in Figure 1 The encoder 100 also generally performs video decoding as part of encoding video data.

[0034] Specifically, the input to the decoder includes a bitstream that may be generated by video encoder 100. The bitstream is first entropy decoded (230) to obtain transform coefficients, motion vectors, and other decoded information. The picture segmentation information indicates how the picture is segmented. Thus, the decoder can divide (235) the picture according to the decoded picture segmentation information. The transform coefficients are inverse quantized (240) and inverse transformed (250) to decode the prediction residual. The decoded prediction residual is combined (255) with the prediction block to reconstruct the image block. The prediction block can be obtained (270) from intra prediction (260) or motion compensation prediction (i.e., inter prediction) (275). An in-loop filter (265) is applied to the reconstructed image. The filtered image is stored in the reference picture buffer (280).

[0035] The decoded picture may further undergo post - decoding processing (285), for example, an inverse color transformation (e.g., a conversion from YCbCr 4:2:0 to RGB 4:4:4) or performing an inverse remapping of the remapping process performed in the pre - encoding processing (101). The post - decoding processing may use metadata derived in the pre - encoding processing and signaled in the bitstream.

[0036] Figure 3 A block diagram showing an example of a system in which various aspects and embodiments are implemented is presented. System 1000 may be implemented as a device including various components described below and configured to perform one or more aspects described in this document. Examples of such devices include, but are not limited to, various electronic devices such as personal computers, laptop computers, smart phones, tablet computers, digital multimedia set - top boxes, digital television receivers, personal video recording systems, connected household appliances, and servers. The elements of system 1000 may be implemented singly or in combination in a single integrated circuit (IC), multiple ICs, and / or discrete components. For example, in at least one embodiment, the processing and encoder / decoder elements of system 1000 are distributed across multiple ICs and / or discrete components. In various embodiments, system 1000 is communicatively coupled to one or more other systems or other electronic devices via, for example, a communication bus or through dedicated input and / or output ports. In various embodiments, system 1000 is configured to implement one or more aspects described in this document.

[0037] System 1000 includes at least one processor 1010, which is configured to execute instructions loaded therein for implementing, for example, various aspects described in this document. Processor 1010 may include embedded memory, input - output interfaces, and various other circuits known in the art. System 1000 includes at least one memory 1020 (e.g., a volatile memory device and / or a non - volatile memory device). System 1000 includes a storage device 1040, which may include non - volatile memory and / or volatile memory, including but not limited to electrically erasable programmable read - only memory (EEPROM), read - only memory (ROM), programmable read - only memory (PROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, disk drives, and / or optical disk drives. As a non - limiting example, storage device 1040 may include internal storage devices, attached storage devices (including removable and non - removable storage devices), and / or network - accessible storage devices.

[0038] System 1000 includes an encoder / decoder module 1030, which is configured to process data, for example, to provide encoded video or decoded video, and the encoder / decoder module 1030 may include its own processor and memory. The encoder / decoder module 1030 represents one or more modules that may be included in a device to perform encoding and / or decoding functions. As is known, a device may include one or both of an encoding and a decoding module. Additionally, the encoder / decoder module 1030 may be implemented as a separate element of system 1000 or may be incorporated within processor 1010 as a combination of hardware and software known to those skilled in the art.

[0039] The program code to be loaded onto processor 1010 or encoder / decoder 1030 to perform the various aspects described in this document may be stored in storage device 1040 and subsequently loaded onto memory 1020 for execution by processor 1010. According to various embodiments, one or more of processor 1010, memory 1020, storage device 1040, and encoder / decoder module 1030 may store one or more of the various items during the execution of the processes described in this document. These stored items may include, but are not limited to, input video, decoded video or portions of decoded video, bitstreams, matrices, variables, and intermediate or final results from the processing of equations, formulas, operations, and operation logic.

[0040] In some embodiments, the memory within processor 1010 and / or encoder / decoder module 1030 is used to store instructions and provide a working memory for the processing required during encoding or decoding. However, in other embodiments, memory external to the processing device (e.g., the processing device may be processor 1010 or encoder / decoder module 1030) is used for one or more of these functions. The external memory may be memory 1020 and / or storage device 1040, for example, dynamic volatile memory and / or non-volatile flash memory. In several embodiments, external non-volatile flash memory is used to store, for example, the operating system of a television. In at least one embodiment, fast external dynamic volatile memory such as RAM is used as a working memory for video encoding and decoding operations, such as MPEG-2 (MPEG refers to the Moving Picture Experts Group, MPEG-2 is also known as ISO / IEC 13818, and 13818-1 is also known as H.222, and 13818-2 is also known as H.262), HEVC (HEVC refers to High Efficiency Video Coding, also known as H.265 and MPEG-H Part 2), or VVC (Versatile Video Coding, a new standard developed by JVET, the Joint Video Experts Team).

[0041] As shown in block 1130, input to the components of system 1000 can be provided through various input devices. Such input devices include, but are not limited to: (i) an RF section that receives RF signals transmitted, for example, over the air by a broadcaster; (ii) component (COMP) input terminals (or a set of COMP input terminals); (iii) universal serial bus (USB) input terminals; and / or (iv) high-definition multimedia interface (HDMI) input terminals. Figure 3 Other examples not shown include composite video.

[0042] In various embodiments, the input devices of block 1130 have corresponding input processing elements associated therewith that are known in the art. For example, the RF section can be associated with elements adapted to: (i) select a desired frequency (also referred to as selecting a signal, or band-limiting a signal to a band); (ii) down-convert the selected signal; (iii) again band-limit to a narrower band to select a signal band that can be referred to as a channel, for example, in some embodiments; (iv) demodulate the down-converted and band-limited signal; (v) perform error correction; and (vi) de-multiplex to select a desired data packet stream. The RF section of various embodiments includes one or more elements that perform these functions, such as, for example, a frequency selector, a signal selector, a band limiter, a channel selector, a filter, a down-converter, a demodulator, an error corrector, and a de-multiplexer. The RF section can include a tuner that performs various of these functions, including, for example, down-converting the received signal to a lower frequency (e.g., an intermediate frequency or near-baseband frequency) or baseband. In one set-top box embodiment, the RF section and its associated input processing elements receive RF signals transmitted over a wired (e.g., cable) medium and perform frequency selection by filtering, down-converting, and again filtering to a desired band. Various embodiments reorder the above (and other) elements, remove some of these elements, and / or add other elements that perform similar or different functions. For example, adding elements can include inserting elements between existing elements, such as, for example, inserting an amplifier and an analog-to-digital converter. In various embodiments, the RF section includes an antenna.

[0043] Additionally, the USB and / or HDMI terminal may include a respective interface processor for connecting the system 1000 to other electronic devices via a USB and / or HDMI connection. It should be understood that various aspects of the input processing, such as Reed-Solomon error correction, may be implemented, as needed, within, for example, a separate input processing IC or processor 1010. Similarly, aspects of the USB or HDMI interface processing may be implemented, as needed, within a separate interface IC or within the processor 1010. The demodulated, error-corrected, and demultiplexed streams are provided to various processing elements, including, for example, the processor 1010 and the encoder / decoder 1030, which operate in conjunction with memory and storage elements to process the data stream as needed for presentation on an output device.

[0044] The various elements of the system 1000 may be provided within an integrated housing. Within the integrated housing, the various elements may be interconnected and data may be transmitted therebetween using suitable connection arrangements such as internal buses (including Inter-Integrated Circuit (I2C) buses), wiring, and printed circuit boards known in the art.

[0045] The system 1000 includes a communication interface 1050 that enables communication with other devices via a communication channel 1060. The communication interface 1050 may include, but is not limited to, a transceiver configured to transmit and receive data via the communication channel 1060. The communication interface 1050 may include, but is not limited to, a modem or a network card, and the communication channel 1060 may be implemented, for example, within a wired and / or wireless medium.

[0046] In various embodiments, a wireless network, such as a Wi-Fi network, for example, IEEE 802.11 (IEEE refers to the Institute of Electrical and Electronics Engineers), is used to stream or otherwise provide data to the system 1000. The Wi-Fi signals of these embodiments are received via the communication channel 1060 and the communication interface 1050 suitable for Wi-Fi communication. The communication channel 1060 of these embodiments is typically connected to an access point or router that provides access to an external network, including the Internet, to allow streaming applications and other over-the-top communications. Other embodiments use a set-top box to provide streaming data to the system 1000, which passes the data via an HDMI connection of the input block 1130. Still other embodiments use an RF connection of the input block 1130 to provide streaming data to the system 1000. As described above, various embodiments provide data in a non-streaming manner. Additionally, various embodiments use wireless networks other than Wi-Fi, such as cellular networks or Bluetooth networks.

[0047] System 1000 can provide output signals to various output devices, which include a display 1100, a speaker 1110, and other peripheral devices 1120. The display 1100 of various embodiments includes, for example, one or more of a touch screen display, an organic light emitting diode (OLED) display, a curved display, and / or a foldable display. The display 1100 can be used for a television, a tablet computer, a laptop computer, a cellular phone (mobile phone), or other devices. The display 1100 can also be integrated with other components (e.g., as in a smart phone), or applied separately (e.g., an external monitor for a laptop computer). In various examples of the embodiments, the other peripheral devices 1120 include one or more of a stand-alone digital video disc (or digital versatile disc) (DVR, two items), a disc player, a stereo system, and / or a lighting system. Various embodiments use one or more peripheral devices 1120, which provide functions based on the output of the system 1000. For example, the disc player performs the function of playing the output of the system 1000.

[0048] In various embodiments, signaling such as an AV link, consumer electronics control (CEC), or other communication protocols is used to transmit control signals between the system 1000 and the display 1100, the speaker 1110, or the other peripheral devices 1120, which enables device-to-device control with or without user intervention. The output devices can be communicatively coupled to the system 1000 via dedicated connections through corresponding interfaces 1070, 1080, and 1090. Alternatively, the output devices can be connected to the system 1000 using a communication channel 1060 via a communication interface 1050. For example, the display 1100 and the speaker 1110 can be integrated in a single unit in an electronic device (e.g., a television set) together with other components of the system 1000. In various embodiments, for example, the display interface 1070 includes a display driver, such as a timing controller (T Con) chip.

[0049] For example, if the RF portion of the input 1130 is part of a separate set-top box, the display 1100 and the speaker 1110 can alternatively be separated from one or more of the other components. In various embodiments where the display 1100 and the speaker 1110 are external components, the output signals can be provided via a dedicated output connection, which includes, for example, an HDMI port, a USB port, or a COMP output.

[0050] These embodiments can be executed by computer software implemented by the processor 1010, hardware, or a combination of hardware and software. As a non-limiting example, the embodiments can be implemented by one or more integrated circuits. The memory 1020 can be of any type suitable for the technical environment and can be implemented using any appropriate data storage technology, as non-limiting examples, such as optical memory devices, magnetic memory devices, semiconductor-based memory devices, fixed memory, and removable memory. The processor 1010 can be of any type suitable for the technical environment and can include one or more of, as non-limiting examples, a microprocessor, a general-purpose computer, a special-purpose computer, and a processor based on a multi-core architecture.

[0051] Figure 4 Illustrated is, for example, a decoding tree and a coding tree unit (CTU) structure for representing a compressed picture according to the HEVC video compression standard. In video compression standards, such as in HEVC, motion-compensated temporal prediction is employed to exploit the redundancy present between consecutive pictures of a video. A motion vector is associated with each prediction unit (PU) included in a CTU. The set of CTUs of a picture is represented as a decoding tree in the compressed domain. This is a quadtree partitioning of the CTUs, where each leaf is called a coding unit (CU). Each CU is given some intra or inter prediction parameters (prediction information). For this purpose, a CU is spatially split into one or more prediction units (PUs), and each PU is assigned some prediction information. An intra or inter decoding mode is assigned at the CU level. A motion vector is assigned to each PU in HEVC. This motion vector is used for motion-compensated temporal prediction of the considered PU. Thus, in video compression standards similar to HEVC, the motion model that links the prediction block and its reference block exists in translation.

[0052] Figure 5A and Figure 5B Illustrated are 4×4 sub-block-based affine motion vector fields for two and three control points, respectively. A recent addition to video compression standards, such as in the Joint Exploration Model (JEM) developed by the JVET (Joint Video Exploration Team) group and the subsequent Versatile Video Coding (VVC) test model, supports some richer motion models to improve temporal prediction. For this purpose, a PU can be spatially partitioned into sub-PUs, and a richer model can be used to assign a dedicated motion vector to each sub-PU. A CU is no longer partitioned into PUs or TUs, and some motion data is directly assigned to each CU. In this new codec design, a CU can be partitioned into sub-CUs, and a motion vector can be calculated for each sub-CU. One of the new motion models introduced in JEM is the affine model, which consists of using an affine model to represent the motion vectors in a CU.

[0053] Figure 5AThe affine motion field of two control points, also known as the four-parameter affine model, includes the motion vector component values for each position (x, y) within the considered block according to Equation [Equation 1]:

[0054] [Equation 1]

[0055] where (v 0x , v 0y ) and (v 1x , v 1y ) are the so-called control point motion vectors 51A and 52A used to generate the affine motion field. (v 0x , v 0y ) is the motion vector of the upper left corner control point. (v 1x , v 1y ) is the motion vector of the upper right corner control point.

[0056] As Figure 5B shown in, a model with three control points (referred to as a six-parameter affine motion model) can also be used to represent the sub-block based motion field of a given decoding unit. In the case of the six-parameter affine model, the motion field is calculated as in Equation [Equation 2], where (v 0x , v 0y ) is vector 51B, (v 1x , v 1y ) is vector 52B, and (v 2x , v 2y ) is vector 53B.

[0057] [Equation 2]

[0058] In practice, to keep the complexity reasonable, the same motion vector is calculated for the pixels of the 4×4 sub-blocks (sub-CUs) of the considered CU, as Figure 5A and Figure 5B shown in. At the center position of each sub-block, the affine motion vector is calculated based on the control point motion vectors. As a result, the obtained motion vectors are represented with 1 / 16 pixel accuracy. Consequently, the prediction unit (PU) of the decoding unit in the affine mode is constructed by performing motion compensation prediction on each sub-block using its own motion vector.

[0059] In VTM, a CU larger than 8×8 can be predicted in the affine AMVP mode. This is signaled by a flag in the bitstream decoded at the CU level. The generation of the affine motion field for an inter CU involves determining the control point motion vectors (CPMVs), which are obtained by the decoder as the sum of the motion vector difference plus the control point motion vector prediction (CPMVP). The CPMVP is a pair (for a 4-parameter affine model specified by 2 control point motion vectors) or a triplet (for a 6-parameter affine model specified by 3 control point motion vectors) of motion vectors that serves as a predictor for the CPMV of the current CU. The CPMVP of the current CU can be inherited from an affine neighboring CU (as in the affine merge mode). The spatial location from which the inherited CPMVP is retrieved is selected from an ordered list of given candidate positions. If the reference picture of the inherited CPMVP is equal to the reference picture of the current CU, the inherited CPMVP is considered valid.

[0060] In the affine merge mode, the CU-level flag indicates whether the CU in the merge mode employs affine motion compensation. If so, then in JEM, the first available neighboring CU decoded in the affine mode is selected from an ordered list of given candidate positions. Once the first neighboring CU in the affine mode is obtained, the CPMVP of the current CU can be inherited from this affine neighboring CU. Three motion vectors from the upper-left, upper-right, and lower-left corners of the neighboring affine CU are retrieved Based on these three motion vectors, the current CU has two or three CPMVs for its upper-left, upper-right, and / or lower-left corners derived as follows.

[0061] For a CU with a 4-parameter affine mode, two CPMVs of the current CU are derived as follows:

[0062] ο

[0063] ο

[0064] ο

[0065] ο

[0066] For a CU with a 6-parameter affine mode, three control point motion vectors of the current CU are derived as follows:

[0067] ο

[0068] ο

[0069] ο

[0070] ο

[0071] ο

[0072] When obtaining the motion vector of the control point of the current CU and / or a motion field within the current CU is calculated on a 4×4 sub-block basis by the model of Equation [Equation 1] or [Equation 2]. Thus, the affine model can be regarded as a sub-block-based model.

[0073] Due to VTM-3.0, the SbTMVP (sub-block based temporal motion vector predictor, also known as ATMVP) candidates are part of the affine merge list as the first sub-block candidates to be evaluated in the rate-distortion optimization process of the affine merge mode. The SbTMVP predicts the motion vector of an 8×8 sub-block within the current CU by collecting the motion information of the corresponding sub-blocks in the collocated pictures as a regular TMVP predictor.

[0074] Figure 6 An example of updating the list in the HMVP method is shown. Since VTM-3.0, the history-based motion vector prediction (HMVP) has also been introduced. The history-based is to maintain a list consisting of multiple motion information (motion vectors, associated reference frames, BCW indices,...) that have been used for decoding the CUs before the current CU. Whenever a non-affine and non-triangular inter-frame CU (which is a non-sub-block-based block) is decoded, the associated motion information is inserted at the end of the list as a new HMVP candidate. As Figure 6 shown, when inserting a new motion candidate into the table, the constrained FIFO rule is utilized, where a redundancy check is first applied to find whether there is the same HMVP in the table. If found, the same HMVP is removed from the table, and then all HMVP candidates are moved forward, i.e., the index is decreased by 1.

[0075] The HMVP candidates can be used in the construction process of the merge candidate list. The latest several HMVP candidates in the table are sequentially checked and inserted into the candidate list after the TMVP candidates. Pruning is applied to the HMVP candidates to exclude spatial or temporal merge candidates from the sub-block motion candidates (i.e., ATMVP).

[0076] According to this principle, motion information related to an inter-frame CU decoded in a sub-block mode is added to a motion information data list, such as an HMVP list, where the sub-block mode is a block that decodes sub-blocks based on motion vectors determined according to motion information data associated with the block. Items in the list are used to determine motion information data of other blocks in the same current picture. Since the affine and SbTMVP modes use sub-block-based motion compensation, several motion vectors are stored for each CU decoded in the sub-block mode. The position within the CU, i.e., the sub-block, from which the motion vector to be saved in the list is selected may affect the BD-rate performance.

[0077] Currently, motion information from CUs decoded in a sub-block or triangle mode is not considered in HMVP list construction and update. According to this principle, zero, one, or several motion information data determined according to an inter-frame CU decoded in the sub-block mode are inserted into the HMVP list as regular inter-frame motion information.

[0078] The principles of the present invention apply to each sub-block decoding mode. The sub-block decoding mode is not limited to the 3 modes of VTM-5.0: (i) affine AMVP; (ii) affine merge; and (iii) SbTMVP. However, it can be extended to all known sub-block modes, such as planar motion vectors, regression, triangle, FRUC, interleaved affine, etc., and can be extended to all new sub-block modes that may be proposed in the future. It is possible to consider only one of these sub-block modes or a combination of modes (e.g., two affine modes) in HMVP list saving, or a combination of all sub-block modes.

[0079] Figure 7 A method 70 for encoding / decoding a block of a current picture is illustrated by way of example. In step 71, a first block to be encoded / decoded according to the sub-block mode is encoded / decoded. Motion vectors for encoding / decoding the sub-blocks of the first block are determined according to motion information contained in the CU of the first block itself, such as Figure 5A and Figure 5B as shown. The motion information data includes different data for motion prediction, such as motion vectors, associated reference frames, BCW indices, etc.

[0080] According to the principles of the present invention, in step 72, a second set of motion information data is determined as a function of the first motion information data; the number of motion information data determined in this set can be zero, one, or several. The choice of the size of this set is guided by the efficiency of this choice for the encoding / decoding process. It is parameterized and shared by the encoder and decoder. The relevant parameters can be encoded in the header of the video stream.

[0081] In an embodiment, second motion information is constructed using data of first motion information, i.e., motion information associated with a first block. The motion vector of the second motion information is determined as a function of a subset of the motion vectors determined for sub-blocks of the first block. Sub-blocks providing their motion vectors for combination (i.e., the function of the motion vectors) are selected according to their positions within the first block. For example, a single sub-block is selected, e.g., a sub-block centered on the first block (e.g., sub-block [1,1] or [1,2] in the case of a 4×4 sub-block partitioning), and the function is an identity function. Thus, the motion vector of the determined second motion information data is the motion vector of the selected sub-block. In another example, the motion vector of the second motion information data is the average motion vector of two or three or four motion vectors of sub-blocks selected for the position distribution of the sub-blocks within the first block, such as two central sub-blocks or three sub-blocks: two in the upper corners and one in the bottom center, or four sub-blocks in the corners of the first block. The function can be different from the average value. The selection of the function and the selection of the positions of the selected sub-blocks are guided by the efficiency of such selection in the encoding / decoding process. It is parameterized and shared by the encoder and the decoder. The relevant parameters can be encoded in the header of the video stream.

[0082] In another embodiment, the second motion information data includes a motion vector determined as a function of the motion vectors of the first motion information data. For example, for a first block encoded / decoded according to a 4-parameter affine model, the motion vector of the second motion information data can be the weighted average of two control point vectors. According to this embodiment, zero, one, or several motion information data can be determined using different functions and / or different parameters.

[0083] According to a variant applicable to the above two embodiments, the selected function of the motion vector and / or the position of the selected sub-block depends on the size of the first block and / or the size of the sub-block. For example, if the first block is smaller than a given size (e.g., 128×128, 64×64, 32×32, or 16×16), the function is an identity, and the selected sub-block is selected at the center of the first block; but if the first block is larger than the given size, the function is the weighted average of two vectors of sub-blocks in the corners of the first block (the weights depending on the size of the first block).

[0084] According to a variant, the number of determined second motion information data depends on the size of the first block and / or the size of the sub-blocks. For example, if the size of the first block is smaller than a given size, no second motion data is determined. In another example, the sub-blocks of the first block can be divided into regions of a fixed size (e.g., 64×64, 32×32, or 16×16) or regions of an adaptive size (e.g., dividing the sub-block CU into regions of 2, 4, or 8). For these regions, one second motion information data is determined according to one of the above embodiments and variants. For example, if the region has a fixed size of 32×32, then for a 16×32 sub-block CU, only one second motion information data is determined, and for a 64×64 sub-block CU, four second motion information data are determined. In another example, if the region has an adaptive size and the sub-block CU is divided into four regions, then for a 16×32 sub-block CU, each is for 8×16, and for a 64×64 sub-block CU, each is for 32×32, and there are always four motion information data saved.

[0085] The number of determined second motion information data and the choice of different functions used for their determination are guided by the efficiency of this choice in the encoding / decoding process. It is parameterized and shared by the encoder and the decoder. The relevant parameters can be encoded in the header of the video stream. For example, an effective choice can be to determine one motion information data using the motion vector of the central sub-block only for sub-blocks with an area greater than or equal to 256 square pixels.

[0086] In Figure 7 step 73 of method 70, the determined set of motion information data is added to the motion information data list, such as the HMVP list of the VCC test model. This list is defined as: its items are used to determine the motion information data of other blocks of the current picture to be encoded / decoded.

[0087] Various implementations relate to decoding. As used in this application, "decoding" can include, for example, all or part of the process performed on the received encoded sequence in order to produce a final output suitable for display. In various embodiments, such processes include one or more of the processes typically performed by a decoder, such as entropy decoding, inverse quantization, inverse transform, and differential decoding. In various embodiments, such processes also or alternatively include the processes performed by the decoders of the various implementations described in this application, for example, managing the motion information data list to feed Figure 1 and Figure 2 the motion prediction modules 170 and 275.

[0088] As a further example, in one embodiment, "decoding" refers only to entropy decoding, in another embodiment, "decoding" refers only to differential decoding, and in another embodiment, "decoding" refers to a combination of entropy decoding and differential decoding. Whether the phrase "decoding process" is intended to specifically refer to a subset of operations or generally refer to a broader decoding process will be clear based on the context of the specific description and is believed to be well understood by those skilled in the art.

[0089] Various implementations involve encoding. In a manner similar to the above discussion regarding "decoding", "encoding" as used in this application can include, for example, all or part of the process performed on an input video sequence in order to produce an encoded bitstream. In various embodiments, such processes include one or more processes typically performed by an encoder, e.g., segmentation, differential encoding, transformation, quantization, and entropy encoding. In various embodiments, such processes also or alternatively include processes performed by the encoders of the various implementations described in this application, e.g., managing a list of motion information data for feeding into Figure 1 and Figure 2 the motion prediction modules 170 and 275.

[0090] As a further example, in one embodiment, "encoding" refers only to entropy encoding, in another embodiment, "encoding" refers only to differential encoding, and in another embodiment, "encoding" refers to a combination of differential encoding and entropy encoding. Whether the phrase "encoding process" is intended to specifically refer to a subset of operations or generally refer to a broader encoding process will be clear based on the context of the specific description and is believed to be well understood by those skilled in the art.

[0091] Note that the grammatical elements used here are descriptive terms. Thus, they do not exclude the use of other grammatical element names.

[0092] When the drawings are presented as flowcharts, it should be understood that they also provide a block diagram of the corresponding apparatus. Similarly, when the drawings are presented as block diagrams, it should be understood that they also provide a flowchart of the corresponding method / process.

[0093] The implementations and aspects described herein can be implemented, for example, in a method or process, an apparatus, a software program, a data stream, or a signal. Even if discussed only in the context of a single form of implementation (e.g., only as a method), the implementation of the features discussed can be implemented in other forms (e.g., an apparatus or a program). For example, an apparatus can be implemented in appropriate hardware, software, and firmware. The method can be implemented, for example, in a processor, which generally refers to a processing device, including, for example, a computer, a microprocessor, an integrated circuit, or a programmable logic device. A processor also includes a communication device, such as a computer, a cellular phone, a portable / personal digital assistant ("PDA"), and other devices that facilitate information communication between end users.

[0094] References to "one embodiment", "an embodiment", "one implementation", "an implementation", and other variations thereof mean that the specific features, structures, characteristics, etc. described in connection with the embodiment are included in at least one embodiment. Thus, the appearances of the phrases "in one embodiment", "in an embodiment", "in one implementation", "in an implementation", and any other variations thereof throughout this application do not necessarily all refer to the same embodiment.

[0095] In addition, this application may relate to "determining" various information. Determining information may include, for example, one or more of estimating information, calculating information, predicting information, or retrieving information from a memory.

[0096] Furthermore, this application may relate to "accessing" various information. Accessing information may include, for example, one or more of receiving information, retrieving information (e.g., from a memory), storing information, moving information, copying information, calculating information, determining information, predicting information, or estimating information.

[0097] In addition, this application may relate to "receiving" various information. Like "accessing", receiving is intended to be a broad term. Receiving information may include, for example, one or more of accessing information or retrieving information (e.g., from a memory). Furthermore, "receiving" is typically involved in one way or another during operations such as storing information, processing information, sending information, moving information, copying information, erasing information, calculating information, determining information, predicting information, or estimating information.

[0098] It should be understood that in cases such as "A / B", "A and / or B", and "at least one of A and B", the use of any one of the following, namely " / ", "and / or", and "at least one of...", is intended to cover the selection of only the first-listed option (A), or only the second-listed option (B), or the selection of both options (A and B). As a further example, in the cases of "A, B, and / or C" and "at least one of A, B, and C", such wording is intended to include the selection of only the first-listed option (A), or only the second-listed option (B), or only the third-listed option (C), or only the first and second-listed options (A and B), or only the first and third-listed options (A and C), or only the second and third-listed options (B and C), or the selection of all three options (A and B and C). This can be extended to any number of items listed, which will be apparent to those of ordinary skill in the art and related fields.

[0099] In addition, as used herein, the word "signal" particularly refers to indicating something to a corresponding decoder. For example, in some embodiments, an encoder signals a particular one of a plurality of parameters for determining second information data from a sub-block mode encoded block. Thus, in an embodiment, the same parameters are used on the encoder side and the decoder side. Therefore, for example, an encoder can convey (explicit signaling) a particular parameter to a decoder such that the decoder can use the same particular parameter. Conversely, if the decoder already has a particular parameter as well as other parameters, signaling can be used without conveyance (implicit signaling) to simply allow the decoder to know and select the particular parameter. By avoiding the transmission of any actual functions, bit saving is achieved in various embodiments. It should be understood that signaling can be implemented in various ways. For example, in various embodiments, one or more syntax elements, flags, etc. are used to signal information to a corresponding decoder. Although the foregoing relates to the verb form of the word "signal", the word "signal" can also be used as a noun herein.

[0100] As will be appreciated by those skilled in the art, implementations can generate various signals that can be formatted to carry information such as can be stored or transmitted. This information can include, for example, instructions for performing a method or data generated by one of the described implementations. For example, a signal can be formatted to carry a bitstream of the described embodiments. Such a signal can be formatted as, for example, an electromagnetic wave (e.g., using the radio frequency portion of the spectrum) or a baseband signal. Formatting can include, for example, encoding a data stream and modulating a carrier with the encoded data stream. The information carried by the signal can be, for example, analog or digital information. As is known, signals can be transmitted over various different wired or wireless links. The signal can be stored on a processor-readable medium.

Claims

1. A method, the method comprising: - Decoding sub - blocks of a first block of a current picture, the sub - blocks being decoded based on at least two first motion vectors determined according to at least two first motion information data associated with the first block; - Determining a second set of motion information data as a function of the at least two first motion information data, wherein the second motion information data includes a second motion vector determined as a function of the at least two first motion vectors, the at least two first motion vectors being included in the at least two first motion information data; And - Adding the second set of motion information data to a history - based motion vector prediction (HMVP) list, items of the HMVP list being used to determine motion information data of a second block of the current picture.

2. The method according to claim 1, wherein the second motion information data includes a motion vector determined as a function of the at least two first motion vectors of a selected sub - block of the first block, the selected sub - block being at a given position within the first block.

3. The method according to claim 1 or 2, wherein the function of the at least two first motion vectors depends on the size of the first block or the size of the sub - block.

4. The method according to claim 1, wherein the second set of motion information data includes a number of items, the number of items depending on the size of the first block or the size of the sub - block.

5. An apparatus for decoding a block of a current picture, the apparatus comprising a processor configured to: - Decoding sub - blocks of a first block of a current picture, the sub - blocks being decoded based on at least two first motion vectors determined according to at least two first motion information data associated with the first block; - Determining a second set of motion information data as a function of the at least two first motion information data, wherein the second motion information data includes a second motion vector determined as a function of the at least two first motion vectors, the at least two first motion vectors being included in the at least two first motion information data; And - Adding the second set of motion information data to a history - based motion vector prediction (HMVP) list, items of the HMVP list being used to determine motion information data of a second block of the current picture.

6. The apparatus according to claim 5, wherein the second motion information data includes a motion vector determined as a function of the at least two first motion vectors of a selected sub - block of the first block, the selected sub - block being at a given position within the first block.

7. The apparatus according to claim 5 or 6, wherein the function of the at least two first motion vectors depends on the size of the first block or the size of the sub - block.

8. The apparatus according to claim 5, wherein the second set of motion information data includes a number of items, the number of items depending on the size of the first block or the size of the sub - block.

9. An equipment, the equipment comprising: The apparatus according to any one of claims 5 to 8; And At least one of the following: (i) an antenna configured to receive a signal including decoded blocks of a picture; (ii) a band limiter configured to limit the received signal to a band including the decoded blocks of the picture; and (iii) a display configured to display an output representing the decoded blocks of the picture.

10. A method comprising: - encoding sub-blocks of a first block of a current picture, the sub-blocks being encoded based on at least two first motion vectors determined according to at least two first motion information data associated with the first block; - determining a second set of motion information data as a function of the at least two first motion information data, wherein the second motion information data includes second motion vectors determined as a function of the at least two first motion vectors, the at least two first motion vectors being included in the at least two first motion information data; and - adding the second set of motion information data to a history-based motion vector prediction (HMVP) list, items of the HMVP list being used to determine motion information data for a second block of the current picture.

11. The method according to claim 10, wherein the second motion information data includes motion vectors determined as a function of the at least two first motion vectors of a selected sub-block of the first block, the selected sub-block being at a given position within the first block.

12. The method according to claim 10 or 11, wherein the function of the at least two first motion vectors depends on the size of the first block or the size of the sub-block.

13. The method according to claim 10, wherein the second set of motion information data includes a number of items, the number of items depending on the size of the first block or the size of the sub-block.

14. An apparatus for encoding a block of a current picture, the apparatus including a processor configured to: - encode sub-blocks of a first block of a current picture, the sub-blocks being encoded based on at least two first motion vectors determined according to at least two motion information data associated with the first block; - determine a second set of motion information data as a function of the at least two first motion information data, wherein the second motion information data includes second motion vectors determined as a function of the at least two first motion vectors, the at least two first motion vectors being included in the at least two first motion information data; and - add the second set of motion information data to a history-based motion vector prediction (HMVP) list, items of the HMVP list being used to determine motion information data for a second block of the current picture.

15. The apparatus according to claim 14, wherein the second motion information data includes motion vectors determined as a function of the at least two first motion vectors of a selected sub-block of the first block, the selected sub-block being at a given position within the first block.

16. The apparatus according to claim 14 or 15, wherein the function of the at least two first motion vectors depends on the size of the first block or the size of the sub-block.

17. The apparatus according to claim 14, wherein the second set of motion information data includes a number of items, the number of items depending on the size of the first block or the size of the sub-block.

Citation Information

Patent Citations

  • Method and apparatus for affine inter prediction for video coding system

    US20190028731A1