Motion refinement by bidirectional matching for affine motion compensation in video coding
By employing bidirectional matching-based motion refinement, the video coding method improves the accuracy of motion vectors for affine motion prediction modes, addressing signaling overhead and enhancing coding efficiency.
Patent Information
- Application Number
- JP2023577623
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-06-17
- Filing Date
- 2022-06-16
- Publication Date
- 2025-06-11
- Estimated Expiration
- 2042-06-16
AI Technical Summary
Existing video coding standards, such as VVC and AVS3, face challenges in accurately estimating motion vectors for affine motion prediction modes, leading to signaling overhead and reduced coding efficiency.
The implementation of a video coding method that performs bidirectional matching-based motion refinement at the block level to iteratively update initial motion vectors, improving accuracy and reducing signaling overhead.
This approach enhances the accuracy of motion information for affine merge mode, achieving higher coding efficiency by refining motion vectors at both the block and sub-block levels.
Smart Images

Figure 0007691529000028 
Figure 0007691529000029 
Figure 0007691529000030
Abstract
Description
Technical Field
[0001] This application relates to video coding and compression. More specifically, this application relates to a video processing system and method for motion refinement in video.
Background Art
[0002] Digital video is supported by various electronic devices such as digital televisions, laptop or desktop computers, tablet computers, digital cameras, digital recording devices, digital media players, home game consoles, smartphones, video teleconferencing devices, video streaming devices, etc. Electronic devices transmit and receive digital video data through a communication network, or communicate in other ways, and / or store digital video data in a storage device. Since the bandwidth capacity of the communication network is limited and the memory resources of the storage device are limited, before the video data is communicated or stored, the video data may be compressed according to one or more video coding standards by video coding. Examples of video coding standards include Versatile Video Coding (VVC), Joint Exploration test Model (JEM), High-Efficiency Video Coding (HEVC / H.265), Advanced Video Coding (AVC / H.264), Moving Picture Expert Group (MPEG) coding, and the like. Video coding generally utilizes prediction methods (e.g., inter prediction, intra prediction, etc.) that utilize the redundancy inherent in video data. The purpose of video coding is to compress video data into a form that uses a lower bit rate while avoiding or minimizing the degradation of video quality.
Summary of the Invention
[0003] The implementation of the present disclosure provides a video coding method for motion refinement in video. The video coding method may include determining, by one or more processors, an initial motion vector for a video block of a video frame from the video. The video coding method may further include determining, by one or more processors, a matching target based on a weighted combination of a first reference block from a first reference frame in the video and a second reference block from a second reference frame in the video. The video coding method may further include performing, by one or more processors, a bidirectional matching-based motion refinement process at the block level to iteratively update the initial motion vector based on the matching target until a refined motion vector of the video block is obtained. The video coding method may further include refining, by one or more processors, a motion vector for each sub-block within the video block by using the refined motion vector of the video block as a starting point of the motion vector for the sub-blocks. Refining the motion vector at the sub-block level applies an affine motion model of the video block.
[0004] The implementation of the present disclosure also provides a video coding apparatus for motion refinement in video. The video coding apparatus may include a memory and one or more processors. The memory may be configured to store at least one video frame of the video. The video frame includes at least one video block. The one or more processors may be configured to determine an initial motion vector of the video block. The one or more processors may be configured to determine a matching target based on a weighted combination of a first reference block from a first reference frame in the video and a second reference block from a second reference frame in the video. The one or more processors may further be configured to perform a bidirectional matching-based motion refinement process at the block level to iteratively update the initial motion vector based on the matching target until a refined motion vector of the video block is obtained. The one or more processors may be configured to use the refined motion vector of the video block as a starting point of the motion vector for the sub-blocks and refine the motion vector for each sub-block within the video block. The one or more processors may apply an affine motion model of the video block to refine the motion vector at the sub-block level.
[0005] Also, when the implementation of the present disclosure is executed by one or more processors, it provides a non-transitory computer-readable storage medium storing instructions for causing the one or more processors to execute a video coding method for motion refinement in a video. The video coding method may include determining an initial motion vector of a video block of a video frame from the video based on a merge list of video blocks. The video coding method may further include determining a matching target based on a weighted combination of a first reference block from a first reference frame in the video and a second reference block from a second reference frame in the video. The video coding method may further include performing a bidirectional matching-based motion refinement process at the block level to iteratively update the initial motion vector based on the matching target until a refined motion vector of the video block is obtained. The video coding method may further include refining the motion vector for each sub-block within the video block using the refined motion vector of the video block as a starting point of the motion vector for the sub-blocks. Refining the motion vector at the sub-block level applies an affine motion model of the video block. The video coding method may further include generating a bitstream including a merge index for identifying the initial motion vector from the merge list, a first reference index for identifying the first reference frame, and a second reference index for identifying the second reference frame. The bitstream is stored in a non-transitory computer-readable storage medium.
[0006] It should be understood that both the above summary and the following detailed description are for illustrative purposes only and do not limit the present disclosure.
[0007] The accompanying drawings incorporated herein and constituting a part of this specification illustrate examples consistent with the present disclosure and, together with the specification, help to explain the principles of the present disclosure.
Brief Description of the Drawings
[0008]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 4C
Figure 4D
Figure 4E
Figure 5A
Figure 5B
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
DETAILED DESCRIPTION OF THE INVENTION
[0009] Here, specific implementations will be referred to in detail. Those examples are shown in the accompanying drawings. In the following detailed description, many non-limiting specific details are described to help understand the subject matter presented herein. However, it will be apparent to those skilled in the art that various alternatives can be used without departing from the scope of the claims, and the subject matter can be implemented without such specific details. For example, it will be apparent to those skilled in the art that the subject matter presented herein can be implemented on many types of electronic devices with digital video capabilities.
[0010] It should be explained that terms such as "first", "second", etc. used in the specification, the claims of the present disclosure, and the accompanying drawings are used to distinguish objects and are not used to describe a specific order or sequence. The data used in this way is interchangeable under appropriate conditions. Therefore, it should be understood that the embodiments of the present disclosure described herein can be implemented in an order other than that shown in the accompanying drawings or described in the present disclosure.
[0011] In the current Versatile Video Coding (VVC) standard and the Third Generation Audio Video Coding standard (AVS3), the motion information of the current coding block in a video decoder is inherited from spatially or temporally adjacent blocks in the form of a merge mode candidate index, or is derived based on the explicit signaling of the estimated motion information transmitted from the video encoder. However, the explicit signaling of the estimated motion information may result in signaling overhead. On the other hand, applying the merge mode motion vector (MV) can reduce the signaling overhead, but the merge mode MV may be copied only from adjacent blocks and thus may have low accuracy.
[0012] A video processing system and method consistent with the present disclosure are disclosed herein to improve the accuracy of motion vector estimation for the affine motion prediction mode used in both the VVC and AVS3 standards. Since bidirectional matching is a motion refinement method that does not require extra signaling, the systems and methods disclosed herein can apply bidirectional matching to improve the accuracy of the motion information for the affine merge mode and achieve higher coding efficiency. For example, various video coding techniques (including merge mode, affine mode, bidirectional matching, etc.) can be combined and applied to the systems and methods disclosed herein to enhance the motion information at both the block level and the sub-block level.
[0013] The systems and methods disclosed herein that are consistent with the present disclosure can improve the affine merge mode by applying bidirectional matching to refine the motion information of video blocks. Specifically, the systems and methods disclosed herein use a merge mode to derive an initial motion vector for a video block, determine a matching target for the video block, and perform a bidirectional matching-based motion refinement process at the video block level to repeatedly update the initial motion vector until a refined motion vector for the video block is obtained. For example, when bidirectional matching is applied, the initial motion vector is first derived for the video block as a starting point (e.g., a starting motion vector), and then the starting motion vector is repeatedly updated to obtain a refined motion vector with the minimum matching cost. The refined motion vector with the minimum matching cost can be selected as the motion vector of the video block at the video block level. Subsequently, the refined motion vector at the video block level can be used as a new starting point for further refining the motion information of sub-blocks at the sub-block level in the affine mode.
[0014] The affine merge mode described herein that conforms to the present disclosure may also be referred to as a combination of a merge mode and an affine mode. The merge mode can be an inter-coding mode used in video compression. In the merge mode, the motion vector of an adjacent video block is inherited for the currently encoded or decoded video block. For example, the merge mode causes the currently encoded video block to inherit the motion vector of a given neighbor. In another example, an index value may be used to identify a particular neighbor from which the currently encoded video block inherits its motion vector. A neighbor can be a video block that is spatially adjacent in the same video frame (e.g., an upper, upper-right, left, or lower-left video block), or a video block that is in the same location in a temporally adjacent video frame. The merge mode that conforms to the present disclosure can be used to determine the initial motion vector of the currently encoded video block (e.g., as a starting point for motion refinement). Regarding the affine mode, an affine motion model can be applied to inter-prediction. With reference to FIGS. 5A through 5B, the affine mode will be described in more detail below.
[0015] The affine mode design in the VVC standard that conforms to the present disclosure can be used as an exemplary implementation of an affine motion prediction mode to facilitate the description of the present disclosure. It is contemplated that the systems and methods disclosed herein can also apply different designs of affine motion prediction modes or other coding tools having the same or similar design spirit.
[0016] FIG. 1 is a block diagram showing an exemplary system 10 for encoding and decoding video blocks in parallel according to some implementations of the present disclosure. As shown in FIG. 1, system 10 includes a source device 12 that generates and encodes video data to be later decoded by a destination device 14. The source device 12 and the destination device 14 can each be any of a variety of electronic devices, including a desktop or laptop computer, a tablet computer, a smartphone, a set-top box, a digital television, a camera, a display device, a digital media player, a home game console, a video streaming device, and the like. In some implementations, the source device 12 and the destination device 14 have a wireless communication function.
[0017] In some implementations, the destination device 14 can receive the encoded video data to be decoded via a link 16. The link 16 can include any type of communication medium or device capable of transferring the encoded video data from the source device 12 to the destination device 14. In one example, the link 16 can include a communication medium that enables the source device 12 to directly transmit the encoded video data to the destination device 14 in real time. The encoded video data can be modulated according to a communication standard such as a wireless communication protocol and transmitted to the destination device 14. The communication medium can include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network such as a local area network or a wide area network, or part of a global network such as the Internet. The communication medium can include a router, a switch, a base station, or any other device useful for facilitating communication from the source device 12 to the destination device 14.
[0018] In other embodiments, the encoded video data may be transmitted from the output interface 22 to the storage device 32. The encoded video data within the storage device 32 may then be accessed by the destination device 14 via the input interface 28. The storage device 32 may include any of various distributed or local access data storage media, such as a hard drive, Blu-ray Disc (registered trademark), Digital Versatile Disc (DVD), Compact Disc Read-Only Memories (CD-ROM), flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. In yet another example, the storage device 32 may correspond to a file server or another intermediate storage device capable of storing the encoded video data generated by the source device 12. The destination device 14 may access the stored video data from the storage device 32 by streaming or downloading. The file server may be any type of computer capable of storing the encoded video data and transmitting the encoded video data to the destination device 14. Exemplary file servers include a web server (e.g., for a website), a File Transfer Protocol (FTP) server, a Network Attached Storage (NAS) device, or a local disk drive. The destination device 14 may access the encoded video data via any standard data connection, including a wireless channel (e.g., a Wireless Fidelity (Wi-Fi) connection), a wired connection (e.g., a Digital Subscriber Line (DSL), cable modem, etc.), or any combination of these suitable for accessing the encoded video data stored on the file server.The transmission of the encoded video data from the storage device 32 can be a streaming transmission, a download transmission, or a combination of both.
[0019] As shown in FIG. 1, the source device 12 includes a video source 18, a video encoder 20, and an output interface 22. The video source 18 can include, for example, a video capture device such as a video camera, a video archive including previously captured video, a video feed interface that receives video data from a video content provider, and / or a computer graphics system that generates computer graphics data as source video, or a combination of these sources. As an example, when the video source 18 is a video camera of a security monitoring system, the source device 12 and the destination device 14 can include a camera phone or a videophone. However, the implementations described in this disclosure are generally applicable to video coding and applicable to wireless and / or wired applications.
[0020] The captured, pre-captured, or computer-generated video can be encoded by the video encoder 20. The encoded video data may be directly transmitted to the destination device 14 via the output interface 22 of the source device 12. The encoded video data may be stored in the storage device 32 for later access by the destination device 14 or other devices for decoding and / or playback. The output interface 22 may further include a modem and / or a transmitter.
[0021] The destination device 14 includes an input interface 28, a video decoder 30, and a display device 34. The input interface 28 includes a receiver and / or a modem and can receive encoded video data via link 16. The encoded video data communicated via link 16 or provided on storage device 32 can include various syntax elements generated by video encoder 20 for use when video decoder 30 decodes the video data. Such syntax elements can be included within the encoded video data, transmitted over a communication medium, stored on a storage medium, or stored on a file server.
[0022] In some implementations, the destination device 14 may include a display device 34, which may be a display device integrally formed with the destination device 14 or an external display device configured to communicate with the destination device 14. The display device 34 displays the decoded video data for the user and can be any of various display devices such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or other types of display devices.
[0023] The video encoder 20 and the video decoder 30 are operable according to proprietary or industry standards such as VVC, HEVC, MPEG-4, Part 10, AVC, or extensions of these standards. It should be understood that the present disclosure is not limited to specific video encoding / decoding standards and is applicable to other video encoding / decoding standards as well. The video encoder 20 of the source device 12 is generally considered configurable to encode video data according to any of these current or future standards. Similarly, the video decoder 30 of the destination device 14 is generally considered configurable to decode video data according to any of these current or future standards.
[0024] Video encoder 20 and video decoder 30 can each be implemented as any of various suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When implemented partially in software, the electronic device stores software-oriented instructions in a suitable non-transitory computer-readable medium and executes the instructions in hardware using one or more processors for performing the video encoding / decoding operations disclosed in the present disclosure. Each of video encoder 20 and video decoder 30 may be included in one or more encoders or decoders, and any of these encoders or decoders may be integrated as part of a codec within their respective devices.
[0025] FIG. 2 is a block diagram showing an exemplary video encoder 20 according to some implementations described in the present application. Video encoder 20 can perform intra and inter prediction coding of video blocks within a video frame. Intra prediction coding relies on spatial prediction to reduce or remove spatial redundancy in video data within a given video frame or picture. Inter prediction coding relies on temporal prediction to reduce or remove temporal redundancy in video data within adjacent video frames or pictures of a video sequence. Note that the term "frame" may be used as a synonym for the term "image" or "picture" in the field of video coding.
[0026] As shown in FIG. 2, the video encoder 20 includes a video data memory 40, a prediction processing unit 41, a decoded picture buffer (DPB) 64, a summer 50, a conversion processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 further includes a motion estimation unit 42, a motion compensation unit 44, a splitting unit 45, an intra prediction processing unit 46, and an intra block copy (BC) unit 48. In some implementations, the video encoder 20 further includes an inverse quantization unit 58, an inverse conversion processing unit 60, and a summer 62 for video block reconstruction. An in-loop filter 63, such as a deblocking filter, that filters block boundaries to remove block artifacts from the reconstructed video data, may be disposed between the summer 62 and the DPB 64. Another in-loop filter, such as a sample adaptive offset (SAO) filter and / or an adaptive in-loop filter (ALF), may be used in addition to the deblocking filter to filter the output of the summer 62. In some embodiments, the in-loop filter may be omitted, and the decoded video blocks may be directly supplied to the DPB 64 by the summer 62. The video encoder 20 may take the form of a fixed or programmable hardware unit, or may be divided into one or more of the illustrated fixed or programmable hardware units.
[0027] The video data memory 40 can store video data to be encoded by components of the video encoder 20. The video data in the video data memory 40 may be obtained from the video source 18, for example, as shown in FIG. 1. The DPB 64 is a buffer that stores reference video data (for example, a reference frame or picture) used when the video encoder 20 encodes video data (for example, in an intra or inter prediction coding mode). The video data memory 40 and the DPB 64 may be formed by any of various memory devices. In various embodiments, the video data memory 40 may be on-chip with other components of the video encoder 20 or off-chip with respect to those components.
[0028] As shown in FIG. 2, after receiving video data, the splitting unit 45 in the prediction processing unit 41 splits the video data into video blocks. The above splitting may further include splitting the video frame into slices, tiles (e.g., a set of video blocks), or other larger coding units (CUs) according to a predefined splitting structure such as a quad tree (QT) structure associated with the video data. The video frame may be regarded as, or may be regarded as, a two-dimensional array or matrix of samples having sample values. Samples in the array are also called pixels or pels. The number of samples in the horizontal and vertical directions (or axes) of the array or picture defines the size and / or resolution of the video frame. The video frame can be split into a plurality of video blocks, for example, using QT splitting. A video block, although smaller in size than the video frame, may be regarded as, or may be regarded as, a two-dimensional array or matrix of samples having sample values. The number of samples in the horizontal and vertical directions (or axes) of the video block defines the size of the video block. The video block may be further split into one or more block splits or sub-blocks (the blocks can be reformed) by repeatedly using, for example, QT splitting, binary tree (BT) splitting, triple tree (TT) splitting, or any combination thereof. It should be noted that the term "block" or "video block" used here may be a part of a frame or picture, particularly a rectangular (square or non-square) portion. For example, with respect to HEVC and VVC, a block or video block may be a coding tree unit (CTU), a coding unit (CU), a prediction unit (PU), or a transform unit (TU), and / or a corresponding block, such as a coding tree block (CTB), a coding block (CB), a prediction block (PB), or a transform block (TB), or may correspond to these.As an alternative or in addition, the block or video block may be or correspond to sub-blocks such as CTB, CB, PB, TB, etc.
[0029] Based on the error results (e.g., coding rate and distortion level), the prediction processing unit 41 may select one of a plurality of prediction coding modes for the current video block, such as one of a plurality of intra prediction coding modes or one of a plurality of inter prediction coding modes. The prediction processing unit 41 may provide the resulting intra or inter prediction coded block (e.g., prediction block) to the summer 50 to generate a residual block, or may provide it to the summer 62 to reconstruct a coding block for subsequent use as part of a reference frame. The prediction processing unit 41 also provides syntax elements such as motion vectors, intra mode indicators, partition information, and other such syntax information to the entropy coding unit 56.
[0030] To select an appropriate intra prediction coding mode for the current video block, the intra prediction processing unit 46 within the prediction processing unit 41 may perform intra prediction coding of the current video block on one or more adjacent blocks within the same frame as the current block to be coded, providing spatial prediction. The motion estimation unit 42 and motion compensation unit 44 within the prediction processing unit 41 perform inter prediction coding of the current video block on one or more prediction blocks within one or more reference frames, providing temporal prediction. The video encoder 20 may execute a plurality of coding paths, for example, to select an appropriate coding mode for each block of video data.
[0031] In some embodiments, the motion estimation unit 42 determines an inter prediction mode for the current video frame by generating a motion vector indicating the displacement of a video block in the current video frame relative to a predicted block in a reference frame according to a predetermined pattern within a series of video frames. The motion estimation performed by the motion estimation unit 42 may be a process of generating a motion vector, which may estimate the motion of the video block. For example, the motion vector can indicate the displacement of a video block in the current video frame or picture relative to a predicted block in a reference frame. The predetermined pattern can specify video frames in the sequence as P frames or B frames. The intra BC unit 48 can determine a vector for intra BC coding, such as a block vector, in a manner similar to the determination of the motion vector for inter prediction by the motion estimation unit 42, and the motion estimation unit 42 can also be used to determine the block vector.
[0032] The predicted block of the video block may be or correspond to a block of the reference frame, or a reference block, that is considered to highly match the video block coded in terms of pixel difference, which may be determined by the sum of absolute differences (SAD), the sum of square differences (SSD), or other difference metrics. In some embodiments, the video encoder 20 can calculate the values of the sub-integer pixel positions of the reference frames stored in the DPB 64. For example, the video encoder 20 may interpolate the values of the 1 / 4 pixel position, 1 / 8 pixel position, or other fractional pixel positions of the reference frame. Accordingly, the motion estimation unit 42 may perform motion search for full pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.
[0033] The motion estimation unit 42 calculates the motion vector of a video block in an inter-predicted coded frame by comparing the position of the video block with the position of a predicted block in a reference frame selected from a first reference frame list (List0) or a second reference frame list (List1) that each identify one or more reference frames stored in the DPB 64. The motion estimation unit 42 transmits the calculated motion vector to the motion compensation unit 44 and then to the entropy encoding unit 56.
[0034] Motion compensation performed by the motion compensation unit 44 may include fetching or generating a predicted block based on the motion vector determined by the motion estimation unit 42. Upon receiving the motion vector of the current video block, the motion compensation unit 44 may find a predicted block in one of the reference frame lists to which the motion vector is directed, search for the predicted block from the DPB 64, and forward the predicted block to the summer 50. The summer 50 then generates a residual block of pixel difference values by subtracting the pixel values of the predicted block provided by the motion compensation unit 44 from the pixel values of the coded current video block. The pixel difference values for generating the residual block may include luminance or chrominance difference components, or both. Also, the motion compensation unit 44 may generate syntax elements associated with the video blocks of the video frame for use when the video decoder 30 decodes the video blocks of the video frame. The syntax elements can include, for example, syntax elements that define the motion vectors used to identify the predicted blocks, any flags indicating the prediction mode, or any other syntax information described herein. Note that the motion estimation unit 42 and the motion compensation unit 44, which are separately illustrated in FIG. 2 for conceptual purposes, may be integrally formed.
[0035] In some embodiments, the intra BC unit 48 can generate vectors and fetch prediction blocks in a manner similar to the method described above in relation to the motion estimation unit 42 and the motion compensation unit 44, where the prediction blocks are in the same frame as the current block being coded, and the vectors are called block vectors as opposed to motion vectors. In particular, the intra BC unit 48 may determine the intra prediction mode to be used for coding the current block. In some examples, the intra BC unit 48 may code the current block using various intra prediction modes, for example, between separate coding passes, and test their performance through rate-distortion analysis. The intra BC unit 48 can then select an appropriate intra prediction mode to use from among the various tested intra prediction modes and generate an intra mode indicator accordingly. For example, the intra BC unit 48 can calculate rate-distortion values using rate-distortion analysis for the various tested intra prediction modes and select, as the appropriate intra prediction mode to use, the intra prediction mode having the optimal rate-distortion characteristics from among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the coded block and the original uncoded block that was coded to generate the coded block, as well as the bit rate (i.e., the number of bits) used to generate the coded block. The intra BC unit 48 may calculate a ratio from the distortions and rates of the various coded blocks and determine which intra prediction mode exhibits the optimal rate-distortion value for the block.
[0036] In other embodiments, the intra BC unit 48 may use all or part of the motion estimation unit 42 and the motion compensation unit 44 to perform such functions for intra BC prediction according to the implementations described herein. In any case, for intra block copy, the predicted block may be a block that is considered to highly match the block to be coded with respect to pixel differences that can be determined by SAD, SSD, or other difference metrics, and the identification of the predicted block may include the calculation of sub-integer pixel position values.
[0037] Regardless of whether the predicted block is from the same frame by intra prediction or from a different frame by inter prediction, the video encoder 20 can generate a residual block and form pixel difference values by subtracting the pixel values of the predicted block from the pixel values of the current video block being coded. The pixel difference values for generating the residual block may include both the differences of the luminance component and the chrominance component.
[0038] The intra prediction processing unit 46 may intra predict the current video block as an alternative to the inter prediction performed by the motion estimation unit 42 and the motion compensation unit 44, or the intra block copy prediction performed by the intra BC unit 48, as described above. In particular, the intra prediction processing unit 46 may determine the intra prediction mode to be used for encoding the current block. For example, the intra prediction processing unit 46 may use various intra prediction modes to encode the current block, for example, between separate encoding paths, and the intra prediction processing unit 46 (or in some embodiments, the mode selection unit) may select an appropriate intra prediction mode to use from the tested intra prediction modes. The intra prediction processing unit 46 may provide information indicating the selected intra prediction mode for the block to the entropy encoding unit 56. The entropy encoding unit 56 may encode the information indicating the selected intra prediction mode in the bitstream.
[0039] After the prediction processing unit 41 determines the prediction block of the current video block by either inter prediction or intra prediction, the summer 50 generates a residual block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and is provided to the transformation processing unit 52. The transformation processing unit 52 transforms the residual video data into transform coefficients using a transformation such as a discrete cosine transform (DCT) or a conceptually similar transformation.
[0040] The transformation processing unit 52 may send the obtained transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameter. In some embodiments, the quantization unit 54 may then perform a scan of the matrix containing the quantized transform coefficients. Alternatively, the entropy coding unit 56 may perform the scan.
[0041] Following quantization, entropy coding unit 56 can use an entropy coding technique, such as Context Adaptive Variable Length Coding (CAVLC), Context Adaptive Binary Arithmetic Coding (CABAC), Syntax-based context-adaptive Binary Arithmetic Coding (SBAC), Probability Interval Partitioning Entropy (PIPE) coding, or another entropy coding technique or technology, to encode the quantized transform coefficients into a video bitstream. The encoded bitstream may then be transmitted to video decoder 30 as shown in FIG. 1, or may be archived in storage device 32 as shown in FIG. 1 for later transmission to or retrieval by video decoder 30. Entropy coding unit 56 may also use an entropy coding technique to encode the motion vectors and other syntax elements of the current video frame being coded.
[0042] Inverse quantization unit 58 and inverse transform processing unit 60 each apply inverse quantization and inverse transform to reconstruct the residual block in the pixel domain to generate a reference block for prediction of other video blocks. The reconstructed residual block may be generated. As described above, motion compensation unit 44 can generate a motion compensation prediction block from one or more reference blocks of the frames stored in DPB 64. Motion compensation unit 44 can also apply one or more interpolation filters to the prediction block to calculate sub-integer pixel values for use in motion estimation.
[0043] Summer 62 generates a reference block for adding the reconstructed residual block to the motion compensation prediction block generated by the motion compensation unit 44 and storing it in the DPB 64. Then, the reference block may be used as a prediction block by the intra BC unit 48, the motion estimation unit 42, and the motion compensation unit 44 to inter-predict another video block in a subsequent video frame.
[0044] FIG. 3 is a block diagram showing an exemplary video decoder 30 according to some implementations of the present application. The video decoder 30 includes a video data memory 79, an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, a summer 90, and a DPB 92. The prediction processing unit 81 further includes a motion compensation unit 82, an intra prediction unit 84, and an intra BC unit 85. The video decoder 30 can perform a decoding process that generally reverses the encoding process described above for the video encoder 20 in relation to FIG. 2. For example, the motion compensation unit 82 can generate prediction data based on the motion vector received from the entropy decoding unit 80, while the intra prediction unit 84 can generate prediction data based on the intra prediction mode indicator received from the entropy decoding unit 80.
[0045] In some embodiments, the units of the video decoder 30 may be tasked to perform the implementations of the present application. Also, in some embodiments, the implementations of the present disclosure may be divided among one or more units of the video decoder 30. For example, the intra BC unit 85 can perform the implementations of the present application alone or in combination with other units of the video decoder 30 such as the motion compensation unit 82, the intra prediction unit 84, and the entropy decoding unit 80. In some embodiments, the video decoder 30 may not include the intra BC unit 85, and the functions of the intra BC unit 85 may be performed by other components of the prediction processing unit 81 such as the motion compensation unit 82.
[0046] The video data memory 79 can store video data such as an encoded video bitstream that is decoded by other components of the video decoder 30. The video data stored in the video data memory 79 may be obtained, for example, from the storage device 32, may be obtained from a local video source such as a camera through wired or wireless network communication of the video data, or may be obtained by accessing a physical data storage medium (such as a flash drive or a hard disk). The video data memory 79 may include a Coded Picture Buffer (CPB) that stores encoded video data from the encoded video bitstream. The DPB 92 of the video decoder 30 stores reference video data for use in decoding video data by the video decoder 30 (for example, in an intra or inter prediction coding mode). The video data memory 79 and the DPB 92 may be formed by any of various memory devices such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. For the sake of explanation, the video data memory 79 and the DPB 92 are shown as two separate components of the video decoder 30 in FIG. 3. However, it will be apparent to those skilled in the art that the video data memory 79 and the DPB 92 may be provided by the same memory device or separate memory devices. In some embodiments, the video data memory 79 may be on-chip with other components of the video decoder 30 or may be off-chip with respect to those components.
[0047] During decoding, video decoder 30 receives an encoded video bitstream that represents video blocks of an encoded video frame and associated syntax elements. Video decoder 30 can receive syntax elements at the video frame level and / or at the video block level. Entropy decoding unit 80 of video decoder 30 decodes the bitstream using entropy decoding techniques and can obtain quantization coefficients, motion vectors or intra prediction mode indicators, and other syntax elements. Entropy decoding unit 80 then transfers the motion vector or intra prediction mode indicator, and other syntax elements to prediction processing unit 81.
[0048] When a video frame is coded as an intra prediction coded (e.g., I) frame, or is coded for an intra coding prediction block within another type of frame, intra prediction unit 84 of prediction processing unit 81 can generate prediction data for video blocks of the current video frame based on the signaled intra prediction mode and reference data from blocks decoded prior to the current frame.
[0049] When a video frame is coded as an inter prediction coded (i.e., B or P) frame, motion compensation unit 82 of prediction processing unit 81 generates one or more prediction blocks for video blocks of the current video frame based on the motion vectors and other syntax elements received from entropy decoding unit 80. Each of the prediction blocks may be generated from a reference frame within one of the reference frame lists. Video decoder 30 can construct reference frame lists such as List0 and List1 using default construction techniques based on the reference frames stored in DPB 92.
[0050] In some embodiments, when a video block is coded according to the intra BC mode described herein, the intra BC unit 85 of the prediction processing unit 81 generates a prediction block of the current video block based on the block vector and other syntax elements received from the entropy decoding unit 80. The prediction block may be within the reconstructed area of the same picture as the current video block processed by the video encoder 20.
[0051] The motion compensation unit 82 and / or the intra BC unit 85 determines prediction information of the video blocks of the current video frame by analyzing the motion vectors and other syntax elements, and then uses the prediction information to generate a prediction block of the current video block being decoded. For example, the motion compensation unit 82 uses some of the received syntax elements to determine the prediction mode (e.g., intra or inter prediction) used to code the video blocks of the video frame, the inter prediction frame type (e.g., B or P), the construction information of one or more reference frame lists of the frame, the motion vectors of each inter prediction coded video block of the frame, the inter prediction status of each inter prediction coded video block of the frame, and other information to decode the video blocks in the current video frame.
[0052] Similarly, the intra BC unit 85 may use some of the received syntax elements, such as flags, to determine that the current video block is in the intra BC mode, the video blocks of the frame are within the reconstructed area, the construction information to be stored in the DPB 92, the block vectors of each intra BC predicted video block of the frame, the intra BC prediction status of each intra BC predicted video block of the frame, and other information to determine that it has been predicted, and decode the video blocks in the current video frame.
[0053] The motion compensation unit 82 may perform interpolation using an interpolation filter used by the video encoder 20 during the encoding of video blocks, and calculate interpolation values for sub-integer pixels of a reference block. In this case, the motion compensation unit 82 can determine the interpolation filter used by the video encoder 20 from the received syntax elements and generate a prediction block using the interpolation filter.
[0054] The inverse quantization unit 86 inverse quantizes the quantized transform coefficients provided in the bitstream and entropy decoded by the entropy decoding unit 80 using the same quantization parameters calculated by the video encoder 20 for each video block in the video frame to determine the degree of quantization. The inverse transform processing unit 88 applies an inverse transform, e.g., an inverse DCT, an inverse integer transform, or a conceptually similar inverse transform process, to the transform coefficients to reconstruct the residual block in the pixel domain.
[0055] After the motion compensation unit 82 or the intra BC unit 85 generates a prediction block of the current video block based on vectors and other syntax elements, the summer 90 reconstructs the decoded video block of the current video block by summing the residual block from the inverse transform processing unit 88 and the corresponding prediction block generated by the motion compensation unit 82 and the intra BC unit 85. The decoded video block is also referred to as the reconstructed block of the current video block. An in-loop filter 91, such as a deblocking filter, an SAO filter, and / or an ALF, may be located between the summer 90 and the DPB 92 to further process the decoded video block. In some embodiments, the in-loop filter 91 may be omitted, and the decoded video block may be directly supplied from the summer 90 to the DPB 92. The decoded video blocks within a given frame are then stored in the DPB 92, and the DPB 92 stores the reference frames used for subsequent motion compensation of the next video block. The DPB 92, or a memory device separate from the DPB 92, may store the decoded video for later presentation on a display device such as the display device 34 of FIG. 1.
[0056] In a typical video coding process (including, for example, video encoding and video decoding processes), a video sequence typically includes a set of ordered frames or pictures. Each frame can include three sample arrays denoted as SL, SCb, and SCr. SL is a two-dimensional array of luminance samples. SCb is a two-dimensional array of Cb chrominance samples. SCr is a two-dimensional array of Cr chrominance samples. In other examples, a frame may be monochromatic and thus include only one two-dimensional array of luminance samples.
[0057] As shown in FIG. 4A, the video encoder 20 (or more specifically, the splitting unit 45) generates an encoded representation of a frame by first splitting the frame into a set of CTUs. A video frame can include an integer number of CTUs arranged in a raster scan order from left to right and top to bottom. Each CTU is the largest logical coding unit, and the width and height of the CTU are signaled by the video encoder 20 in the sequence parameter set such that all CTUs in the video sequence have the same size, which can be one of 128×128, 64×64, 32×32, 16×16. However, it should be noted that the CTUs in the present disclosure are not necessarily limited to a specific size. As shown in FIG. 4B, each CTU may include a CTB of one luminance sample, corresponding coding tree blocks of two chrominance samples, and syntax elements used to code the samples of the coding tree blocks. The syntax elements describe the properties of different types of units of the coded pixel block and how the video sequence can be reconstructed by the video decoder 30, and include inter or intra prediction, intra prediction mode, motion vectors, and other parameters. In a monochrome picture or a picture having three separate color planes, the CTU may include a single coding tree block and syntax elements used to code the samples of the coding tree block. The coding tree block may be a block of samples of N×N.
[0058] To achieve better performance, the video encoder 20 may recursively perform tree partitioning, such as binary tree partitioning, ternary tree partitioning, quad tree partitioning, or a combination thereof, on the coding tree blocks of the CTU, and divide the CTU into smaller CUs. As shown in FIG. 4C, a 64×64 CTU 400 is first divided into four smaller CUs, each of the smaller CUs having a block size of 32×32. Among the four smaller CUs, CU 410 and CU 420 are each divided into four CUs of 16×16 block size. Two 16×16 CUs 430 and 440 are further divided into four CUs of 8×8 block size each. FIG. 4D shows a quad tree data structure indicating the final result of the partitioning process of CTU 400 as shown in FIG. 4C, and each leaf node of the quad tree corresponds to one CU of each size from 32×32 to 8×8. Similar to the CTU shown in FIG. 4B, each CU may include a CB of luminance samples, corresponding coding blocks of chrominance samples of two frames of the same size, and syntax elements used to code the samples of the coding blocks. In a monochrome picture or a picture having three separate color planes, the CU may include a single coding block and a syntax structure used to code the samples of the coding block. The quad tree partitioning shown in FIGS. 4C and 4D is for illustrative purposes only, and it should be noted that one CTU can be divided into multiple CUs to adapt to various local characteristics based on quad / ternary / binary tree partitioning. In a multi-type tree structure, one CTU may be divided by a quad tree structure, and the leaf CUs of each quad tree may be further divided by binary and ternary tree structures. As shown in FIG. 4E, there are multiple possible partitioning types for a coding block having a width W and a height H, namely, four-way partitioning, vertical binary partitioning, horizontal binary partitioning, vertical ternary partitioning, vertically extended ternary partitioning, horizontal ternary partitioning, and horizontally extended ternary partitioning.
[0059] In some implementation forms, the video encoder 20 can further divide the coding block of the CU into one or more M×N PBs. A PB is a block of rectangular (square or non-square) samples to which the same inter or intra prediction is applied. The PU of the CU may include a PB of luminance samples, corresponding PBs of two chrominance samples, and syntax elements used to predict the PB. In a monochromatic picture or a picture having three separate color planes, the PU may include a single PB and a syntax structure used to predict the PB. The video encoder 20 can generate prediction luminance, Cb, and Cr blocks for the luminance, Cb, and Cr PBs of each PU of the CU.
[0060] The video encoder 20 can generate a prediction block of the PU using intra prediction or inter prediction. When the video encoder 20 uses intra prediction to generate a prediction block of the PU, the video encoder 20 can generate a prediction block of the PU based on the decoded samples of the frame associated with the PU. When the video encoder 20 uses inter prediction to generate a prediction block of the PU, the video encoder 20 can generate a prediction block of the PU based on the decoded samples of one or more frames other than the frame associated with the PU.
[0061] After the video encoder 20 generates prediction luminance, Cb, and Cr blocks for one or more PUs of a CU, the video encoder 20 can generate a luminance residual block of the CU by subtracting the predicted luminance block from the original luminance coding block of the CU. Thus, each sample of the luminance residual block of the CU indicates the difference between the luminance sample in one of the predicted luminance blocks of the CU and the corresponding sample in the original luminance coding block of the CU. Similarly, the video encoder 20 can generate a Cb residual block and a Cr residual block of the CU, respectively. Thus, each sample of the Cb residual block of the CU indicates the difference between the Cb sample in one of the predicted Cb blocks of the CU and the corresponding sample in the original Cb coding block of the CU, and each sample of the Cr residual block of the CU may indicate the difference between the Cr sample in one of the predicted Cr blocks of the CU and the corresponding sample in the original Cr coding block of the CU.
[0062] Furthermore, as shown in FIG. 4C, the video encoder 20 can use quadtree partitioning to decompose the luminance, Cb, and Cr residual blocks of a CU into one or more luminance, Cb, and Cr transform blocks. A transform block may include a block of samples in a rectangle (square or non-square) to which the same transform is applied. A TU of a CU may include a transform block of luminance samples, corresponding transform blocks of two chrominance samples, and syntax elements used to transform the transform block samples. Thus, each TU of a CU may be associated with a luminance transform block, a Cb transform block, and a Cr transform block. In some embodiments, the luminance transform block associated with a TU may be a sub-block of the luminance residual block of the CU. The Cb transform block may be a sub-block of the Cb residual block of the CU. The Cr transform block may be a sub-block of the Cr residual block of the CU. In a monochrome picture or a picture having three separate color planes, a TU may include a single transform block and a syntax structure used to transform the samples of the transform block.
[0063] The video encoder 20 can apply one or more transforms to the luminance transform blocks of the TUs to generate the luminance coefficient blocks of the TUs. The coefficient blocks may be two-dimensional arrays of transform coefficients. The transform coefficients may be scalar quantities. The video encoder 20 can apply one or more transforms to the Cb transform blocks of the TUs to generate the Cb coefficient blocks of the TUs. The video encoder 20 can apply one or more transforms to the Cr transform blocks of the TUs to generate the Cr coefficient blocks of the TUs.
[0064] After generating a coefficient block (e.g., a luminance coefficient block, a Cb coefficient block, or a Cr coefficient block), the video encoder 20 can quantize the coefficient block. Quantization generally refers to the process of quantizing the transform coefficients to minimize the amount of data used to represent the transform coefficients and provide further compression. After the video encoder 20 quantizes the coefficient block, the video encoder 20 can apply an entropy coding technique to encode the syntax elements indicating the quantized transform coefficients. For example, the video encoder 20 can perform CABAC on the syntax elements indicating the quantized transform coefficients. Finally, the video encoder 20 can output a bitstream including a sequence of bits forming the coded frame and the representation of the associated data, and the bitstream is stored in the storage device 32 or transmitted to the destination device 14.
[0065] After receiving the bitstream generated by the video encoder 20, the video decoder 30 can analyze the bitstream and obtain syntax elements from the bitstream. The video decoder 30 can reconstruct a frame of video data based at least in part on the syntax elements obtained from the bitstream. The process of reconstructing the video data generally reverses the encoding process executed by the video encoder 20. For example, the video decoder 30 can perform an inverse transform on the coefficient block associated with the TU of the current CU to reconstruct the residual block associated with the TU of the current CU. The video decoder 30 further reconstructs the coding block of the current CU by adding the samples of the prediction block of the current CU's PU to the corresponding samples of the transform block of the current CU's TU. After reconstructing the coding blocks of each CU in the frame, the video decoder 30 can reconstruct the frame.
[0066] As described above, video coding mainly uses two modes, namely, intra prediction (or intra-frame prediction) and inter prediction (or inter-frame prediction), to achieve video compression. Note that intra-block copy (IBC) can be regarded as either intra prediction or a third mode. Among the two modes, inter prediction contributes more to coding efficiency than intra prediction because it uses a motion vector to predict the current video block from a reference video block.
[0067] However, as video data capture technology improves and the video block size for retaining video data details becomes finer, the amount of data required to represent the motion vectors of the current frame has increased significantly. One way to overcome this problem is to take advantage of the fact that groups of adjacent CUs in both the spatial and temporal domains not only have similar video data for prediction purposes, but also have similar motion vectors between these adjacent CUs. Therefore, the motion information of spatially adjacent CUs and / or CUs at the same location temporally, also called the "Motion Vector Predictor (MVP)" of the current CU, can be used as an approximation of the motion information (e.g., motion vectors) of the current CU by exploring its spatial and temporal correlations.
[0068] Instead of encoding the actual motion vectors of the current CU into the video bitstream (e.g., as described above in relation to FIG. 2, where the actual motion vectors are determined by the motion estimation unit 42), the motion vector predictor of the current CU is subtracted from the actual motion vectors of the current CU to generate the Motion Vector Difference (MVD) of the current CU. By doing so, it becomes unnecessary to encode the motion vectors determined by the motion estimation unit 42 for each CU of the frame into the video bitstream, and the amount of data used to represent the motion information within the video bitstream can be significantly reduced.
[0069] Similar to the process of selecting a predicted block within a reference frame during inter - frame prediction of a code block, a set of rules can be adopted by both the video encoder 20 and the video decoder 30 to construct a motion vector candidate list (also referred to as a "merge list") for the current CU using potential candidate motion vectors related to the current CU and spatially adjacent CUs and / or CUs at the same temporal location, and then select one member from the motion vector candidate list as the motion vector predictor for the current CU. By doing so, it is not necessary to transmit the motion vector candidate list itself from the video encoder 20 to the video decoder 30, and the index of the selected motion vector predictor within the motion vector candidate list is sufficient for both the video encoder 20 and the video decoder 30 to use the same motion vector predictor within the motion vector candidate list for encoding and decoding the current CU. Therefore, only the index of the selected motion vector predictor needs to be sent from the video encoder 20 to the video decoder 30.
[0070] A brief description of the affine mode is provided herein with reference to FIGS. 5A - 5B. In HEVC, only the translational motion model is applied to motion - compensated prediction. In the real world, various types of motion such as zoom - in, zoom - out, rotation, perspective motion, and other irregular motions may exist. In the VVC and AVS3 standards, affine motion - compensated prediction can be applied by signaling a flag for each inter - coding block to indicate whether the translational motion model or the affine motion model is applied to inter - prediction. In some implementations, one of two affine modes (e.g., the 4 - parameter affine motion model shown in FIG. 5A or the 6 - parameter affine motion model shown in FIG. 5B) may be selected and applied to the affine - coded video block.
[0071] The 4-parameter affine motion model shown in FIG. 5A includes the following affine parameters: namely, two parameters for translational movement in the horizontal and vertical directions respectively, one parameter for zoom movement, and one parameter for rotational movement in both the horizontal and vertical directions. In this model, the horizontal zoom parameter can be made equivalent to the vertical zoom parameter, and the horizontal rotation parameter can be made equivalent to the vertical rotation parameter. To achieve better adaptation of the motion vector and the affine parameters, the affine parameters of this model can be encoded by two motion vectors (referred to as control point motion vectors (CPMV)) located at two control points (e.g., the upper left corner and the upper right corner) of the current video block. As shown in FIG. 5A, the affine motion field of the video block (e.g., the motion vectors of the video block) can be described by two CPMVs, V 0 and V1. Based on the control point motion, the motion field of the affine-encoded sub-block having a position (x, y) within the video block can be derived by the following equation (1).
Equation
[0072] In Equation (1), v x and v y represent the x-component and y-component respectively at the (x, y) position of the motion vector of the affine-encoded sub-block. w represents the width of the video block. v 0x and v 0y represent the x-component and y-component of CPMV V 0 respectively. v 1x and v 1y represent the x-component and y-component of CPMV V 1 respectively.
[0073] The 6-parameter affine motion model shown in FIG. 5B has the following affine parameters, namely, two parameters for translation in the horizontal and vertical directions respectively, two parameters for zoom and rotation motions in the horizontal direction, and two more parameters for zoom and rotation motions in the vertical direction. The 6-parameter affine motion model can be encoded by three CPMVs at three control points. As shown in FIG. 5B, the three control points of the 6-parameter affine video block are at the upper left corner, upper right corner, and lower left corner of the video block, and are associated with CPMV V 0 , V 1 , V 2 . The motion at the upper left control point is related to translation, the motion at the upper right control point is related to rotation and zoom motions in the horizontal direction, and the motion at the lower left control point is related to rotation and zoom motions in the vertical direction. Compared with the 4-parameter affine motion model, the rotation and zoom motions in the horizontal direction of the 6-parameter affine motion model may not be the same as the rotation and zoom motions in the vertical direction. The motion vector (v x , v y ) of each sub-block at the position (x, y) of the video block can be derived as follows using three CPMVs at three control points.
Equation
[0074] In Equation (2), v x and v y represent the x-component and y-component respectively at the (x, y) position of the motion vector of the affine-encoded sub-block. w and h are the width and height of the video block. v 0x and v 0y represent the x-component and y-component of CPMV V 0 respectively. v 1x and v 1y represent the x-component and y-component of CPMV V 1 respectively. v 2x and v 2y represent the x-component and y-component of CPMV V 2represents the x-component and y-component.
[0075] FIG. 6 is a diagram illustrating exemplary bidirectional matching according to some implementations of the present disclosure. In the area of video coding, bidirectional matching is a technique in which the motion information of a currently encoded video block is not signaled to the decoder side but is derived at the decoder side. When bidirectional matching is used in the motion derivation process, an initial motion vector can be derived for the entire video block first. Specifically, the merge list of the video block can be checked, and a candidate motion vector from the merge list that leads to the minimum matching cost among all candidate motion vectors in the merge list can be selected as the starting point. Next, a local search centered on the starting point within the search range can be performed, and the motion vector with the minimum matching cost within the search range can be used as the motion vector for the entire video block. Subsequently, using the motion vector of the entire video block as a new starting point, the motion information can be further refined at the sub-block level. For example, a plurality of CPMVs can be derived for the entire video block, and then, based on the above formula (1) or (2), by applying the CPMV at the video block level, the motion vector at the sub-block level can be derived.
[0076] As shown in FIG. 6, by finding two best-match reference blocks 604, 606 along the motion trajectory of the video block from two different reference frames, the motion information of the video block 602 in the video frame can be derived using bidirectional matching. Assuming a continuous motion trajectory, the motion vectors MV0 and MV1 pointing to the two reference blocks 604, 606 can be proportional to the temporal distances of the reference frames with respect to the video frame (e.g., TD0 and TD1), respectively. As a special case, when the video frame is temporally between two reference frames and the temporal distances from the video frame to the two reference frames are the same (e.g., TD0 = TD1), the motion vector derived from bidirectional matching becomes a mirror-based bidirectional motion vector.
[0077] FIG. 7 is a block diagram showing an exemplary process 700 for motion refinement by bidirectional matching for affine motion compensation, according to some implementations of the present disclosure. In some implementations, process 700 may be performed by a prediction processing unit 41 (e.g., including a motion estimation unit 42, a motion compensation unit 44, etc.) of video encoder 20, or a prediction processing unit 81 (e.g., including a motion compensation unit 82) of video decoder 30. In some implementations, process 700 may be performed by a video processor (e.g., processor 1120 as shown in FIG. 11) on the encoder side or the decoder side. For illustrative purposes only, the following description of process 700 is provided with respect to a video processor.
[0078] To encode or decode a video block from a video frame of a video, the video processor may perform an initial motion vector estimation 702 to generate an initial motion vector 704 for the video block. For example, the video processor can determine the initial motion vector 704 for the video block based on the merge list of the video block. Specifically, the merge list of the video block can be checked, and a candidate motion vector from the merge list that leads to the minimum matching cost among all candidate motion vectors in the merge list can be selected as the initial motion vector 704.
[0079] The video processor may perform bidirectional matching-based motion refinement processing 706 at the video block level to iteratively update the initial motion vector 704 until a refined motion vector 714 of the video block is obtained. The initial motion vector 704 can be used as a starting point (e.g., starting motion vector) for the bidirectional matching-based motion refinement processing 706. When iterative updates around the starting motion vector are performed, the matching cost (e.g., bidirectional matching cost) between the current prediction of the video block and the matching target can be calculated iteratively to guide the gradual update of the starting motion vector of the video block. In some implementations, the matching cost between the current prediction of the video block and the matching target can be calculated based on a matching cost function. The matching cost function can be the sum of absolute differences (SAD), the mean removed SAD (MRSAD) obtained by removing SAD, the sum of square differences (SSD), or other appropriate difference metrics between the current prediction of the video block and the matching target.
[0080] When the video block is coded in affine mode, the initial motion vector 704 may include one or more initial CPMVs at one or more control points of the video block. The refined motion vector 714 may include one or more refined CPMVs at one or more control points.
[0081] First, in the bidirectional matching-based motion refinement process 706, the video processor may execute a matching target determination operation 708 to determine a matching target for iterative update of motion information. For example, referring to FIG. 8, the video processor may determine a first reference block Ref0 and a second reference block Ref1 from a first reference frame 802 and a second reference frame 804 of the video, respectively, based on the initial motion vector 704. The video processor may determine the matching target based on a weighted combination of the first reference block Ref0 and the second reference block Ref1. For example, the matching target may be equal to the weighted sum of Ref0 and Ref1 (e.g., matching target = w0 * Ref0 + w1 * Ref1, where w0 and w1 represent the weights of Ref0 and Ref1, respectively).
[0082] In some implementations, the intercoding mode (e.g., merge mode) disclosed herein may be a bidirectional prediction indicating that two different lists of reference frames (e.g., List0 and List1) are used to identify two predictions of a video block. For example, List0 may include a list of reference frames preceding the video block, and List1 may include a list of reference frames following the video block. Ref0 may be a List0 prediction from the first reference frame 802 based on the initial motion vector 704. Ref1 may be a List1 prediction from the second reference frame 804 based on the initial motion vector 704. The matching target may be a weighted sum of the List0 prediction and the List1 prediction derived based on the initial motion vector 704.
[0083] Alternatively, the matching target can also be the weighted combination of the List0 and List1 predictions plus the corresponding prediction residuals associated with the List0 and List1 predictions. In this case, the matching target can be the weighted combination of the List0 reconstruction and the List1 reconstruction. For example, List0 reconstruction = List0 prediction + List0 prediction residual, List1 reconstruction = List1 prediction + List1 prediction residual, and matching target = w0 * List0 reconstruction + w1 * List1 reconstruction.
[0084] In some implementations, the weights w0 and w1 may reuse the same values derived on the encoder side for normal weighted bidirectional prediction (e.g., CU-level weighted bidirectional prediction). Alternatively, the weights w0 and w1 may have predetermined values. For example, w0 = w1 = 1 / 2. In another example, w0 = 1 and w1 = 0, or w0 = 0 and w1 = 1. When one of the weights w0 and w1 is 0, the bidirectional matching becomes unidirectional motion vector refinement instead of bidirectional motion vector refinement.
[0085] In some implementations, the weights w0 and w1 may have values with different signs. For example, w0 = 1, w1 = -1. In this case, the bidirectional prediction difference can be used to calculate the matching cost. Specifically, the bidirectional prediction difference generated before the starting motion vector is updated and the bidirectional prediction difference generated after the starting motion vector is updated are calculated to determine the matching cost.
[0086] Referring again to the bidirectional matching-based motion refinement process 706 of FIG. 7, the video processor may repeatedly execute the motion refinement operation 710 and the motion vector update operation 712 until the refined motion vector 714 is generated for the video block. For example, the video processor can use the initial motion vector 704 to initialize the intermediate motion vector for the video block and determine the motion refinement of the intermediate motion vector based on the matching target. The video processor can update the intermediate motion vector based on the motion refinement. The intermediate motion vector can represent the motion vector of the video block while performing the bidirectional matching-based motion refinement process 706. Next, the video processor may determine whether a predetermined iteration stop condition is satisfied. If the predetermined iteration stop condition is satisfied, the video processor may determine the intermediate motion vector as the refined motion vector 714. On the other hand, if the predetermined iteration stop condition is not satisfied, the video processor may repeatedly determine the motion refinement for the intermediate motion vector and continue to update the intermediate motion vector based on the motion refinement until the predetermined iteration stop condition is satisfied.
[0087] In some implementations, the predetermined iteration stop condition may be satisfied when the intermediate motion vector converges. Alternatively, the predetermined iteration stop condition may be satisfied when the total number of iterations meets a predetermined threshold (e.g., the total number of iterations reaches a predetermined upper limit).
[0088] In some implementations, the motion refinement of the intermediate motion vector can be determined through a calculation-based derivation, a search-based derivation, or a combination of a calculation-based derivation and a search-based derivation. A first exemplary process in which a calculation-based derivation is used to determine motion refinement, a second exemplary process in which a search-based derivation is used to determine motion refinement, and a third exemplary process in which a combination of a calculation-based derivation and a search-based derivation is used to determine motion refinement are provided below.
[0089] In a first exemplary process to which a computation-based derivation is applied, the video processor can determine a current prediction of a video block based on an intermediate motion vector. For example, the video processor can determine a third reference block (Ref2) and a fourth reference block (Ref3) from a first reference frame 802 and a second reference frame 804 respectively based on the intermediate motion vector. The video processor can determine the current prediction of the video block based on a weighted combination of the third reference block Ref2 and the fourth reference block Ref3 (e.g., current prediction = w2*Ref2 + w3*Ref3, where w2 and w3 represent the weights of Ref2 and Ref3 respectively). In some embodiments, the third reference block Ref2 and the fourth reference block Ref3 can be the intermediate List0 prediction and the intermediate List1 prediction of the video block respectively. The intermediate List0 prediction and the intermediate List1 prediction can be the List0 prediction and the List1 prediction of the video block based on the intermediate motion vector respectively. In some implementations, w2 and w3 may be equal to w0 and w1 respectively. Alternatively, w2 and w3 may have different values from w0 and w1 respectively.
[0090] The video processor may determine a temporary motion model between the current prediction of the video block and the matching target, and derive motion refinement for the intermediate motion vector based on the temporary motion model. For example, the temporary motion model can be used for the motion refinement calculation described later. In some implementations, before the bidirectional matching-based motion refinement process 706 is executed, the affine motion model of the video block may be a four-parameter affine motion model (having two CPMVs) or a six-parameter affine motion model (having three CPMVs). When bidirectional matching is utilized, the temporary motion model between the current prediction and the matching target may be linear or non-linear, and it can be represented by a two-parameter (linear), four-parameter (non-linear), or six-parameter (non-linear) motion model.
[0091] In some embodiments, the temporary motion model may have the same number of parameters as the affine motion model of the video block. For example, the temporary motion model is a 6-parameter motion model, and the affine motion model is also a 6-parameter affine motion model. In other embodiments, the temporary motion model is a 4-parameter motion model, and the affine motion model is also a 4-parameter affine motion model. Alternatively, the temporary motion model may have a different number of parameters from the affine motion model of the video block. For example, while the affine motion model of the video block is a 6-parameter affine motion model, the temporary motion model is a 2-parameter motion model or a 4-parameter motion model. In other embodiments, while the affine motion model of the video block is a 4-parameter affine motion model, the temporary motion model is a 2-parameter motion model or a 6-parameter motion model.
[0092] For example, the affine motion model may be a 6-parameter affine motion model with three control points {(v 0x ,v 0y ), (v 1x ,v 1y ), (v 2x ,v 2y )} having three CPMVs. The motion refinement of the intermediate motion vector (e.g., the motion refinement of the three CPMVs) can be represented by {(dv 0x ,dv 0y ), (dv 1x ,dv 1y ), (dv 2x ,dv 2y )}. The matching target luminance signal can be represented as l(i,j) and is associated with the matching target. The predicted luminance signal can be represented as l' k (i,j) and is associated with the current prediction of the video block. The spatial gradients g x (i,j) and g y (i,j) can be derived by applying the Sobel filter to the predicted signal l' k (i,j) in the horizontal and vertical directions, respectively. Using a 6-parameter temporary motion model, the motion refinement for each CPMV can be derived as follows.
Number
[0093] In the above formula (3), (dv x (x,y), dv y (x,y)) represents the delta motion refinement of CPMV, a and b represent the delta translation parameters, c and d represent the horizontal delta zoom and rotation parameters, and e and f represent the vertical delta zoom and rotation parameters.
[0094] The coordinates of the upper left, upper right, and lower left control points {(v 0x , v 0y ), (v 1x , v 1y ), (v 2x , v 2y )} are (0,0), (w0), and (0,h) respectively. w and h represent the width and height of the video block. Based on the above formula (3), the motion refinements of the three CPMVs at the three control points can be derived from the following formulas (4) to (6) for their respective coordinates.
Number
Number
Number
[0095] Based on the optical flow equation, the relationship between the change in luminance and the spatial gradient and temporal motion can be formulated as the following formula (7).
Number
[0096] dv x (i,j) and dv ySubstituting (i,j) into Equation (3), Equation (8) for the parameter set (a,b,c,d,e,f) is obtained as follows:
Equation
[0097] Since all samples within the video block satisfy Equation (8), the parameter set (a,b,c,d,e,f) of Equation (8) can be solved by the least squares error method. Then, the motion refinement at the three control points {(v 0x , v 0y ), (v 1x , v 1y ), (v 2x , v 2y )} can be solved from Equations (4) to (6) and can be rounded to a specific accuracy (e.g., 1 / 16 pel). Using the above calculation process iteratively, the CPMV at the three control points is refined until convergence when the parameter set (a,b,c,d,e,f) is all zero or the total number of iterations satisfies a predetermined upper limit.
[0098] In another embodiment, the temporary motion model may be a four-parameter motion model. For the motion refinement of each CPMV, the four-parameter motion model can be expressed using the following Equation (9).
Equation
[0099] The coordinates of the upper left and upper right control points {(v 0x , v 0y ), (v 1x , v 1y )} can be (0,0) and (w0) respectively. Based on the above Equation (9), the delta motion refinement of the CPMV at the two control points can be derived as the following Equations (10) and (11) for their respective coordinates.
Equation
[0100] dv in Equation (7) x (i, j) and dv y Substituting (i, j) into Equation (9), Equation (12) for the parameter set (a, b, c, d) is obtained as follows: [Mathematics]
[0101] Similar to the above Equation (8), the parameter set (a, b, c, d) of Equation (12) can be solved by the least squares method by considering all samples in the video block.
[0102] In yet another embodiment, the temporary motion model may be a two-parameter motion model. For the motion refinement of each CPMV, c = d = e = f = 0 (e.g., according to the above Equation (3)). And the two-parameter temporary motion model can be expressed as follows. [Mathematics]
[0103] As shown in the above Equation (13), the motion refinement (e.g., delta motion refinement at any control point) is the same for any CPMV. dv in Equation (7) x (i, j) and dv y Substituting (i, j) into Equation (13), Equation (14) for the parameter set (a, b) is obtained as follows: [Mathematics]
[0104] Similar to the above Equation (8), the parameter set (a, b) of Equation (14) can be solved by the least squares method by considering all samples in the video block.
[0105] After obtaining motion refinement according to the above formula (3), (9), or (13), the video processor can update the intermediate motion vector using the derived motion refinement and obtain a refined motion vector 714 based on the affine motion model of the video block. For example, each CPMV can be updated using the following formula.
Number
[0106] In the above formula (15),
Number
Number
[0107] For example, when the temporary motion model between the current prediction and the matching target is a 2-parameter motion model, the motion refinement can be derived using the above formula (13). That is, for each of the CPMVs, dv x (x,y)=a and dv y(x, y) = b. Further, according to the above formula (15), when the affine motion model of the video block is a 6-parameter affine motion model, three refined CPMVs at three control points with coordinates (0, 0), (w0), and (0, h) of the 6-parameter affine motion model can be derived as follows.
Number
Number
Number
[0108] In another embodiment, when the temporary motion model between the current prediction and the matching target is a 4-parameter motion model, motion refinement can be derived using the above formula (9). That is, for each of the CPMVs, dv x (x, y) = c * x - d * y + a and dv y (x, y) = d * x + c * y + b. Further, according to the above formula (15), when the affine motion model of the video block is a 6-parameter affine motion model, three refined CPMVs at three control points with coordinates (0, 0), (w0), and (0, h) of the 6-parameter affine motion model can be derived as follows.
Number
Number
Number
[0109] Furthermore, in another embodiment, when the temporary motion model between the current prediction and the matching target is a 6-parameter motion model, motion refinement can be derived using the above formula (3). That is, dvx (x, y) = c * x - d * y + a and dv y (x, y) = e * x + f * y + b. Further, according to the above formula (15), when the affine motion model of the video block is a six-parameter affine motion model, three refined CPMVs at three control points with coordinates (0, 0), (w0), and (0, h) of the six-parameter affine motion model can be derived as follows.
Number
Number
Number
[0110] In a second exemplary process where search-based derivation is applied to derive motion refinement, the video processor can repeatedly apply an incremental change (e.g., +1 or -1) to the intermediate motion vector of each control point in the horizontal and / or vertical directions. The corresponding change in the intermediate motion vector leading to a smaller matching cost is retained for each control point until a refined motion vector is obtained, and can be set as a new starting point for the next search.
[0111] For example, the intermediate motion vector may be bidirectional and include a first motion vector for List0 (e.g., called the L0 motion vector) and a second motion vector for List1 (e.g., called the L1 motion vector). Progressive refinement of the L0 and L1 motion vectors may be performed separately by fixing the matching target and iteratively updating the current prediction of the video block using the updated L0 and / or L1 motion vectors. To reduce the complexity of the process, refinement can be performed simultaneously for both the L0 and L1 motion vectors by using the same amount of refinement in the reverse direction for the L0 and L1 motion vectors. For example, the updated L0 and L1 motion vectors can be determined using the following equation (25).
Number
[0112] In the above equation (25), v 0 and v 1 represent the motion vectors of L0 and L1 respectively, v 0 ' and v 1 ' represent the updated motion vectors of L0 and L1 respectively, Δ and -Δ represent the motion refinements applied to List0 and List1 in the reverse direction respectively, and k represents a scaling factor that can be used to consider the temporal distance. For example, k may be determined based on the ratio of the first temporal distance between the video frame and the first reference frame and the second temporal distance between the video frame and the second reference frame.
[0113] In some implementation forms, in order to repeatedly update the intermediate motion vectors of each control point, the video processor can first generate a first corrected motion vector based on the intermediate motion vectors within a predetermined search range and the changes in the first motion vectors. For example, the first corrected motion vector can be made equal to the sum of the intermediate motion vector and the change in the first motion vector, and the change in the first motion vector can be the incremental change in the intermediate motion vector. The video processor can determine whether to assign the change in the first motion vector as motion refinement for the intermediate motion vector based on (a) the matching cost associated with the intermediate motion vector and (b) the current matching cost associated with the first corrected motion vector.
[0114] For example, the video processor can determine the prediction of a video block based on the intermediate motion vector and determine the matching cost associated with the intermediate motion vector based on the matching target and the prediction of the video block. The prediction of the video block can be a weighted combination of the List0 prediction and the List1 prediction of the video block based on the intermediate motion vector. The matching cost can be determined based on any matching cost function disclosed herein. Similarly, the video processor can determine the current prediction of the video block based on the first corrected motion vector and also determine the current matching cost associated with the first corrected motion vector based on the matching target and the current prediction of the video block. The current prediction of the video block can be a weighted combination of the List0 prediction and the List1 prediction of the video block based on the first corrected motion vector.
[0115] When the current matching cost associated with the first corrected motion vector is smaller than the matching cost associated with the intermediate motion vector, the video processor can derive the motion refinement as the change in the first motion vector. As a result, the intermediate motion vector can be updated to the first corrected motion vector (e.g., intermediate motion vector = intermediate motion vector + change in the first motion vector).
[0116] If the current matching cost associated with the first corrected motion vector is equal to or greater than the matching cost associated with the intermediate motion vector, the video processor can determine not to assign a change in the first motion vector as motion refinement. Instead, the video processor can generate a second corrected motion vector based on the intermediate motion vector and a change in a second motion vector within a predetermined search range (e.g., second corrected motion vector = intermediate motion vector + change in second motion vector). The video processor can determine whether to assign the change in the second motion vector as motion refinement based on the matching cost associated with the intermediate motion vector and another current matching cost associated with the second corrected motion vector. If the other current matching cost associated with the second corrected motion vector is less than the matching cost associated with the intermediate motion vector, the video processor can derive the motion refinement as the change in the second motion vector. And the intermediate motion vector can be updated to be the second corrected motion vector. If the other current matching cost associated with the second corrected motion vector is equal to or greater than the matching cost associated with the intermediate motion vector, the video processor can determine not to assign the change in the second motion vector as motion refinement.
[0117] By performing a similar operation, the video processor can repeatedly update the intermediate motion vector until a predetermined iteration stop condition is satisfied. For example, the predetermined iteration stop condition can be satisfied when a motion vector change available within a predetermined search range is detected and processed, or when the total number of iterations satisfies a predetermined upper limit. When the predetermined iteration stop condition is satisfied, the video processor can determine the refined motion vector 714 of the video block as the intermediate motion vector.
[0118] In a third exemplary process for determining motion refinement using a combination of a computation-based derivation and a search-based derivation, first, a computation-based derivation is used to quickly refine an intermediate motion vector, and then, according to the search-based derivation, further refinement can be provided to obtain a refined motion vector 714. Specifically, the video processor determines the motion refinement of the intermediate motion vector by a computation-based derivation and updates the intermediate motion vector based on the motion refinement determined by the computation-based derivation. Next, the video processor determines the motion refinement of the intermediate motion vector again by a search-based derivation and updates the intermediate motion vector again based on the motion refinement determined by the search-based derivation. As a result, a refined motion vector 714 for the video block is obtained.
[0119] After obtaining the refined motion vector 714 of the video block, the video processor may execute a sub-block motion vector refinement process 716 to generate a motion vector 718 for each sub-block within the video block. Specifically, the video processor can refine the motion vector of each sub-block within the video block by using the refined motion vector 714 of the video block as the starting point of the motion vector of the sub-block. The video processor can apply an affine motion model of the video block to refine the motion vector at the sub-block level. For example, the video processor can obtain a plurality of refined CPMVs for the video block through the above-described bidirectional matching-based motion refinement process 706, and then apply the above formula (1) or (2) according to whether the affine motion model is a four-parameter model or a six-parameter affine motion model, and derive the motion vector of each sub-block using the refined CPMV.
[0120] Exemplary application conditions for the bidirectional matching-based motion refinement process 706 that conform to the present disclosure are provided herein. Specifically, for bidirectional matching, a motion trajectory is assumed. However, if the motion trajectory is not linear, bidirectional matching may not be used to derive a reliable motion vector. For example, bidirectional matching does not function well with complex motions such as rotation, zoom, and warping. To derive a more reliable motion refinement, specific application conditions may be determined to limit the overuse of bidirectional matching.
[0121] In some implementations, the bidirectional matching-based motion refinement process 706 is applied only when two reference frames are on two different sides of the current video frame (e.g., when one reference frame is before the current video frame and the other reference frame is after the current video frame). In some implementations, when two reference frames are on the same side of the current video frame (e.g., when two reference frames are before the current video frame, or when two reference frames are after the current video frame), and the temporal distance between the two reference frames meets a predefined threshold (e.g., when the temporal distance is less than (or greater than) a predefined value), the bidirectional matching-based motion refinement process 706 can be applied. In some implementations, when two reference frames are on two different sides of the current video frame and the first temporal distance between one of the reference frames and the current video frame is the same as the second temporal distance between the other reference frame and the current video frame, the bidirectional matching-based motion refinement process 706 can be applied.
[0122] In accordance with the present disclosure, the bidirectional matching-based motion refinement process 706 can be applied to affine motion at the block level, while the sub-block motion vector refinement 716 can be applied to regular motion at the sub-block level. This is because only normal motion is included at the sub-block level, while affine motion (e.g., zoom in / zoom out, rotation, or perspective motion, etc.) can also be included at the block level. In some implementations, the regular motion can be equivalent to a two-parameter affine motion model.
[0123] FIG. 9 is a flowchart of an exemplary method 900 for motion refinement in a video according to some implementations of the present disclosure. Method 900 may be performed by a video processor associated with video encoder 20 or video decoder 30 and may include steps 902 through 906, as described below. Some steps may be optional in order to implement the disclosure provided herein. Additionally, some operations may be performed simultaneously or in an order different from that shown in FIG. 9.
[0124] In step 902, the video processor may determine an initial motion vector for a video block of a video frame from the video.
[0125] In step 903, the video processor may determine the object to be matched based on the weighted combination of a first reference block from a first reference frame in the video and a second reference block from a second reference frame in the video. For example, the video processor may determine a first reference block and a second reference block from the first reference frame and the second reference frame of the video respectively based on the first motion vector. The video processor may determine a first weight for the first reference block and a second weight for the second reference block respectively. The video processor may use the first and second weights to determine the weighted combination of the first reference block and the second reference block. The video processor may determine the object to be matched based on the weighted combination of the first and second reference blocks.
[0126] In some implementations, the first weight and the second weight may be the same as the corresponding weights derived on the encoder side for normal weighted bi-prediction. For example, normal weighted bi-prediction may have a weight for List0 prediction and a weight for List1 prediction. The first and second weights may be equal to the weight for List0 prediction and the weight for List1 prediction respectively. Alternatively, the first and second weights may have predetermined values. For example, each of the first and second weights may be 0.5. In another example, the first weight may be 0 and the second weight may be 1. Or, the first weight may be 1 and the second weight may be 0.
[0127] In step 904, the video processor may perform block-level bidirectional matching-based motion refinement processing to iteratively update the initial motion vector based on the object to be matched until a refined motion vector of the video block is obtained. For example, the video processor may use the initial motion vector to initialize an intermediate motion vector, determine the motion refinement of the intermediate motion vector based on the object to be matched, and update the intermediate motion vector based on the motion refinement. The video processor may determine whether a predetermined iteration stop condition is satisfied.
[0128] In response to the satisfaction of a predetermined iteration stop condition, the video processor may determine the intermediate motion vector as a refined motion vector. In response to the non-satisfaction of the predetermined iteration stop condition, the video processor may repeatedly determine motion refinement for the intermediate motion vector until the predetermined iteration stop condition is satisfied, and may continue to update the intermediate motion vector based on the motion refinement.
[0129] In some implementations, motion refinement can be determined via a computation-based derivation, a search-based derivation, or a combination of a computation-based derivation and a search-based derivation.
[0130] In step 906, the video processor can refine the motion vector of each sub-block within the video block by using the refined motion vector of the video block as the starting point of the motion vector of the sub-blocks. The video processor can apply an affine motion model of the video block to refine the motion vector at the sub-block level.
[0131] FIG. 10 is a flowchart of another exemplary method 1000 for motion refinement in video according to some implementations of the present disclosure. Method 1000 may be implemented by a video processor associated with video encoder 20 or video decoder 30 and may include steps 1002 through 1016, as described below. Some steps may be optional in order to implement the disclosure provided herein. Additionally, some operations may be performed simultaneously or in an order different from that shown in FIG. 10.
[0132] In step 1002, the video processor may determine an initial motion vector of a video block of a video frame from the video based on a merge list of the video block.
[0133] In step 1004, the video processor may determine a matching target from the first reference frame and the second reference frame of the video based on the initial motion vector.
[0134] In step 1006, the video processor may use the initial motion vector to initialize the intermediate motion vector of the video block.
[0135] In step 1008, the video processor may determine the motion refinement of the intermediate motion vector based on the matching target.
[0136] In step 1010, the video processor may update the intermediate motion vector based on the motion refinement.
[0137] In step 1012, the video processor may determine whether a predetermined iteration stop condition is satisfied. In response to the establishment of the predetermined iteration stop condition, method 1000 may proceed to step 1014. Otherwise, method 1000 may return to step 1008.
[0138] In step 1014, the video processor may determine the intermediate motion vector as the refined motion vector for the video block.
[0139] In step 1016, the video processor may generate a bitstream including a merge index for identifying the initial motion vector from the merge list, a first reference index for identifying the first reference frame, and a second reference index for identifying the second reference frame.
[0140] FIG. 11 shows an arithmetic environment 1110 coupled to a user interface 1150 according to some embodiments of the present disclosure. The arithmetic environment 1110 can be part of a data processing server. The arithmetic environment 1110 includes a processor 1120, a memory 1130, and an input / output (I / O) interface 1140.
[0141] Processor 1120 typically controls the overall operation of the computing environment 1110, such as operations related to display, data collection, data communication, and image processing. Processor 1120 may include one or more processors for executing instructions to perform all or some of the steps of the methods described above. Further, Processor 1120 may include one or more modules to facilitate the interaction between Processor 1120 and other components. Processor 1120 may be a central processing unit (CPU), a microprocessor, a single-chip machine, a graphical processing unit (GPU), or the like.
[0142] Memory 1130 is configured to store various types of data to support the operation of the computing environment 1110. Memory 1130 may include a predetermined software 1132. Examples of such data include instructions for any application or method operating in the computing environment 1110, video data sets, image data, and the like. Memory 1130 can be implemented by using any type of volatile or non-volatile memory device, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, optical disk, or a combination thereof.
[0143] The I / O interface 1140 provides an interface between the processor 1120 and peripheral interface modules such as a keyboard, click wheel, buttons, and the like. The buttons may include, but are not limited to, a home button, a scan start button, and a scan stop button. The I / O interface 1140 can be coupled to an encoder and a decoder.
[0144] In some implementation forms, in order to execute the above method, for example, a non-transitory computer-readable storage medium including a plurality of programs executable by a processor 1120 in a computing environment 1110 is also provided in a memory 1130. Alternatively, the non-transitory computer-readable storage medium may store, for example, a bitstream or a data stream including encoded video information (for example, video information including one or more syntax elements) generated by an encoder (for example, the video encoder 20 in FIG. 2) using the above-described encoding method used by a decoder (for example, the video decoder 30 in FIG. 3) when decoding video data. The non-transitory computer-readable storage medium may be, for example, a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, or the like.
[0145] In some implementation forms, a computing device including one or more processors (for example, processor 1120) and a non-transitory computer-readable storage medium or a memory 1130 storing a plurality of programs executable by the one or more processors is also provided, and the one or more processors are configured to execute the above method when executing the plurality of programs.
[0146] In some implementation forms, in order to execute the above method, for example, a computer program product including a plurality of programs executable by a processor 1120 in a computing environment 1110 is also provided in a memory 1130. For example, the computer program product may include a non-transitory computer-readable storage medium.
[0147] In some implementation forms, in order to execute the above method, the computing environment 1110 may be implemented using one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), FPGAs, GPUs, controllers, microcontrollers, microprocessors, or other electronic components.
[0148] The description of the present disclosure is presented for illustrative purposes and is not intended to be exhaustive or to limit the present disclosure. Many modifications, changes, and alternative implementations will be apparent to those skilled in the art who benefit from the teachings shown in the foregoing description and the related drawings.
[0149] Unless otherwise specified, the order of the steps of the methods according to the present disclosure is merely intended to be illustrative, and the steps of the methods according to the present disclosure are not limited to the specifically described order and may be changed according to actual conditions. Furthermore, at least one of the steps of the methods according to the present disclosure may be adjusted, combined, or deleted according to actual requirements.
[0150] The examples are selected and described to explain the principles of the disclosure and to enable those skilled in the art to understand the disclosure regarding various implementations and to best utilize the underlying principles and various implementations with various modifications for the specific uses intended. Therefore, it should be understood that the scope of the present disclosure is not limited to the specific examples of the disclosed implementations, and that modifications and other implementations are intended to be included within the scope of the present disclosure.
Claims
1. Determining, by one or more processors, an initial motion vector of a video block of a video frame from a video; Determining, by the one or more processors, a matching target based on a weighted combination of a first reference block from a first reference frame in the video and a second reference block from a second reference frame in the video; Performing, by the one or more processors, a bidirectional matching-based motion refinement process at a block level to iteratively update the initial motion vector based on the matching target until a refined motion vector of the video block is obtained; Refining, by the one or more processors, a motion vector of each sub-block within the video block by using the refined motion vector of the video block as a starting point of the motion vector of the sub-block, wherein refining the motion vector at the sub-block level is to apply an affine motion model of the video block to refine the motion vector; A video coding method for motion refinement in a video, comprising the above.
2. Determining the matching target further comprises: Determining a first weight for the first reference block and a second weight for the second reference block respectively; Determining a weighted combination of the first reference block and the second reference block by using the first weight and the second weight; The method according to claim 1, further comprising the above.
3. The method according to claim 2, wherein the first weight and the second weight are the same as corresponding weights derived on the encoder side for weighted bi-prediction, or the first weight and the second weight have predetermined values.
4. Performing the bidirectional matching-based motion refinement process further comprises: Using the initial motion vector to initialize an intermediate motion vector; Determining a motion refinement of the intermediate motion vector based on the matching target; Updating the intermediate motion vector based on the motion refinement; The method according to claim 1, further comprising the above.
5. Performing the bidirectional matching-based motion refinement process further comprises: Determining whether a predetermined iteration stop condition is satisfied; In response to the satisfaction of the predetermined iteration stop condition, determining the intermediate motion vector as the refined motion vector, or In response to the non-satisfaction of the predetermined iteration stop condition, repeatedly determining the motion refinement for the intermediate motion vector until the predetermined iteration stop condition is satisfied, and continuously updating the intermediate motion vector based on the motion refinement; The method according to claim 4, further comprising.
6. The method according to claim 5, wherein the motion refinement is determined by a calculation-based derivation, a search-based derivation, or a combination of the calculation-based derivation and the search-based derivation.
7. The motion refinement is determined by the calculation-based derivation, and determining the motion refinement of the intermediate motion vector includes: Determining a current prediction of the video block based on the intermediate motion vector; Determining a temporary motion model between the current prediction and the matching target, wherein the temporary motion model is used for calculating the motion refinement; Calculating the motion refinement of the intermediate motion vector based on the temporary motion model; The method according to claim 6, further comprising.
8. The method according to claim 7, wherein the predetermined iteration stop condition is satisfied when the intermediate motion vector converges or the total number of iterations satisfies a predetermined threshold.
9. The method according to claim 7, wherein the total number of parameters of the temporary motion model is equal to the total number of parameters of the affine motion model, or the total number of parameters of the temporary motion model is different from the total number of parameters of the affine motion model.
10. The motion refinement is determined by the search-based derivation, and determining the motion refinement of the intermediate motion vector includes: Generating a first corrected motion vector based on the intermediate motion vector within a predetermined search range and a change in the first motion vector; Determining whether to assign the change in the first motion vector as the motion refinement based on a matching cost associated with the intermediate motion vector and a current matching cost associated with the first corrected motion vector; The method according to claim 6, further comprising.
11. The current matching cost associated with the first corrected motion vector is: Determining a current prediction of the video block based on the first corrected motion vector, Determine the current matching cost associated with the first modified motion vector based on the matching target and the current prediction of the video block The method according to claim 10, determined thereby
12. Further comprising deriving the motion refinement as a change in the first motion vector such that the intermediate motion vector is updated to the first modified motion vector in response to the current matching cost associated with the first modified motion vector being less than the matching cost associated with the intermediate motion vector, or In response to the current matching cost associated with the first modified motion vector being equal to or greater than the matching cost associated with the intermediate motion vector, Do not assign the change in the first modified motion vector as the motion refinement, Generate a second modified motion vector based on the intermediate motion vector and a change in a second motion vector within the predetermined search range, Determine whether to assign the change in the second motion vector as the motion refinement based on the matching cost associated with the intermediate motion vector and another current matching cost associated with the second modified motion vector The method according to claim 10, further comprising this
13. The method according to claim 10, wherein the predetermined iteration stop condition is satisfied when a change in a motion vector available within the predetermined search range is detected and processed, or when the total number of iterations satisfies a predetermined threshold
14. The motion refinement is determined by a combination of the calculation-based derivation and the search-based derivation, and performing the bidirectional matching-based motion refinement process Determine the motion refinement of the intermediate motion vector by a calculation-based derivation based on the matching target, Update the intermediate motion vector based on the motion refinement determined by the calculation-based derivation, Determine the motion refinement of the intermediate motion vector again by a search-based derivation based on the matching target, Update the intermediate motion vector again based on the motion refinement determined by the search-based derivation The method according to claim 6, including this
15. The bidirectional matching-based motion refinement process is One of the first reference frame and the second reference frame is in front of the video frame, and the other of the first reference frame and the second reference frame is behind the video frame, or Both the first reference frame and the second reference frame are in front of or behind the video frame, and the temporal distance between the first reference frame and the second reference frame satisfies a predetermined threshold, The method according to claim 1, which is executed to obtain the refined motion vector when one of the conditions is satisfied.
16. A memory configured to store at least one video frame of a video, the video frame including at least one video block, and One or more processors for executing the method according to any one of claims 1 to 15 A video coding device for motion refinement in a video, comprising:
17. A method of storing a bitstream, comprising: generating a bitstream by executing the video coding method according to any one of claims 1 to 15; and storing the generated bitstream.
18. A method for receiving a bitstream to be decoded by the video coding device according to claim 16, or transmitting a bitstream generated by the video coding device.
Citation Information
Patent Citations
Image decoding device, image decoding method, and program
JP2021002725A
Method and device for video signal processing
US20200154127A1