Decoder-Side Motion Vector Refinement (DMVR) Inter-Frame Prediction Using a Shared Interpolation Filter and Reference Pixels
By using shared interpolation filters and reference data to perform inter-frame prediction and refine motion vectors in video decoding technology, the problem of large data volume and heavy processing burden during processing of high-quality video data in the prior art is solved, and the bit rate reduction and video quality improvement are achieved.
Patent Information
- Application Number
- CN202380026512.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-03-16
- Filing Date
- 2023-02-20
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-02-20
AI Technical Summary
When processing high-quality video data, existing video decoding technology faces the problem of large amount of data and heavy processing burden, which leads to excessive burden on communication networks and equipment.
Inter-prediction is performed using shared interpolation filters and reference data, and the prediction accuracy of video data blocks is improved by refining the motion vector and reducing the amount of data.
Through the improved inter prediction method, the bit rate of video data is reduced, the video quality is improved, and the burden on communication networks and devices is reduced.
Smart Images

Figure CN118844062B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to video coding (e.g., including encoding and / or decoding of video data). For example, aspects of the present disclosure relate to systems and techniques for performing inter prediction (e.g., decoder-side motion vector refinement (DMVR)) using a shared interpolation filter and reference data. Background Art
[0002] Many devices and systems allow video data to be processed and output for consumption. Digital video data includes a large amount of data to meet the needs of consumers and video providers. For example, consumers of video data expect the highest quality video, with high fidelity, high resolution, high frame rate, etc. As a result, the large amount of video data required to meet these needs places a burden on communication networks and devices that process and store video data.
[0003] Various video coding techniques can be used to compress video data. Video coding is performed according to one or more video coding standards. For example, video coding standards include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), MPEG-2 Part 2 coding (MPEG stands for Moving Picture Experts Group), and so on, as well as proprietary video codecs / formats, such as AOMedia Video 1 (AV1) developed by the Alliance for Open Media. Video coding typically utilizes prediction methods (e.g., inter prediction, intra prediction, etc.), which utilize the redundancy present in video images or sequences. The goal of video coding techniques is to compress video data into a form that uses a lower bitrate while avoiding or minimizing degradation of video quality. As evolving video services become available, there is a need for coding techniques with better coding efficiency. Summary of the Invention
[0004] In some examples, systems and techniques for improved inter prediction using a shared interpolation filter are described. According to at least one illustrative example, a method of processing video data is provided. The method includes: obtaining a reference data block for predicting a video data block; determining one or more refined motion vectors based on the reference data block using an inter prediction processing path; and performing inter prediction for the video data block using the inter prediction processing path, where the inter prediction is based on the reference data block and the one or more refined motion vectors.
[0005] In another example, there is provided an apparatus for processing video data, the apparatus including at least one memory (e.g., configured to store data such as virtual content data, one or more images, etc.) and at least one processor (e.g., implemented in circuitry) coupled to the at least one memory. The at least one processor is configured to and may: obtain a reference data block for predicting a video data block; determine one or more refined motion vectors based on the reference data block using an inter-frame prediction processing path; and perform inter-frame prediction for the video data block using the inter-frame prediction processing path, where the inter-frame prediction is based on the reference data block and the one or more refined motion vectors.
[0006] In another example, there is provided a non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: obtain a reference data block for predicting a video data block; determine one or more refined motion vectors based on the reference data block using an inter-frame prediction processing path; and perform inter-frame prediction for the video data block using the inter-frame prediction processing path, where the inter-frame prediction is based on the reference data block and the one or more refined motion vectors.
[0007] In another example, there is provided an apparatus for processing video data. The apparatus includes: means for obtaining a reference data block for predicting a video data block; means for determining one or more refined motion vectors based on the reference data block using an inter-frame prediction processing path; and means for performing inter-frame prediction for the video data block using the inter-frame prediction processing path, where the inter-frame prediction is based on the reference data block and the one or more refined motion vectors.
[0008] In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: determining 2-tap horizontal interpolation based on the reference data block using a first interpolation filter; and determining 2-tap vertical interpolation based on the reference data block using a second interpolation filter.
[0009] In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: determining a sum of absolute differences (SAD) based on the 2-tap horizontal interpolation and the 2-tap vertical interpolation; and generating one or more refined motion vectors using the SAD and one or more original motion vectors associated with the video data block.
[0010] In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: determining 8-tap horizontal interpolation based on the reference data block and a first refined motion vector using a first interpolation filter; determining 8-tap vertical interpolation based on the reference data block and a second refined motion vector using a second interpolation filter; and generating a plurality of inter-frame prediction pixels using the 8-tap horizontal interpolation and the 8-tap vertical interpolation.
[0011] In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: determining a weighted prediction using 8-tap horizontal interpolation and 8-tap vertical interpolation, wherein the weighted prediction is determined based on the sum of the 8-tap horizontal interpolation and the 8-tap vertical interpolation; and generating the plurality of inter-frame predicted pixels based on the weighted prediction.
[0012] In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: generating one or more refined motion vectors based on the sum of absolute differences (SAD) determined for 2-tap horizontal interpolation and 2-tap vertical interpolation; wherein the SAD is determined based on the sum of the negative value of the 2-tap vertical interpolation and the 2-tap horizontal interpolation.
[0013] In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: using arithmetic logic to determine the weighted prediction and to determine the SAD.
[0014] In some aspects, the inter-frame prediction processing path includes a first interpolation filter and a second interpolation filter. In some aspects, the first interpolation filter is a 2-tap 32x2 horizontal interpolation filter. In some aspects, the second interpolation filter is a 2-tap 32x2 vertical interpolation filter.
[0015] In some aspects, to determine one or more refined motion vectors, one or more of the methods, apparatuses, and computer-readable media described above further include using the inter-frame prediction processing path as a motion vector refinement path. In some aspects, to perform inter-frame prediction for a video data block, one or more of the methods, apparatuses, and computer-readable media described above further include using the inter-frame prediction processing path as a pixel prediction path.
[0016] In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: using the inter-frame prediction processing path as a motion vector refinement path by configuring the first interpolation filter to perform 2-tap 20x2 horizontal interpolation and configuring the second interpolation filter to perform 2-tap 20x2 vertical interpolation; and using the inter-frame prediction processing path as a pixel prediction path by configuring the first interpolation filter to perform 8-tap 8x2 horizontal interpolation and configuring the second interpolation filter to perform 8-tap 8x2 vertical interpolation.
[0017] In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: generating an encoded video bitstream including one or more pictures, at least one of the one or more pictures being based on inter-frame prediction performed for a video data block.
[0018] In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: sending the encoded video bitstream to a decoding device, where the encoded video bitstream is sent together with signaling information.
[0019] In some aspects, store the encoded video bitstream.
[0020] In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: obtaining one or more encoded pictures, where at least one of the one or more encoded pictures includes video data blocks; and decoding the video data blocks from at least one of the encoded pictures.
[0021] In some aspects, one or more of the methods, apparatuses, and computer-readable media described above further include: decoding the video data blocks from at least one of the encoded pictures includes reconstructing the video data blocks.
[0022] In some aspects, the apparatus is one of the following or a part of one of the following: a mobile device (e.g., a mobile phone or a so-called "smartphone" or other mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a vehicle (a car, a truck, etc. or a component or system of a car, a truck, etc.), a personal computer, a laptop computer, a server computer, a robotic device, or other devices. In some aspects, the apparatus includes radio detection and ranging (radar) for capturing radio frequency (RF) signals. In some aspects, the apparatus includes one or more light detection and ranging (LIDAR) sensors, radar sensors, or other light-based sensors for capturing light-based (e.g., optical frequency) signals. In some aspects, the apparatus includes one camera or multiple cameras for capturing one or more images. In some aspects, the apparatus further includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the above apparatus may include one or more sensors that can be used to determine the position of the apparatus, the state of the apparatus (e.g., temperature, humidity level, and / or other states), and / or for other purposes.
[0023] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood in reference to the appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.
[0024] The foregoing and other features and embodiments will become more apparent when referring to the following specification, claims, and drawings. Brief Description of the Drawings
[0025] The exemplary embodiments of the present application are described in detail below with reference to the following accompanying drawings:
[0026] Figure 1 is a block diagram illustrating an example of an encoding device and a decoding device according to some examples;
[0027] Figure 2 is a block diagram illustrating an example of an inter-pixel prediction data path with decoder-side motion vector refinement (DMVR) capability according to some examples;
[0028] Figure 3 is a block diagram illustrating an example of an improved inter-pixel prediction data path with DMVR capability according to some examples;
[0029] Figure 4 is a flowchart illustrating an example of a process for processing audio data according to some examples;
[0030] Figure 5 is a block diagram illustrating an example video encoding device according to some examples; and
[0031] Figure 6 is a block diagram illustrating an example video decoding device according to some examples. Detailed Description
[0032] Certain aspects and embodiments of the present disclosure are provided below. Some of these aspects and embodiments may be applied independently, and some of them may be applied in combination, which will be obvious to those skilled in the art. In the following description, specific details are set forth for the purpose of explanation to provide a thorough understanding of the embodiments of the present application. However, it will be obvious that the various embodiments may be practiced without these specific details. The accompanying drawings and description are not intended to be restrictive.
[0033] The following description only provides example embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. On the contrary, the following description of the exemplary embodiments will provide those skilled in the art with an enabling description for implementing the exemplary embodiments. It should be understood that various changes can be made to the functions and arrangements of the elements without departing from the essence and scope of the present application as set forth in the appended claims.
[0034] Digital video data can include a large amount of data, especially as the demand for high-quality video data continues to grow. For example, consumers of video data generally expect increasingly high-quality video, with high fidelity, high resolution, high frame rate, etc. However, the large amount of video data required to meet such demands may impose a significant burden on communication networks as well as devices for processing and storing video data.
[0035] Video decoding devices implement video compression techniques to efficiently encode and decode video data. Video compression techniques can include applying different prediction modes, including spatial prediction (e.g., intra-frame prediction or intra-prediction), temporal prediction (e.g., inter-frame prediction or inter-prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques for reducing or removing redundancy inherent in a video sequence. A video encoder can divide each picture of an original video sequence into a plurality of rectangular regions, which are referred to as video blocks or coding units (described in more detail below). These video blocks can be encoded using a specific prediction mode.
[0036] Video blocks can be divided into one or more groups of smaller blocks in one or more ways. Blocks can include coding tree blocks, prediction blocks, transform blocks, and / or other suitable blocks. Unless otherwise specified, a reference to "block" generally refers to such video blocks (e.g., coding tree blocks, coding blocks, prediction blocks, transform blocks, or other appropriate blocks or sub-blocks as would be understood by a person of ordinary skill in the art). Additionally, each of these blocks can also be interchangeably referred to herein as a "unit" (e.g., coding tree unit (CTU), coding unit, prediction unit (PU), coding unit (CU), transform unit (TU), etc.). In some cases, a unit can indicate a coding logic unit encoded in a bitstream, while a block can indicate a portion targeted by a process in a video frame buffer.
[0037] For an inter-frame prediction mode, a video encoder can search for a block similar to the block being encoded in a frame (or picture) located at another temporal position, which is referred to as a reference frame or reference picture. The video encoder can limit the search to a certain spatial displacement from the block to be encoded. A two-dimensional (2D) motion vector including a horizontal displacement component and a vertical displacement component can be used to locate the best match. For an intra-frame prediction mode, a video encoder can use spatial prediction techniques to form a prediction block based on data from previously encoded neighboring blocks within the same picture.
[0038] A video encoder can determine a prediction error. For example, the prediction can be determined as the difference between the pixel values in the block being encoded and a predicted block. The prediction error may also be referred to as a residual. The video encoder can also apply a transform to the prediction error (e.g., a discrete cosine transform (DCT) or other suitable transform) to generate transform coefficients. After the transform, the video encoder can quantize the transform coefficients. The quantized transform coefficients and motion vectors can be represented using syntax elements and, together with control information, form a decoded representation of the video sequence. In some instances, the video encoder can entropy code the quantized transform coefficients and / or syntax elements to further reduce the number of bits required for their representation.
[0039] After entropy decoding and dequantizing the received bitstream, the video decoder can use the syntax elements and control information discussed above to construct prediction data (e.g., a predicted block) for decoding the current frame. For example, the video decoder can add the predicted block and the compressed prediction error. The video decoder can determine the compressed prediction error by weighting the transform basis functions using the quantized coefficients. The difference between the reconstructed frame and the original frame is referred to as the reconstruction error.
[0040] Video coding can be performed according to a specific video coding standard. Examples of video coding standards include, but are not limited to, ITU-T H.261, ISO / IEC MPEG-1 Video, ITU-T H.262 or ISO / IEC MPEG-2 Video, ITU-T H.263, ISO / IEC MPEG-4 Video, Advanced Video Coding (AVC) or ITU-T H.264 (including its Scalable Video Coding (SVC) and Multi-View Video Coding (MVC) extensions), High Efficiency Video Coding (HEVC) or ITU-T H.265 (including its Range and Screen Content Coding, 3D Video Coding (3D-HEVC), Multi-View (MV-HEVC) and Scalable (SHVC) extensions), Versatile Video Coding (VVC) or ITU-T H.266 and its extensions, VP9, Alliance for Open Media (AOMedia) Video 1 (AV1), Essential Video Coding (EVC), and so on.
[0041] As noted above, a video encoder may divide each picture of an original video sequence into one or more smaller blocks or rectangular regions, which may then be encoded using, for example, inter-prediction (or inter-frame prediction) to remove the temporal redundancy inherent in the original video sequence. If a block is encoded in an inter-prediction mode, a prediction block is formed based on previously encoded and reconstructed blocks, which are available in both the video encoder and the video decoder for forming prediction references. For example, instead of directly encoding the pixel values of each current (e.g., currently encoded or currently decoded) block, the video encoder may perform inter-prediction by searching a previously encoded frame for a block similar to the current block. These previously encoded frames serve as reference frames. When a matching or similar block is found, the block may be encoded with a motion vector that points to the location of the matching block in the reference frame.
[0042] When the video encoder successfully finds a matching block that is similar but not identical to the current block, the video encoder may compute the difference between the matching block and the current block. These residual values may be used as a prediction error, which is transformed and provided to the video decoder. Based on the motion vector (e.g., pointing to the matching block in the reference frame) and the prediction error, the video decoder may perform inter-prediction to recover the original pixel values of the current block.
[0043] The process of determining the motion vector is referred to as motion estimation. As described above, the video encoder may perform inter-prediction by using motion estimation to determine the motion vector, where the motion vector may be used to encode the location of the matching block in the reference frame. Accurate motion vectors and / or motion information may improve the inter-prediction performance at the decoder side. More accurate motion information may also be associated with an increased bitrate or otherwise occupy a larger percentage of the bitstream between the video encoder and the video decoder. In the HEVC standard and the VVC standard, a merge mode is used to reduce the bitrate required for motion vector signaling. In some examples, the merge mode may generate inaccurate motion vectors, which may result in imprecise inter-prediction.
[0044] In some cases, the video decoder may perform inter-prediction to improve decoding performance. For example, the video decoder may perform decoder-side motion vector refinement (DMVR) to refine an initial motion vector (e.g., the video decoder may perform DMVR to refine an initial motion vector from the merge mode). The video decoder may perform DMVR to generate a new motion vector or refine the motion vector with increased precision relative to the precision of the initial motion vector from the merge mode.
[0045] In an illustrative example, a video decoder may perform DMVR based on analyzing one or more motion prediction candidates from spatially or temporally adjacent blocks. The result of DMVR is a refined motion vector having some offset from the original motion vector (e.g., the merged mode initial motion vector received at the video decoder). In some cases, the video decoder may perform DMVR based on a search performed around the original motion vector on L0 and L1 reference frames. The search may be based on computing a distortion between L0 candidate samples and L1 candidate samples. In some cases, the sum of absolute differences (SAD) may be used as the distortion metric. In some examples, the refined motion vector may be derived based on the offset having the lowest SAD. The offset may be an integer offset. In some cases, DMVR may be performed at a maximum integer offset (e.g., the maximum search integer offset equal to two).
[0046] In the current VVC standard, the DMVR search region is larger than the block region (e.g., larger than the predicted block region). For example, according to the VVC standard, a video decoder may perform DMVR by using a 2-tap interpolation filter to obtain predicted pixel data values for the DMVR search region. The 2-tap interpolation is performed in each prediction direction L0 and L1, and the predicted pixel values for the DMVR search region are generated using the original merged candidate motion vectors. Since the DMVR search region depends on pixel data predicted using the original motion vector, the refined motion vector prediction pipeline at the video decoder has a dependency on the pixel prediction pipeline.
[0047] In some examples of hardware implementations of DMVR pixel - to - pixel prediction, in addition to the pixel prediction path, a separate motion vector refinement path is provided. Such methods may be associated with increased latency due to performing separate pre - fetching of reference data for the motion vector refinement path and the pixel prediction path. Additional die area is required to support the two separate prediction paths. In some cases, out - of - order prediction may be performed in an effort to reduce the latency impact of multiple pre - fetches. A large amount of static random access memory (SRAM) may be required to support out - of - order prediction, which increases power consumption and die size.
[0048] As described in more detail herein, systems, apparatuses, methods, and computer-readable media for providing improved inter-frame prediction are described (collectively referred to as "systems and techniques"). For example, the systems and techniques may perform inter-frame prediction using shared reference data prefetching and a shared interpolation data path for DMVR SAD and pixel prediction. In some cases, the shared reference data prefetching may obtain (e.g., prefetch) reference data or pixel blocks that are utilized by the shared interpolation data path to perform decoder-side motion vector refinement (DMVR) and perform pixel prediction. For example, the shared reference data prefetching may be at least partially based on common or identical reference data that is prefetched by the shared interpolation data path and used to perform different processing tasks. The shared interpolation data path may be used to perform multiple (e.g., two or more) processing tasks, where the multiple processing tasks utilize some (or all) of the same hardware components and / or processing units. For example, the shared interpolation data path may use some or all of the same interpolation filters to perform DMVR and perform pixel prediction, where the shared interpolation data path performs DMVR and pixel prediction separately. For example, the shared interpolation data path may use the same hardware components and / or processing units to sequentially perform DMVR and pixel prediction. In some cases, the shared interpolation data path may first use the same or shared hardware components and / or processing units to perform DMVR and determine a refined motion vector; the refined motion vector may then be looped back to the input of the shared interpolation path and subsequently used to perform pixel prediction using the same or shared hardware components and / or processing units that were previously used to perform DMVR. In some examples, the shared interpolation data path may use a set of one or more interpolation filters to implement both the DMVR data path or configuration and the pixel prediction data path or configuration. For example, the set of one or more interpolation filters included in the shared interpolation data path may be adjusted or configured between implementing the DMVR data path and implementing the pixel prediction data path.
[0049] According to some aspects, systems and techniques can include using shared reference data prefetching, the size of which is based on the maximum possible offset between a refined motion vector and an original motion vector (e.g., from which the refined motion vector is generated). In some examples, a shared reference data prefetch can be obtained for a given block (e.g., PU, CU, CTU, etc.) for which inter-frame prediction will be performed. When the block (e.g., PU, CU, etc.) is a DMVR block (e.g., DMVR PU, etc.), decoder-side motion vector refinement can be first performed to generate one or more refined motion vectors. The refined motion vectors can be generated based on the shared reference data prefetch and one or more original motion vectors associated with the block (e.g., PU, CU, etc.). The refined motion vectors can then be used together with lower-layer block (e.g., PU, CU, etc.) data (e.g., including a portion of the shared reference data prefetch) to perform inter-pixel prediction for the block. A PU will be used as an example of a block herein. However, the systems and techniques described herein can be used for other blocks (e.g., CU, CTU, TU, etc.).
[0050] In some examples, a shared interpolation data path can be used to perform DMVR or otherwise generate refined motion vectors. The same shared interpolation data path can be used to perform inter-pixel prediction. For example, the shared interpolation data path can be configured to perform 2-tap interpolation and implement a motion vector refinement path. The shared interpolation data path can also be configured to perform 8-tap interpolation and implement an inter-pixel prediction path. When the PU is a DMVR PU, the shared interpolation data path can be used to perform motion vector refinement and inter-pixel prediction in a continuous manner. In some examples, the shared interpolation data path can generate inter-pixel prediction without performing unordered prediction and / or PU reordering. When the PU is a non-DMVR PU, the shared interpolation data path can be used to perform only inter-pixel prediction.
[0051] Other details regarding the systems and techniques will be described with respect to the drawings.
[0052] Figure 1FIG. 0 is a block diagram illustrating an example of a system 100 that includes an encoding device 104 and a decoding device 112. The encoding device 104 may be part of a source device, and the decoding device 112 may be part of a receiving device. The source device and / or the receiving device may include an electronic device, such as a mobile or fixed telephone handset (e.g., a smart phone, a cellular phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the source device and the receiving device may include one or more wireless transceivers for wireless communication. The decoding techniques described herein are applicable to video decoding in a variety of multimedia applications, including streaming video transmission (e.g., over the Internet), television broadcast or transmission, encoding digital video for storage on a data storage medium, decoding digital video stored on a data storage medium, or other applications. As used herein, the term decoding may refer to encoding and / or decoding. In some examples, the system 100 may support unidirectional or bidirectional video transmission to support applications such as video conferencing, video streaming, video playback, video broadcast, gaming, and / or video telephony.
[0053] The encoding device 104 (or encoder) may be used to encode video data using a video decoding standard, format, codec, or protocol to generate an encoded video bitstream. Examples of video decoding standards and formats, codecs include ITU-T H.261, ISO / IEC MPEG-1 Video, ITU-T H.262 or ISO / IEC MPEG-2 Video, ITU-T H.263, ISO / IEC MPEG-4 Video, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) (including its scalable video coding (SVC) and multi-view video coding (MVC) extensions), High Efficiency Video Coding (HEVC) or ITU-T H.265 and Versatile Video Coding (VVC) or ITU-T H.266. There are various extensions to HEVC for handling multi-layer video coding, including range and screen content coding extensions, 3D video coding (3D-HEVC) and multi-view extensions (MV-HEVC) and scalable extensions (SHVC). HEVC and its extensions have been developed by the Joint Collaborative Team on Video Coding (JCT-VC) as well as the Joint Collaborative Team on 3D Video Coding Extensions (JCT-3V) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG). VP9, AOMedia Video 1 (AV1) developed by the Alliance for Open Media (AOMedia), and Essential Video Coding (EVC) are other video decoding standards to which the techniques described herein may be applied.
[0054] VVC is the latest video coding standard developed by the Joint Video Exploration Team (JVET) of ITU-T and ISO / IEC to at least partially achieve high compression capabilities beyond HEVC for a wide range of applications. The VVC specification was completed in July 2020 and released by both ITU-T and ISO / IEC. The VVC specification specifies the standard bitstream and picture formats, high-level syntax (HLS) and decoding unit-level syntax, parsing process, decoding process, etc. VVC also specifies profile / tier / level (PTL) constraints, byte stream format, hypothetical reference decoder, and supplementary enhancement information (SEI) in annexes.
[0055] The systems and techniques described herein can be applied to any of the existing video codecs (e.g., VVC, HEVC, AVC, or other suitable existing video codecs), and / or can be effective coding tools for any video coding standard being developed and / or future video coding standards. For example, video codecs such as VVC, HEVC, AVC, and / or their extensions can be used to perform the examples described herein. However, the techniques and systems described herein can also be applicable to other coding standards, codecs, or formats, such as MPEG, JPEG (or other coding standards for still images), VP9, AV1, their extensions, or other suitable coding standards that are already available or not yet available or not yet developed. For example, in some examples, the encoding device 104 and / or the decoding device 112 can operate according to a proprietary video codec / format (such as AV1, an extension of AVI, and / or a successor to AV1 (e.g., AV2)) or other proprietary format or industry standard. Thus, although the techniques and systems described herein may be described with reference to a particular video coding standard, those of ordinary skill in the art will understand that the description should not be construed as being applicable only to that particular standard.
[0056] Refer to Figure 1 , the video source 102 can provide video data to the encoding device 104. The video source 102 can be part of the source device or can be part of a device other than the source device. The video source 102 can include a video capture device (e.g., a camera, a camera phone, a video phone, etc.), a video archive including stored video, a video server or content provider providing video data, a video feed interface receiving video from a video server or content provider, a computer graphics system for generating computer graphics video data, a combination of such sources, or any other suitable video source.
[0057] Video data from video source 102 can include one or more input pictures or frames. A picture or frame is a static image, which in some cases is part of a video. In some examples, the data from video source 102 can be a static image that is not part of a video. In HEVC, VVC, and other video coding specifications, a video sequence can include a series of pictures. A picture can include three sample arrays, which are represented as SL, SCb, and SCr. SL is a two-dimensional array of luminance samples, SCb is a two-dimensional array of Cb chrominance samples, and SCr is a two-dimensional array of Cr chrominance samples. Chrominance samples may also be referred to herein as "chroma" samples. A pixel can refer to all three components (luminance sample and chrominance samples) at a given position in an array of a picture. In other cases, a picture can be monochromatic and can include only an array of luminance samples, in which case the terms pixel and sample can be used interchangeably. Regarding the example techniques described herein that refer to individual samples for illustrative purposes, the same techniques can be applied to pixels (e.g., all three sample components at a given position in an array of a picture). Regarding the example techniques described herein that refer to pixels (e.g., all three sample components at a given position in an array of a picture) for illustrative purposes, the same techniques can be applied to individual samples.
[0058] The encoder engine 106 (or encoder) of the encoding device 104 encodes video data to generate an encoded video bitstream. In some examples, the encoded video bitstream (or "video bitstream" or "bitstream") is a series of one or more decoded video sequences. A decoded video sequence (CVS) includes a series of access units (AUs), which starts from an AU having a random access point picture in the base layer and having certain attributes, until the next AU having a random access point picture in the base layer and having certain attributes and does not include the next AU. For example, certain attributes of the random access point picture that starts the CVS may include a RASL flag (e.g., NoRaslOutputFlag) equal to 1. Otherwise, a random access point picture (having a RASL flag equal to 0) does not start a CVS. An access unit (AU) includes one or more decoded pictures and control information corresponding to the decoded pictures sharing the same output time. The decoded slices of a picture are encapsulated as data units at the bitstream level, and the data units are called network abstraction layer (NAL) units. For example, an HEVC video bitstream may include one or more CVSs, which include NAL units. Each NAL unit in the NAL units has a NAL unit header. In one example, the header is one byte for H.264 / AVC (except for multi-layer extensions) and two bytes for HEVC. The syntax elements in the NAL unit header take specified bits and are thus visible to all kinds of systems and transport layers (such as transport stream, real-time transport (RTP) protocol, file format, etc.).
[0059] There are two types of NAL units in the HEVC standard, including video coding layer (VCL) NAL units and non-VCL NAL units. VCL NAL units include a slice or a segment of decoded picture data (described below), and non-VCL NAL units include control information related to one or more decoded pictures. In some cases, NAL units may be referred to as packets. An HEVC AU includes a VCL NAL unit containing decoded picture data and a non-VCL NAL unit (if any) corresponding to the decoded picture data.
[0060] A NAL unit may include a bit sequence that forms a decoded representation of video data (e.g., an encoded video bitstream, a CVS of the bitstream, etc.), such as a decoded representation of a picture in the video. The encoder engine 106 generates a decoded representation of a picture by splitting each picture into a plurality of slices. A slice is independent of other slices such that the information in the slice can be decoded without relying on data from other slices within the same picture. A slice includes one or more slice segments, which include independent slice segments and (if present) one or more dependent slice segments that depend on a previous slice segment. The slice is split into coding tree blocks (CTBs) of luminance samples and chrominance samples. The CTB of luminance samples and one or more CTBs of chrominance samples, together with the syntax for the samples, are referred to as a coding tree unit (CTU). A CTU may also be referred to as a "tree block" or "largest coding unit" (LCU). The CTU is the basic processing unit for HEVC encoding. A CTU can be split into a plurality of coding units (CUs) of different sizes. A CU contains an array of luminance samples and an array of chrominance samples called coding blocks (CBs).
[0061] The luminance CB and the chrominance CB can be further split into prediction blocks (PBs). A PB is a block of samples of a luminance component or a chrominance component that uses the same motion parameters for inter-frame prediction or intra-block copy prediction (when available or enabled). The luminance PB and one or more chrominance PBs, together with the associated syntax, form a prediction unit (PU). For inter-frame prediction, a set of motion parameters (e.g., one or more motion vectors, reference indices, etc.) is signaled in the bitstream for each PU, and for inter-frame prediction of the luminance PB and one or more chrominance PBs. The motion parameters may also be referred to as motion information. A CB can also be split into one or more transform blocks (TBs). A TB represents a square block of samples of a color component to which a residual transform (e.g., in some cases, the same two-dimensional transform) is applied to decode a prediction residual signal. A transform unit (TU) represents the TBs of luminance samples and chrominance samples and the corresponding syntax elements.
[0062] The size of the CU corresponds to the size of the coding mode and can be square-shaped. For example, the size of the CU can be 8x8 samples, 16x16 samples, 32x32 samples, 64x64 samples, or any other suitable size up to the size of the corresponding CTU. The phrase "NxN" is used herein to refer to the pixel dimensions of a video block in the vertical and horizontal dimensions (e.g., 8 pixels x 8 pixels). The pixels in a block can be arranged in rows and columns. In some examples, the block may not have the same number of pixels in the horizontal direction as in the vertical direction. The syntax data associated with the CU can describe, for example, splitting the CU into one or more PUs. The splitting pattern may vary depending on whether the CU is coded in an intra prediction mode or an inter prediction mode. The PU can be split into a non-square shape. The syntax data associated with the CU can also describe, for example, splitting the CU into one or more TUs according to the CTU. The TU can be square or non-square shaped.
[0063] According to the HEVC standard, transform units (TUs) can be used to perform transforms. The TUs can vary for different CUs. The size of the TU can be set based on the size of the PU within a given CU. The TU can have the same size as the PU or be smaller than the PU. In some examples, a quadtree structure called a residual quadtree (RQT) can be used to subdivide the residual samples corresponding to the CU into smaller units. The leaf nodes of the RQT can correspond to the TUs. The pixel differences associated with the TUs can be transformed to produce transform coefficients. The transform coefficients can be quantized by the encoder engine 106.
[0064] Once the pictures of the video data are segmented into CUs, the encoder engine 106 uses prediction modes to predict each PU. The prediction unit or prediction block is subtracted from the original video data to obtain a residual (described below). For each CU, the prediction mode can be signaled in the bitstream using syntax data. The prediction mode can include intra prediction (or picture-in prediction) or inter prediction (or picture-to-picture prediction). Intra prediction exploits the correlation between spatially adjacent samples within a picture. For example, using intra prediction, each PU is predicted from adjacent image data in the same picture using, for example, DC prediction to find the average value for the PU, plane prediction to fit a planar surface to the PU, directional prediction to infer from adjacent data, or any other suitable prediction type. Inter prediction uses the temporal correlation between pictures to derive a motion-compensated prediction for a block of image samples. For example, using inter prediction, each PU is predicted from image data in one or more reference pictures (before or after the current picture in output order). For example, a decision can be made at the CU level as to whether to use picture-to-picture prediction or picture-in prediction to code a picture region.
[0065] The encoder engine 106 and the decoder engine 116 (described in more detail below) may be configured to operate according to VVC. According to VVC, a video decoder (such as the encoder engine 106 and / or the decoder engine 116) divides a picture into multiple coding tree units (CTUs) (where the CTB of the luminance samples and one or more CTBs of the chrominance samples together with the syntax for the samples are referred to as a CTU). The video decoder may divide the CTU according to a tree structure, such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure eliminates the concept of multiple partition types, such as the separation between the CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels, including a first level that is divided according to quadtree partitioning and a second level that is divided according to binary tree partitioning. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary tree correspond to coding units (CUs).
[0066] In the MTT partitioning structure, quadtree partitioning, binary tree partitioning, and one or more types of ternary tree partitioning may be used to partition a block. Ternary tree partitioning is a partitioning in which a block is split into three sub-blocks. In some examples, ternary tree partitioning divides a block into three sub-blocks without dividing the original block through the center. The partitioning types in the MTT (e.g., quadtree, binary tree, and ternary tree) may be symmetric or asymmetric.
[0067] When operating according to the AV1 codec, the encoding device 104 and the decoding device 112 may be configured to decode video data in units of blocks. In AV1, the largest decoding block that can be processed is called a superblock. In AV1, a superblock may be 128x128 luminance samples or 64x64 luminance samples. However, in subsequent video decoding formats (e.g., AV2), a superblock may be defined by a different (e.g., larger) luminance sample size. In some examples, a superblock is the top level of a block quadtree. The encoding device 104 may further divide a superblock into smaller decoding blocks. The encoding device 104 may divide the superblock and other decoding blocks into smaller blocks using square or non-square partitioning. Non-square blocks may include N / 2xN blocks, NxN / 2 blocks, N / 4xN blocks, and NxN / 4 blocks. The encoding device 104 and the decoding device 112 may perform separate prediction and transform processing on each decoding block.
[0068] AV1 also defines tiles of video data. A tile is a rectangular array of superblocks that can be decoded independently of other tiles. That is, the encoding device 104 and the decoding device 112 can encode and decode the decoded blocks within a tile respectively without using the video data from other tiles. However, the encoding device 104 and the decoding device 112 can perform filtering across tile boundaries. The size of the tiles can be uniform or non-uniform. Tile-based decoding can enable parallel processing and / or multithreading implemented by the encoder and decoder.
[0069] In some examples, the encoding device 104 and the decoding device 112 can use a single QTBT or MTT structure to represent each of the luminance and chrominance components, while in other examples, the video decoder can use two or more QTBT or MTT structures, such as one QTBT or MTT structure for the luminance component and another QTBT or MTT structure for the two chrominance components (or two QTBTs and / or MTTs for the respective chrominance components).
[0070] The encoding device 104 and the decoding device 112 can be configured to use the quadtree partitioning, QTBT partitioning, MTT partitioning, or other partitioning structures of each HEVC.
[0071] In some examples, a slice type is assigned to one or more slices of a picture. The slice types include I slices, P slices, and B slices. An I slice (intra-frame, independently decodable) is a slice of a picture that is decoded only by intra-frame prediction and can therefore be decoded independently because an I slice only requires intra-frame data to predict any prediction unit or prediction block of the slice. A P slice (unidirectional predicted frame) is a slice of a picture that can be decoded using intra-frame prediction and unidirectional inter-frame prediction. Each prediction unit or prediction block within a P slice is decoded using intra-frame prediction or inter-frame prediction. When inter-frame prediction is applied, the prediction unit or prediction block is predicted by only one reference picture, and thus the reference samples are only from one reference region of one frame. A B slice (bi-directionally predicted frame) is a slice of a picture that can be decoded using intra-frame prediction and inter-frame prediction (e.g., either bi-directional prediction or unidirectional prediction). The prediction units or prediction blocks of a B slice can be predicted bi-directionally from two reference pictures, where each picture contributes to one reference region and the sample sets of the two reference regions are weighted (e.g., with equal weights or with different weights) to generate the prediction signal of the bi-directional prediction block. As described above, the slices of a picture are decoded independently. In some cases, a picture can be decoded as only one slice.
[0072] As noted above, in - picture prediction exploits the correlation between spatially - adjacent samples within a picture. There are a variety of intra - prediction modes (also referred to as “intra - modes”). In some examples, the intra - prediction of a luminance block includes 35 modes, which include a planar mode, a DC mode, and 33 angular modes (e.g., diagonal intra - prediction modes and angular modes adjacent to diagonal intra - prediction modes). The encoding device 104 and / or the decoding device 112 can select, for each block, a prediction mode that minimizes the residual between the predicted block and the block to be encoded (e.g., based on the sum of absolute errors (SAE), sum of absolute differences (SAD), sum of absolute transform differences (SATD), or other measures of similarity). For example, the SAE can be calculated by obtaining the absolute difference between each pixel (or sample) in the block to be encoded and the corresponding pixel (or sample) in the prediction block for comparison. The differences of the pixels (or samples) are summed to produce a measure of block similarity, such as the L1 norm of the difference image, the Manhattan distance between two image blocks, or other calculations. Using SAE as an example, the SAE for each prediction using each intra - prediction mode in the intra - prediction modes indicates the magnitude of the prediction error. The intra - prediction mode with the best match to the actual current block is given by the intra - prediction mode that gives the minimum SAE.
[0073] Indexing for the 35 modes of intra - prediction is shown in Table 1 below. In other examples, more intra - modes can be defined, including prediction angles that may not yet be represented by the 33 angular modes. In other examples, the prediction angles associated with the angular modes can be different from those used in HEVC.
[0074]
[0075]
[0076] Table 1 - Specification of Intra Prediction Modes and Associated Names
[0077] Inter - picture prediction uses the temporal correlation between pictures to derive motion - compensated predictions for blocks of image samples. Using a translational motion model, the position of a block in a previously decoded picture (reference picture) is indicated by a motion vector (Δx, Δy), where Δx specifies the horizontal displacement of the reference block relative to the position of the current block, and Δy specifies the vertical displacement of the reference block relative to the position of the current block. In some cases, the motion vector (Δx, Δy) can have integer sampling accuracy (also known as integer accuracy), in which case the motion vector points to the integer - pixel grid (or integer - pixel sampling grid) of the reference frame. In some cases, the motion vector (Δx, Δy) can have fractional sampling accuracy (also known as fractional - pixel accuracy or non - integer accuracy) to more accurately capture the movement of underlying objects, not limited to the integer - pixel grid of the reference frame. The accuracy of the motion vector can be represented by the quantization level of the motion vector. For example, the quantization level can be integer accuracy (e.g., 1 pixel) or fractional - pixel accuracy (e.g., 1 / 4 pixel, 1 / 2 pixel, or other sub - pixel values). When the corresponding motion vector has fractional sampling accuracy, interpolation is applied to the reference picture to derive the prediction signal. For example, samples available at integer positions can be filtered (e.g., using one or more interpolation filters) to estimate the values at fractional positions. The previously decoded reference pictures are indicated by the reference index (refIdx) of the reference picture list. The motion vector and the reference index can be referred to as motion parameters. Two types of inter - picture prediction can be performed, including uni - directional prediction and bi - directional prediction.
[0078] In the case of inter - frame prediction using bi - directional prediction (also known as bi - directional inter - frame prediction), two sets of motion parameters (Δx 0 , y 0 , refIdx 0 and Δx 1 , y 1 , refIdx 1 ) are used to generate two motion - compensated predictions (from the same reference picture or possibly from different reference pictures). For example, in the case of bi - directional prediction, each predicted block uses two motion - compensated prediction signals, and B prediction units are generated. The two motion - compensated predictions are combined to obtain the final motion - compensated prediction. For example, the two motion - compensated predictions can be combined by averaging. In another example, weighted prediction can be used, in which case different weights can be applied to each motion - compensated prediction. The reference pictures that can be used in bi - directional prediction are stored in two separate lists, denoted as list 0 and list 1 respectively. A motion - estimation process can be used at the encoding device 104 to derive the motion parameters.
[0079] In the case of inter - frame prediction using uni - directional prediction (also known as uni - directional inter - frame prediction), a set of motion parameters (Δx 0 , y0 , refIdx 0 ) Generate motion-compensated predictions from reference pictures. For example, in the case of uni-directional prediction, each prediction block uses at most one motion-compensated prediction signal, and P prediction units are generated.
[0080] A PU may include data related to the prediction process (e.g., motion parameters or other suitable data). For example, when a PU is encoded using intra prediction, the PU may include data describing the intra prediction mode used for the PU. As another example, when a PU is encoded using inter prediction, the PU may include data defining the motion vector for the PU. The data defining the motion vector for the PU may describe, for example, the horizontal component (Δx) of the motion vector, the vertical component (Δy) of the motion vector, the resolution for the motion vector (e.g., integer precision, quarter-pixel precision, or eighth-pixel precision), the reference picture to which the motion vector points, the reference index, the list of reference pictures for the motion vector (e.g., list 0, list 1, or list C), or any combination thereof.
[0081] AV1 includes two general techniques for encoding and decoding decoding blocks of video data. These two general techniques are intra prediction (e.g., intra prediction or spatial prediction) and inter prediction (e.g., inter prediction or temporal prediction). In the context of AV1, when using an intra prediction mode to predict a block of the current frame of video data, the encoding device 104 and the decoding device 112 do not use video data from other frames of the video data. For most intra prediction modes, the video encoding device 104 encodes a block of the current frame based on the difference between the sample values in the current block and the predicted values generated from reference samples in the same frame. The video encoding device 104 determines the predicted values generated from the reference samples based on the intra prediction mode.
[0082] After performing prediction using intra prediction and / or inter prediction, the encoding device 104 may perform transformation and quantization. For example, after prediction, the encoder engine 106 may compute a residual value corresponding to the PU. The residual value may include the pixel difference between the current pixel block (PU) being decoded and the prediction block used to predict the current block (e.g., the predicted version of the current block). For example, after generating a prediction block (e.g., using inter prediction or intra prediction), the encoder engine 106 may generate a residual block by subtracting the prediction block generated by the prediction unit from the current block. The residual block includes a set of pixel differences that quantize the difference between the pixel values of the current block and the pixel values of the prediction block. In some examples, the residual block may be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of pixel values). In such examples, the residual block is a two-dimensional representation of pixel values.
[0083] A block transform is used to transform any residual data that may remain after performing prediction. The block transform can be based on a discrete cosine transform, a discrete sine transform, an integer transform, a wavelet transform, other suitable transform functions, or any combination thereof. In some cases, one or more block transforms (e.g., of size 32x32, 16x16, 8x8, 4x4, or other suitable size) can be applied to the residual data in each CU. In some examples, TUs can be used for the transform and quantization processes implemented by the encoder engine 106. A given CU having one or more PUs can also include one or more TUs. As described in further detail below, the residual values can be transformed into transform coefficients using the block transform, and the TUs can be used for quantization and scanning to produce serialized transform coefficients for entropy coding.
[0084] In some examples, after performing intra prediction or inter prediction decoding using the PUs of a CU, the encoder engine 106 can compute the residual data for the TUs of the CU. The PUs can include pixel data in the spatial domain (or pixel domain). The TUs can include coefficients in the transform domain after applying the block transform. As previously noted, the residual data can correspond to the pixel differences between the pixels in the unencoded picture and the predicted values corresponding to the PUs. The encoder engine 106 can form TUs that include the residual data for the CU and can transform the TUs to produce the transform coefficients for the CU.
[0085] The encoder engine 106 can perform quantization of the transform coefficients. Quantization provides further compression by quantizing the transform coefficients to reduce the amount of data used to represent the coefficients. For example, quantization can reduce the bit depth associated with some or all of the coefficients. In one example, a coefficient having a value of n bits can be rounded down to a value of m bits during quantization, where n is greater than m.
[0086] Once quantization has been performed, the decoded video bitstream includes the quantized transform coefficients, prediction information (e.g., prediction mode, motion vectors, block vectors, etc.), segmentation information, and any other suitable data (such as other syntax data). Different elements of the decoded video bitstream can be entropy encoded by the encoder engine 106. In some examples, the encoder engine 106 can scan the quantized transform coefficients using a predefined scan order to produce a serialized vector that can be entropy encoded. In some examples, the encoder engine 106 can perform adaptive scanning. After scanning the quantized transform coefficients to form a vector (e.g., a one-dimensional vector), the encoder engine 106 can entropy encode the vector. For example, the encoder engine 106 can use context-adaptive variable length decoding, context-adaptive binary arithmetic decoding, syntax-based context-adaptive binary arithmetic decoding, probability interval segmentation entropy decoding, or another suitable entropy coding technique.
[0087] The output terminal 110 of the encoding device 104 can transmit NAL units constituting the encoded video bitstream data to the decoding device 112 of the receiving device over the communication link 120. The input terminal 114 of the decoding device 112 can receive the NAL units. The communication link 120 can include a channel provided by a wireless network, a wired network, or a combination of a wired network and a wireless network. The wireless network can include any wireless interface or combination of wireless interfaces, and can include any suitable wireless network (e.g., the Internet or other wide area network, packet-based network, WiFi TM , radio frequency (RF), UWB, WiFi Direct, cellular, Long Term Evolution (LTE), WiMax TM , etc.). The wired network can include any wired interface (e.g., fiber optic, Ethernet, powerline Ethernet, coaxial cable-based Ethernet, digital subscriber line (DSL), etc.). The wired network and / or wireless network can be implemented using various equipment (such as base stations, routers, access points, bridges, gateways, switches, etc.). The encoded video bitstream data can be modulated according to a communication standard (such as a wireless communication protocol) and sent to the receiving device.
[0088] In some examples, the encoding device 104 may store the encoded video bitstream data in the storage device 108. The output terminal 110 may retrieve the encoded video bitstream data from the encoder engine 106 or from the storage device 108. The storage device 108 may include any of a variety of distributed or locally accessible data storage media. For example, the storage device 108 may include a hard disk drive, a storage disk, a flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing the encoded video data. The storage device 108 may also include a decoded picture buffer (DPB) for storing reference pictures used in inter-frame prediction. In a further example, the storage device 108 may correspond to a file server or another intermediate storage device that can store the encoded video generated by the source device. In such cases, the receiving device including the decoding device 112 may access the stored video data from the storage device via streaming or downloading. The file server may be any type of server capable of storing the encoded video data and sending the encoded video data to the receiving device. Example file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The receiving device may access the encoded video data through any standard data connection, including an Internet connection, and may include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device 108 may be streaming, download transmission, or a combination thereof.
[0089] The input terminal 114 of the decoding device 112 receives the encoded video bitstream data and may provide the video bitstream data to the decoder engine 116, or provide it to the storage device 118 for later use by the decoder engine 116. For example, the storage device 118 may include a DPB for storing reference pictures used in inter-frame prediction. The receiving device including the decoding device 112 may receive the encoded video data to be decoded via the storage device 108. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and sent to the receiving device. The communication medium for transmitting the encoded video data may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful for facilitating communication from the source device to the receiving device.
[0090] The decoder engine 116 can decode the encoded video bitstream data by entropy decoding (e.g., using an entropy decoder) and extracting elements of one or more encoded video sequences that make up the encoded video data. The decoder engine 116 can rescale the encoded video bitstream data and perform an inverse transform on it. The residual data is passed to the prediction stage of the decoder engine 116. The decoder engine 116 predicts pixel blocks (e.g., PUs). In some examples, the prediction is added to the output of the inverse transform (residual data).
[0091] The decoding device 112 can output the decoded video to a video destination device 122, which may include a display or other output device for presenting the decoded video data to a consumer of the content. In some aspects, the video destination device 122 can be part of a receiving device that includes the decoding device 112. In some aspects, the video destination device 122 can be part of a separate device that is different from the receiving device.
[0092] In some examples, the video encoding device 104 and / or the video decoding device 112 can be integrated with an audio encoding device and an audio decoding device, respectively. The video encoding device 104 and / or the video decoding device 112 can also include other hardware or software necessary to implement the encoding techniques described above, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. The video encoding device 104 and the video decoding device 112 can be integrated as part of a combined encoder / decoder (codec) in the respective devices. Examples of specific details of the encoding device 104 are described below with reference to FIG. 8. Examples of specific details of the decoding device 112 are described below with reference to FIG. 9.
[0093] In Figure 1The example system shown is an illustrative example that can be used herein. Techniques for processing video data using the techniques described herein can be performed by any digital video encoding and / or decoding device. Although generally the techniques of the present disclosure are performed by a video encoding device or a video decoding device, these techniques can also be performed by a combined video encoder-decoder (commonly referred to as a "codec"). In addition, the techniques of the present disclosure can also be performed by a video preprocessor. The source device and the receiving device are merely examples of such decoding devices, where the source device generates the decoded video data for transmission to the receiving device. In some examples, the source device and the receiving device can operate in a substantially symmetric manner such that each of these devices includes video encoding and decoding components. Thus, the example system can support unidirectional or bidirectional video transmission between video devices, e.g., for video streaming, video playback, video broadcasting, or video telephony.
[0094] Extensions to the HEVC standard include multi-view video coding extensions (referred to as MV-HEVC) and scalable video coding extensions (referred to as SHVC). The MV-HEVC and SHVC extensions share the concept of hierarchical coding, where different layers are included in the encoded video bitstream. Each layer in the decoded video sequence is addressed by a unique layer identifier (ID). The layer ID can be present in the header of the NAL unit to identify the layer associated with the NAL unit. In MV-HEVC, different layers can represent different views of the same scene in the video bitstream. In SHVC, different scalable layers are provided that represent the video bitstream at different spatial resolutions (or picture resolutions) or at different reconstruction fidelities. The scalable layers can include a base layer (where the layer ID = 0) and one or more enhancement layers (where the layer ID = 1, 2, … n). The base layer can conform to the profile of the first version of HEVC and represents the lowest available layer in the bitstream. Compared to the base layer, the enhancement layers have increased spatial resolution, temporal resolution, or frame rate and / or reconstruction fidelity (or quality). The enhancement layers are hierarchically organized and may (or may not) depend on the lower layers. In some examples, different layers can be decoded using a single standard codec (e.g., encoding all layers using HEVC, SHVC, or other coding standards). In some examples, different layers can be decoded using a multi-standard codec. For example, the base layer can be decoded using AVC, while the SHVC and / or MV-HEVC extensions to the HEVC standard can be used to decode one or more enhancement layers.
[0095] Typically, a layer includes a set of VCL NAL units and a corresponding set of non-VCL NAL units. A specific layer ID value is assigned to the NAL units. Layers can be hierarchical in the sense that a layer can be dependent on a lower layer. A layer set refers to a self - contained set of layers represented within a bitstream, meaning that the layers within a layer set can be dependent on other layers within that layer set during the decoding process, but not on any other layers to be decoded. Thus, the layers within a layer set can form an independent bitstream representing video content. The set of layers within a layer set can be obtained from another bitstream through the operation of a sub - bitstream extraction process. A layer set can correspond to the set of layers to be decoded when a decoder wants to operate according to certain parameters.
[0096] As previously described, the HEVC bitstream includes groups of NAL units, including VCL NAL units and non - VCL NAL units. The VCL NAL units include the decoded picture data that forms the decoded video bitstream. For example, the bit sequence forming the decoded video bitstream is present in the VCL NAL units. In addition to other information, the non - VCL NAL units can also contain parameter sets with high - level information related to the encoded video bitstream. For example, the parameter sets can include a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), and a Picture Parameter Set (PPS). Examples of the objectives of the parameter sets include bit - rate efficiency, error - tolerance, and providing a system - layer interface. Each slice references a single active PPS, SPS, and VPS to access the information that the decoding device 112 can use to decode the slice. An identifier (ID) can be decoded for each parameter set, including the VPS ID, SPS ID, and PPS ID. The SPS includes the SPS ID and the VPS ID. The PPS includes the PPS ID and the SPS ID. Each slice header includes the PPS ID. Using the ID, the valid parameter sets can be identified for a given slice.
[0097] The PPS includes information that applies to all slices in a given picture. In some examples, all slices in a picture reference the same PPS. Slices in different pictures may also reference the same PPS. The SPS includes information that applies to all pictures in the same decoded video sequence (CVS) or bitstream. As previously described, a decoded video sequence is a series of access units (AUs) that starts at a random access point picture (e.g., an instantaneous decoding reference (IDR) picture or a broken link access (BLA) picture, or other suitable random access point picture) in the base layer and having certain properties (described above), and continues until the next AU (or the end of the bitstream) in the base layer that has a random access point picture and has certain properties, and does not include that next AU. The information in the SPS may not change with pictures within the decoded video sequence. Pictures in the decoded video sequence may use the same SPS. The VPS includes information that applies to all layers within the decoded video sequence or bitstream. The VPS includes a syntax structure having syntax elements that apply to the entire decoded video sequence. In some embodiments, the VPS, SPS, or PPS may be sent in-band with the encoded bitstream. In some embodiments, the VPS, SPS, or PPS may be sent out-of-band in a separate transmission from the NAL unit containing the decoded video data.
[0098] The present disclosure may generally relate to "signaling" certain information, such as syntax elements. The term "signaling" generally may refer to the conveyance of the values of syntax elements and / or other data used to decode the encoded video data. For example, the video encoding device 104 may signal the value of a syntax element in the bitstream. Generally speaking, signaling refers to generating a value in the bitstream. As noted above, the video source 102 may transmit the bitstream to the video destination device 122 substantially in real time or not in real time (such as may occur when storing the syntax elements to the storage device 108 for later retrieval by the video destination device 122).
[0099] The video bitstream may also include Supplemental Enhancement Information (SEI) messages. For example, an SEI NAL unit may be part of the video bitstream. In some cases, the SEI message may contain information that is not required for the decoding process. For example, the information in the SEI message may not be necessary for the decoder to decode the video pictures of the bitstream, but the decoder may use this information to improve the display or processing of the picture (e.g., the decoded output). The information in the SEI message may be embedded metadata. In an illustrative example, the information in the SEI message may be used by decoder-side entities to improve the visibility of the content. In some instances, certain application standards may indicate the presence of such SEI messages in the bitstream, such that quality improvements can be brought to all devices compliant with the application standard (e.g., carrying frame-packing SEI messages for frame-compatible planar stereoscopic 3DTV video formats, where an SEI message is carried for each frame of the video; handling recovery point SEI messages; using pull-down scan rectangle SEI messages in DVB; and many other examples).
[0100] As noted above, a video encoder (e.g., encoding device 104) and / or a video decoder (e.g., decoding device 112) may perform decoder-side motion vector refinement (DMVR) based on analyzing one or more motion prediction candidates from spatially or temporally adjacent blocks. The result of DMVR is a refined motion vector having some offset from the original motion vector (e.g., the merge mode initial motion vector). In some cases, the video encoder and / or decoder may perform DMVR based on a search performed around the original motion vector on L0 and L1 reference frames. The search may include calculating the distortion between L0 candidate samples and L1 candidate samples, where the sum of absolute differences (SAD) is used as the distortion metric. In some examples, the refined motion vector may be derived based on identifying the integer offset with the lowest SAD. In some cases, the video encoder and / or video decoder may perform DMVR using a maximum integer offset of 2. In the current VVC standard, the DMVR search region is larger than the block region (e.g., larger than the prediction block region). For example, according to the VVC standard, the video decoder may perform DMVR by using a 2-tap interpolation filter to obtain the predicted pixel data values of the DMVR search region. The 2-tap interpolation is performed in each prediction direction L0 and L1, and the predicted pixel values of the DMVR search region are generated using the original merge candidate motion vectors. Since the DMVR search region depends on the pixel data predicted using the original motion vector, the refined motion vector prediction pipeline at the video decoder has a dependency on the pixel prediction pipeline.
[0101] Figure 2Example diagram 200 illustrating an inter-pixel prediction data path with DMVR capabilities. An inter-pixel prediction data path with DMVR capabilities (also referred to as a "pipeline") may include a reference prefetch block 210 that receives an inter-frame prediction prediction unit (PU) command stream as input, where the inter-frame prediction prediction unit (PU) command stream is indicated as "inter-frame prediction PU_CMD[original]" in Figure 2 . The inter-frame prediction PU commands may each include an original motion vector associated with the PU. For example, the original motion vector may be obtained or received from a motion vector prediction block. In some cases, the original motion vector included in the inter-frame prediction PU commands may be a merged mode motion vector (e.g., generated by a video encoder). In some examples, the inter-frame prediction mode is bi-directional prediction in the L0 and L1 prediction directions. The inter-frame prediction PU commands may include a first motion vector MVL0 for the L0 prediction direction and a second motion vector MVL1 for the L1 prediction direction. Each prediction block (e.g., PU) may have a block size SbWxSbH, where SbW represents the sub-block width and SbH represents the sub-block height. According to the VVC standard, the prediction block (e.g., PU) size may be 16x16, 16x8, or 8x16. The following discussion refers to an example based on a 16x16 prediction unit, but it should be noted that one or more different block sizes may be utilized, as mentioned above.
[0102] The inter-frame prediction PU commands may include a combination of DMVR PUs and non-DMVR PUs. As will be explained in more depth below, DMVR may be performed to generate a refined motion vector for DMVR PUs prior to inter-pixel prediction (e.g., performing inter-pixel prediction using the refined motion vector), while non-DMVR PUs may perform inter-pixel prediction directly (e.g., performing inter-pixel prediction using the original motion vector).
[0103] As illustrated, the inter-pixel prediction pipeline 200 with DMVR capabilities includes a motion vector refinement path and a pixel prediction path. The motion vector refinement path (also referred to as the "DMVR SAD path") includes at least 2-tap interpolation filters 252 and 254, an SAD block 256, and a DMVR block 230. In some examples, the 2-tap interpolation filter 252 may be a 2-tap horizontal interpolation filter, and the 2-tap interpolation filter 254 may be a 2-tap vertical interpolation filter.
[0104] The pixel prediction path includes at least 8-tap interpolation filters 242 and 244, a weighted prediction block 246, and a reconstruction interface 270. In some examples, the 8-tap interpolation filter 242 can be an 8-tap horizontal interpolation filter, and the 8-tap interpolation filter 244 can be an 8-tap vertical interpolation filter. The DMVR PU can first be provided to the motion vector refinement path and then loop back to the input of the pixel prediction path. The non-DMVR PU can be directly provided to the pixel prediction path without going through the motion vector refinement path.
[0105] As previously mentioned, the reference prefetch block 210 receives the inter-frame prediction PU command stream as input. In response to receiving an inter-frame prediction PU command, the reference prefetch block 210 can communicate with a direct memory access (DMA) controller or interface 212 to perform or request reference data prefetching for the inter-frame prediction PU command. In some examples, the size of the reference data prefetch can be based on the type of the inter-frame prediction PU command (e.g., DMVR, non-DMVR) and / or the currently processed component of the inter-frame prediction PU command (e.g., luminance, chrominance). In an illustrative example, the reference prefetch block 210 first obtains a 16x16 luminance DMVR prediction unit, followed by the 16x16 chrominance component of the prediction unit, followed by a sequence of non-DMVR E prediction units and their corresponding chrominance prediction units. This sequence of prediction units will be discussed in turn below.
[0106] Starting with the example of a 16x16 luminance DMVR PU, the reference prefetch block 210 initiates a prefetch request or a prefetch communication with the DMA 212 (e.g., indicated as "prefetch" in Figure 2 ), thereby requesting 24x24 reference data blocks for the L0 direction and for the L1 direction. In some cases, the reference prefetch block 210 requests 24x24 reference data blocks because the 2-tap interpolation filters 252 and 254 operate over an extended range of ±2 with respect to the 16x16 PU input (e.g., resulting in 20x20) and the SAD block 256 introduces a further extension of ±2 (e.g., resulting in 24x24). After sending the prefetch request to the DMA 212, the reference prefetch block 210 can provide an indication or identifier of the specific PU to which the prefetch request was sent to the DMA 212 to the data path scheduler 220. Then, the reference prefetch block 210 can continue to pipeline the prefetch request to the DMA 212 for incoming subsequent prediction units after the 16x16 luminance DMVR PU (e.g., in this example, which is chrominance or non-DMVR).
[0107] The DMA 212 can obtain the requested data based on a prefetch request received from the reference prefetch block 210 and can store the obtained data in a cache and / or a first-in-first-out (FIFO) buffer. Since the DMA 212 can be associated with a relatively long latency between receiving the prefetch request and subsequently delivering the obtained item or data to the cache / FIFO, the reference prefetch block 210 can time the sending of the prefetch request such that the DMA 212 loads the requested data into the cache / FIFO before the data path scheduler 220 sends a reference data read request to the DMA 212. In some examples, the DMA 212 can send an indication that prefetching has been completed for a given block or PU to the data path scheduler 220.
[0108] For a 16x16 luma DMVR prediction unit, the data path scheduler 220 reads 24x24 reference data for SAD calculation for DMVR from the DMA 212, as previously discussed. The reference data read operation is indicated as "reference read" in Figure 2 Then, the data path scheduler 220 can schedule the reference data into the motion vector refinement path. For example, in the motion vector refinement path, the 24x24 reference data for the luma DMVR prediction unit is provided to a first interpolation filter 252 (shown here as a 2-tap horizontal 20x2 interpolation filter) and a second interpolation filter 254 (shown here as a 2-tap vertical 20x2 interpolation filter).
[0109] Using the first 2-tap interpolation filter 252 and the second 2-tap interpolation filter 254, the motion vector refinement path generates a 20x20 block for the 16x16 luma DMVR prediction unit. For a given input of an SbWxSbH luma DMVR prediction unit, the motion vector refinement path can apply the interpolation filters 252 and 254 to generate an interpolated block of size (SbW + 4)x(SbH + 4). For example, recalling that the VVC specification provides 16x16, 16x8, or 8x16 luma size blocks for DMVR inter prediction, the interpolation filters 252 and 254 can additionally be used to produce interpolated blocks of size 20x12 (e.g., for a 16x8 luma DMVR input) or 12x20 (e.g., for an 8x16 luma DMVR input). In some examples, the motion vector refinement path uses the first interpolation filter 252 and the second interpolation filter 254 to generate a first 20x20 array DPred0 for the L0 prediction direction and a second 20x20 array DPred1 for the L1 prediction direction.
[0110] The SAD block 256 generates a SAD (e.g., sum of absolute differences) array based on receiving DPred0 and DPred1 as inputs. For example, the SAD block 256 may generate a DSAD of size 24x24.
[0111] Based on the DSAD generated by the SAD block 256, the DMVR block 230 refines the original motion vectors MVL0 and MVL1 associated with the 16x16 luma DMVR prediction unit provided as an input to the motion vector refinement path. In some examples, the original motion vectors MVL0 and MVL1 are refined to generate new motion vectors DMVL0 and DMVL1, which in some cases may be generated to have an offset of no more than ±2 pixels in the x and y directions respectively relative to the original motion vectors MVL0 and MVL1.
[0112] The output of the DMVR block 230 is the output of the motion vector refinement path (e.g., two refined motion vectors DMVL0 and DMVL1), which is indicated as "Inter - frame prediction PU_CMD[refined MV]" in Figure 2 As illustrated, the refined motion vectors are provided to the reference pre - fetch block 210. In some examples, the original 16x16 luma DMVR prediction unit can then be looped back to the reference pre - fetch block 210 and combined with the refined motion vectors generated for the original 16x16 luma DMVR prediction unit. Then, the original 16x16 luma DMVR block and its refined motion vectors can be treated as a normal (e.g., non - DMVR) prediction unit using the refined motion vectors instead of the original motion vectors, and are scheduled into the pixel prediction paths 242, 244, 246, 270.
[0113] The pixel prediction path can be used to perform inter - frame prediction for an input PU. Inter - frame prediction can be performed in the same way for a DMVR PU (e.g., having refined motion vectors generated by the refined motion vector path) and for a non - DMVR PU (e.g., still having its original motion vector). The pixel prediction path may include a first 8 - tap interpolation filter 242 (shown here as an 8 - tap horizontal 8x2 interpolation filter) and a second 8 - tap interpolation filter 244 (shown here as an 8 - tap vertical 8x2 interpolation filter). As previously mentioned, the pixel prediction path can operate on the original input size SbWxSbH of each PU. The pixel prediction path can apply the first 8 - tap interpolation filter 242 and the second 8 - tap interpolation filter 244 to generate PredL0 and PredL1 blocks of size SbWxSbH for prediction directions L0 and L1 respectively.
[0114] In the context of the previous example of the 16x16 luminance DMVR block that is looped back to the reference prefetch block 210 and input to the pixel prediction path in combination with the refined motion vectors DMVL0 and DMVL1 generated by the motion vector refinement path, the pixel prediction path uses the refined motion vectors to generate the PredL0 and PredL1 blocks. When the pixel prediction path receives a non-DMVR block as input, the PredL0 and PredL1 blocks are generated using the original motion vectors MVL0 and MVL1 (e.g., using two 8-tap interpolation filters 242 and 244).
[0115] Then, the final prediction block for the inter-frame prediction PU is determined as the combination of PredL0 and PredL1 and is a block of size SbWxSbH. For example, a 16x16 luminance DVMR block will produce a final prediction luminance block of size 16x16. The corresponding 16x16 chrominance block associated with this luminance block will produce a final prediction chrominance block also of size 16x16, but the chrominance block can be processed using only the pixel prediction path (e.g., the chrominance block can be directly provided to the pixel prediction path even if its corresponding luminance block is provided to the motion vector refinement path). In some examples, the prediction of the chrominance components PredCb and PredCr can be obtained in the same manner as described above for the luminance component using the pixel prediction path but using 4-tap interpolation filters instead of the 8-tap interpolation filters 242 and 244, and using the original motion vectors MVL0 and MVL1 instead of the refined motion vectors DMVL0 and DMVL1.
[0116] In one illustrative example, the final prediction block for the inter-frame prediction PU can be determined as the combination of PredL0 and PredL1 using a weighted prediction block 246 (shown here as an 8x2 weighted prediction). The weighted prediction 246 can perform a weighted prediction that is calculated as the weighted average of L0 and L1.
[0117] As Figure 2As illustrated, the pixel prediction path further includes a reconstruction interface 270 that receives the weighted prediction from the weighted prediction block 246 and communicates with the PU reordering SRAM 260. The PU reordering SRAM 260 is provided at the inter-frame prediction pixel reconstruction interface to accommodate the out-of-order processing of the PUs following the luminance DMVR PU. Recalling the example currently being discussed with reference to a sequence in which the reference prefetch block 210 first fetches a 16x16 luminance DMVR block, followed by the corresponding 16x16 chrominance (non-DMVR) block and / or additional non-DMVR blocks, in some examples, out-of-order pixel prediction is performed. For example, out-of-order pixel prediction can occur when a DMVR block (e.g., a 16x16 luminance DMVR block) is provided to the motion vector refinement path; when motion vector refinement is performed, the reference prefetch block 210 and / or the data path scheduler 220 can obtain the next non-DMVR block (e.g., in the input stream of the inter-frame prediction PU command) and provide the next non-DMVR block to the pixel prediction path. The pixel prediction path then processes the next non-DMVR block in parallel with the current DMVR luminance block.
[0118] In some examples, using the motion vector refinement path to calculate the refined motion vector for the luminance or other DMVR blocks can introduce out-of-order pixel prediction. The aforementioned PU reordering SRAM 260 can be used to reverse the out-of-order pixel prediction, and the aforementioned PU reordering SRAM can store or otherwise buffer any pixel prediction results output in an out-of-order manner by the pixel prediction path. For example, the PU reordering SRAM 260 can store or buffer the out-of-order pixel prediction generated for the 16x16 chrominance non-DMVR block while the corresponding 16x16 luminance DMVR block is being processed through the pixel prediction pipeline. The reconstruction interface 270 can then reorder or resequence the out-of-order pixel prediction stored in the PU reordering SRAM 260 to thereby generate a properly ordered output of the inter-frame prediction pixel values (e.g., the output of the inter-frame prediction pixels). In some cases, the size of the PU reordering SRAM 260 can depend on the prefetch latency processed by the refined motion vector prediction unit. For example, this can be the prefetch latency associated with cycling the DMVR PU and its calculated refined motion vector back to the reference prefetch block 210, and / or the prefetch latency associated with the reference prefetch block 210 performing an additional (e.g., second) reference data prefetch from the DMA 212. The second reference data prefetched from the DMA 212 can correspond to the reference data required to process the DMVR PU through the pixel prediction path; the first reference data prefetched from the DMA 212 can correspond to the reference data required to process the DMVR PU through the motion vector refinement path previously.
[0119] Figure 3Example diagram 300 illustrates an improved DMVR inter-pixel prediction data path. As illustrated, the improved DMVR inter-pixel prediction data path includes at least a first shared interpolation filter 382 and a second shared interpolation filter 384. In one illustrative example, the first shared interpolation filter 382 and the second shared interpolation filter 384 may be provided as a shared 2-tap 32x2 horizontal interpolation filter and a shared 2-tap 32x2 vertical interpolation filter, respectively. In some examples, the first shared interpolation filter 382 and the second shared interpolation filter 384 may be included in a shared DMVR-SAD and pixel prediction path (also referred to as a shared inter-frame prediction processing path). The shared DMVR-SAD and pixel prediction path (e.g., the shared inter-frame prediction processing path) includes the first shared interpolation filter 382, the second shared interpolation filter 384, a shared weighted prediction and SAD block 386, and a DMVR block 330. As will be explained in more depth below, in some examples, the shared weighted prediction and SAD block 386 may include one or more shared arithmetic units, shared logic units, shared arithmetic logic units, etc.
[0120] In one illustrative example, the DMVR block 330 may be the same as or similar to the Figure 2 DMVR block 230. In one illustrative example, the 2-tap 32x2 shared horizontal interpolation filter 382 may be used to provide a 2-tap 20x2 horizontal interpolation filter 252 and an 8-tap 8x2 horizontal interpolation filter 242, both of which are illustrated in Figure 2 . The 2-tap 32x2 shared vertical interpolation filter 384 may be used to provide a 2-tap 20x2 vertical interpolation filter 254 and an 8-tap 8x2 vertical interpolation filter 244, both of which are illustrated in Figure 2 .
[0121] For example, the first 2-tap 32x2 shared interpolation filter 382 and the second 2-tap 32x2 shared interpolation filter 384 may each include a plurality of interpolation logic units that may be reconfigured or reallocated to perform 2-tap 20x2 interpolation and 8-tap 8x2 interpolation as needed. For example, the 2-tap 32x2 shared interpolation filters 382 and 384 may each include 2 * 32 = 64 interpolation logic units. To perform the 2-tap 20x2 interpolation associated with the motion vector refinement path, 24 of the interpolation logic units may remain inactive (e.g., 64 - 2 * 20 = 24). To perform the 8-tap 8x2 interpolation associated with the pixel prediction path, all 64 interpolation logic units may be active, reconfigured from the 2-tap 32x2 configuration to the 8-tap 8x2 configuration (e.g., 2 * 32 = 64 = 8 * 8).
[0122] By providing a first 2-tap 32x2 shared interpolation filter 382 and a second 2-tap 32x2 shared interpolation filter 384, the improved DMVR pixel prediction architecture 300 can be selectively configured to perform motion vector refinement processing tasks (e.g., for DMVR PUs) and pixel prediction processing tasks (e.g., for all PUs, DMVR or non-DMVR) using a single processing path or pipeline, which tasks are performed using Figure 2 two separate processing paths in the method of. By transforming between an 8-tap 8x2 interpolation filter and a 2-tap 32x2 interpolation filter, the two shared interpolation filters 382 and 384 can be reused for dual functionality to implement motion vector refinement processing and pixel prediction processing in a single pipeline.
[0123] The shared interpolation filter logic can improve the inter-frame prediction performance and efficiency, for example, by reducing the die size and / or providing better alignment with hardware implementation limitations that limit the number of pixels that can be interpolated in a single clock cycle. For example, the improved DMVR pixel prediction architecture 300 can eliminate the dedicated 2-tap horizontal and vertical filters used in the dedicated DMVR interpolation path for SAD (e.g., Figure 2 the dedicated 2-tap horizontal interpolation filter 252 and 2-tap vertical interpolation filter 254 illustrated). The improved DMVR pixel prediction architecture 300 can perform inter-frame prediction using DMVR with lower power consumption (due to fewer hardware, logic, circuit, etc. components), smaller die size, and lower silicon cost.
[0124] In one illustrative example, the improved DMVR pixel prediction architecture 300 can perform inter-frame prediction using DMVR without a reordering step, because the improved DMVR pixel prediction architecture 300 can avoid unordered pixel prediction. For example, a 16x16 luminance DMVR block can be processed by a shared pipeline first configured to perform motion vector refinement (e.g., the shared interpolation filters 382 and 384 are configured to perform 2-tap 20x2 horizontal interpolation and 2-tap 20x2 vertical interpolation, respectively). The interpolation outputs of the 20x2 horizontal and vertical interpolations can include the same 20x20 DPred0 and DPred1 arrays previously described with respect to Figure 2 Then, the DPred0 and DPred1 arrays can be provided to the shared weighted prediction and SAD block 386. When
[0125] the shared pipeline of is configured to perform motion vector refinement, the shared weighted prediction and SAD block 386 can be configured to calculate the SAD array DSAD from DPred0 and DPred1 (e.g., as described with respect to Figure 3 and Figure 2as described). The shared weighted prediction and SAD block 386 can implement the shared logic between the weighted prediction operation (which takes the sum of the inputs) and the SAD operation (which takes the difference of the inputs) by negating one of the input values provided to the shared block 386. For example, negating one of DPred0 and DPred1 at the shared weighted prediction and SAD block 386 and then summing the two values is equivalent to taking the difference between DPred0 and DPred1 (and SAD is the sum of absolute differences). In some examples, the shared weighted prediction and SAD block 386 can include one or more shared arithmetic units, shared logic units, shared arithmetic logic units, etc., for summing the inputs of the weighted prediction operation and summing the inputs of the SAD operation (e.g., where one is negated).
[0126] In one illustrative example, the improved DMVR pixel prediction architecture 300 can avoid out-of-order pixel prediction by configuring the first shared interpolation filter 382 and the second shared interpolation filter 384 to perform Figure 2 the 2-tap 20x2 interpolation of the interpolation filters 252 and 254 and by configuring the shared weighted prediction and SAD block 386 to perform Figure 2 the SAD calculation of the SAD block 256. Returning to the example of the 16x16 luminance DMVR PU, the reference prefetch block 310, the data path scheduler 320, and the DMA 312 (which can be the same or similar to Figure 2 the corresponding components therein) can first provide the 16x16 luminance DMVR PU to the shared processing pipeline (e.g., the shared inter-frame prediction processing path) of 382, 384, 386, 330. The DMVR block 330 generates a refined motion vector for the 16x16 luminance DMVR PU and returns it to the data path scheduler 320. Then, the data path scheduler 320 can configure the shared inter-frame prediction processing path to perform pixel prediction based on the refined motion vector. Since the motion vector refinement and pixel prediction are performed sequentially and using the same shared hardware and processing pipeline, out-of-order pixel prediction can be avoided, and the Figure 2 PU reordering SRAM 260 (e.g., which can include a large amount of SRAM) can be removed.
[0127] In some examples, the improved DMVR pixel prediction architecture 300 can use shared reference data prefetching between motion vector refinement operations and pixel prediction operations. For example, the reference prefetch block 310 can perform a single reference data prefetch (e.g., from the DMA 312), which is large enough for both the DMVR SAD configuration of the shared inter-frame prediction processing path of the architecture 300 and for the subsequent pixel prediction configuration of the shared inter-frame prediction processing path of the architecture 300. By sharing the reference data prefetch between successive motion vector refinement and pixel prediction operations, the prefetch latency caused by the DMA 312 can be reduced (e.g., halved) relative to Figure 2 the prefetch latency of the illustrated system.
[0128] In one illustrative example, the reference prefetch block 310 can request or prefetch a 27x27 reference data block from the DMA 312 (e.g., an increase from the 24x24 prefetch reference data block size of the example discussed Figure 2 . Since the maximum change of the refined motion vector from the original motion vector can be limited to a range of -2 to +2 pixels, the 27x27 reference data prefetch block size can be large enough for both the initial motion vector refinement (e.g., which utilizes a 24x24 block size) and the subsequent pixel prediction using the refined motion vector (e.g., since the refined motion vector can be limited to a range of ±2).
[0129] As illustrated, the output of the refined motion vector (e.g., generated by the DMVR block 330) can be directly provided to the data path scheduler 320 because no additional reference data prefetching is performed between the motion vector refinement processing and the subsequent pixel prediction processing performed using the shared inter-frame prediction processing path (e.g., the same 27x27 reference data prefetch associated with the motion vector refinement processing will be reused by the data path scheduler 320 to schedule the pixel prediction processing).
[0130] Figure 4 is a flow diagram illustrating an example of a process 400 for processing image and / or video data. At block 402, the process 400 includes obtaining a reference data block for predicting a video data block. For example, one or more of the illustrated reference prefetch block 310 and / or the direct memory access (DMA) controller 312 can be used to obtain the reference data block. In some cases, the reference data block can be obtained based on an inter-frame prediction PU command received at the reference prefetch block (e.g., Figure 3 the illustrated reference prefetch block 310). The reference prefetch block (e.g., Figure 3 the illustrated reference prefetch block 310) can use the inter-frame prediction PU command to generate a prefetch request and send it to the DMA controller (e.g., Figure 3 the illustrated reference prefetch block 310) can use the inter-frame prediction PU command to generate a prefetch request and send it to the DMA controller (e.g., Figure 3the illustrated DMA controller 312). In some examples, the reference data block may be obtained by or provided to a data path scheduler block, such as Figure 3 the illustrated data path scheduler block 320.
[0131] At block 404, process 400 includes determining one or more refined motion vectors based on a reference data block using a shared inter-frame prediction processing path. For example, the shared inter-frame prediction processing path may include Figure 3 one or more of the illustrated first shared interpolation filter 382 and second shared interpolation filter 384. In some examples, the shared inter-frame prediction processing path may additionally or alternatively include Figure 3 the illustrated shared weighted prediction (e.g., WPRED) and sum of absolute differences (e.g., SAD) block 386. In some examples, process 400 includes determining one or more refined motion vectors by using the inter-frame prediction processing path as a motion vector refinement path. For example, it may be possible to use the inter-frame prediction processing path as a motion vector refinement path by configuring a first interpolation filter (e.g., Figure 3 the illustrated first shared interpolation filter 382) to perform 2-tap 20x2 horizontal interpolation and configuring a second interpolation filter (e.g., second shared interpolation filter 384) to perform 2-tap 20x2 vertical interpolation. In some aspects, process 400 includes performing inter-frame prediction for a video data block by using the inter-frame prediction processing path as a pixel prediction path. For example, it may be possible to use the inter-frame prediction processing path as a pixel prediction path by configuring a first interpolation filter (e.g., Figure 3 the illustrated first shared interpolation filter 382) to perform 8-tap 8x2 horizontal interpolation and configuring a second interpolation filter (e.g., second shared interpolation filter 384) to perform 8-tap 8x2 vertical interpolation.
[0132] In some examples, one or more refined motion vectors may be determined by performing 2-tap horizontal interpolation based on a reference data block using a first interpolation filter. The first interpolation filter may be the same as or similar to Figure 3 the illustrated first shared interpolation filter 382, and / or may be the same as or similar to Figure 2 the illustrated 2-tap horizontal interpolation filter 252. In some examples, one or more refined motion vectors may be determined by performing 2-tap vertical interpolation based on a reference data block using a second interpolation filter. The second interpolation filter may be the same as or similar to Figure 3 the illustrated second shared interpolation filter 384, and / or may be the same as or similar to Figure 2 the illustrated 2-tap vertical interpolation filter 254.
[0133] In some examples, a shared inter - prediction processing path can be used to determine the sum of absolute differences (SAD) based on 2 - tap horizontal interpolation and 2 - tap vertical interpolation. For example, the shared inter - prediction processing path can include Figure 3 the illustrated shared WPRED and SAD block 386. Figure 3 The illustrated shared WPRED and SAD block 386 can be used to determine weighted prediction and determine SAD using the same arithmetic logic. In some cases, process 400 can further include generating one or more refined motion vectors using the SAD and one or more original motion vectors associated with a video data block. In some aspects, process 400 can include generating one or more refined motion vectors based on the SAD determined for 2 - tap horizontal interpolation and 2 - tap vertical interpolation, where the SAD is determined based on the sum of the negative value of the 2 - tap vertical interpolation and the 2 - tap horizontal interpolation. For example, the SAD can be determined by multiplying one of the 2 - tap horizontal interpolation or 2 - tap vertical interpolation by - 1 and determining the sum of the negative value and the remaining one of the 2 - tap horizontal interpolation or 2 - tap vertical interpolation.
[0134] At block 406, process 400 includes performing inter - prediction for a video data block using the shared inter - prediction processing path. The inter - prediction can be based on a reference data block and one or more refined motion vectors. For example, the shared inter - prediction processing path can include Figure 3 one or more of the illustrated first shared interpolation filter 382 and second shared interpolation filter 384. For example, the shared inter - prediction processing path can include a 2 - tap 32x2 horizontal interpolation filter as the first interpolation filter (e.g., Figure 3 the illustrated first shared interpolation filter 382). In some examples, the shared inter - prediction processing path can include a 2 - tap 32x2 vertical interpolation as the second interpolation filter (e.g., Figure 3 the illustrated second shared interpolation filter 384). In some examples, the shared inter - prediction processing path can additionally or alternatively include Figure 3 the illustrated shared weighted prediction (e.g., WPRED) and sum of absolute differences (e.g., SAD) block 386. In some examples, inter - prediction can be performed by using the shared inter - prediction processing path to perform pixel prediction using the refined motion vectors. The refined motion vectors can be generated by using the shared inter - prediction processing path to perform DMVR and generate refined motion vectors.
[0135] In some examples, performing inter - prediction for a video data block can include determining an 8 - tap horizontal interpolation based on a reference data block and a first refined motion vector using a first interpolation filter. The first interpolation filter can be the same as or similar to Figure 3 the illustrated first shared interpolation filter 382, and / or can be the same asFigure 2 The illustrated 8-tap horizontal interpolation filter 242 is the same as or similar to. In some examples, performing inter-frame prediction for a video data block may include using a second interpolation filter to determine 8-tap vertical interpolation based on a reference data block and a second refined motion vector. The second interpolation filter may be the same as or similar to Figure 3 the illustrated second shared interpolation filter 384, and / or may be the same as or similar to Figure 2 the illustrated 8-tap vertical interpolation filter 244. In some examples, 8-tap horizontal interpolation and 8-tap vertical interpolation may be used to generate multiple inter-frame prediction pixels.
[0136] For example, 8-tap horizontal interpolation and 8-tap vertical interpolation may be used to determine a weighted prediction, where the weighted prediction is determined based on the sum of the 8-tap horizontal interpolation and the 8-tap vertical interpolation. In some examples, Figure 3 the illustrated shared weighted prediction and SAD block 386 may be used to determine the weighted prediction. In some aspects, process 400 may include generating multiple inter-frame prediction pixels based on the weighted prediction.
[0137] In some aspects, process 400 includes generating an encoded video bitstream including one or more pictures, at least one of the one or more pictures being based on inter-frame prediction performed for a video data block. For example, the encoded video bitstream may be generated by Figure 3 the illustrated reconstruction interface 370 and / or may be generated based on inter-frame prediction pixels output by Figure 3 the illustrated reconstruction interface 370. In some aspects, process 400 includes sending the encoded video bitstream to a decoding device, the encoded video bitstream being sent together with signaling information. In some aspects, process 400 includes storing the encoded video bitstream. In some aspects, process 400 includes obtaining one or more encoded pictures, at least one of the one or more encoded pictures including a video data block. In some examples, process 400 includes decoding a video data block from at least one encoded picture. In some aspects, process 400 includes decoding a video data block from at least one encoded picture by reconstructing the video data block.
[0138] In some examples, process 400 may be performed by a decoding device (e.g., Figure 1 and Figure 6 the decoding device 112). In some cases, process 400 may be performed by an encoding device (e.g., Figure 1 and Figure 5 the encoding device 104).
[0139] In some embodiments, the processes (or methods) described herein may be performed by a computing device or apparatus (such as Figure 1The system 100) shown is executed. For example, it can be executed by Figure 1 and Figure 5 the encoding device 104 shown, by another video source side device or video transmitting device, by Figure 1 and Figure 6 the decoding device 112 shown, and / or by another client side device such as a player device, a display, or any other client side device. In some cases, a computing device or apparatus may include a processor, a microprocessor, a microcomputer, or other components of a device configured to perform the steps of the processes described herein. In some examples, a computing device or apparatus may include a camera configured to capture video data (e.g., a video sequence) including video frames. In some examples, the camera or other capture device that captures the video data is separate from the computing device, in which case the computing device receives or obtains the captured video data. The computing device may also include a network interface configured to communicate the video data. The network interface may be configured to communicate data based on the Internet Protocol (IP) or other types of data. In some examples, a computing device or apparatus may include a display for displaying output video content (such as samples of pictures of a video bitstream).
[0140] These processes can be described with respect to logical flowcharts, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally speaking, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform a particular function or implement a particular data type. The operations are not to be construed as limiting in the order in which they are described, and any number of the described operations can be combined in any order and / or in parallel to implement the process.
[0141] Furthermore, these processes can be executed under the control of one or more computer systems configured with executable instructions and can be implemented in hardware as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed jointly on one or more processors, or a combination thereof. As noted above, the code can be stored on a computer-readable or machine-readable storage medium, e.g., in the form of a computer program including a plurality of instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.
[0142] The decoding techniques discussed herein can be implemented in an exemplary video encoding and decoding system (e.g., system 100). In some examples, a system includes a source device that provides encoded video data to be decoded later by a destination device. Specifically, the source device provides the video data to the destination device via a computer-readable medium. The source device and the destination device can include any of a variety of devices, including desktop computers, notebooks (i.e., laptops) computers, tablet computers, set-top boxes, telephone handheld devices (such as so-called "smart" phones, so-called "smart" tablets), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, and the like. In some cases, the source device and the destination device can be equipped for wireless communication.
[0143] The destination device can receive the encoded video data to be decoded via a computer-readable medium. The computer-readable medium can include any type of medium or device capable of moving the encoded video data from the source device to the destination device. In one example, the computer-readable medium can include a communication medium that enables the source device to send the encoded video data directly in real time to the destination device. The encoded video data can be modulated according to a communication standard (such as a wireless communication protocol) and sent to the destination device. The communication medium can include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network such as: a local area network, a wide area network, or a global network such as the Internet. The communication medium can include routers, switches, base stations, or any other equipment that can be useful for facilitating communication from the source device to the destination device.
[0144] In some examples, the encoded data can be output from the output interface to a storage device. Similarly, the encoded data can be accessed from the storage device via the input interface. The storage device can include any data storage medium among various distributed or locally accessible data storage media, such as a hard disk drive, a Blu-ray disc, a DVD, a CD-ROM, a flash memory, a volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. In another example, the storage device can correspond to a file server or another intermediate storage device that can store the encoded video generated by the source device. The target device can access the stored video data from the storage device via streaming or downloading. The file server can be any type of server capable of storing the encoded video data and sending the encoded video data to the target device. Example file servers include web servers (e.g., for websites), FTP servers, network-attached storage (NAS) devices, or local disk drives. The target device can access the encoded video data through any standard data connection, including an Internet connection. This can include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on the file server. The sending of the encoded video data from the storage device can be streaming, download sending, or a combination of them.
[0145] The techniques of the present disclosure are not necessarily limited to wireless applications or settings. These techniques can be applied to video coding to support any of various multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission (such as HTTP-based Dynamic Adaptive Streaming over HTTP (DASH)), digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, the system can be configured to support unidirectional or bidirectional video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0146] In one example, the source device includes a video source, a video encoder, and an output interface. The target device can include an input interface, a video decoder, and a display device. The video encoder of the source device can be configured to apply the techniques disclosed herein. In other examples, the source device and the target device can include other components or arrangements. For example, the source device can receive video data from an external video source such as an external camera. Similarly, the target device can interface with an external display device instead of including an integrated display device.
[0147] The above example systems are merely examples. Techniques for simultaneously processing video data can be performed by any digital video encoding and / or decoding device. Although generally the techniques of the present disclosure are performed by a video encoding device, these techniques can also be performed by a video encoder / decoder (commonly referred to as a “codec”). Additionally, the techniques of the present disclosure can also be performed by a video preprocessor. The source device and the destination device are merely examples of such decoding devices, where the source device generates the decoded video data for transmission to the destination device. In some examples, the source device and the destination device can operate in a substantially symmetric manner such that each of these devices includes video encoding and decoding components. Thus, the example systems can support one-way or two-way video transmission between video devices, e.g., for video streaming, video playback, video broadcast, or video telephony.
[0148] The video source can include a video capture device such as a camera, a video archival unit including previously captured video, and / or a video feed interface for receiving video from a video content provider. As an additional alternative, the video source can generate computer graphics-based data as the source video, or generate a combination of live video, archived video, and computer-generated video. In some cases, if the video source is a camera, the source device and the destination device can form a so-called camera phone or video phone. However, as mentioned above, the techniques described in the present disclosure are generally applicable to video decoding and can be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video can be encoded by a video encoder. The encoded video information can then be output by an output interface to a computer-readable medium.
[0149] As noted, the computer-readable medium can include a transient medium (such as a wireless broadcast or a wired network transmission) or a storage medium (i.e., a non-transient storage medium) such as a hard disk, a flash drive, a compact disc, a digital video disc, a Blu-ray disc, or other computer-readable media. In some examples, a network server (not shown) can receive the encoded video data from the source device and provide the encoded video data to the destination device, e.g., via a network transmission. Similarly, a computing device of a media production facility (such as a disc stamping facility) can receive the encoded video data from the source device and produce a disc containing the encoded video data. Thus, in various examples, the computer-readable medium can be understood to include one or more computer-readable media in various forms.
[0150] The input interface of the target device receives information from a computer-readable medium. The information of the computer-readable medium may include syntax information defined by a video encoder and also used by a video decoder, and the syntax information includes syntax elements that describe the characteristics and / or processing of blocks and other decoded units (e.g., groups of pictures (GOPs)). A display device displays the decoded video data to a user and may include any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device. Various embodiments of the present application have been described.
[0151] Specific details of the encoding device 104 and the decoding device 112 are shown respectively in Figure 5 and Figure 6 are shown. Figure 5 is a block diagram illustrating an example encoding device 104 that can implement one or more of the techniques described in the present disclosure. For example, the encoding device 104 can generate the syntax structures described herein (e.g., the syntax structures of VPS, SPS, PPS, or other syntax elements). The encoding device 104 can perform intra prediction and inter prediction decoding of video blocks within a video slice. As previously described, intra decoding relies at least in part on spatial prediction to reduce or remove spatial redundancy within a given video frame or picture. Inter decoding relies at least in part on temporal prediction to reduce or remove temporal redundancy within adjacent or surrounding frames of a video sequence. The intra mode (I mode) can refer to any of several space-based compression modes. The inter modes (such as unidirectional prediction (P mode) or bidirectional prediction (B mode)) can refer to any of several time-based compression modes.
[0152] The encoding device 104 includes a partitioning unit 35, a prediction processing unit 41, a filter unit 63, a picture memory 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 includes a motion estimation unit 42, a motion compensation unit 44, and an intra prediction processing unit 46. For video block reconstruction, the encoding device 104 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62. The filter unit 63 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although the filter unit 63 is shown as an in-loop filter in Figure 5 it can be implemented as a post-loop filter in other configurations. A post-processing device 57 can perform additional processing on the encoded video data generated by the encoding device 104. In some examples, the techniques of the present disclosure can be implemented by the encoding device 104. However, in other cases, one or more of the techniques of the present disclosure can be implemented by the post-processing device 57.
[0153] As Figure 5 shown, the encoding device 104 receives video data, and the splitting unit 35 splits the data into video blocks. The splitting may also include splitting into slices, slice segments, tiles, or other larger units as well as video block splitting (e.g., according to the quadtree structure of LCU and CU). The encoding device 104 generally shows components for encoding video blocks within a video slice to be encoded. A slice may be divided into multiple video blocks (and possibly divided into sets of video blocks called tiles). The prediction processing unit 41 may select one of multiple possible decoding modes for the current video block based on error results (e.g., decoding rate and distortion level, etc.), such as one of multiple intra-prediction decoding modes or one of multiple inter-prediction decoding modes. The prediction processing unit 41 may provide the resulting intra-coded or inter-coded block to the adder 50 to generate residual block data and to the adder 62 to reconstruct the encoded block for use as a reference picture.
[0154] The intra-prediction processing unit 46 within the prediction processing unit 41 may perform intra-prediction decoding of the current video block with respect to one or more adjacent blocks in the same frame or slice as the current block to be decoded to provide spatial compression. The motion estimation unit 42 and the motion compensation unit 44 within the prediction processing unit 41 perform inter-prediction decoding of the current video block with respect to one or more prediction blocks in one or more reference pictures to provide temporal compression.
[0155] The motion estimation unit 42 may be configured to determine the inter-prediction mode of a video slice according to a predetermined pattern of the video sequence. The predetermined pattern may specify video slices in the sequence as P slices, B slices, or GPB slices. The motion estimation unit 42 and the motion compensation unit 44 may be highly integrated but are shown separately for conceptual purposes. The motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector, which estimates the motion of a video block. For example, the motion vector may indicate the displacement of a prediction unit (PU) of a video block within the current video frame or picture relative to a prediction block within a reference picture.
[0156] A prediction block is a block that is found to closely match the PU of the video block to be decoded in terms of pixel difference, which can be determined by the sum of absolute differences (SAD), the sum of squared differences (SSD), or other difference metrics. In some examples, the encoding device 104 may calculate the values at sub-integer pixel positions of the reference picture stored in the picture memory 64. For example, the encoding device 104 may interpolate the values at quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference picture. Thus, the motion estimation unit 42 may perform a motion search with respect to full pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.
[0157] The motion estimation unit 42 calculates the motion vector of the PU of the video block in the slice decoded inter - frame by comparing the position of the PU of the video block in the slice decoded inter - frame with the position of the prediction block in the reference picture. The reference picture can be selected from the first reference picture list (list 0) or the second reference picture list (list 1), where each identifies one or more reference pictures stored in the picture memory 64. The motion estimation unit 42 transmits the calculated motion vector to the entropy coding unit 56 and the motion compensation unit 44.
[0158] The motion compensation performed by the motion compensation unit 44 may involve extracting or generating a prediction block based on the motion vector determined by the motion estimation, and thus may perform interpolation at sub - pixel accuracy. After receiving the motion vector of the PU of the current video block, the motion compensation unit 44 may locate the prediction block pointed to by the motion vector in the reference picture list. The encoding device 104 forms a residual video block by subtracting the pixel values of the prediction block from the pixel values of the currently decoded video block, thereby forming a pixel difference. The pixel difference forms the residual data of the block and may include both a luminance difference component and a chrominance difference component. The summer 50 represents one or more components that perform this subtraction operation. The motion compensation unit 44 may also generate syntax elements associated with the video block and the video slice for use by the decoding device 112 when decoding the video blocks of the video slice.
[0159] The intra - prediction processing unit 46 may perform intra - prediction on the current block as an alternative to the inter - prediction performed by the motion estimation unit 42 and the motion compensation unit 44, as described above. Specifically, the intra - prediction processing unit 46 may determine the intra - prediction mode for encoding the current block. In some examples, the intra - prediction processing unit 46 may, for example, use various intra - prediction modes to encode the current block during a separate encoding process, and the intra - prediction processing unit 46 may select an appropriate intra - prediction mode to use from the tested modes. For example, the intra - prediction processing unit 46 may calculate rate - distortion values using rate - distortion analysis for various tested intra - prediction modes, and may select the intra - prediction mode with the best rate - distortion characteristics among the tested modes. Rate - distortion analysis generally determines the amount of distortion (or error) between the encoded block and the original un - encoded block that was encoded to produce the encoded block, as well as the bit rate (i.e., the number of bits) used to produce the encoded block. The intra - prediction processing unit 46 may calculate a ratio based on the distortion and rate of various encoded blocks to determine which intra - prediction mode exhibits the best rate - distortion value for the block.
[0160] In any case, after selecting an intra prediction mode for a block, the intra prediction processing unit 46 may provide information indicating the selected intra prediction mode for the block to the entropy coding unit 56. The entropy coding unit 56 may encode the information indicating the selected intra prediction mode. The encoding device 104 may include in the transmitted bitstream configuration data a definition of the coding context for various blocks and an indication of the most probable intra prediction mode, intra prediction mode index table, and modified intra prediction mode index table for each of the contexts. The bitstream configuration data may include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also referred to as codeword mapping tables).
[0161] After the prediction processing unit 41 generates a prediction block for a current video block via inter prediction or intra prediction, the encoding device 104 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and applied to the transform processing unit 52. The transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform such as a discrete cosine transform (DCT) or a conceptually similar transform. The transform processing unit 52 may convert the residual video data from the pixel domain to a transform domain, such as a frequency domain.
[0162] The transform processing unit 52 may transfer the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, the quantization unit 54 may then perform a scan of the matrix including the quantized transform coefficients. Alternatively, the entropy coding unit 56 may perform the scan.
[0163] After quantization, the entropy coding unit 56 performs entropy coding on the quantized transform coefficients. For example, the entropy coding unit 56 may perform context - adaptive variable - length coding (CAVLC), context - adaptive binary arithmetic coding (CABAC), syntax - based context - adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding technique. After entropy coding by the entropy coding unit 56, the encoded bitstream may be sent to the decoding device 112, or archived for later transmission or retrieval by the decoding device 112. The entropy coding unit 56 may also perform entropy coding on the motion vectors and other syntax elements of the current video slice being decoded.
[0164] The inverse quantization unit 58 and the inverse transform processing unit 60 respectively apply inverse quantization and inverse transform to reconstruct the residual block in the pixel domain for later use as a reference block of the reference picture. The motion compensation unit 44 can calculate the reference block by adding the residual block to a predicted block of one of the reference pictures in the reference picture list. The motion compensation unit 44 can also apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer pixel values for use in motion estimation. The adder 62 adds the reconstructed residual block to the motion-compensated predicted block generated by the motion compensation unit 44 to generate a reference block for storage in the picture memory 64. The reference block can be used by the motion estimation unit 42 and the motion compensation unit 44 as a reference block for inter-frame prediction of blocks in subsequent video frames or pictures.
[0165] In this way, Figure 5 the encoding device 104 represents an example of a video encoder configured to perform the techniques described herein. For example, the encoding device 104 can perform any of the techniques described herein, including the processes described herein. In some cases, some of the techniques of the present disclosure can also be implemented by a post-processing device 57.
[0166] Figure 6 is a block diagram illustrating an example decoding device 112. The decoding device 112 includes an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, an adder 90, a filter unit 91, and a picture memory 92. The prediction processing unit 81 includes a motion compensation unit 82 and an intra-prediction processing unit 84. In some examples, the decoding device 112 can perform a decoding pass that is substantially inverse to the encoding pass described with respect to the encoding device 104 from Figure 5 the encoding device 104.
[0167] During the decoding process, the decoding device 112 receives an encoded video bitstream representing video blocks and associated syntax elements of an encoded video slice transmitted by the encoding device 104. In some embodiments, the decoding device 112 can receive the encoded video bitstream from the encoding device 104. In some embodiments, the decoding device 112 can receive the encoded video bitstream from a network entity 79 such as a server, a media-aware network element (MANE), a video editor / splicer, or other such device configured to implement one or more of the above techniques. The network entity 79 may or may not include the encoding device 104. Some of the techniques described in the present disclosure can be implemented by the network entity 79 before the network entity 79 sends the encoded video bitstream to the decoding device 112. In some video decoding systems, the network entity 79 and the decoding device 112 can be part of separate devices, while in other examples, the functionality described with respect to the network entity 79 can be performed by the same device including the decoding device 112.
[0168] The entropy decoding unit 80 of the decoding device 112 performs entropy decoding on the bitstream to generate the quantized coefficients, motion vectors, and other syntax elements. The entropy decoding unit 80 forwards the motion vectors and other syntax elements to the prediction processing unit 81. The decoding device 112 may receive syntax elements at the video slice level and / or the video block level. The entropy decoding unit 80 may process and parse both fixed-length syntax elements and variable-length syntax elements in one or more parameter sets (such as VPS, SPS, and PPS).
[0169] When a video slice is decoded as an intra-coded (I) slice, the intra prediction processing unit 84 of the prediction processing unit 81 may generate prediction data for the video blocks of the current video slice based on the signalized intra prediction mode and the data from the previously decoded blocks of the current frame or picture. When a video frame is decoded as an inter-coded (i.e., B, P, or GPB) slice, the motion compensation unit 82 of the prediction processing unit 81 generates a prediction block for the video blocks of the current video slice based on the motion vectors and other syntax elements received from the entropy decoding unit 80. The prediction block may be generated from one of the reference pictures in the reference picture list. The decoding device 112 may construct the reference frame lists (list 0 and list 1) using a default construction technique based on the reference pictures stored in the picture memory 92.
[0170] The motion compensation unit 82 determines the prediction information for the video blocks of the current video slice by parsing the motion vectors and other syntax elements, and uses the prediction information to generate a prediction block for the current video block being decoded. For example, the motion compensation unit 82 may use one or more syntax elements in the parameter set to determine the prediction mode (e.g., intra prediction or inter prediction) for decoding the video blocks of the video slice, the inter prediction slice type (e.g., B slice, P slice, or GPB slice), the construction information of one or more reference picture lists of the slice, the motion vectors of each inter-coded video block of the slice, the inter prediction state of each inter-coded video block of the slice, and other information for decoding the video blocks in the current video slice.
[0171] The motion compensation unit 82 may also perform interpolation based on an interpolation filter. The motion compensation unit 82 may use the interpolation filter used by the encoding device 104 during the encoding of the video block to calculate the interpolated values of the sub-integer pixels of the reference block. In this case, the motion compensation unit 82 may determine the interpolation filter used by the encoding device 104 from the received syntax elements, and may use the interpolation filter to generate the prediction block.
[0172] The inverse quantization unit 86 inverse quantizes or dequantizes the quantized transform coefficients provided in the bitstream and decoded by the entropy decoding unit 80. The inverse quantization process may include determining the degree of quantization using the quantization parameter calculated by the encoding device 104 for each video block in the video slice, and correspondingly determining the degree of inverse quantization to be applied. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT or other suitable inverse transform), an inverse integer transform, or a conceptually similar inverse transform process to the transform coefficients to generate a residual block in the pixel domain.
[0173] After the motion compensation unit 82 generates a prediction block for the current video block based on the motion vector and other syntax elements, the decoding device 112 forms the decoded video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82. The adder 90 represents one or more components that perform this summation operation. If desired, a loop filter (either in the decoding loop or after the decoding loop) may also be used to smooth pixel transitions or otherwise improve video quality. The filter unit 91 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although the filter unit 91 is shown as an in-loop filter in Figure 6 , in other configurations, the filter unit 91 may be implemented as a post-loop filter. Then, the decoded video blocks in a given frame or picture are stored in the picture memory 92, which stores reference pictures for subsequent motion compensation. The picture memory 92 also stores the decoded video for later presentation on a display device (such as Figure 1 the video destination device 122 shown).
[0174] In this way, Figure 6 the decoding device 112 represents an example of a video decoder configured to perform the techniques described herein. For example, the decoding device 112 may perform any of the techniques described herein, including the processes described herein.
[0175] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. The computer-readable medium may include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagated wirelessly or over a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, memory, or memory devices. The computer-readable medium may have code and / or machine-executable instructions stored thereon, which may represent a process, function, subroutine, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program statements. By passing and / or receiving information, data, arguments, parameters, or memory contents, a code segment may be coupled to another code segment or hardware circuit. The information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means, including memory sharing, message passing, token passing, network transmission, etc.
[0176] In some embodiments, computer-readable storage devices, media, and memories may include wired or wireless signals including bitstreams, etc. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as power consumption, carrier signals, electromagnetic waves, and signals themselves.
[0177] Specific details are provided in the foregoing description to provide a thorough understanding of the embodiments and examples provided herein. However, one of ordinary skill in the art will understand that the embodiments may be practiced without these specific details. For clarity, in some instances, the present technology may be presented as including separate functional blocks, including functional blocks containing devices, device components, steps or routines in a method embodied in software or a combination of hardware and software. Additional components other than those shown in the figures and / or described herein may be used. For example, circuits, systems, networks, processes, and other components may be shown in block diagram form as components to avoid obscuring these embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments.
[0178] Individual embodiments may be described above as processes or methods depicted as flowcharts, flow diagrams, data flow diagrams, structure diagrams, or block diagrams. Although a flowchart may depict operations as a sequential process, many of the operations may be performed in parallel or concurrently. In addition, the order of the operations may be rearranged. A process is terminated when the operations of the process are completed, but the process may have additional steps not included in the figure. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, the termination of the process may correspond to the function returning to the calling function or the main function.
[0179] The processes and methods according to the above examples can be implemented using computer-executable instructions stored or otherwise obtained from a computer-readable medium. These instructions can include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a certain function or group of functions. Portions of the computer resources used can be accessed via a network. The computer-executable instructions can be, for example, binary, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, the information used, and / or the information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices with non-volatile memory, networked storage devices, etc.
[0180] Devices implementing the processes and methods according to these disclosures can include hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof, and can take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., a computer program product) for performing the necessary tasks can be stored in a computer-readable or machine-readable medium. A processor can execute the necessary tasks. Typical examples of form factors include laptop computers, smart phones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein can also be embodied in peripheral devices or plug-in cards. By further example, such functionality can also be implemented on a circuit board among different chips or different processes executed on a single device.
[0181] Instructions, the medium for conveying such instructions, the computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.
[0182] In the foregoing description, aspects of the present application have been described with reference to specific embodiments of the present application, but those skilled in the art will recognize that the present application is not limited thereto. Although exemplary embodiments of the present application have been described in detail herein, it is to be understood that the inventive concept may be embodied and employed in other various ways, and the appended claims are intended to be construed to cover such variations, unless limited by the prior art. The various features and aspects of the above application may be used alone or in combination. Additionally, the embodiments may be used in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that in alternative embodiments, the methods may be performed in a different order than that described.
[0183] Those of ordinary skill in the art will understand that the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced with less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively, without departing from the scope of the specification.
[0184] In cases where a component is described as “configured to” perform certain operations, such a configuration may be achieved, for example, by designing electronic circuitry or other hardware to perform the operations, by programming a programmable electronic circuit (e.g., a microprocessor or other suitable electronic circuit) to perform the operations, or any combination thereof.
[0185] The phrase “coupled to” means that any component is directly or indirectly physically connected to another component, and / or any component directly or indirectly communicates with another component (e.g., is connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0186] Claim language or other language reciting “at least one of” a set and / or “one or more of” a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting “at least one of A and B” means A, B, or A and B. In another example, claim language reciting “at least one of A, B, and C” means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language “at least one of” a set and / or “one or more of” a set does not limit the set to the items listed in the set. For example, claim language reciting “at least one of A and B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.
[0187] The various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or any combination thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0188] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general purpose computer, a wireless communication device handset, or an integrated circuit device having multiple uses, including applications in a wireless communication device handset and other devices. Any feature described as a module or component may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be at least partially implemented by a computer-readable data storage medium comprising program code, including instructions that, when executed, perform one or more of the above-described methods. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. Additionally or alternatively, the techniques may be at least partially implemented by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0189] The program code can be executed by a processor, which can include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor can be a microprocessor, but in an alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, as used herein, the term "processor" can refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within a dedicated software module or hardware module configured for encoding and decoding, or incorporated into a combined video encoder-decoder (CODEC).
[0190] Exemplary aspects of the present disclosure include:
[0191] Aspect 1. An apparatus for processing video data, the apparatus comprising: at least one memory; and at least one processor, the at least one processor coupled to the at least one memory and configured to: obtain a reference data block for predicting a video data block; determine one or more refined motion vectors based on the reference data block using an inter-frame prediction processing path; and perform inter-frame prediction for the video data block using the inter-frame prediction processing path, wherein the inter-frame prediction is based on the reference data block and the one or more refined motion vectors.
[0192] Aspect 2. The apparatus according to aspect 1, wherein, to determine the one or more refined motion vectors, the at least one processor is configured to: determine a 2-tap horizontal interpolation based on the reference data block using a first interpolation filter; and determine a 2-tap vertical interpolation based on the reference data block using a second interpolation filter.
[0193] Aspect 3. The apparatus according to any one of aspects 1 to 2, wherein the at least one processor is configured to: determine a sum of absolute differences (SAD) based on the 2-tap horizontal interpolation and the 2-tap vertical interpolation; and generate the one or more refined motion vectors using the SAD and one or more original motion vectors associated with the video data block.
[0194] Aspect 4. The apparatus according to any one of Aspects 1 to 3, wherein, in order to perform inter-frame prediction on the video data block, the at least one processor is configured to: determine an 8-tap horizontal interpolation based on the reference data block and a first refined motion vector using the first interpolation filter; determine an 8-tap vertical interpolation based on the reference data block and a second refined motion vector using the second interpolation filter; and generate a plurality of inter-frame prediction pixels using the 8-tap horizontal interpolation and the 8-tap vertical interpolation.
[0195] Aspect 5. The apparatus according to any one of Aspects 1 to 4, wherein the at least one processor is configured to: determine a weighted prediction using the 8-tap horizontal interpolation and the 8-tap vertical interpolation, wherein the weighted prediction is determined based on the sum of the 8-tap horizontal interpolation and the 8-tap vertical interpolation; and generate the plurality of inter-frame prediction pixels based on the weighted prediction.
[0196] Aspect 6. The apparatus according to any one of Aspects 1 to 5, wherein the at least one processor is configured to: generate the one or more refined motion vectors based on the sum of absolute differences (SAD) determined for the 2-tap horizontal interpolation and the 2-tap vertical interpolation; wherein the SAD is determined based on the sum of the negative value of the 2-tap vertical interpolation and the 2-tap horizontal interpolation.
[0197] Aspect 7. The apparatus according to any one of Aspects 1 to 6, wherein the at least one processor is configured to: use arithmetic logic to determine the weighted prediction and determine the SAD.
[0198] Aspect 8. The apparatus according to any one of Aspects 1 to 7, wherein: the inter-frame prediction processing path includes the first interpolation filter and the second interpolation filter; the first interpolation filter is a 2-tap 32x2 horizontal interpolation filter; and the second interpolation filter is a 2-tap 32x2 vertical interpolation filter.
[0199] Aspect 9. The apparatus according to any one of Aspects 1 to 8, wherein: in order to determine the one or more refined motion vectors, the at least one processor is configured to use the inter-frame prediction processing path as a motion vector refinement path; and in order to perform inter-frame prediction on the video data block, the at least one processor is configured to use the inter-frame prediction processing path as a pixel prediction path.
[0200] Aspect 10. The apparatus according to any one of Aspects 1 to 9, wherein the at least one processor is configured to: use the inter-frame prediction processing path as a motion vector refinement path by configuring the first interpolation filter to perform 2-tap 20x2 horizontal interpolation and configuring the second interpolation filter to perform 2-tap 20x2 vertical interpolation; and use the inter-frame prediction processing path as a pixel prediction path by configuring the first interpolation filter to perform 8-tap 8x2 horizontal interpolation and configuring the second interpolation filter to perform 8-tap 8x2 vertical interpolation.
[0201] Aspect 11. The apparatus according to any one of Aspects 1 to 10, wherein the at least one processor is configured to: generate an encoded video bitstream including one or more pictures, at least one of the one or more pictures being based on the inter-frame prediction performed on the video data block.
[0202] Aspect 12. The apparatus according to any one of Aspects 1 to 11, wherein the at least one processor is configured to: send the encoded video bitstream to a decoding device, the encoded video bitstream being sent together with signaling information.
[0203] Aspect 13. The apparatus according to any one of Aspects 1 to 12, wherein the at least one processor is configured to: store the encoded video bitstream.
[0204] Aspect 14. The apparatus according to any one of Aspects 1 to 13, wherein the at least one processor is configured to: obtain one or more encoded pictures, at least one of the one or more encoded pictures including the video data block; and decode the video data block from the at least one encoded picture.
[0205] Aspect 15. The apparatus according to any one of Aspects 1 to 14, wherein, in order to decode the video data block from the at least one encoded picture, the at least one processor is configured to reconstruct the video data block.
[0206] Aspect 16. A method of processing video data, the method comprising: obtaining a reference data block for predicting a video data block; determining one or more refined motion vectors based on the reference data block using an inter-frame prediction processing path; and performing inter-frame prediction on the video data block using the inter-frame prediction processing path, wherein the inter-frame prediction is based on the reference data block and the one or more refined motion vectors.
[0207] Aspect 17. The method according to aspect 16, wherein determining the one or more refined motion vectors includes: determining a 2-tap horizontal interpolation based on the reference data block using a first interpolation filter; and determining a 2-tap vertical interpolation based on the reference data block using a second interpolation filter.
[0208] Aspect 18. The method according to any one of aspects 16 to 17, the method further comprising: determining a sum of absolute differences (SAD) based on the 2-tap horizontal interpolation and the 2-tap vertical interpolation; and generating the one or more refined motion vectors using the SAD and one or more original motion vectors associated with the video data block.
[0209] Aspect 19. The method according to any one of aspects 16 to 18, wherein performing inter-frame prediction for the video data block includes: determining an 8-tap horizontal interpolation based on the reference data block and a first refined motion vector using the first interpolation filter; determining an 8-tap vertical interpolation based on the reference data block and a second refined motion vector using the second interpolation filter; and generating a plurality of inter-frame prediction pixels using the 8-tap horizontal interpolation and the 8-tap vertical interpolation.
[0210] Aspect 20. The method according to any one of aspects 16 to 19, the method further comprising: determining a weighted prediction using the 8-tap horizontal interpolation and the 8-tap vertical interpolation, wherein the weighted prediction is determined based on the sum of the 8-tap horizontal interpolation and the 8-tap vertical interpolation; and generating the plurality of inter-frame prediction pixels based on the weighted prediction.
[0211] Aspect 21. The method according to any one of aspects 16 to 20, the method further comprising: generating the one or more refined motion vectors based on the sum of absolute differences (SAD) determined for the 2-tap horizontal interpolation and the 2-tap vertical interpolation; wherein the SAD is determined based on the sum of the negative value of the 2-tap vertical interpolation and the 2-tap horizontal interpolation.
[0212] Aspect 22. The method according to any one of aspects 16 to 21, the method further comprising: using arithmetic logic to determine the weighted prediction and determine the SAD.
[0213] Aspect 23. The method according to any one of aspects 16 to 22, wherein: the inter-frame prediction processing path includes the first interpolation filter and the second interpolation filter; the first interpolation filter is a 2-tap 32x2 horizontal interpolation filter; and the second interpolation filter is a 2-tap 32x2 vertical interpolation filter.
[0214] Aspect 24. The method according to any one of aspects 16 to 23, wherein: determining the one or more refined motion vectors includes using the inter-frame prediction processing path as a motion vector refinement path; and performing inter-frame prediction on the video data block includes using the inter-frame prediction processing path as a pixel prediction path.
[0215] Aspect 25. The method according to any one of aspects 16 to 24, the method further comprising: using the inter-frame prediction processing path as a motion vector refinement path by configuring the first interpolation filter to perform 2-tap 20x2 horizontal interpolation and configuring the second interpolation filter to perform 2-tap 20x2 vertical interpolation; and using the inter-frame prediction processing path as a pixel prediction path by configuring the first interpolation filter to perform 8-tap 8x2 horizontal interpolation and configuring the second interpolation filter to perform 8-tap 8x2 vertical interpolation.
[0216] Aspect 26. The method according to any one of aspects 16 to 25, the method further comprising: generating an encoded video bitstream including one or more pictures, at least one of the one or more pictures being based on the inter-frame prediction performed on the video data block.
[0217] Aspect 27. The method according to any one of aspects 16 to 26, the method further comprising: sending the encoded video bitstream to a decoding device, the encoded video bitstream being sent together with signaling information.
[0218] Aspect 28. The method according to any one of aspects 16 to 27, the method further comprising: storing the encoded video bitstream.
[0219] Aspect 29. The method according to any one of aspects 16 to 28, the method further comprising: obtaining one or more encoded pictures, at least one of the one or more encoded pictures including the video data block; and decoding the video data block from the at least one encoded picture.
[0220] Aspect 30. The method according to any one of aspects 16 to 29, wherein decoding the video data block from the at least one encoded picture includes reconstructing the video data block.
Claims
1. An apparatus for processing video data, the apparatus comprises: at least one memory; and at least one processor, the at least one processor being coupled to the at least one memory and configured to: obtain a reference data block for predicting a video data block; determine one or more refined motion vectors based on the reference data block using a first configuration of a plurality of interpolation logic units of a first interpolation filter and a second interpolation filter for implementing a shared inter-frame prediction processing path, wherein in the first configuration, the first interpolation filter is configured to determine a 2-tap horizontal interpolation based on the reference data block, and the second interpolation filter is configured to determine a 2-tap vertical interpolation based on the reference data block; perform inter-frame prediction for the video data block using a second configuration of the plurality of interpolation logic units of the shared inter-frame prediction processing path, wherein the inter-frame prediction is based on the reference data block and the one or more refined motion vectors, and wherein in the second configuration, the first interpolation filter is configured to determine an 8-tap horizontal interpolation based on the reference data block and a first refined motion vector, and the second interpolation filter is configured to determine an 8-tap vertical interpolation based on the reference data block and a second refined motion vector; and generate a plurality of inter-frame prediction pixels using the 8-tap horizontal interpolation and the 8-tap vertical interpolation.
2. The apparatus according to claim 1, wherein the at least one processor is configured to: determine a sum of absolute differences SAD based on the 2-tap horizontal interpolation and the 2-tap vertical interpolation; and generate the one or more refined motion vectors using the SAD and one or more original motion vectors associated with the video data block.
3. The apparatus according to claim 1, wherein the at least one processor is configured to: determine a weighted prediction using the 8-tap horizontal interpolation and the 8-tap vertical interpolation, wherein the weighted prediction is determined based on a sum of the 8-tap horizontal interpolation and the 8-tap vertical interpolation; and generate the plurality of inter-frame prediction pixels based on the weighted prediction.
4. The apparatus according to claim 3, wherein the at least one processor is configured to: generate the one or more refined motion vectors based on a sum of absolute differences SAD determined for the 2-tap horizontal interpolation and the 2-tap vertical interpolation; wherein the SAD is determined based on a sum of a negative value of the 2-tap vertical interpolation and the 2-tap horizontal interpolation.
5. The apparatus according to claim 4, wherein the at least one processor is configured to: use arithmetic logic to determine the weighted prediction and determine the SAD.
6. The apparatus according to claim 1, wherein: the inter-frame prediction processing path includes the first interpolation filter and the second interpolation filter; the first interpolation filter is a 2-tap 32x2 horizontal interpolation filter; and the second interpolation filter is a 2-tap 32x2 vertical interpolation filter.
7. The apparatus according to claim 6, wherein: To determine the one or more refined motion vectors, the at least one processor is configured to use the inter prediction processing path as a motion vector refinement path; and To perform inter prediction for the video data block, the at least one processor is configured to use the inter prediction processing path as a pixel prediction path.
8. The apparatus according to claim 7, wherein the at least one processor is configured to:[[]] Use the inter prediction processing path as a motion vector refinement path by configuring the first interpolation filter to perform 2-tap 20x2 horizontal interpolation and configuring the second interpolation filter to perform 2-tap 20x2 vertical interpolation; and Use the inter prediction processing path as a pixel prediction path by configuring the first interpolation filter to perform 8-tap 8x2 horizontal interpolation and configuring the second interpolation filter to perform 8-tap 8x2 vertical interpolation.
9. The apparatus according to claim 1, wherein the at least one processor is configured to:[[]] Generate an encoded video bitstream including one or more pictures, at least one of the one or more pictures being based on the inter prediction performed for the video data block.
10. The apparatus according to claim 9, wherein the at least one processor is configured to:[[]] Send the encoded video bitstream to a decoding device, the encoded video bitstream being sent together with signaling information.
11. The apparatus according to claim 9, wherein the at least one processor is configured to:[[]] Store the encoded video bitstream.
12. The apparatus according to claim 1, wherein the at least one processor is configured to:[[]] Obtain one or more encoded pictures, at least one of the one or more encoded pictures including the video data block; and Decode the video data block from the at least one encoded picture.
13. The apparatus according to claim 12,[[]] wherein,[[]] To decode the video data block from the at least one encoded picture, the at least one processor is configured to reconstruct the video data block.
14. A method of processing video data, the method comprises:[[]] Obtain a reference data block for predicting a video data block; Determine one or more refined motion vectors based on the reference data block using a first configuration of a plurality of interpolation logic units of a first interpolation filter and a second interpolation filter for implementing a shared inter prediction processing path, wherein in the first configuration, the first interpolation filter is configured to determine 2-tap horizontal interpolation based on the reference data block, and the second interpolation filter is configured to determine 2-tap vertical interpolation based on the reference data block; A second configuration of the plurality of interpolation logic units using the shared inter-frame prediction processing path performs inter-frame prediction on the video data block, wherein the inter-frame prediction is based on the reference data block and the one or more refined motion vectors, and wherein in the second configuration, the first interpolation filter is configured to determine an 8-tap horizontal interpolation based on the reference data block and a first refined motion vector, and the second interpolation filter is configured to determine an 8-tap vertical interpolation based on the reference data block and a second refined motion vector; and Generate a plurality of inter-frame predicted pixels using the 8-tap horizontal interpolation and the 8-tap vertical interpolation.
15. The method according to claim 14, the method further comprises: Determine the sum of absolute differences SAD based on the 2-tap horizontal interpolation and the 2-tap vertical interpolation; and Generate the one or more refined motion vectors using the SAD and one or more original motion vectors associated with the video data block.
16. The method according to claim 14, the method further comprises: Determine a weighted prediction using the 8-tap horizontal interpolation and the 8-tap vertical interpolation, wherein the weighted prediction is determined based on the sum of the 8-tap horizontal interpolation and the 8-tap vertical interpolation; and Generate the plurality of inter-frame predicted pixels based on the weighted prediction.
17. The method according to claim 16, the method further comprises: Generate the one or more refined motion vectors based on the sum of absolute differences SAD determined for the 2-tap horizontal interpolation and the 2-tap vertical interpolation; wherein the SAD is determined based on the sum of the negative value of the 2-tap vertical interpolation and the 2-tap horizontal interpolation.
18. The method according to claim 17, the method further comprises: Use arithmetic logic to determine the weighted prediction and determine the SAD.
19. The method according to claim 14, wherein: The inter-frame prediction processing path includes the first interpolation filter and the second interpolation filter; The first interpolation filter is a 2-tap 32x2 horizontal interpolation filter; and The second interpolation filter is a 2-tap 32x2 vertical interpolation filter.
20. The method according to claim 19, wherein: Determining the one or more refined motion vectors includes using the inter-frame prediction processing path as a motion vector refinement path; and Performing inter-frame prediction on the video data block includes using the inter-frame prediction processing path as a pixel prediction path.
21. The method according to claim 20, the method further comprises: Use the inter-frame prediction processing path as a motion vector refinement path by configuring the first interpolation filter to perform 2-tap 20x2 horizontal interpolation and configuring the second interpolation filter to perform 2-tap 20x2 vertical interpolation; and Use the inter-frame prediction processing path as a pixel prediction path by configuring the first interpolation filter to perform 8-tap 8x2 horizontal interpolation and configuring the second interpolation filter to perform 8-tap 8x2 vertical interpolation.
22. The method according to claim 14, the method further comprises: generating an encoded video bitstream including one or more pictures, at least one of the one or more pictures being based on the inter prediction performed on the video data block.
23. The method according to claim 22, the method further comprises: sending the encoded video bitstream to a decoding device, the encoded video bitstream being sent together with signaling information.
24. The method according to claim 22, the method further comprises: storing the encoded video bitstream.
25. The method according to claim 14, the method further comprises: obtaining one or more encoded pictures, at least one of the one or more encoded pictures including the video data block; and decoding the video data block from the at least one encoded picture.
26. The method according to claim 25, wherein decoding the video data block from the at least one encoded picture comprises reconstructing the video data block.