Block-level collocated motion field projection for video coding
By optimizing block-level processing and hardware buffers, the problems of excessive read/write bandwidth and buffer size in juxtaposed motion field projection are solved, thus improving the efficiency of video decoding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QUALCOMM INC
- Filing Date
- 2023-07-07
- Publication Date
- 2026-04-24
AI Technical Summary
Existing video decoding technologies suffer from excessively high read/write bandwidth and buffer size requirements when processing juxtaposed motion field projections, which affects the decoding efficiency of video data.
By adopting a block-level processing approach, the motion vectors of the juxtaposed reference frames are projected into a block-level motion field, and a block-level buffer is used in the hardware for storage and decoding, which reduces the number of read/write operations and the buffer size.
It effectively reduces the number of clock cycles for juxtaposed motion vector projection, lowers read/write bandwidth requirements, and improves the decoding efficiency of video data.
Smart Images

Figure CN119698834B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates generally to video encoding and decoding. For example, aspects of this disclosure include improved video decoding techniques associated with juxtaposed sports field projections. Background Technology
[0002] Digital video capabilities can be incorporated into a wide variety of devices, including digital televisions, digital live broadcast systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, tablets, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite wireless phones (so-called "smartphones"), video conferencing equipment, video streaming devices, and more. These devices allow video data to be processed and output for consumption. Digital video data comprises vast amounts of data to meet the needs of both consumers and video providers. For example, consumers of video data expect the highest quality video, with high fidelity, high resolution, and high frame rates. As a result, the large amounts of video data required to meet these needs place a burden on the communication networks and devices that process and store the video data.
[0003] Digital video devices implement video decoding techniques to compress video data. Video decoding is performed according to one or more video decoding standards or formats. Examples of video decoding standards or formats include Universal Video Decoding (VVC), High-Efficiency Video Decoding (HEVC), Advanced Video Decoding (AVC), MPEG-2 Part 2 Decoding (MPEG stands for Moving Picture Experts Group), and proprietary video codecs / formats such as AOMedia Video 1 (AV1) developed by the Alliance for Open Media. Video decoding typically utilizes predictive methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that take advantage of redundancy present in video images or sequences. The goal of video decoding technology is to compress video data into a form using a lower bitrate while avoiding or minimizing video quality degradation. As more and more video services become available, there is a need for decoding technologies with better decoding efficiency. Summary of the Invention
[0004] In some examples, systems and techniques for block-level juxtaposed motion field projection are described. For example, these systems and techniques read motion vectors associated with one or more juxtaposed reference frames and project them into a block-level motion field. The one or more juxtaposed reference frames may be associated with a current (e.g., currently decoded) frame of video data. Each corresponding block-level motion field may be associated with a block included in that current frame of video data. The block-level motion field and the block included in that current frame of video data may have the same relative position relative to the juxtaposed reference frames and the current block of video data, respectively. The projected block-level motion field may be stored in one or more buffers. The current block of video data may be decoded concurrently with the motion field associated with it for one or more subsequent blocks of video data.
[0005] According to at least one exemplary example, an apparatus for processing video data is provided, the apparatus including at least one memory (e.g., the at least one memory configured to store data such as virtual content data, one or more images, etc.) and at least one processor coupled to the at least one memory (e.g., the at least one processor is implemented in a circuit). The at least one processor is configured to: obtain one or more first sets of juxtaposed motion vector data associated with a first block of video data included in a current frame of the video data; project each of the one or more first sets of juxtaposed motion vector data into a first projection motion field associated with a first buffer; obtain one or more second sets of juxtaposed motion vector data associated with a second block of video data included in the current frame of the video data; project each of the one or more second sets of juxtaposed motion vector data into the first projection motion field associated with the first buffer; and decode the first block of video data based on the first projection motion field associated with the first buffer, based on each of the one or more first sets of juxtaposed motion vector data and each of the one or more second sets of juxtaposed motion vector data projected into the first projection motion field associated with the first buffer.
[0006] In another example, a method for processing video data is provided, the method comprising: obtaining one or more first sets of juxtaposed motion vector data associated with a first block of video data included in a current frame of the video data; projecting each of the one or more first sets of juxtaposed motion vector data into a first projection motion field associated with a first buffer; obtaining one or more second sets of juxtaposed motion vector data associated with a second block of video data included in the current frame of the video data; projecting each of the one or more second sets of juxtaposed motion vector data into the first projection motion field associated with the first buffer; and decoding the first block of video data based on the first projection motion field associated with the first buffer, based on each of the one or more first sets of juxtaposed motion vector data and each of the one or more second sets of juxtaposed motion vector data projected into the first projection motion field associated with the first buffer.
[0007] In another example, a non-transitory computer-readable medium is provided having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to: obtain one or more first sets of juxtaposed motion vector data associated with a first block of video data included in a current frame of video data; project each of the one or more first sets of juxtaposed motion vector data into a first projection motion field associated with a first buffer; obtain one or more second sets of juxtaposed motion vector data associated with a second block of video data included in the current frame of video data; project each of the one or more second sets of juxtaposed motion vector data into the first projection motion field associated with the first buffer; and decode the first block of video data based on the first projection motion field associated with the first buffer, based on each of the one or more first sets of juxtaposed motion vector data and each of the one or more second sets of juxtaposed motion vector data projected into the first projection motion field associated with the first buffer.
[0008] In another example, an apparatus for processing video data is provided, the apparatus comprising: means for obtaining one or more first sets of juxtaposed motion vector data associated with a first block of video data included in a current frame of the video data; means for projecting each of the one or more first sets of juxtaposed motion vector data into a first projection motion field associated with a first buffer; means for obtaining one or more second sets of juxtaposed motion vector data associated with a second block of video data included in the current frame of the video data; means for projecting each of the one or more second sets of juxtaposed motion vector data into the first projection motion field associated with the first buffer; and means for decoding the first block of video data based on the first projection motion field associated with the first buffer, based on the projection of each of the one or more first sets of juxtaposed motion vector data and each of the one or more second sets of juxtaposed motion vector data into the first projection motion field associated with the first buffer.
[0009] In some aspects, one or more of the devices described herein are, are part of, and / or include: mobile devices or wireless communication devices (e.g., mobile phones or other mobile devices), extended reality (XR) devices or systems (e.g., virtual reality (VR) devices, augmented reality (AR) devices, or mixed reality (MR) devices), wearable devices (e.g., network-connected watches or other wearable devices), cameras, personal computers, laptop computers, vehicles or computing devices or components of vehicles, server computers or server equipment, other devices, or combinations thereof. In some aspects, the device includes one or more cameras for capturing one or more images. In some aspects, the device also includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the aforementioned devices may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyroscope testers, one or more accelerometers, any combination thereof, and / or other sensors).
[0010] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to define the scope of the claimed subject matter. This subject matter should be understood with reference to the appropriate portions of the entire specification, any or all drawings, and each claim.
[0011] The foregoing and other features and aspects will become more apparent from the following description, claims and accompanying drawings. Attached Figure Description
[0012] The exemplary aspects of this application are described in detail below with reference to the following figures:
[0013] Figure 1 This is a block diagram illustrating examples of encoding and decoding devices according to some examples of this disclosure;
[0014] Figure 2A This is a conceptual diagram illustrating example space adjacent motion vector candidates for merging patterns, according to some examples of this disclosure;
[0015] Figure 2B This is a conceptual diagram illustrating example space neighboring motion vector candidates for an advanced motion vector prediction (AMVP) pattern, according to some examples of this disclosure;
[0016] Figure 3A This is a conceptual diagram illustrating example time motion vector predictor (TMVP) candidates according to some examples of this disclosure;
[0017] Figure 3B This is a conceptual diagram illustrating examples of motion vector scaling according to some examples of this disclosure;
[0018] Figure 4A This is a conceptual diagram illustrating examples of neighboring samples of the current decoding unit for estimating motion compensation parameters of the current decoding unit, according to some examples of this disclosure;
[0019] Figure 4B This is a conceptual diagram illustrating an example of a neighboring sample of a reference block used to estimate motion compensation parameters of the current decoding unit, according to some examples of this disclosure;
[0020] Figure 5 Examples of aspects of scaling motion vectors for time-merging candidates used in processing blocks are illustrated according to some examples of this disclosure;
[0021] Figure 6A Examples of juxtaposed motion vector projections are shown according to some examples of this disclosure;
[0022] Figure 6B Examples of available types of juxtaposed reference frames according to some examples of this disclosure;
[0023] Figures 7A to 7D Examples of overriding juxtaposed motion vector projections and existing motion vector field projections according to some examples of this disclosure;
[0024] Figure 8A Examples of frame-level raster scan sequences that can be used to perform juxtaposed motion vector projections, according to some examples of this disclosure;
[0025] Figure 8BExample diagrams illustrating a system for performing juxtaposed motion vector projection based on frame-level raster scanning, according to some examples of this disclosure;
[0026] Figure 9A Examples of block-level raster scan sequences that can be used to perform juxtaposed motion vector projections, according to some examples of this disclosure;
[0027] Figure 9B Example diagrams illustrating a system for performing juxtaposed motion vector projection based on block-level processing, according to some examples of this disclosure;
[0028] Figure 10 This is an illustration illustrating an example of a block-based juxtaposed motion vector projection using a sliding motion vector buffer, according to some examples of this disclosure;
[0029] Figure 11 This is an illustration illustrating an example of a block-based juxtaposed motion vector projection using a circular motion vector buffer according to some examples of this disclosure;
[0030] Figure 12 This is an illustration illustrating an example of block-based juxtaposed motion vector projection using a ping-pong motion vector buffer according to some examples of this disclosure;
[0031] Figure 13 This is an illustration illustrating an example of block-based juxtaposed motion vector projection using a sliding motion vector buffer, wherein the projection is performed in reverse N order and the decoding is performed in Z order, according to some examples of this disclosure.
[0032] Figure 14 This is an illustration illustrating an example of a row-based index used to overwrite a conflict associated with a block-based juxtaposed motion vector projection, according to some examples of this disclosure;
[0033] Figure 15 This is a flowchart illustrating an example process for video decoding using block-based juxtaposed motion vector projection, according to some examples of this disclosure;
[0034] Figure 16 This is a block diagram illustrating an example video encoding device according to some examples of this disclosure; and
[0035] Figure 17 This is a block diagram illustrating an example video decoding device according to some examples of this disclosure. Detailed Implementation
[0036] Certain aspects and facets of this disclosure are provided below. Some of these aspects and facets may be applied independently, and some may be applied in combination, as will be apparent to those skilled in the art. Specific details are set forth in the following description for purposes of explanation to provide a thorough understanding of the aspects of this application. However, it will be apparent that aspects may be practiced without these specific details. The accompanying drawings and descriptions are not intended to be limiting.
[0037] The following description provides only exemplary aspects and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of the exemplary aspects will provide those skilled in the art with a description that can be used to implement the exemplary aspects. It should be understood that various changes can be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.
[0038] Video decoding devices (e.g., encoding devices, decoding devices, or combined encoder-decoder devices) implement video compression techniques to efficiently decode (e.g., encode and / or decode) video data. Video compression techniques may include applying different prediction modes, including spatial prediction (e.g., intra-frame prediction or intra-prediction), temporal prediction (e.g., inter-frame prediction or inter-prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques for reducing or removing inherent redundancy in a video sequence. A video encoder may divide each frame of the original video sequence into rectangular regions, referred to as video blocks or decoding units (described in more detail below). These video blocks may be encoded using specific prediction modes.
[0039] Video blocks can be divided into one or more smaller blocks in one or more ways. Blocks may include decode tree blocks, prediction blocks, transform blocks, or other suitable blocks. Unless otherwise specified, the reference to “block” generally refers to such video blocks (e.g., decode tree blocks, decode blocks, prediction blocks, transform blocks, or other suitable blocks or sub-blocks, as will be understood by one of ordinary skill in the art). Furthermore, each of these blocks may also be interchangeably referred to herein as a “unit” (e.g., decode tree unit (CTU), decode unit, prediction unit (PU), transform unit (TU), etc.). In some cases, a unit may refer to a decoded logic unit encoded in the bitstream, while a block may refer to a portion of the process targeted in the video frame buffer.
[0040] For inter-frame prediction mode, the video encoder searches for blocks similar to the block to be encoded in a frame (or picture) located at another time position (called a reference frame or reference picture). The video encoder can restrict the search to a certain spatial displacement from the block to be encoded. The best match can be located using two-dimensional (2D) motion vectors that include horizontal and vertical displacement components. For intra-frame prediction mode, the video encoder can use spatial prediction techniques to form the prediction block based on data from previously encoded adjacent blocks within the same picture.
[0041] In some cases, juxtaposed motion vector projection can be performed based on some (or all) of the time motion vectors in the set of time motion vectors determined relative to the current frame and its reference frames projected onto the starting frame. In some cases, a motion vector projection that can be located within and / or outside the given frame x can be generated using the time motion vectors determined for a given frame with index x. In some examples, motion vector projection can be invalidated (e.g., rejected or not selected) if the location of the motion vector projection is not included in one of the following: the given frame from which the motion vector projection originates, the horizontally adjacent frame to the left of the given frame, and the horizontally adjacent frame to the right of the given frame. In some examples, juxtaposed motion vector projection can be performed for up to three juxtaposed reference frames (referred to as REF0, REF1, and REF2). Based on having the same projection end position, the motion vector associated with REF0 can be projected first and can be overwritten by the motion vectors projected later with REF1 and / or REF2.
[0042] In some examples, frame-level raster scan order can be used to perform juxtaposed motion vector projection. For example, for a given frame of video data, all REF0 motion vectors can be projected in frame-by-frame raster scan order, followed by the projection of all REF1 and REF2 motion vectors (e.g., sequentially, and each in frame-by-frame raster scan order). Performing juxtaposed motion vector projection by processing three juxtaposed reference frames in frame-level raster scan order can be associated with four read operations and four write operations on the juxtaposed data. For example, the four read operations may include three read operations for reading REF0, REF1, and REF2 motion vector data (e.g., before performing the corresponding motion vector projection for each reference frame) and a fourth read operation for reading the final contents of the projected motion vector buffer. The four write operations may include three write operations for writing the REF0, REF1, and REF2 motion vector projections to the projected motion vector buffer and a fourth write operation for writing the final motion vector prediction of the current frame to memory (e.g., based on the combined motion vector projection field stored in the projected motion vector buffer).
[0043] It is necessary to reduce the number of clock cycles used to perform juxtaposition motion vector projection on one or more (e.g., up to three) juxtaposition reference frames. It is necessary to reduce the read / write bandwidth associated with decoding frames of video data using juxtaposition motion vector projection. It is also necessary to reduce the size (e.g., storage capacity) of one or more buffers associated with the projection motion vectors determined during the performance of juxtaposition motion vector projection.
[0044] This document describes systems, apparatuses, processes (also referred to as methods), and computer-readable media (collectively referred to herein as "systems and techniques") for performing juxtaposed motion vector projection using block-level processing. In some aspects, the methods described herein can be used to perform instantaneous (e.g., real-time or near real-time) juxtaposed motion vector projection according to a block-based processing order. In some examples, the systems and techniques can be used to read and project motion vectors from reference frames based on block-level scanning. In some aspects, the current decoded frame of video data may be associated with three juxtaposed reference frames (e.g., REF0, REF1, REF2). The associated motion vector data included in the three juxtaposed reference frames can be read and projected at the 64×64 block level. For example, the current decoded frame of video data can be divided into multiple 64×64-sized blocks, and motion vectors can be obtained from the three juxtaposed reference frames (e.g., REF0, REF1, REF2) using the same 64×64 block size. In some cases, the associated motion vector data included in the juxtaposed reference frames can be read and projected at the 128×128 block level. In some examples, the system and techniques can avoid writing the projected motion vectors to DDR or other memory based on the block-based processing order described herein for performing juxtaposed motion vector projection.
[0045] In some aspects, the system and technology may include one or more MV projection buffers implemented in hardware, wherein the size of the one or more MV projection buffers is set to be three times the block level size (e.g., for a 64×64 block level, the size of the one or more MV projection buffers may be set to store three 64×64 blocks; for a 128×128 block level, the size of the one or more MV projection buffers may be set to store three 128×128 blocks).
[0046] In some examples, a given block of video data can be decoded in parallel with a juxtaposed motion vector projection operation performed on one or more additional blocks and / or subsequent blocks of video data. For example, additional blocks and / or subsequent blocks of video data may be horizontally adjacent to the right side of the currently decoded block of video data (e.g., where two blocks are included in the same frame of video data). In some cases, a given block of video data can be decoded based on determining that MV projection has been performed on the given block, where decoding of subsequent blocks of video data is performed independently of the MV projection operation.
[0047] The techniques described herein can be implemented using one or more decoding devices (including one or more encoding devices, decoding devices, or combinations of encoding-decoding devices). Decoding devices can be implemented using one or more of the following: player devices (such as mobile devices), extended reality (XR) devices, vehicles or computing systems of vehicles, server devices or systems (e.g., distributed server systems including multiple servers, single server devices or systems, etc.), or other devices or systems.
[0048] The systems and techniques described herein can be applied to any existing video codec, any video codec under development, and / or any future video decoding standard, including but not limited to High Efficiency Video Decoding (HEVC), Advanced Video Decoding (AVC), Universal Video Decoding (VVC), VP9, AOMedia Video 1 (AV1) format / codec, and / or other existing, under development, or pending video decoding standards. The systems and techniques described herein can improve the operation of communication systems and devices within the system by enhancing the performance of video data transmission by devices with improved compression and associated improved video quality based on improved motion vector selection from adaptive bilateral matching as described herein.
[0049] Various aspects of the system and technology will be described with reference to the accompanying drawings.
[0050] Figure 1 This is a block diagram illustrating an example of a system 100 including an encoding device 104 and a decoding device 112. The encoding device 104 may be part of a source device, and the decoding device 112 may be part of a receiving device. The source device and / or receiving device may include electronic devices such as mobile or landline phones (e.g., smartphones, cellular phones, etc.), desktop computers, laptops or notebook computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, Internet Protocol (IP) cameras, or any other suitable electronic devices. In some examples, the source device and receiving device may include one or more wireless transceivers for wireless communication. The decoding techniques described herein are applicable to video decoding in a variety of multimedia applications, including streaming video (e.g., via the Internet), television broadcasting or transmission, encoding digital video for storage on a data storage medium, decoding digital video stored on a data storage medium, or other applications. As used herein, the term decoding may refer to encoding and / or decoding. In some examples, system 100 may support one-way or two-way video transmission to support applications such as video conferencing, video streaming, video playback, video broadcasting, gaming, and / or video telephony.
[0051] Encoding device 104 (or encoder) can be used to encode video data using video decoding standards, formats, codecs, or protocols to generate an encoded video bitstream. Examples of video decoding standards and formats / codecs include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) (including its Scalable Video Decoding (SVC) and Multi-View Video Decoding (MVC) extensions), High-Efficiency Video Decoding (HEVC) or ITU-T H.265, and Universal Video Decoding (VVC) or ITU-T H.266. Various extensions to HEVC are available for multi-layer video decoding, including range and screen content decoding extensions, 3D video decoding (3D-HEVC), multi-view extensions (MV-HEVC), and scalable extensions (SHVC). HEVC and its extensions have been developed by the Joint Collaborative Team for Video Decoding (JCT-VC) and the Joint Collaborative Team for the Development of 3D Video Decoding Extensions of the ITU-T Video Decoding Experts Group (VCEG) and the ISO / IEC Animation Experts Group (MPEG). VP9, AOMedia Video 1 (AV1), and Basic Video Decoding (EVC), developed by the Alliance for Open Media (AOMedia), are other video decoding standards to which the technologies described herein can be applied.
[0052] The techniques described herein can be applied to any of existing video codecs (e.g., High Efficiency Video Decoding (HEVC), Advanced Video Decoding (AVC), or other suitable existing video codecs), and / or can be efficient decoding tools for any video decoding standard under development and / or future video decoding standards (e.g., VVC and / or other video decoding standards under development or to be developed). For example, the examples described herein can be performed using video codecs such as VVC, HEVC, AVC, and / or their extensions. However, the techniques and systems described herein can also be applied to other decoding standards, codecs, or formats such as MPEG, JPEG (or other decoding standards for still images), VP9, AV1, their extensions, or other suitable decoding standards that are already available, not yet available, or not yet developed. For example, in some examples, encoding device 104 and / or decoding device 112 can operate according to proprietary video codecs / formats (such as AV1, extensions of AVI, and / or successors to AV1 (e.g., AV2)) or other proprietary formats or industry standards. Therefore, although the techniques and systems described herein may be described with reference to specific video decoding standards, it will be understood by those skilled in the art that the description should not be construed as applicable only to that particular standard.
[0053] refer to Figure 1 Video source 102 can provide video data to encoding device 104. Video source 102 can be part of a source device or part of a device other than a source device. Video source 102 may include video capture devices (e.g., cameras, camera phones, video phones, etc.), video archives containing stored video, video servers or content providers that provide video data, video feed interfaces that receive video from video servers or content providers, computer graphics systems for generating computer graphics video data, combinations of such sources, or any other suitable video source.
[0054] Video data from video source 102 may include one or more input pictures or frames. A picture or frame is a still image, which in some cases is part of the video. In some examples, the data from video source 102 may be a still image that is not part of the video. In HEVC, VVC, and other video decoding specifications, a video sequence may include a series of pictures. A picture may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luminance samples, SCb is a two-dimensional array of Cb chrominance samples, and SCr is a two-dimensional array of Cr chrominance samples. Chrominance samples may also be referred to herein as “chroma” samples. A pixel may refer to all three components (luminance and chrominance samples) at a given location in the array of pictures. In other instances, a picture may be monochrome and may consist only of an array of luminance samples; in this case, the terms pixel and sample are used interchangeably. The same techniques described herein, which reference various samples for illustrative purposes, can be applied to pixels (e.g., all three sample components at a given location in the array of pictures). The example techniques described herein, which refer to pixels (e.g., all three sample components at a given location in an array of images) for illustrative purposes, can be applied to individual samples.
[0055] Encoding device 104's encoder engine 106 (or encoder) encodes the video data to generate an encoded video bitstream. In some examples, the encoded video bitstream (or "video bitstream" or "bitstream") is a series of one or more decoded video sequences. The decoded video sequence (CVS) includes a series of access units (AUs) that begin with an AU that has a random access point picture in the base layer and possesses certain attributes, and continue until the next AU that has a random access point picture in the base layer and possesses certain attributes, but does not include that next AU. For example, certain attributes of the random access point picture that begins the CVS may include a RASL flag equal to 1 (e.g., NoRaslOutputFlag). Otherwise, a random access point picture (with a RASL flag equal to 0) does not begin the CVS. An access unit (AU) includes one or more decoded pictures and corresponding control information for decoded pictures sharing the same output time. Decoded slices of pictures are encapsulated into data units at the bitstream level, which are called Network Abstraction Layer (NAL) units. For example, an HEVC video bitstream may include one or more CVSs, which include NAL units. Each NAL unit has a NAL unit header. In one example, this header is one byte for H.264 / AVC (except for multi-layer extensions) and two bytes for HEVC. The syntax elements in the NAL unit header use specified bits and are therefore visible to all kinds of systems and transport layers, such as transport streams, Real-Time Transport (RTP) protocols, file formats, etc.
[0056] The HEVC standard contains two types of NAL units: Video Decoding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units contain decoded picture data that forms the decoded video bitstream. For example, the bit sequence that forms the decoded video bitstream is present in a VCL NAL unit. A VCL NAL unit may include a slice or fragment of the decoded picture data (described below), and non-VCL NAL units contain control information relating to one or more decoded pictures. In some cases, NAL units may be referred to as packets. A HEVC AU includes VCL NAL units containing decoded picture data and non-VCL NAL units (if any) corresponding to the decoded picture data. Among other information, non-VCL NAL units may also contain a set of parameters with high-level information relating to the encoded video bitstream. For example, the parameter set may include a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), and a Picture Parameter Set (PPS). In some cases, each slice or other portion of the bitstream may reference a single valid PPS, SPS, and / or VPS to allow the decoding device 112 to access information that can be used to decode the slice or other portion of the bitstream.
[0057] NAL units can include bit sequences that form a decoded representation of video data (e.g., an encoded video bitstream, a CVS of a bitstream, etc.), such as a decoded representation of images in a video. Encoder engine 106 generates a decoded representation of an image by dividing each image into multiple slices. Slices are independent of other slices, allowing information within that slice to be decoded without relying on data from other slices within the same image. A slice includes one or more segments, comprising independent segments and (if present) one or more dependent segments that depend on previous segments.
[0058] In HEVC, the slice is then divided into decoder tree blocks (CTBs) for luma and chroma samples. One or more CTBs for luma samples and one CTB for chroma samples, along with the syntax used for the samples, are called decoder tree units (CTUs). CTUs are also referred to as "tree blocks" or "maximum decoder units" (LCUs). The CTU is the basic processing unit used for HEVC encoding. A CTU can be subdivided into multiple decoder units (CUs) of different sizes. Each CU contains an array of luma samples and an array of chroma samples called a decoder block (CB).
[0059] Luminance CBs and chrominance CBs can be further subdivided into prediction blocks (PBs). A PB is a sample block of the luminance or chrominance component that uses the same motion parameters for inter-frame prediction or intra-block copy (IBC) prediction (when available or enabled). A luminance PB and one or more chrominance PBs, along with their associated syntax, form a prediction unit (PU). For inter-frame prediction, the set of motion parameters (e.g., one or more motion vectors, reference indices, etc.) is signaled in the bitstream for each PU, as well as for inter-frame prediction of the luminance PB and one or more chrominance PBs. Motion parameters can also be referred to as motion information. CBs can also be subdivided into one or more transform blocks (TBs). A TB represents a square block of samples of the chrominance component, to which a residual transform (e.g., in some cases, the same two-dimensional transform) is applied to decode the prediction residual signal. A transform unit (TU) represents the TB of the luminance and chrominance samples and the corresponding syntax elements. Transform decoding is described in more detail below.
[0060] The size of a CU corresponds to the size of the decoding mode and can be square. For example, the size of a CU can be 8×8 samples, 16×16 samples, 32×32 samples, 64×64 samples, or any other suitable size up to the size of the corresponding CTU. The phrase “N×N” is used herein to refer to the pixel size of a video block in both the vertical and horizontal dimensions (e.g., 8 pixels × 8 pixels). Pixels in a block can be arranged in rows and columns. In some implementations, a block may not have the same number of pixels horizontally as it does vertically. The syntax data associated with a CU can describe, for example, the segmentation of the CU into one or more PUs. The segmentation mode may vary depending on whether the CU is coded using intra-frame prediction mode or inter-frame prediction mode. PUs can be segmented into non-square shapes. The syntax data associated with a CU can also describe, for example, the segmentation of the CU into one or more TUs according to the CTU. TUs can be square or non-square.
[0061] According to the HEVC standard, transform units (TUs) can be used to perform transforms. TUs can vary for different core cells (CUs). The size of the TU can be set based on the size of the pixel unit (PU) within a given CU. A TU can have the same size as or smaller than the PU. In some examples, a quadtree structure called a residual quadtree (RQT) can be used to subdivide the residual samples corresponding to the CU into smaller units. The leaf nodes of the RQT can correspond to TUs. The pixel differences associated with the TU can be transformed to produce transform coefficients. The transform coefficients can then be quantized by the encoder engine 106.
[0062] Once the video data is segmented into Units (CUs), the encoder engine 106 uses a prediction mode to predict each Processing Unit (PU). The prediction unit or block is then subtracted from the original video data to obtain the residual (described below). For each CU, the prediction mode can be signaled within the bitstream using syntax data. Prediction modes can include intra-frame prediction (or intra-picture prediction) or inter-frame prediction (or inter-picture prediction). Intra-frame prediction utilizes the correlation between spatially adjacent samples within a picture. For example, using intra-frame prediction, each PU is predicted from neighboring image data in the same picture using, for example, DC prediction to find the average value for the PU, planar prediction to adapt a planar surface to the PU, orientation prediction to infer from neighboring data, or any other suitable prediction type. Inter-frame prediction uses temporal correlations between pictures to derive motion-compensated predictions for blocks of image samples. For example, using inter-frame prediction, each PU is predicted from image data in one or more reference pictures (in the output order before or after the current picture) using motion-compensated prediction. For example, a decision can be made at the CU level whether to use inter-picture prediction or intra-picture prediction to decode a picture region.
[0063] Encoder engine 106 and decoder engine 116 (described in more detail below) can be configured to operate according to VVC. According to VVC, the video decoder (such as encoder engine 106 and / or decoder engine 116) segments the image into multiple decoder tree units (CTUs) (where one or more CTBs of luminance samples and chrominance samples, together with the syntax used for the samples, are referred to as CTUs). The video decoder can segment CTUs according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure removes the concept of multiple partition types, such as the separation between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels: a first level of segmentation based on quadtree segmentation and a second level of segmentation based on binary tree segmentation. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to decoder units (CUs).
[0064] In the MTT partitioning structure, blocks can be partitioned using quadtree partitioning, binary tree partitioning, and one or more types of ternary tree partitioning. A ternary tree partitioning is a partition where a block is split into three sub-blocks. In some examples, a ternary tree partitioning divides a block into three sub-blocks without partitioning the original block through a center. The partitioning type in MTT (e.g., quadtree, binary tree, and ternary tree) can be symmetric or asymmetric.
[0065] When operating according to the AV1 codec, encoding device 104 and decoding device 112 can be configured to decode video data in blocks. In AV1, the largest decoded block that can be processed is called a superblock. In AV1, a superblock can be 128×128 luminance samples or 64×64 luminance samples. However, in subsequent video decoding formats (e.g., AV2), superblocks can be defined by different (e.g., larger) luminance sample sizes. In some examples, the superblock is the top level of a block quadtree. Encoding device 104 can further divide the superblock into smaller decoded blocks. Encoding device 104 can use square or non-square partitioning to divide the superblock and other decoded blocks into smaller blocks. Non-square blocks can include N / 2×N blocks, N×N / 2 blocks, N / 4×N blocks, and N×N / 4 blocks. Encoding device 104 and decoding device 112 can perform separate prediction and transform processing on each decoded block within the decoded block.
[0066] AV1 also defines tiles for video data. A tile is a rectangular array of superblocks that can be decoded independently of other tiles. That is, the encoding device 104 and the decoding device 112 can encode and decode the decoding blocks within a tile separately without using video data from other tiles. However, the encoding device 104 and the decoding device 112 can perform filtering across tile boundaries. The size of the tiles can be uniform or non-uniform. Tile-based decoding enables parallel processing and / or multithreading in the encoder and decoder implementations.
[0067] In some examples, the video decoder may use a single QTBT or MTT structure to represent each of the luminance and chrominance components, while in other examples, the video decoder may use two or more QTBT or MTT structures, such as one QTBT or MTT structure for the luminance component and another QTBT or MTT structure for the two chrominance components (or two QTBT and / or MTT structures for the respective chrominance components).
[0068] The video decoder can be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, superblock partitioning, or other partitioning structures.
[0069] In some examples, a slice type is assigned to one or more slices of an image. Slice types include intra-frame decoded slices (I-slices), inter-frame decoded P-slices, and inter-frame decoded B-slices. An I-slice (an intra-frame decoded frame, independently decodeable) is a slice of an image that is decoded only by intra-frame prediction and is therefore independently decodeable because an I-slice only requires intra-frame data to predict any prediction unit or prediction block of the slice. A P-slice (a one-way prediction frame) is a slice of an image that can be decoded using both intra-frame prediction and one-way inter-frame prediction. Each prediction unit or prediction block within a P-slice is decoded using either intra-frame or inter-frame prediction. When inter-frame prediction is applied, the prediction unit or prediction block is predicted by only one reference image, and therefore the reference sample comes from only one reference region within a frame. A B-slice (a two-way prediction frame) is a slice of an image that can be decoded using both intra-frame and inter-frame prediction (e.g., either two-way or one-way prediction). Prediction units or blocks of a B-slice can be bidirectionally predicted from two reference images, where each image contributes to a reference region and the sample sets of the two reference regions are weighted (e.g., with equal weights or with different weights) to generate the prediction signal for the bidirectional prediction block. As described above, slices of an image are decoded independently. In some cases, an image can be decoded into only one slice.
[0070] As mentioned above, intra-image prediction leverages the correlation between spatially adjacent samples within an image. Several intra-prediction modes exist (also referred to as "intra-frame modes"). In some examples, intra-frame prediction for a luma patch includes 35 modes, comprising a planar mode, a DC mode, and 33 angular modes (e.g., a diagonal intra-prediction mode and angular modes adjacent to the diagonal intra-prediction mode). These 35 intra-prediction modes can be indexed. In other examples, more intra-frame modes can be defined, including prediction angles that may not yet be represented by the 33 angular modes. In other examples, the prediction angles associated with angular modes may differ from those used in HEVC.
[0071] Inter-image prediction uses temporal correlations between images to derive motion-compensated predictions for image sample patches. Using a translational motion model, the position of a patch in a previously decoded image (reference image) is indicated by a motion vector (Δx, Δy), where Δx specifies the horizontal displacement of the reference patch relative to the current patch's position, and Δy specifies the vertical displacement of the reference patch relative to the current patch's position. In some cases, the motion vector (Δx, Δy) may have integer sample accuracy (also known as integer accuracy), in which case the motion vector points to an integer pixel grid (or integer pixel sample grid) of the reference frame. In other cases, the motion vector (Δx, Δy) may have fractional sample accuracy (also known as fractional pixel accuracy or non-integer accuracy) to more accurately capture the movement of underlying objects, not limited to an integer pixel grid of the reference frame.
[0072] The accuracy of a motion vector can be represented by its quantization level. For example, the quantization level can be integer accuracy (e.g., 1 pixel) or fractional pixel accuracy (e.g., ...). 1 / 4 pixels, 1 ( / 2 pixel or other sub-pixel value). When the corresponding motion vector has fractional sample accuracy, interpolation is applied to the reference image to derive the predicted signal. For example, samples available at integer positions can be filtered (e.g., using one or more interpolation filters) to estimate values at fractional positions. The previously decoded reference image is indicated by a reference index (refIdx) in the list of reference images. The motion vector and reference index can be referred to as motion parameters. Two types of inter-image prediction can be performed, including one-way prediction and two-way prediction.
[0073] In the case of inter-frame prediction using bidirectional prediction (also known as bidirectional inter-frame prediction), two sets of motion parameters (Δx0, y0, refIdx0 and Δx1, y1, refIdx1) are used to generate two motion-compensated predictions (from the same reference image or possibly from different reference images). For example, in the case of bidirectional prediction, each prediction block uses two motion-compensated prediction signals and generates B prediction units. The two motion-compensated predictions are then combined to obtain the final motion-compensated prediction. For example, the two motion-compensated predictions can be combined by averaging. In another example, weighted prediction can be used, in which case different weights can be applied to each motion-compensated prediction. The reference images that can be used in bidirectional prediction are stored in two separate lists, denoted as list 0 and list 1. The motion parameters can be derived at the encoder using a motion estimation process.
[0074] In the case of inter-frame prediction using unidirectional prediction (also known as one-way inter-frame prediction), motion-compensated predictions are generated from a reference image using a set of motion parameters (Δx0, y0, refIdx0). For example, in the case of unidirectional prediction, each prediction block uses at most one motion-compensated prediction signal and generates P prediction units.
[0075] The PU (Programming Component) may include data related to the prediction process (e.g., motion parameters or other suitable data). For example, when the PU uses intra-frame prediction for encoding, it may include data describing the intra-frame prediction mode used by the PU. As another example, when the PU uses inter-frame prediction for encoding, it may include data defining the motion vectors used by the PU. The data defining the motion vectors used by the PU may describe, for example, the horizontal component (Δx) of the motion vector, the vertical component (Δy) of the motion vector, the resolution used for the motion vector (e.g., integer precision, quarter-pixel precision, or eighth-pixel precision), the reference picture to which the motion vector points, the reference index, a list of reference pictures used for the motion vector (e.g., list 0, list 1, or list C), or any combination thereof.
[0076] AV1 includes two common techniques for encoding and decoding blocks of video data. These two common techniques are intra-frame prediction (e.g., intra-frame prediction or spatial prediction) and inter-frame prediction (e.g., inter-frame prediction or temporal prediction). In the context of AV1, when using intra-frame prediction modes to predict blocks of video data for the current frame, the encoding device 104 and the decoding device 112 do not use video data from other frames of the video data. For most intra-frame prediction modes, the video encoding device 104 encodes the blocks of the current frame based on the difference between sample values in the current block and predicted values generated from reference samples in the same frame. The video encoding device 104 determines the predicted values generated from the reference samples based on the intra-frame prediction mode.
[0077] After performing prediction using intra-frame prediction and / or inter-frame prediction, encoding device 104 may perform transform and quantization. For example, after prediction, encoder engine 106 may compute a residual value corresponding to the PU. The residual value may include the pixel difference between the current pixel block (PU) being decoded and the prediction block used to predict the current block (e.g., a predicted version of the current block). For example, after generating a prediction block (e.g., publishing inter-frame prediction or intra-frame prediction), encoder engine 106 may generate a residual block by subtracting the prediction block generated by the prediction unit from the current block. The residual block includes a set of pixel differences that quantize the differences between the pixel values of the current block and the pixel values of the prediction block. In some examples, the residual block may be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of pixel values). In such examples, the residual block is a two-dimensional representation of the pixel values.
[0078] Block transforms are used to transform any residual data that may remain after prediction is performed. These block transforms can be based on discrete cosine transforms, discrete sine transforms, integer transforms, wavelet transforms, other suitable transform functions, or any combination thereof. In some cases, one or more block transforms (e.g., of size 32×32, 16×16, 8×8, 4×4, or other suitable sizes) can be applied to the residual data in each CU. In some aspects, TUs can be used for transform and quantization processes implemented by encoder engine 106. A given CU with one or more PUs may also include one or more TUs. As described further in detail below, residual values can be transformed into transform coefficients using block transforms, and then quantized and scanned using TUs to produce serialized transform coefficients for entropy decoding.
[0079] In some aspects, after intra-frame prediction decoding or inter-frame prediction decoding using the PU of the CU, the encoder engine 106 can compute residual data for the TU of the CU. The PU may include pixel data in the spatial domain (or pixel domain). The TU may include coefficients in the transform domain after applying the block transform. As previously noted, the residual data may correspond to the pixel difference between a pixel in the uncoded image and the predicted value corresponding to the PU. The encoder engine 106 may form a TU including the residual data for the CU, and may then transform the TU to produce transform coefficients for the CU.
[0080] The encoder engine 106 can perform quantization on the transform coefficients. Quantization provides further compression by reducing the amount of data used to represent the coefficients. For example, quantization can reduce the bit depth associated with some or all of the coefficients. In one example, a coefficient with an n-bit value can be rounded down to an m-bit value during quantization, where n is greater than m.
[0081] Once quantization is performed, the decoded video bitstream includes the quantized transform coefficients, prediction information (e.g., prediction modes, motion vectors, block vectors, etc.), segmentation information, and any other suitable data (such as other syntax data). The different elements of the decoded video bitstream can then be entropy-coded by encoder engine 106. In some examples, encoder engine 106 may scan the quantized transform coefficients using a predefined scan order to produce a serialized vector that can be entropy-coded. In some examples, encoder engine 106 may perform an adaptive scan. After scanning the quantized transform coefficients to form a vector (e.g., a one-dimensional vector), encoder engine 106 may entropy-code that vector. For example, encoder engine 106 may use context-adaptive variable-length decoding, context-adaptive binary arithmetic decoding, syntax-based context-adaptive binary arithmetic decoding, probabilistic interval segmentation entropy decoding, or another suitable entropy coding technique.
[0082] The output 110 of encoding device 104 can transmit NAL units constituting the encoded video bitstream data to decoding device 112 of the receiving device via communication link 120. The input 114 of decoding device 112 can receive the NAL units. Communication link 120 may include a channel provided by a wireless network, a wired network, or a combination of wired and wireless networks. The wireless network may include any wireless interface or combination of wireless interfaces and may include any suitable wireless network (e.g., the Internet or other wide area networks, packet-based networks, WiFi™, radio frequency (RF), ultra-wideband (UWB), WiFi Direct, cellular, LTE, WiMax™, etc.). The wired network may include any wired interface (e.g., fiber optic, Ethernet, powerline Ethernet, coaxial cable-based Ethernet, digital signal line (DSL), etc.). Wired and / or wireless networks can be implemented using various equipment (such as base stations, routers, access points, bridges, gateways, switches, etc.). The encoded video bitstream data may be modulated according to communication standards (such as wireless communication protocols) and transmitted to the receiving device.
[0083] In some examples, encoding device 104 may store encoded video bitstream data in storage device 108. Output 110 may retrieve encoded video bitstream data from encoder engine 106 or from storage device 108. Storage device 108 may include any of a variety of distributed or locally accessed data storage media. For example, storage device 108 may include hard disk drives, storage disks, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. Storage device 108 may also include a decoded picture buffer (DPB) for storing reference pictures used in inter-frame prediction. In further examples, storage device 108 may correspond to a file server or another intermediate storage device that may store encoded video generated by a source device. In such cases, a receiving device including decoding device 112 may access the stored video data from the storage device via streaming or downloading. The file server may be any type of server capable of storing encoded video data and sending encoded video data to the receiving device. Example file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The receiving device can access the encoded video data via any standard data connection, including an internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing encoded video data stored on a file server. Transmission of the encoded video data from storage device 108 can be streaming, downloading, or a combination thereof.
[0084] Input 114 of decoding device 112 receives encoded video bitstream data and can provide the video bitstream data to decoder engine 116 or to storage device 118 for later use by decoder engine 116. For example, storage device 118 may include a decoded picture buffer (DPB) for storing reference pictures used in inter-frame prediction. A receiving device including decoding device 112 can receive the encoded video data to be decoded via storage device 108. The encoded video data may be modulated according to communication standards (such as wireless communication protocols) and transmitted to the receiving device. The communication medium used to transmit the encoded video data may include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, wide area network, or global network (such as the Internet). The communication medium may include routers, switches, base stations, or any other equipment that may be useful for facilitating communication from the source device to the receiving device.
[0085] Decoder engine 116 decodes the encoded video bitstream data by entropy decoding (e.g., using an entropy decoder) and extracting elements of one or more decoded video sequences that constitute the encoded video data. Decoder engine 116 can then rescale the encoded video bitstream data and perform an inverse transform on it. The residual data is then passed to the prediction stage of decoder engine 116. Decoder engine 116 then predicts pixel blocks (e.g., PUs). In some examples, the prediction is added to the output of the inverse transform (the residual data).
[0086] Video decoding device 112 can output decoded video to video destination device 122, which may include a display or other output device for displaying the decoded video data to a consumer of the content. In some aspects, video destination device 122 may be part of a receiving device that includes decoding device 112. In some aspects, video destination device 122 may be part of a separate device, distinct from the receiving device.
[0087] In some respects, the video encoding device 104 and / or the video decoding device 112 may be integrated with the audio encoding device and the audio decoding device, respectively. The video encoding device 104 and / or the video decoding device 112 may also include other hardware or software necessary for implementing the decoding techniques described above, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic components, software, hardware, firmware, or any combination thereof. The video encoding device 104 and the video decoding device 112 may be integrated as part of a combined encoder / decoder (codec) in the respective device.
[0088] exist Figure 1 The example system shown is an illustrative example that can be used herein. The techniques used to process video data using the techniques described herein can be implemented by any digital video encoding and / or decoding device. Although the techniques disclosed herein are generally implemented by video encoding or video decoding devices, these techniques can also be implemented by a combined video encoder-decoder (often referred to as a “codec”). Furthermore, the techniques disclosed herein can also be implemented by a video preprocessor. The source device and receiving device are merely examples of such decoding devices, where the source device generates decoded video data for transmission to the receiving device. In some examples, the source device and receiving device can operate in a substantially symmetrical manner, such that each of these devices includes video encoding and decoding components. Therefore, the example system can support one-way or two-way video transmission between video devices, for example, for video streaming, video playback, video broadcasting, or video telephony.
[0089] This disclosure may generally relate to "signaling" certain information (such as syntax elements). The term "signaling" can generally refer to the communication of values and / or other data of syntax elements used to decode encoded video data. For example, video encoding device 104 may signal the values of syntax elements in a bitstream. Generally, signaling refers to generating values in a bitstream. As noted above, video source 102 may transmit the bitstream to video destination device 122 substantially in real time or not in real time (such as when syntax elements are stored in storage device 108 for later retrieval by video destination device 122).
[0090] Video bitstreams may also include Supplemental Enhancement Information (SEI) messages. For example, an SEI NAL unit can be part of the video bitstream. In some cases, SEI messages may contain information not required by the decoding process. For example, the information in the SEI message may not be necessary for the decoder to decode the video images in the bitstream, but the decoder can use the information to improve the display or processing of the images (e.g., the decoded output). The information in the SEI message can be embedded metadata. In an exemplary example, the information in the SEI message can be used by a decoder-side entity to improve the visibility of the content. In some instances, certain application standards may mandate the presence of such SEI messages in the bitstream, resulting in quality improvements for all devices conforming to the application standard (e.g., carrying frame encapsulation SEI messages for frame-compatible planar stereoscopic 3DTV video formats, where an SEI message is carried for each frame of the video; processing recovery point SEI messages; using pull-scan rectangle SEI messages in DVB; and many other examples).
[0091] As described above, for each block, a set of motion information (also referred to as motion parameters in this document) may be available. The set of motion information contains motion information for both the forward and backward prediction directions. The forward and backward prediction directions are the two prediction directions of a bidirectional prediction mode, in which case the terms "forward" and "backward" do not necessarily have geometric meaning. Instead, "forward" and "backward" correspond to the reference image list 0 (RefPicList0 or L0) and reference image list 1 (RefPicList1 or L1) for the current image. In some examples, when only one reference image list is available for an image or slice, only RefPicList0 is available and the motion information for each block of the slice is always forward.
[0092] In some cases, motion vectors, along with their reference indices, are used in the decoding process (e.g., motion compensation). Such motion vectors with associated reference indices are represented as a single prediction set of motion information. For each prediction direction, the motion information may include a reference index and a motion vector. In some cases, for simplicity, it may be assumed that the motion vectors themselves are referenced by an associated reference index. The reference index is used to identify reference images in the current list of reference images (RefPicList0 or RefPicList1). The motion vector has horizontal and vertical components that provide the offset from the coordinate position in the current image to the coordinates in the reference image identified by the reference index. For example, the reference index may indicate a specific reference image that should be used for a block in the current image, and the motion vector may indicate the position of the best-matching block (the block that best matches the current block) in the reference image.
[0093] In H.264 / AVC, each inter-frame macroblock (MB) can be partitioned in four different ways: one 16×16 MB partition; two 16×8 MB partitions; two 8×16 MB partitions; and four 8×8 MB partitions. Different MB partitions within a single MB can have different reference index values (RefPicList0 or RefPicList1) for each direction. In some cases, when an MB is not partitioned into four 8×8 MB partitions, it may have only one motion vector for each MB partition in each direction. In some cases, when an MB is partitioned into four 8×8 MB partitions, each 8×8 MB partition can be further subdivided into sub-blocks, in which case each sub-block can have a different motion vector in each direction. In some examples, there are four different ways to obtain sub-blocks from 8×8 MB partitions: one 8×8 sub-block; two 8×4 sub-blocks; two 4×8 sub-blocks; and four 4×4 sub-blocks. Each sub-block can have a different motion vector in each direction. Therefore, motion vectors exist at a level equal to or higher than that of the sub-blocks.
[0094] In AVC, a time-direct mode can be enabled at the MB level or MB segment level for use in skip and / or direct modes in B slices. For each MB partition, a motion vector is derived using the motion vector of the block co-located in RefPicList1[0] of the current block with the current MB partition. Each motion vector in the co-located block is scaled based on the POC distance.
[0095] Direct spatial mode can also be performed in AVC. For example, in AVC, direct mode can also predict motion information from spatial neighbors.
[0096] As mentioned above, in HEVC, the largest decoding unit in a slice is called a decode tree block (CTB). A CTB contains a quadtree whose nodes are decoder units. In the HEVC master profile, the size of the CTB can range from 16×16 to 64×64. In some cases, an 8×8 CTB size is supported. Decoding units (CUs) can have the same size as the CTB and can be as small as 8×8. In some cases, each decoding unit decodes in a single mode. When a CU is inter-decoded, it can be further divided into 2 or 4 prediction units (PUs) or can become a single PU without applying further division. When two PUs exist within a single CU, it can be a rectangle of half the size or a rectangle containing CUs. 1 / 4 or 3 Two rectangles of size 4. When the CU undergoes inter-frame decoding, a motion information set exists for each PU. Additionally, each PU is decoded using a unique inter-frame prediction mode to derive the motion information set.
[0097] For example, for motion prediction in HEVC, there are two inter-frame prediction modes: a merge mode for prediction units (PUs) and an Advanced Motion Vector Prediction (AMVP) mode. Special cases considered for merging are skipped. In AMVP or merge mode, a motion vector (MV) candidate list is maintained for multiple motion vector predictors. The motion vector for the current PU is generated by selecting a candidate from the MV candidate list, along with a reference index in the merge mode. In some examples, one or more scaling window offsets, along with the stored motion vectors, may be included in the MV candidate list.
[0098] In an example where the MV candidate list is used for block motion prediction, the MV candidate list can be constructed separately by the encoding and decoding devices. For example, the MV candidate list can be generated by the encoding device when the block is encoded, and by the decoding device when the block is decoded. Information related to motion information candidates in the MV candidate list (e.g., information related to one or more motion vectors, information related to one or more LIC flags in the MV candidate list that can be stored in some cases, and / or other information) can be signaled between the encoding and decoding devices. For example, in merge mode, the index value to the stored motion information candidate can be signaled from the encoding device to the decoding device (e.g., in syntax structures such as Picture Parameter Set (PPS), Sequence Parameter Set (SPS), Video Parameter Set (VPS), slice header, Supplemental Enhancement Information (SEI) messages and / or other signaling transmitted in or separately from the video bitstream). The decoding device can construct the MV candidate list and use the signaled references or indexes to obtain one or more motion information candidates from the constructed MV candidate list for motion compensation prediction. For example, decoding device 112 can construct a list of motion vector (MV) candidates and use motion vectors from indexed positions (and in some cases, LIC flags) for motion prediction of blocks. In the case of AMVP mode, in addition to reference or index, the difference or residual value emitted by the signal can be used as an increment. For example, in AMVP mode, decoding device can construct one or more lists of MV candidates and apply the increment value to one or more motion information candidates obtained using the index value emitted by the signal in performing motion compensation prediction of the block.
[0099] In some examples, the MV candidate list contains up to five candidates for the merge mode and two candidates for the AMVP mode. In other examples, different numbers of candidates may be included in the MV candidate list for the merge mode and / or AMVP mode. The merge candidates may contain a set of motion information. For example, the set of motion information may include motion vectors corresponding to both the reference image list (list 0 and list 1) and the reference index. If the merge candidate is identified by the merge index, the reference image and associated motion vector for the prediction of the current block are determined. However, in AMVP mode for each potential prediction direction from list 0 or list 1, the reference index along with the MVP index needs to be explicitly signaled to the MV candidate list, since the AMVP candidate only contains motion vectors. In AMVP mode, the predicted motion vectors can be further refined.
[0100] As seen above, the merged candidate corresponds to the entire set of motion information, while the AMVP candidate contains only a single motion vector for a specific prediction direction and reference index. Candidates for both modes are similarly derived from the same spatially and temporally adjacent blocks.
[0101] In some examples, the merge mode allows inter-frame prediction PUs to inherit one or more identical motion vectors, prediction directions, and one or more reference picture indices from inter-frame prediction PUs that include motion data locations selected from a set of spatially adjacent motion data locations and two temporally co-located motion data locations. For AMVP mode, one or more motion vectors of the PU can be predictively decoded relative to one or more motion vector predictors (MVPs) from an AMVP candidate list constructed by the encoder and / or decoder. In some instances, for unidirectional inter-frame prediction of the PU, the encoder and / or decoder can generate a single AMVP candidate list. In some instances, for bidirectional prediction of the PU, the encoder and / or decoder can generate two AMVP candidate lists, one using motion data from spatially and temporally adjacent PUs in the forward prediction direction, and one using motion data from spatially and temporally adjacent PUs in the backward prediction direction.
[0102] Candidates for both modes can be derived from spatially and / or temporally adjacent blocks. For example, Figure 2A and Figure 2B This includes a conceptual graph illustrating adjacent candidates in the instance space. Figure 2A Examples of spatially adjacent motion vector (MV) candidates for merging patterns are shown. Figure 2B An example of spatial neighbor motion vector (MV) candidates for AMVP mode is shown. Spatial MV candidates are derived from neighboring blocks for a specific PU (PU0), but the method for generating candidates from blocks differs for merge mode and AMVP mode.
[0103] In merging mode, the encoder can form a list of merging candidates by considering merging candidates from various motion data locations. For example, such as Figure 2A As shown, regarding Figure 2A Using spatially adjacent motion data positions indicated by numbers 0 to 4, up to five spatial MV candidates can be derived. The MV candidates can be ordered in the merged candidate list in the order shown by numbers 0 to 4. For example, positions and orders can include: left position (0), above position (1), right upper position (2), left lower position (3), and left upper position (4). Figure 2A In this context, block 200 includes PU0 202 and PU1 204. In some examples, when the video decoder decodes the motion information of PU0 202 using a merging mode, the video decoder may add motion information from spatial neighbor blocks 210, 212, 214, 216, and 218 to the candidate list in the order described above.
[0104] exist Figure 2BIn the AVMP mode shown, adjacent blocks are divided into two groups: the left group, including blocks 0 and 1; and the upper group, including blocks 2, 3, and 4. Figure 2B In this context, blocks 0, 1, 2, 3, and 4 are labeled as blocks 230, 232, 234, 236, and 238, respectively. Here, block 220 includes PU0 222 and PU1 224, and blocks 230, 232, 234, 236, and 238 represent the spatial neighbors of PU0 222. For each group, potential candidates in neighboring blocks that refer to the same reference picture indicated by a reference index emitted by a signal have the highest priority to be selected to form the final candidates for that group. All neighboring blocks may not contain motion vectors pointing to the same reference picture. Therefore, if no such candidate can be found, the first available candidate will be scaled to form the final candidate, thus compensating for temporal distance differences.
[0105] Figure 3A and Figure 3B This includes a concept diagram illustrating time motion vector prediction. Figure 3A Example CU 300 includes PU0 302 and PU1 304. PU0 302 includes a center block 310 for PU0 302 and a bottom block 306 to the right side of PU0 302. Figure 3A An external block 308 is also shown that can predict motion information from motion information of PU0 302, as discussed below. Figure 3B An example is the current image 342 of the current block 326 for which motion information is to be predicted. Figure 3B Also illustrated are the juxtaposed image 330 to the current image 342 (including the juxtaposed block 324 to the current block 326), the current reference image 340, and the juxtaposed reference image 332. The juxtaposed motion vector 320, which is used as the motion information of the current block 326, is used to predict the juxtaposed block 324.
[0106] The video decoder can add a Temporal Motion Vector Predictor (TMVP) candidate (e.g., TMVP candidate 322) (if enabled and available) to the MV candidate list after any spatial motion vector candidate. The motion vector derivation process for the TMVP candidate is the same for both merge mode and AMVP mode. However, in some instances, the target reference index of the TMVP candidate in merge mode is always set to zero.
[0107] The main block location used for TMVP candidate derivation is the bottom right block 306, which is placed outside PU 304, as shown below. Figure 3A As shown, this is to compensate for the bias in the upper and left blocks used to generate spatially adjacent candidates. However, if block 306 is located outside the current CTB (or LCU) row (e.g., as...), Figure 3A(as illustrated by block 308 in the example) or if the motion information of block 306 is unavailable, then the central block 310 of PU 302 is used to replace the block.
[0108] refer to Figure 3B The motion vector of TMVP candidate 322 can be derived from the juxtaposition block 324 of juxtaposed image 330, indicated at the slice level. Similar to the temporal direct mode in AVC, the motion vector of the TMVP candidate can undergo motion vector scaling, which is performed to compensate for the distance difference between the current image 342 and the current reference image 340, and between the juxtaposed image 330 and the juxtaposed reference image 332. That is, motion vector 320 can be scaled to produce TMVP candidate 322 based on the distance difference between the current image (e.g., current image 342) and the current reference image (e.g., current reference image 340), and between the juxtaposed image (e.g., juxtaposed image 330) and the juxtaposed reference image (e.g., juxtaposed reference image 332).
[0109] Other aspects of motion prediction are covered in the HEVC standard and / or other standards, formats, or codecs. For example, several other aspects are covered in merge mode and AMVP mode. One aspect includes motion vector scaling. Regarding motion vector scaling, it can be assumed that the value of a motion vector is proportional to the distance between the images during rendering time. Motion vectors associate two images, a reference image and the image containing the motion vector (i.e., the containing image). When predicting another motion vector using one motion vector, the distance between the containing image and the reference image is calculated based on the Picture Order Count (POC) value.
[0110] For a motion vector to be predicted, its associated containing image and reference image can be different. Therefore, a new distance (based on POC) is calculated. Furthermore, the motion vector can be scaled based on these two POC distances. For spatially adjacent candidates, the containing images of two motion vectors are the same, while the reference images are different. In HEVC, motion vector scaling is applied to both TMVP and AMVP for spatially and temporally adjacent candidates.
[0111] Another aspect of motion prediction involves the generation of artificial motion vector candidates. For example, if the list of motion vector candidates is incomplete, artificial motion vector candidates are generated and inserted at the end of the list until all candidates are obtained. In the merging mode, there are two types of artificial motion vector candidates: combined candidates derived only for B-slices; and zero candidates used only for AMVP if the first type does not provide enough artificial candidates. For each pair of candidates that is already in the candidate list and has the necessary motion information, a bidirectional combined motion vector candidate is derived by combining the motion vector of the first candidate from the image in reference list 0 and the motion vector of the second candidate from the image in reference list 1.
[0112] In some implementations, a pruning process can be performed when a new candidate is added to or inserted into the MV candidate list. For example, in some cases, MV candidates from different blocks may include the same information. In this case, storing duplicate motion information for multiple MV candidates in the MV candidate list can lead to redundancy and reduced efficiency. In some examples, the pruning process can eliminate or minimize redundancy in the MV candidate list. For example, the pruning process may include comparing potential MV candidates to be added to the MV candidate list with MV candidates already stored in the MV candidate list. In an exemplary example, the horizontal displacement (ΔX) and vertical displacement (Δy) of the stored motion vectors (indicating the position of the reference block relative to the current block) can be compared with the horizontal displacement (Δx) and vertical displacement (Δy) of the motion vectors of potential candidates. If the comparison reveals that the motion vector of a potential candidate does not match any of one or more of the stored motion vectors, the potential candidate is not considered a candidate to be pruned and can be added to the MV candidate list. If a match is found based on the comparison, the potential MV candidate is not added to the MV candidate list, thus avoiding the insertion of duplicate candidates. In some cases, to reduce complexity, only a limited number of comparisons are performed during the pruning process instead of comparing each potential MV candidate with all existing candidates.
[0113] In some decoding schemes such as HEVC, weighted prediction (WP) is supported, where a scaling factor (denoted by a), a shift number (denoted by s), and an offset (denoted by b) are used in motion compensation. Assuming the pixel value at position (x, y) in the reference image is p(x, y), then p'(x, y) = ((a*p(x, y) + (1 << (s-1))) >> s) + b, instead of p(x, y) being used as the predicted value in motion compensation.
[0114] When WP is enabled, a signal is used to indicate whether WP is applicable to a reference image for each current slice. If WP is applied to a reference image, the WP parameter set (i.e., a, s, and b) is passed to the decoder and used for motion compensation from the reference image. In some examples, to flexibly enable / disable WP for the luma and chroma components, the WP flag and WP parameters are signaled for the luma and chroma components respectively. In WP, the same set of WP parameters is used for all pixels in a reference image.
[0115] Figure 4AThis is an illustration illustrating an example of neighboring reconstructed samples of the current block 402 and neighboring samples of the reference block 404 for unidirectional inter-frame prediction. Motion vector MV 410 can be decoded for the current block 402, where MV 410 may include a reference index to a list of reference images and / or other motion information for identifying the reference block 404. For example, MV may include horizontal and vertical components providing offsets from coordinate positions in the current image to coordinates in a reference image identified by the reference index. Figure 4B This is an illustration of an example of neighboring reconstructed samples of the current block 422 and neighboring samples of the first reference block 424 and the second reference block 426 used for bidirectional inter-frame prediction. In this case, two motion vectors MV0 and MV1 can be decoded for the current block 422 to identify the first reference block 424 and the second reference block 426, respectively.
[0116] As previously explained, OBMC is an example motion compensation technique that can be implemented for motion compensation. OBMC can increase prediction accuracy and avoid block artifacts. In OBMC, a prediction can be a weighted sum of multiple predictions or include a weighted sum of multiple predictions. In some cases, blocks can be larger in each dimension and can overlap with adjacent blocks in a quadrant manner. Therefore, each pixel can belong to multiple blocks. For example, in some exemplary cases, each pixel can belong to 4 blocks. In this scheme, OBMC can achieve four predictions for each pixel, totaling a weighted average.
[0117] In some cases, specific syntax at the CU level can be used to enable and disable OBMCs. In some examples, there are two orientation modes in the OBMC (e.g., top, left, right, bottom, or below), including a CU boundary OBMC mode and a sub-block boundary OBMC mode. When using the CU boundary OBMC mode, the original prediction block using the current CU MV and another prediction block using an adjacent CU MV (e.g., an "OBMC block") are mixed. In some examples, the top left sub-block in the CU (e.g., the first or leftmost sub-block in the first / top row of the CU) has both a top OBMC block and a left OBMC block, and other topmost sub-blocks (e.g., other sub-blocks in the first / top row of the CU) may only have a top OBMC block. Other leftmost sub-blocks (e.g., the sub-block in the first column of the CU on the left side of the CU) may only have a left OBMC block.
[0118] Subblock boundary OBMC mode can be enabled when sub-CU decoding tools that allow different MVs based on subblocks are enabled in the current CU (e.g., affine motion compensation prediction, advanced temporal motion vector prediction (ATMVP), etc.). In subblock boundary OBMC mode, individual OBMC blocks using the MVs of connected adjacent subblocks can be blended with the original prediction block using the MV of the current subblock. In some examples, individual OBMC blocks using the MVs of connected adjacent subblocks can be blended in parallel with the original prediction block using the MV of the current subblock in subblock boundary OBMC mode, as further described herein. In other examples, individual OBMC blocks using the MVs of connected adjacent subblocks can be blended sequentially with the original prediction block using the MV of the current subblock in subblock boundary mode. In some cases, CU boundary OBMC mode can be performed before subblock boundary OBMC mode, and the predefined blending order of subblock boundary OBMC mode can include top, left, bottom, and right.
[0119] The prediction of the MV based on neighboring sub-blocks N (e.g., the sub-blocks above, to the left, below, and to the right of the current sub-block) can be represented as P. N And the prediction based on the MV of the current sub-block can be expressed as P C When sub-block N contains the same motion information as the current sub-block, the original prediction block can be mixed with the prediction block without relying on the MV of sub-block N. In some cases, P can be... N The 4 rows / columns of the sample and P C Mixing the same samples in the same way.
[0120] Figure 5 Examples of motion vector scaling 500 for time-merging candidates used in a processing block are illustrated according to some examples of this disclosure. In some examples, time-merging candidate derivation can be performed, wherein a merge candidate (e.g., a time-merging candidate) is added to a merge candidate list. Time-merging candidate derivation can be performed based on the scaled motion vector. The scaled motion vector can be derived based on the juxtaposed CUs included in the juxtaposed reference images. The list of reference images to be used for deriving the juxtaposed CUs can be explicitly signaled in the slice header.
[0121] For example, Figure 5 The current image 510 and the juxtaposed image 530 are depicted and can be associated with the current reference image 515 and the juxtaposed reference image 535, respectively. Figure 5 The current CU 512 (e.g., associated with the current image 510) and the juxtaposed CU (e.g., associated with the juxtaposed image 530) are also depicted. In some examples, it can be as follows: Figure 5 The example illustrates the derivation or acquisition of scaled motion vectors for time merging candidate derivations. For example, Figure 5Dashed line 511 is depicted, which is scaled from the motion vector of the juxtaposed CU 532 using Picture Order Count (POC) distances tb and td. In some examples, tb is the POC difference between the current reference picture 515 and the current picture 510, while td is the POC difference between the juxtaposed reference picture 535 and the juxtaposed picture 530. The reference picture index of the temporal merge candidate can be set to zero.
[0122] Figure 6A Examples of juxtaposed motion vector projections 610 and 620 used in a processing block according to some examples of this disclosure are illustrated. Figure 6B Examples of various reference frames that can be used as temporal juxtaposition reference frames for juxtaposition motion vector projection are shown.
[0123] The AV1 codec supports up to seven reference frames for inter-frame prediction. For example, as Figure 6B As illustrated, AV1 can support a set of reference frames 630, which includes some (or all) of the following reference frames: LAST, LAST2, LAST3, GOLDEN, BWDREF, ALTREF2, and ALTREF. This set of seven reference frames 630 can be used to perform inter-frame prediction. In some cases, each reference frame in the set of reference frames 630 can be associated with one or more motion vectors. In some aspects, up to three reference frames (e.g., Figure 6B Three of the seven illustrated reference frames can be used as time juxtaposition reference frames. In one illustrative example, up to three reference frames that can be used as time juxtaposition reference frames can be selected from a set of five reference frames. The five reference frames selected from the set of up to three time juxtaposition reference frames may include... Figure 6B The illustrated reference frames are LAST, LAST2, BWDREF, ALTREF2, and ALTREF. In some aspects, selection process 632 can be used to select up to three temporally juxtaposed reference frames, wherein the selected temporally juxtaposed reference frames are associated with a set of temporal motion vectors 642. As previously described, the motion vectors of the juxtaposed reference frames (e.g., temporal motion vectors 642) can be used to perform juxtaposed motion vector projection.
[0124] Figure 6AExamples of juxtaposed motion vector projections 610 and 620 are shown. In some cases, juxtaposed motion vector projection can be performed based on some (or all) of the time motion vectors 642 projected relative to the current frame and its reference frame. For example, one or more current frame offsets cur_offsets can be used to project one or more time motion vectors 642 associated with the starting frame 624 (e.g., a reference frame of the current frame 622) relative to the current frame 622. One or more reference frame offsets ref_offset[rf] can be used to additionally project one or more time motion vectors 642 associated with the starting frame 624 relative to the reference frame 626 (e.g., a reference frame of the starting frame 624).
[0125] An example of juxtaposed motion vector projection 610 presents a simplified representation of juxtaposed motion vector projection 620 (e.g., the current frame 622 may correspond to or be the same as block 612, the starting frame 624 may correspond to or be the same as block 614, and the reference frame 626 may correspond to or be the same as block 616). In an illustrative example, blocks 612, 614, and 616 are 64 × 64 in size. In some aspects, a projection can be valid when the projection of a given time motion vector in time motion vector 642 lies within the block to which the projection is associated, or when the projection lies within one of the horizontally adjacent blocks. For example, when projecting time motion vector 642 based on block 614, a given projection can be valid if it lies within block 614 (e.g., the block to which the projection is associated) or if it lies within block 612 or block 616 (e.g., two horizontally adjacent blocks relative to block 614).
[0126] In some examples, juxtaposed motion vector projection can be performed for up to three juxtaposed reference frames (referred to as REF0, REF1, and REF2). The motion vector associated with REF0 (e.g., obtained from or included in the set of temporal motion vectors 642) can be projected first. In an exemplary example, for a 64×64 block size, the motion vector associated with REF0 can be projected over the entire frame in an 8×8 raster scan order. A valid projection of REF0 MV can lie within the 64×64 REF0 block or within one of the 64×64 blocks horizontally adjacent to the REF0 block. Projections outside one of the three 64×64 blocks can be considered invalid. After determining the valid projection of REF0 MV, the same procedure described above can be used to project the motion vectors associated with REF1 and REF2. In some respects, the projections of the MVs (e.g., a set of 642 time motion vectors) associated with REF0, REF1, and REF2 can be performed sequentially such that if two projections both go to the same location within one of three adjacent 64×64 blocks, the MV projection associated with REF1 will overwrite the MV projection associated with REF0, and if two projections both go to the same location, the MV projection with REF2 will overwrite the MV projection associated with REF1. In some cases, sequentially projecting REF0, REF1, and REF2 can be used to provide an MV projection associated with REF2, which has the highest relative priority, and an MV projection associated with REF0, which has the lowest relative priority, in cases where a subsequent projection overwrites any MV projection already present at the same location.
[0127] Figures 7A to 7D An example of a juxtaposed motion vector projection performed based on two juxtaposed reference frames (e.g., LAST_FRAME and ALT_FRAME). Figure 7A An example of motion vector references included in the LAST_FRAME is shown (e.g., before performing motion vector projection). Arrows derived from each motion vector reference in the LAST_FRAME can represent scaled motion vectors that can be used to perform juxtaposed motion vector projection against the LAST_FRAME. The end of each arrow indicates the block location to which the corresponding motion vector reference within the LAST_FRAME will be projected (e.g., based on the scaled motion vector determined for the corresponding motion vector reference within the LAST_FRAME).
[0128] Figure 7B Example based on Figure 7AThe scaled motion vectors depicted in the diagram are used to generate a motion vector field projection for LAST_FRAME. Block locations with no associated motion vector projection (e.g., empty block locations) can be set to zero or considered invalid. In some aspects, juxtaposed motion vector projections can be performed on LAST_FRAME in raster scan order (e.g., as shown in the diagram). Figure 7A exemplified) to generate Figure 7B The motion vector field is projected. If the motion vector is projected onto an empty or unoccupied position within the motion vector field projection, the projected motion vector can be written to the projected position. For example, processing LAST_FRAME in raster scan order results in MV ref_frame 702 being processed first. Based on the scaled motion vector associated with MV ref_frame 702, MV ref_frame 702 is projected and written to the currently empty position 730. Raster scan order processing can then proceed to MV ref_frame 704. Based on the scaled motion vector associated with MV ref_frame 704, MV ref_frame 704 is also projected to the same position 730, which is now occupied by the projected MV ref_frame 702. Because MV ref_frame 704 is projected at a later point in the raster scan order than MV ref_frame 702, conflicts between multiple MV ref_frames projected to the same position 730 can be resolved by allowing the later projection of MV ref_frame 704 to overwrite the earlier projection of MV ref_frame 702. For example... Figure 7B As illustrated, in the final motion vector field projection determined for LAST_FRAME, position 730 is occupied by the projected MV ref_frame 704.
[0129] Figure 7C This is a second example illustrating the motion vector reference and scaled motion vectors associated with the second juxtaposed reference frame ALT_FRAME. This can be compared with the projection above regarding LAST_FRAME (e.g., as...). Figure 7A In a manner similar to or the one illustrated, the motion vector reference included in the ALT_FRAME is projected based on the associated scaled motion vector to determine the motion vector field projection for the LAST_FRAME (e.g., as shown in the example). Figure 7B exemplified).
[0130] Can Figure 7CThe motion vector references depicted in the LAST_FRAME (e.g., MV ref_frames included in the ALT_FRAME) are projected onto the motion vector field projection previously determined for the LAST_FRAME, wherein the motion vector projection associated with the ALT_FRAME overwrites any existing projections associated with the LAST_FRAME located at the same position within the motion vector field projection. Figure 7D The example shown is the projection of the resulting motion vector field determined by LAST_FRAME and ALT_FRAME.
[0131] In some examples, it can be (for example, such as) Figure 7C The MV ref_frame 714 included in the ALT_FRAME depicted is projected to position 730. As previously mentioned, the MV ref_frames included in the ALT_FRAME can be projected to the position determined by the LAST_FRAME and... Figure 7B On the motion vector field projection shown in the example. In the LAST_FRAME motion vector field projection, position 730 is occupied by LAST_FRAME MV ref_frame 704.
[0132] When ALT_FRAME MV ref_frame 730 is projected to the same position 730 (e.g., during juxtaposed motion vector projection of ALT_FRAME), ALT_FRAME MV ref_frame 730 can override LAST_FRAME MVref_frame 704. In an exemplary example, motion vector projection of ALT_FRAME can be performed in raster scan order (e.g., as described above). Based on the raster scan order, ALT_FRAME MV ref_frame 716 is subsequently projected to the same position 730 currently occupied by ALT_FRAME MV ref_frame 714. The later-projected ALT_FRAME MV ref_frame 716 can override the earlier-projected ALT_FRAME MV ref_frame 716.
[0133] like Figure 7D As illustrated, the resulting motion vector field projection determined for LAST_FRAME and ALT_FRAME only includes the later projection of ALT_FRAME MV ref_frame 716 at position 730. In some respects, when a third juxtaposed reference frame is used to perform the juxtaposed motion vector projection, it can be related to the above regarding... Figures 7A to 7D The third juxtaposed reference frame is projected onto the method described in the same or similar manner. Figure 7DIn the illustrated LAST_FRAME+ALT_FRAME motion vector field projection, position conflicts (e.g., MV projections to occupied positions within the motion vector field projection) are resolved by overwriting earlier projections with later projections.
[0134] Figure 8A and Figure 8B Examples of juxtaposed motion vector projection based on frame-level processing are shown. For example, Figure 8A An example is provided for a frame-level raster scan order 810 that can be used to perform juxtaposition motion vector projection for the currently decoded frame 802. This can be based on a first set of MVs 812 associated with a first juxtaposition reference frame (e.g., in...). Figure 8A The second set of MV 814 associated with the second juxtaposed reference frame (e.g., in 'Frame0 MV') is indicated in the middle. Figure 8A The middle indicator is 'Frame1 MV') and the third set of MV 816 associated with the third juxtaposed reference frame (e.g., in Figure 8A The instruction in the middle is 'Frame2 MV') to perform the juxtaposition of motion vector projection.
[0135] In an exemplary example, the aforementioned first, second, and third juxtaposed reference frames may be associated with or obtained based on the current decoded frame 802. This can be achieved by first reading a first set of MV 812 from memory and then using... Figure 8A The frame-level raster scan sequence depicted in the text projects each MV in the first set of MV 812 to perform a juxtaposed motion vector projection for the currently decoded frame 802. For example, it can be compared with the above regarding... Figures 6A to 7D The first set of MV 812 is projected in the same or similar manner as the described juxtaposed motion vector projection. In some aspects, the projected motion vectors (e.g., motion vector field projection) determined based on the first set of MV 812 projected in frame-level raster scan order can be written to or otherwise stored in a buffer or memory associated with a decoding device used to decode the current frame 802.
[0136] For example, Figure 8BAn example diagram illustrates a system 820 that implements juxtaposed motion vector projection with a frame-level raster scan order 810. REF0 frames MV 812 are read from DDR (e.g., memory) 840 and used to generate REF0 frame MV projection 852. In some respects, the REF0 frames MV 812 and the first set of MV 812 can be the same. The REF0 MV 812 can be read from memory 840 and projected sequentially using the frame-level raster scan order 810. While projecting each REF0 MV 812, the resulting motion vector projection can be written to a projected MV buffer 825 included in memory 840.
[0137] After all REF0 MVs 812 have been projected and written to the projected MV buffer 825, the example system 820 can then repeat the same process for REF1 MVs 814. REF1 MVs 814 can be used again to determine the REF1 frame MV projection 854 using the frame-level raster scan order 810. Each entry of the REF1 frame MV projection 854 can be written to the projected MV buffer 825, where each REF1 frame MV projection entry overwrites any REF0 MV projection entries already stored in the projected MV buffer 825 at the same location (e.g., by overwriting entries as described above). Figures 7A to 7D (Existing entries as described). After all REF1 MV 814 has been projected and written to the projected MV buffer 825, the example system 820 can then repeat the same process again for REF2 MV 816.
[0138] In some cases, example system 820 performs three consecutive raster scans over the entire frame 802, each for each of the three consecutive motion vector projections (e.g., for REF0 MV 812, REF1 MV 814, and REF2 MV 816) by using frame-level raster scan order 810. The size of the projected MV buffer 825 must be additionally set to store the MV projection entries over the entire frame 802 because frame 802 cannot be decoded until motion vector projections have been completed for all three juxtaposed reference frames (e.g., REF0, REF1, REF2).
[0139] In some respects, performing motion vector projection sequentially for each of the juxtaposed reference frames associated with the currently decoded frame 802 can consume significant computational cycles (e.g., clock cycles of one or more processors) when decoding frame 802 begins. For example, significant clock cycles may be consumed while the example system 802 waits for frame-level raster scan order motion vector projections to be completed sequentially for three juxtaposed reference frames REF0, REF1, and REF2. In some cases, the sequential process of projecting the MVs associated with REF0, REF1, and REF2 is itself sequential with the decoding of frame 802, which can introduce further inefficiencies and / or latency into the decoding of frame 802 based on MV projection.
[0140] like Figure 8B As illustrated, performing juxtaposed motion vector projection by processing three juxtaposed reference frames in frame-level raster scan order can be associated with four read operations and four write operations on the juxtaposed data. For example, the four read operations may include three read operations for reading REF0, REF1, and REF2 MV data (e.g., 812, 814, and 816, respectively) before performing the corresponding MV projection for each reference frame, and a fourth read operation for reading the final contents of the projected MV buffer 825 to perform MV prediction at the MV prediction hardware 850. The four write operations may include three write operations for writing the REF0, REF1, and REF2 motion vector projections to the projected MV buffer 825, and a fourth write operation for writing the final MV prediction for the current frame 802 from the MV prediction hardware 850 to memory 840.
[0141] It is necessary to reduce the number of clock cycles used to perform juxtaposition motion vector projection on one or more (e.g., up to three) juxtaposition reference frames. It is necessary to reduce the read / write bandwidth associated with decoding frames of video data using juxtaposition motion vector projection. It is also necessary to reduce the size (e.g., storage capacity) of one or more buffers associated with the projection motion vectors determined during the performance of juxtaposition motion vector projection.
[0142] The systems and techniques described herein can be used to perform instantaneous juxtaposed motion vector projection according to a block-based processing order. As will be described in more detail below, the systems and techniques can be used to read and project motion vectors from reference frames at the block level. In an exemplary example, in a scene where the current decoded frame of video data is associated with three juxtaposed reference frames (e.g., REF0, REF1, REF2), the associated motion vector data included in the three juxtaposed reference frames can be read and projected at a 64×64 block level. In some aspects, the associated motion vector data included in the juxtaposed reference frames can be read and projected at a 128×128 block level. In some aspects, the systems and techniques can avoid writing the projected motion vectors to DDR or other memory (e.g., as described herein) based on the block-based processing order used to perform juxtaposed motion vector projection. Figure 8B (In the example). In an exemplary example, the system and techniques may include one or more MV projection buffers implemented in hardware, wherein the size of the one or more MV projection buffers is set to be three times the block level size (e.g., for a 64×64 block level, the size of the one or more MV projection buffers may be set to store three 64×64 blocks; for a 128×128 block level, the size of the one or more MV projection buffers may be set to store three 128×128 blocks).
[0143] Figure 9A This is an illustration of an example of block-based (e.g., block-level) processing 910 that can be used to perform juxtaposed motion vector projection more efficiently. For example, the current decoded frame of video data can be divided into multiple blocks, and juxtaposed motion vector projection can be performed for each block of the current decoded frame of video data. Figure 9A In the example of block-based processing, three blocks of video data are illustrated (e.g., 902a, 902b, and 902c). In some cases, more or fewer blocks may be used per decoded frame of video data. Blocks 902a-902c are illustrated as having a size of 64×64, but other block sizes (e.g., such as 128×128) may also be used.
[0144] In an exemplary example, the same block level (e.g., block size) associated with video data blocks 902a-902c can be used to obtain MV information blocks from one or more juxtaposed reference frames. The juxtaposed reference frames may be associated with the currently decoded frame of the video data and may include the information described above. Figure 6B The described example juxtaposition reference frames include one or more example juxtaposition reference frames. In one exemplary example, the one or more juxtaposition reference frames may include a first juxtaposition reference frame REF0, a second juxtaposition reference frame REF1, and a third juxtaposition reference frame REF2.
[0145] The MV information associated with the juxtaposed reference frame can be read and projected at the same block level (e.g., block size) associated with the video data blocks included in the currently decoded frame. For example, if the size of video data blocks 902a-902c is 64×64, the first block 902a of the video data can be decoded based on a 64×64 MV information block read from the juxtaposed reference frame. The 64×64 MV information block obtained for the juxtaposed reference frame can be projected into a single 64×64 motion vector field projection generated for the first block 902a of the video data, and the first block 902a of the video data can be decoded based on a projection that has determined the associated 64×64 motion vector field projection has been completed.
[0146] In an exemplary example, a given block of video data (e.g., the first block 902a) can be decoded in parallel with juxtaposition motion vector projection (MV) operations performed for one or more additional blocks and / or subsequent blocks of video data. For example, the first block 902a can be decoded when juxtaposition MV projection operations have been performed for the first block 902a and the second block 902b (e.g., the first block 902a can be decoded in parallel with juxtaposition MV projection operations performed for the third block 902c). The second block 902b can be decoded when juxtaposition MV projection operations have been performed for the second block 902b and the third block 902c (e.g., the second block 902b can be decoded in parallel with juxtaposition MV projection operations performed for the fourth block). For example, the given block of video data can be decoded based on determining that MV projection has been performed for the given block and its adjacent or neighboring blocks, wherein decoding for subsequent non-neighboring blocks of video data is performed independently of MV projection operations.
[0147] like Figure 9A As illustrated, the system and techniques can perform block-based juxtaposition MV projection for each block (e.g., 902a-902c) based on the raster scan order applied within each block. For example, to decode a 64×64 block 902a of video data, a 64×64 block of MVs associated with each juxtaposition reference frame (e.g., REF0, REF1, REF2) can be obtained, and... Figure 9AThe blocks are projected in the block-level raster scan sequence depicted. The REF0 MV block can be projected in the raster scan sequence to generate a REF0 block projection 952, which can be written to or stored in one or more buffers of the MV projection buffer 970. After the REF0 MV 912 block has been projected into the MV projection buffer 970 (e.g., as a REF1 block projection 952), the REF1 MV 914 block can be projected in the raster scan sequence to generate a REF1 block projection 954, which can also be written to or stored in one or more buffers of the MV projection buffer 970.
[0148] By writing (e.g., projecting) REF1 block projection 954 to (e.g., into) a buffer that already includes REF0 block projection 952 (e.g., included in MV projection buffer 970), the system and techniques can generate a combined MV field projection for REF1 and REF0, as shown regarding Figures 10 to 13 To explain in more detail. In a combined MV field projection for REF1 and REF0, the REF1 MV projection, having the same position as the existing REF0 MV projection, will overwrite the existing REF0 MV projection (e.g., as previously discussed regarding...). Figures 7A to 7D (As described).
[0149] After the block of REF1 MV 914 has been projected into the MV projection buffer 970 (e.g., as REF1 block projection 954), the block of REF2 MV 916 can be projected in raster scan order to generate REF2 block projection 956. REF2 block projection 956 can also be written to or stored in one or more buffers of the MV projection buffer 970, thereby overwriting any REF0 and / or REF1 MV projections existing at the same location as the REF2 MV projection.
[0150] In some aspects, one or more MV buffers included in the MV projection buffer 970 may be implemented as sliding buffers, wherein the current decoded block of video data is decoded from the first buffer of the sliding MV buffer 970. After decoding the current block (e.g., the first block 902a of video data) from the first buffer of the sliding MV buffer 970, the contents of the first buffer may be invalidated (e.g., deleted, cleared, zeroed, etc.). The next block of video data (e.g., the second block 902b of video data) may be decoded based on sliding the sliding MV buffer 970 by one step. For example, after invalidating the contents of the first buffer, the sliding MV buffer 970 may slide one step to the right. In some aspects, the old second buffer becomes the new first buffer from which the next block of video data (e.g., the second block 902b of video data) will be decoded; the old third buffer becomes the new second buffer; and the old first buffer (e.g., now empty after invalidating its contents) becomes the new third buffer.
[0151] In an exemplary example, the first buffer of the sliding MV buffer 970 can be used to obtain an MV field projection for decoding the current decoded block of video data based on determining that the first buffer includes an MV projection associated with the current block of video data and includes MV projections for two horizontally adjacent blocks. For example, when the current decoded block of video data is block 902a, the first buffer in the sliding MV buffer 970 can be used to decode block 902a when the first buffer includes an MV projection associated with block 902a (e.g., an MV projection associated with the current block of video data) and includes an MV projection associated with block 902b (e.g., an MV projection associated with the horizontally adjacent block to the right).
[0152] When the current decoded block of the video data is block 902b, the first buffer in the sliding MV buffer 970 can be used to decode block 902b when the first buffer includes the MV projection associated with block 902b (e.g., the MV projection associated with the current block of the video data), includes the MV projection associated with block 902a (e.g., the MV projection associated with the horizontal neighboring block to the left), and includes the MV projection associated with block 902c (e.g., the MV projection associated with the horizontal neighboring block to the right).
[0153] In some respects, systems and techniques can be used to perform on-the-fly juxtaposition MV projection based on determining the MV field projections of a first buffer in a sliding MV buffer 970, including the current block and two horizontally adjacent blocks. In such an example, the MV field projections stored in the first buffer in the sliding MV buffer 970 can be provided as input to an MV prediction unit 960, which generates a final MV and writes the final MV to memory 940. Based on the final MV determined using the MV prediction unit 960, the current block of video data can be decoded on the fly (e.g., before MV projection has been completed for either the juxtaposition reference frame and / or the currently decoded frame of the video data). As will be discussed below regarding... Figure 10 In a more detailed description, the current block of video data can be decoded in parallel with MV projection performed on subsequent (e.g., undecoded) blocks of video data.
[0154] For example, Figure 10 This is an illustration illustrating an example of block-based juxtaposed MV projection using a sliding MV buffer. The current decoded frame of video data 1002 comprises multiple blocks of video data, shown here as block A (e.g., block 1002a), block B (e.g., block 1002b), block C (e.g., block 1002c), and block D (e.g., block 1002d). In some examples, the block level (e.g., block size) may be 64×64. The sliding MV buffer 1070 includes a first buffer 1071, a second buffer 1072, and a third buffer 1073. In one exemplary example, the size of each buffer 1071 to 1073 included in the sliding MV buffer 1070 may be set to store a 64×64 block of data. As will be described below, the size of each buffer 1071 to 1073 may be set to store a 64×64 block of MV projection.
[0155] Multiple blocks of video data included in the frames of video data 1002 can be processed in raster order or in various other scan orders. For example... Figure 10 As illustrated in the examples, juxtaposed MV projection can be performed on frames of video data 1002 using the projection and decoding order of blocks A, B, C, D, etc. The currently decoded frame of video data 1002 may be associated with one or more juxtaposed reference frames (not shown), as previously described. In some examples, the currently decoded frame of video data 1002 may be associated with up to three juxtaposed reference frames.
[0156] Figure 10The illustrated block-based juxtaposition MV projection 1000 can apply the same 64×64 block-level scan to obtain blocks of the MV from each of the three juxtaposition reference frames associated with the current decoded frame of the video data 1002. For example, 64×64 blocks of the MV included in block A of each of the three juxtaposition reference frames (e.g., REF0, REF1, and REF2) can be obtained and projected into the sliding MV buffer 1070. In an exemplary example, the three 64×64 blocks A of the MV can be projected into the first buffer 1071 and the second buffer 1072 of the sliding MV buffer 1070. Because each of the three buffers 1071 to 1073 of the sliding MV buffer 1070 stores a 64×64 block of the video data, in some respects, each buffer 1071 to 1073 can store a different MV projection field.
[0157] For example, the raster scan sequence of block A can be projected from REF0 to obtain the MV of block A, wherein each given block A REF0MV is projected into a first buffer (e.g., corresponding to the MV projection field determined for decoding block A video data 1002a) or into a second buffer (e.g., corresponding to the MV projection field determined for decoding block B video data 1002b), rather than both. For example, as regarding Figure 6A and Figure 6B As described, a valid MV projection for a given block can be located in the same block or in one of two horizontally adjacent blocks, while an invalid MV projection is located outside these three blocks. Since block A is the first block for which an MV projection is performed, there is no horizontally adjacent block to its left. Therefore, the only valid MV projection location for block A will be either the first buffer 1071 or the second buffer 1072.
[0158] After the REF0 block AMV is projected into the first buffer 1071 and the second buffer 1072, the REF1 block AMV is obtained and is also projected into the first buffer 1071 and the second buffer 1072. The REF1 MV projection projected into the location where the REF0 MV projection already exists will overwrite the existing REF0 MV projection (e.g., as previously mentioned). Figures 7A to 7D (As described). Subsequently, after the REF1 block A MV projection has been completed, the REF2 block A MV is obtained and projected again into the first buffer 1071 and the second buffer 1072, overwriting any existing REF0 or REF1 projections that already exist at the same location as the REF2 MV projection.
[0159] When the MV projection of block A in REF2 has been completed, the first buffer 1071 includes MV projections from the locations in REF0, REF1, and REF2 blocks that are projected into block A. The second buffer 1072 includes MV projections from the locations in the horizontally adjacent blocks to the right of block A in REF0, REF1, and REF2 blocks (e.g., projected into block B). The third buffer 1073 remains empty or otherwise does not include the projection of block A.
[0160] The contents of the first buffer 1071 (e.g., the MV projection field to the block A juxtaposed reference frame MV in block A) cannot yet be used to decode the block A video data 1002a. To decode the block A video data 1002a, a combined MV projection field to the block A juxtaposed reference frame and the block B juxtaposed reference frame in block A is required.
[0161] After the MV projection for block A is completed, the MVs included in block B for each of the three juxtaposed reference frames (e.g., REF0, REF1, and REF2) are obtained and projected into buffers 1071, 1072, and 1073 in the same or similar manner as described above regarding the MV projection for block A. When performing the MV projection for block B, the REF0, REF1, and REF2 MVs projected from block B to block A (e.g., the horizontally adjacent block on the left) are written to the first buffer 1071. The REF0, REF1, and REF2 MVs projected from block B to another location within block B are written to the second buffer 1072. The REF0, REF1, and REF2 MVs projected from block B to block C (e.g., the horizontally adjacent block on the right) are written to the third buffer 1073.
[0162] The first buffer 1071 now includes the MV projection fields of three block A reference frames MV projected into block A and three block B reference frames MV projected into block A (e.g., the first buffer 1071 includes the MV projection fields of block A and block B reference frames MV into block A). The MV projection fields of the first buffer 1071 can be used to decode the block A video data 1002a.
[0163] After decoding the video data 1002a of block A, the MV projection field of block A (e.g., the projections of the block A reference frame MV and the block B reference frame MV onto block A) can be invalidated, cleared, or zeroed. The sliding MV buffer 1070 can then slide one step to the right based on reassigning the old second buffer 1072 to the new first buffer 1071. The old third buffer 1073 can be reassigned to the new second buffer 1072, and the old first buffer 1071 (e.g., now empty) can be reassigned to the new third buffer 1073. In an exemplary example, the contents of the buffers included in the sliding MV buffer 1070 are not modified during buffer reassignment. For example, before sliding the sliding MV buffer 1070, the old second buffer 1072 includes the projections of the block A reference frame MV onto block B and the projections of the block B reference frame MV onto block B. After sliding the sliding MV buffer 1070 and reassigning the old second buffer 1072 to the new first buffer 1071, the new first buffer 1071 still includes the same projection of the block A reference frame MV into block B and the projection of the block B reference frame MV into block B.
[0164] As described above, after sliding the sliding MV buffer 1070 one step to the right, the MV included in block C of each of the three juxtaposed reference frames (e.g., REF0, REF1, and REF2) is obtained and projected into buffers 1071, 1072, and 1073. In an exemplary example, because the sliding MV buffer 1070 slides one step to the right before projecting block C MV, block C MV is projected into the newly reassigned first buffer 1071 (e.g., the old second buffer 1072 from the previous step), the newly reassigned second buffer 1072 (e.g., the old third buffer 1073 from the previous step), and the newly reassigned third buffer 1073 (e.g., the old first buffer 1071 cleared / invalidated in the previous step).
[0165] The new first buffer 1071 now contains the MV projection fields of three block A reference frames projected into block B, three block B reference frames projected into block B, and three block C reference frames projected into block B (e.g., the new first buffer 1071 includes a combined MV projection field of the three reference frames MV of block A, the three reference frames MV of block B, and the three reference frames MV of block C, all of which are projected into block B). The combined MV projection field of the new first buffer 1071 can be used to decode the block B video data 1002b.
[0166] After decoding block B video data 1002b, the MV projection field of block B stored in the new first buffer 1071 can be invalidated, cleared, or zeroed, and the MV buffer 1070 can be slid to the right by one step. After clearing the block B MV projection field in the new first buffer 1071 (e.g., after decoding block B video data 1002b), buffer 1071 can be reassigned as a new third buffer 1073 for decoding block C video data 1002c. The second buffer 1072 can be reassigned as a new first buffer 1071 for decoding block C video data 1002c, and the third buffer 1073 can be reassigned as a new second buffer 1072 for decoding block C video data 1002c. The process can then be repeated for block C video data 1002c, block D video data 1002d, etc., until all blocks included in the multiple blocks used to divide frame 1002 have undergone MV projection (e.g., the process described above regarding decoding blocks A, B, and C using the sliding MV buffer 1070 can be repeated for all remaining 64×64 blocks included in the current decoded frame of video data 1002).
[0167] In an exemplary example, the row index of the projected MV determined for a block (e.g., 64×64) of juxtaposed reference frame MV data can be used to modify the overwriting of existing projected MV entries stored in a sliding MV buffer (e.g., sliding MV buffer 1070). For example, the row index of the projected MV associated with the same position as the existing projected MV can be compared with the row index of the existing projected MV to determine whether an overwriting should be performed. In some aspects, the row index-based comparison of the projected MVs can be used to obtain a final MV projection field generated using the block-level MV projection described previously, but the overwriting of the existing MV projection is implemented as if frame-level raster scan sequence MV projection had been performed.
[0168] For example, Figure 14 Figure 1400 illustrates an example of a process that can be used to modify an MV projection overwrite associated with a block-level MV projection to be the same as an MV projection overwrite associated with a frame-level raster scan sequence MV projection.
[0169] Projecting into the existing frame-level raster scan order MV (e.g., as mentioned above) Figure 8A and Figure 8BIn the described frame-level raster scan order (MV projection), the projection of an MV with a higher row index can always overwrite a previously projected MV with a lower row index. For example, in the frame-level raster scan order method, MVs with higher row indices (e.g., closer to the top of the current decoded frame) are processed after all MVs with lower row indices (e.g., closer to the top of the current decoded frame). Therefore, in the frame-level raster scan order method, MVs with higher row indices can always overwrite previously projected MVs with lower row indices.
[0170] However, when implementing a block-level MV projection method, the MV associated with the second block (e.g., horizontally adjacent to the right of the first block) will overwrite the previously projected MV located in the first block, even if the projected MV from the second block has a lower row index than the previously projected MV located in the first block. For example, Figure 14 An example is shown: a first 64×64 block 1410 and a horizontally adjacent second 64×64 block 1420. Also depicted are associated first 64×64 projection spaces 1415 and second 64×64 projection spaces 1425.
[0171] The MVs of the first projection block 1410 are projected first, for example, using the raster scan order within the first projection block 1410. As illustrated, MVs A, B, and C can each be projected into the first projection space 1415 without overwrite conflicts. MV D is projected into the same location within the first projection space 1415 as the previously projected MV B, resulting in an overwrite conflict that can be assessed according to the system and techniques described herein.
[0172] Since the row index of MV D (e.g., row index = 4) is greater than the row index of MV B (e.g., row index = 3), the projection of MV D can overwrite the existing projection of MV B in the first projection space 1415.
[0173] MV E can then be projected from the first projection block 1410 to the second projection space 1425 without an overwrite conflict. The projection of MV F from the second projection block 1420 to the first projection space 1415 results in an overwrite conflict because MV F is projected to the same location as an existing projection of MV C (e.g., the same location as in the first projection space 1415). In this example, the row index of MV F (e.g., row index = 1) is less than the row index of MV C (e.g., row index = 4), and the overwrite conflict can be resolved by maintaining the existing projection of MV C (e.g., by preventing MV F from overwriting the existing projection of MV C).
[0174] Since MV G is projected to the same location as the existing projection of MV A, another overwrite conflict arises from the projection of MV G from the second projection block 1620 to the first projection space 1615. Here, the row index of MV G (e.g., row index = 4) is greater than the row index of MV A (e.g., row index = 0), and the overwrite conflict is resolved by allowing the projection of MV G to overwrite the existing projection of MV A within the first projection space 1415.
[0175] In some respects, the systems and techniques described herein can implement block-based MV projection buffers as ring buffers (e.g., rather than...). Figure 10 The illustrated sliding MV buffer 1070). For example, Figure 11 This is an illustration of an example of a block-based juxtaposed MV projection 1100 using a circular MV buffer 1170. In this example, the three reference frame MV projections associated with block A (e.g., REF0 block A, REF1 block A, and REF2 block A) can be compared with those described above regarding... Figure 10 The methods described above are respectively projected into the first buffer 1171 and the second buffer 1172. Similarly, the projections of the three reference frames MV associated with block B can be the same as those described above. Figure 10 The methods described are respectively projected into the first buffer 1171, the second buffer 1172 and the third buffer 1173.
[0176] After the MV projection for blocks A and B has been completed, the first buffer 1171a can be used to decode the video data 1102a of block A, as described above. Figure 10 As described. However, after decoding the video data 1102a of block A and invalidating the contents of the first buffer 1171a, the now empty first buffer 1171a is not reassigned to a new third buffer (e.g., the MV buffer 1170 does not slide like the sliding MV buffer 1070). Instead, the MV buffer 1170 can be implemented as a ring buffer.
[0177] As illustrated, the three reference frames MV associated with block C (e.g., REF0 block C, REF1 block C, and REF2 block C) can be projected into the second buffer 1172, the third buffer 1173, and the first buffer 1171 in this order, respectively. The video data 1102b of block B can then be decoded from the second buffer 1172, and the second buffer 1172 can be cleared after successful decoding. Then, the three reference frames MV associated with block D can be projected into the third buffer 1173, the first buffer 1171, and the second buffer 1172 in this order, respectively. The video data 1102c of block C can then be decoded from the third buffer 1173, and the third buffer 1173 can be cleared after successful decoding.
[0178] Subsequently, the three reference frames MV associated with the additional block E (e.g., horizontally adjacent to the right side of block D, not shown) can be projected into the first buffer 1171, the second buffer 1172, and the third buffer 1173 respectively in this order. The video data 1102d of block D can then be decoded from the first buffer 1171, and the first buffer 1171 can be cleared after successful decoding. The above process can then be repeated for the remaining blocks included in the currently decoded frame 1102. In some respects, the above regarding... Figure 14 The described row-index-based overwrite conflict control is still compatible with Figure 11 The illustrated block-level juxtaposed MV projection ring buffer implementation is used together.
[0179] On the other hand, systems and techniques can use ping-pong memory buffers to implement block-based MV projection buffers. For example, Figure 12 This is an illustration of an example of a block-based juxtaposed MV projection 1200 using a ping-pong MV buffer 1270. In this example, the MV projections of three reference frames associated with block A (e.g., REF0 block A, REF1 block A, and REF2 block A) can each be projected into a first ping buffer and a second ping buffer (ping_1 and ping_2, respectively), and can additionally be projected into a first pong buffer and a second pong buffer (pong_1 and pong_2, respectively).
[0180] The three reference frames MV associated with block B can each be projected into a first ping buffer, a second ping buffer, and a third ping buffer (e.g., ping_1, ping_2, and ping_3, respectively), and can additionally be projected into a first pong buffer, a second pong buffer, and a third pong buffer (e.g., pong_1, pong_2, and pong_3, respectively). The video data 1202a of block A can be decoded from the ping_1 buffer.
[0181] Subsequently, the three reference frame MV projections associated with block C can each be projected into the first ping buffer, the second ping buffer, and the third ping buffer (e.g., ping_1, ping_2, and ping_3, respectively), and can additionally be projected into the first pong buffer, the second pong buffer, and the third pong buffer (e.g., pong_1, pong_2, and pong_3, respectively). Block B video data 1202b can be decoded from the pong_2 buffer. After decoding block B video data 1202b, the block A projection can be invalidated (e.g., cleared) from all buffers of the MV buffer 1270 (e.g., the block A projection can be invalidated from all ping buffers and all pong buffers).
[0182] Then, the projections of the three reference frames MV associated with block D can each be projected into the first ping buffer, the second ping buffer, and the third ping buffer (e.g., ping_1, ping_2, and ping_3, respectively), and can additionally be projected into the first pong buffer, the second pong buffer, and the third pong buffer (e.g., pong_1, pong_2, and pong_3, respectively). The block C video data 1202c can be decoded from the ping_3 buffer. After decoding the block C video data 1202c, the block B projection can be invalidated (e.g., cleared) from all buffers of the MV buffer 1270 (e.g., the block B projection can be invalidated from all ping buffers and all pong buffers).
[0183] Then, the three reference frame MV projections associated with the additional block E (e.g., horizontally adjacent to the right side of block D, not shown) can each be projected into a first ping buffer, a second ping buffer, and a third ping buffer (e.g., ping_1, ping_2, and ping_3, respectively), and can additionally be projected into a first pong buffer, a second pong buffer, and a third pong buffer (e.g., pong_1, pong_2, and pong_3, respectively). The block D video data 1202d can be decoded from the pong_1 buffer. After decoding the block D video data 1202d, the block D projection can be invalidated (e.g., cleared) from all buffers of the MV buffer 1270 (e.g., the block D projection can be invalidated from all ping buffers and all pong buffers).
[0184] The above process can then be repeated for the remaining blocks included in the currently decoded frame 1202. In some respects, the above regarding Figure 14 The described row-index-based overwrite conflict control is still compatible with Figure 12 The illustrated block-level juxtaposed MV projection ping-pong MV buffer implementation is used together.
[0185] On the other hand, the system and techniques can implement the block-based MV projection described herein using a block-level scan size of 128×128 (e.g., the opposite of the 64×64 block size mentioned above). For example, as Figure 13 As illustrated in Example Figure 1300, the frame of the currently decoded video data 1302 can be divided into multiple 128×128 blocks. Each 128×128 block of video data may include four 64×64 sub-blocks. For example, frame A video data 1302a may include sub-blocks A1, A2, A3, and A4, each of which is 64×64 in size.
[0186] The three reference frames MV associated with video data 1302a of frame A can be read in size 128×128. The 64×64 A1 and A2 portions of frame A MV can be projected into buffers Up_1 and Low_1 respectively, and also into buffers Up_2 and Low_2 respectively. The 64×64 A3 and A4 portions of frame A MV can be projected into buffers Up_1 and Low_1 respectively; into buffers Up_2 and Low_2 respectively; and into buffers Up_3 and Low_3 respectively.
[0187] The three reference frames MV associated with frame B video data 1302b can be read in size 128×128. The 64×64 B1 and B2 portions of frame B MV can be projected into buffers Up_2 and Low_2 respectively; they can be projected into buffers Up_3 and Low_3 respectively; and they can be projected into buffers Up_4 and Low_4 respectively.
[0188] The first and second buffers (e.g., Up_1, Up_2, Low_1, Low_2) can then be used to decode block A video data 1302a. In some aspects, the projection of the 64×64 portion of the 128×128 block of video data (e.g., frame A video data 1302a, frame B video data 1302b, etc.) can be performed in reverse N order, and the decoding of the 64×64 portion of the 128×128 block of video data can be performed in Z order. In some cases, performing the reverse N-order juxtaposed motion vector projection (e.g., the projection of the 64×64 portion or the projection of two 64×128 portions split from the 128×128 block of video data) can be associated with a 50% reduction in DDR or other memory bandwidth at the vertical tile boundaries. For example, for a 128×128 block of video data, the left-line buffer read at the vertical tile boundaries can be 64. In some examples, the A1 to A4 64×64 sub-blocks of video data 1302a in block A can be decoded in Z order, starting with A1, followed by A3, A2 and A4.
[0189] After decoding the four 64×64 sub-blocks of the 128×128 block A video data 1302a, buffers 1 and 2 can be invalidated, cleared, or zeroed. For example, buffers Up_1, Up_2, Low_1, and Low_2 can be invalidated. Subsequently, 64×64 sub-blocks B3, B4 and C1, C3 can be projected into the high and low buffers of buffers 1 to 4, where the old buffers 3 and 4 are reassigned to the new buffers 1 and 2. The old buffers 1 and 2 (now cleared) can be reassigned to the new buffers 3 and 4, thus in accordance with the above... Figure 10 The sliding buffer is implemented in the same or similar manner as described.
[0190] Figure 15 This is a flowchart illustrating an example of a process 1500 for encoding or decoding (decoding) video data according to various aspects described herein. At block 1502, process 1500 includes obtaining one or more first sets of juxtaposed motion vector data associated with a first block of video data included in the current frame of the video data. For example, the first block of video data may be associated with... Figure 4B The example shown is the current block 422. Figure 10 One or more of the illustrated blocks 1002a to 1002d Figure 11 One or more of the illustrated blocks 1102a to 1102d Figure 12 One or more of the illustrated blocks 1202a to 1202d and / or Figure 13 One or more of the illustrated blocks 1302a to 1302d are the same as or similar to each other. In some cases, the current frame of the video data may be the same as... Figure 5 The current image shown is 510. Figure 6A The current frame 622 is shown in the example. Figure 10 The example frame 1002, Figure 11 The example frame 1102, Figure 12 The illustrated frame 1202 and / or Figure 13 The example frame 1302 is the same as or similar to the example frame.
[0191] In some examples, one or more first sets of juxtaposed motion vector data include three first sets of juxtaposed motion vector data obtained from three corresponding reference frames associated with the current frame of the video data. For example, one or more first sets of juxtaposed motion vector data may include Figure 8A The exemplified Frame0 MV 812, Frame1 MV 814 and Frame2 MV 816 and / or Figure 9A One or more of Frame0 MV 912, Frame1 MV 914, and Frame2 MV 916 are illustrated. In some examples, it can be derived from... Figure 6BThree corresponding reference frames selected from the various types of reference frames 630 illustrated in the example obtain three first sets of juxtaposed motion vector data (e.g., three corresponding reference frames may be obtained for each of Frame0 MV 812 / 912, three corresponding reference frames may be obtained for each of Frame1 MV 814 / 914, and three corresponding reference frames may be obtained for each of Frame2 MV 816 / 916).
[0192] In some cases, the three first sets of juxtaposed motion vector data are positioned identically within the corresponding three reference frames. This position of the three first sets of juxtaposed motion vector data within the corresponding three reference frames may be the same as the position of the first block of video data within the current frame of the video data. In some examples, the first block of video data may be a 64×64 block, a 64×128 block, a 128×64 block, a 128×128 block, etc. In some cases, one or more first sets of juxtaposed motion vector data (e.g., three first sets of juxtaposed motion vector data) may have the same size as the first block of video data.
[0193] At block 1504, process 1500 includes projecting each of the one or more first sets of juxtaposed motion vector data into a first projected motion field associated with a first buffer. For example, each of the one or more first sets of juxtaposed motion vector data may be projected into a first projected motion field associated with a first buffer. Figure 9B The motion vector prediction buffer 970 shown is associated with a first projected motion field. One or more first sets of motion vector data are juxtaposed therewith... Figure 9A and Figure 9B In some cases, the first projected motion field may be the same as or similar to the illustrated Ref-0 frame MV 912. Figure 9B The illustrated Ref0 block projection 952 is the same as or similar.
[0194] In which one or more first sets of motion vector data are juxtaposed with those of... Figure 10 In some examples associated with the first block given by the illustrated block 1002a, the first projected motion field may be the same as or similar to projection 1071. One or more first sets of motion vector data juxtaposed therein are associated with... Figure 11 In some cases associated with the first block given by the illustrated block 1102a, the first projected motion field may be the same as or similar to projection 1171. In which one or more first sets of motion vector data are juxtaposed, and... Figure 12 In some examples associated with the first block given by the illustrated block 1202a, the first projected motion field may be the same as or similar to one or more projections in projection 1270; etc.
[0195] In some examples, each of one or more first sets of projected juxtaposed motion vector data includes splitting each of the one or more first sets of juxtaposed motion vector data into a first 64×128 set of juxtaposed motion vector data and a second 64×128 set of juxtaposed motion vector data (e.g., when the first block size and the juxtaposed motion vector data size are 128×128). In some cases, the first 64×128 set and the second 64×128 set of juxtaposed motion vector data may be non-overlapping. In some aspects, the first 64×128 set and the second 64×128 set of juxtaposed motion vector data may be projected in reverse N-order, for example, as... Figure 13 exemplified.
[0196] At box 1506, process 1500 includes obtaining one or more second sets of juxtaposed motion vector data associated with a second block of video data included in the current frame of the video data. For example, the second block of video data may be associated with... Figure 10 One or more of the illustrated blocks 1002a to 1002d Figure 11 One or more of the illustrated blocks 1102a to 1102d Figure 12 One or more of the illustrated blocks 1202a to 1202d and / or Figure 13 One or more of the illustrated blocks 1302a to 1302d are the same or similar.
[0197] In some cases, the second block of video data may be adjacent to the first block of video data. For example, the first block of video data may be... Figure 10 The example block is 1002a, while the second block of video data can be block 1002b; the first block of video data can be... Figure 11 The example shown is block 1102a, while the second block of video data could be block 1102b; etc. In some examples, the first block of video data could be... Figure 10 The example block is 1002b, while the second block of video data could be block 1002c; the first block of video data could be... Figure 11 The example is block 1102b, while the second block of video data could be block 1002c; etc.
[0198] In some examples, the one or more second sets of juxtaposed motion vector data include three second sets of juxtaposed motion vector data obtained from the corresponding three reference frames associated with the current frame of the video data. For example, the three second sets of juxtaposed motion vector data and the three first sets of juxtaposed motion vector data may be obtained from the same three corresponding reference frames (e.g., the same three corresponding reference frames described above with respect to box 1502 of process 1500).
[0199] In some cases, the three second sets of juxtaposed motion vector data are positioned identically within the corresponding three reference frames. This position of the three second sets of juxtaposed motion vector data within the corresponding three reference frames may be the same as the position of the second block of video data within the current frame of the video data. In some examples, the second block of video data may be a 64×64 block, a 64×128 block, a 128×64 block, a 128×128 block, etc. In some cases, one or more second sets of juxtaposed motion vector data (e.g., three second sets of juxtaposed motion vector data) may have the same size as the second block of video data.
[0200] In some examples, the first block of video data and the second block of video data may have the same size. The one or more first sets of juxtaposed motion vector data may have the same size as the one or more second sets of juxtaposed motion vector data (e.g., both may additionally or alternatively be the same size as the first block and the second block of video data).
[0201] At block 1508, process 1500 includes projecting each of the one or more second sets of juxtaposed motion vector data into the first projected motion field associated with the first buffer. For example, each of the one or more second sets of juxtaposed motion vector data may be projected into a first projected motion field associated with the first buffer. Figure 9B The illustrated motion vector prediction buffer 970 is associated with a first projected motion field. In some aspects, each of one or more second sets of juxtaposed motion vector data can be projected into the same first projected motion field to which each of one or more first sets of juxtaposed motion vector data is projected (e.g., as described above with respect to box 1504 of process 1500). In which one or more first sets of juxtaposed motion vector data are associated with... Figure 9A and Figure 9B In some cases similar to or the Ref-0 frame MV 912 illustrated, one or more second sets of juxtaposed motion vector data can be used with Figure 9A and Figure 9B The illustrated Ref-1 frame MV914 is the same as or similar.
[0202] In which one or more first sets of motion vector data are juxtaposed with those of... Figure 10 The illustrated block 1002a provides some examples associated with the first block, where one or more second sets of motion vector data can be combined with data from... Figure 10The illustrated block 1002b is associated with the second block. As previously described, the first projected motion field may be the same as or similar to projection 1071. At block 1504, the juxtaposed MV associated with block 1002a can be projected into the first projected motion field 1071 (e.g., generating...). Figure 10 (Example 'ProjA'). At box 1508, the juxtaposed MV associated with block 1002b can be projected on top of the existing projection of the first projection motion field 1071 (e.g., generating...). Figure 10 (Examples of 'Proj A and B').
[0203] At block 1510, process 1500 includes decoding the first block of video data based on the first projected motion field associated with the first buffer, projecting each of the one or more first sets of juxtaposed motion vector data and each of the one or more second sets of juxtaposed motion vector data onto the first projected motion field associated with the first buffer. For example, decoding may be based on projecting the juxtaposed MV data associated with block 1002a into buffer 1071 and the juxtaposed MV data associated with block 1002b into buffer 1071. Figure 10 The illustrated block 1002a is decoded, wherein the projection of the juxtaposed MV data associated with block 1002b overwrites the conflicting projection of the juxtaposed MV data associated with block 1002a (e.g., overwriting the projection at the same location within the first projection motion field).
[0204] In some examples, process 1500 may further include projecting each of the one or more first sets of juxtaposed motion vector data into a second projection motion field associated with a second buffer, wherein the second projection motion field is associated with decoding the second block of video data. For example, the juxtaposed MV data associated with block 1002a may be projected into a second projection motion field associated with second buffer 1072, such as... Figure 10 As illustrated, the juxtaposed MV data associated with block 1102a can be projected into a second projected motion field associated with the second buffer 1172, such as... Figure 11 exemplified; etc.
[0205] In some cases, process 1500 may further include projecting each of the one or more second sets of juxtaposed motion vector data into the second projection motion field associated with the second buffer; and projecting each of the one or more second sets of juxtaposed motion vector data into a third projection motion field associated with a third buffer, wherein the third projection motion field is associated with decoding a third block of video data, the third block of video data being located adjacent to the second block of video data. For example, the juxtaposed MV data associated with block 1002b may be projected into the second projection motion field associated with second buffer 1072, and may be projected into the third projection motion field associated with third buffer 1073, such as... Figure 10 As illustrated, the juxtaposed MV data associated with block 1102b can be projected into a second projection motion field associated with the second buffer 1172, and can be projected into a third projection motion field associated with the third buffer 1173, as shown. Figure 11 As illustrated; etc.; In the above example, the third block of video data can be located adjacent to the second block of video data (e.g., the third block of video data can be located near the second block of video data). Figure 10 The example block is 1002c; Figure 11 The illustrated block 1102c; etc. (same or similar).
[0206] In some examples, process 1500 may include obtaining one or more third sets of juxtaposed motion vector data associated with the third block of video data included in the current frame of the video data; and projecting each of the one or more third sets of juxtaposed motion vector data into the second projection motion field associated with the second buffer. For example, it may be possible to obtain... Figure 10 The illustrated third block 1002c is associated with the juxtaposed MV data and projected onto the second projection motion field associated with the second buffer 1072; thus, it is possible to obtain the same data as the second buffer 1072. Figure 11 The illustrated third block 1102c is associated with juxtaposed MV data and projected onto a second projection motion field associated with the second buffer 1172; etc. In some examples, after the juxtaposed MV associated with the first, second, and third blocks has been projected onto the second projection motion field associated with the second buffer, the second block of video data can be decoded based on the second projection motion field associated with the second buffer. For example, after the juxtaposed MV associated with the first block 1002a, the second block 1002b, and the third block 1002c has each been projected onto the second projection motion field associated with the second buffer 1072, the second block of video data can be decoded. Figure 10The illustrated second block 1002b is decoded; after the juxtaposed MVs associated with the first block 1102a, the second block 1102b, and the third block 1102c have each been projected into the second projection motion field associated with the second buffer 1172, the... Figure 11 The second 1102b block shown is decoded; etc.
[0207] In some examples, process 1500 may further include projecting each of the one or more third sets of juxtaposed motion vector data into the third projected motion field associated with the third buffer; and removing the first projected motion field from the first buffer based on successful decoding of the first block of video data. For example, with Figure 10 The juxtaposed MV data associated with the illustrated third block 1002c can be projected into a third projected motion field associated with the third buffer 1073, and the first projected motion field stored in the first buffer 1071 can be cleared or deleted based on successful decoding of the first block 1002a (e.g., based on an earlier projection of the first block 1002a and the second block 1002b into the first buffer 1071, as described above). In some examples, process 1500 may also include projecting each of the one or more third sets of juxtaposed motion vector data into a fourth projected motion field, wherein the fourth projected motion field is written into the first buffer. For example, in deletion Figure 10 After the contents of the illustrated first buffer 1071 are used, the first buffer 1071 can be used to store the fourth projected motion field. In some examples, process 1500 may also include deleting the second projected motion field from the second buffer based on successful decoding of the second block of video data. In some cases, the first buffer, the second buffer, and the third buffer may be included in a sliding buffer. In some examples, the first buffer, the second buffer, and the third buffer may be included in a ring buffer or a ping-pong memory.
[0208] In some specific implementations, the processes (or methods) described herein (including process 1500) may be computed by a computing device or apparatus (such as...) Figure 1 The system 100 shown is executed. For example, these processes can be performed by... Figure 1 and Figure 16 The encoding device 104 shown is derived from another video source-side device or video transmission device, or from... Figure 1 and Figure 17The decoding device 112 shown may be performed by another client-side device (such as a player device, display, or any other client-side device). In some cases, the computing device or apparatus may include one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components configured to perform the steps of process 1500.
[0209] In some examples, a computing device may include a mobile device, a desktop computer, a server computer and / or a server system, or other types of computing devices. Components of a computing device (e.g., one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components) may be implemented in circuitry. For example, a component may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.
[0210] In some examples, the computing device or apparatus may include a camera configured to capture video data (e.g., a video sequence) comprising video frames. In some examples, the camera or other capturing device that captures the video data is separate from the computing device, in which case the computing device receives or acquires the captured video data. The computing device may also include a network interface configured to communicate the video data. The network interface may be configured to communicate Internet Protocol (IP) based data or other types of data. In some examples, the computing device or apparatus may include a display for displaying samples of output video content, such as images of a video bitstream.
[0211] These processes can be described relative to logic flowcharts, whose operations represent sequences of operations that can be implemented by hardware, computer instructions, or combinations thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which operations are described is not intended to be construed as limiting, and any number of described operations can be combined in any order and / or in parallel to implement a process.
[0212] Additionally, these processes can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented by hardware or a combination thereof as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors. As noted above, the code can be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.
[0213] The decoding techniques discussed herein can be implemented in an example video encoding and decoding system (e.g., System 100). In some examples, a system includes a source device that provides encoded video data that will later be decoded by a destination device. Specifically, the source device provides the video data to the destination device via a computer-readable medium. The source and destination devices can include any of a wide variety of devices, including desktop computers, laptops, tablets, set-top boxes, handset phones (such as so-called "smartphones" and "smart tablets"), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, and the like. In some cases, the source and destination devices may be equipped for wireless communication.
[0214] The destination device may receive encoded video data to be decoded via a computer-readable medium. The computer-readable medium may include any type of medium or device capable of moving encoded video data from the source device to the destination device. In one example, the computer-readable medium may include a communication medium enabling the source device to transmit encoded video data directly to the destination device in real time. The encoded video data may be modulated according to communication standards, such as wireless communication protocols, and transmitted to the destination device. The communication medium may include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network (such as the Internet). The communication medium may include routers, switches, base stations, or any other equipment that may be useful in facilitating communication from the source device to the destination device.
[0215] In some examples, encoded data can be output from an output interface to a storage device. Similarly, encoded data can be accessed from a storage device via an input interface. The storage device can include any data storage medium of various distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, the storage device may correspond to a file server or another intermediate storage device that can store encoded video generated by a source device. The destination device can access the stored video data from the storage device via streaming or downloading. The file server can be any type of server capable of storing and sending encoded video data to the destination device. Example file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The destination device can access the encoded video data via any standard data connection, including an internet connection. The data connection can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing encoded video data stored on a file server. The transmission of encoded video data from a storage device can be streaming, downloading, or a combination thereof.
[0216] The technologies disclosed herein are not necessarily limited to wireless applications or setups. These technologies can be applied to video decoding to support any multimedia application in a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission (such as HTTP-based Dynamic Adaptive Streaming (DASH)), digital video encoded onto data storage media, decoding of digital video stored on data storage media, or other applications. In some examples, the system can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0217] In one example, the source device includes a video source, a video encoder, and an output interface. The destination device may include an input interface, a video decoder, and a display device. The video encoder of the source device may be configured to apply the techniques disclosed herein. In other examples, the source and destination devices may include other components or arrangements. For example, the source device may receive video data from an external video source such as an external camera. Similarly, the destination device may interface with an external display device, rather than including an integrated display device.
[0218] The example system described above is merely an example. Techniques for parallel processing of video data can be implemented by any digital video encoding and / or decoding device. While the techniques disclosed herein are generally implemented by video encoding devices, they can also be implemented by video encoders / decoders (often referred to as “codecs”). Furthermore, the techniques disclosed herein can also be implemented by video preprocessors. The source and destination devices are merely examples of such decoding devices, where the source device generates decoded video data to be sent to the destination device. In some examples, the source and destination devices can operate in a substantially symmetrical manner, such that each of these devices includes both video encoding and decoding components. Therefore, the example system can support one-way or two-way video transmission between video devices, for example, for video streaming, video playback, video broadcasting, or video telephony.
[0219] Video sources may include video capture devices, such as cameras, video archives including previously captured video, and / or video feed interfaces for receiving video from video content providers. Alternatively, a video source may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In some cases, if the video source is a camera, the source and destination devices may form a so-called camera phone or video phone. However, as mentioned above, the techniques described in this disclosure are generally applicable to video decoding and can be used in wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video may be encoded by a video encoder. The encoded video information may then be output to a computer-readable medium via an output interface.
[0220] As noted, computer-readable media may include transient media (such as wireless broadcasting or wired network transmission) or storage media (i.e., non-transient storage media), such as hard disks, flash drives, compressed optical discs, digital video optical discs, Blu-ray discs, or other computer-readable media. In some examples, a network server (not shown) may receive encoded video data from a source device and, for example, provide the encoded video data to a destination device via a network. Similarly, computing devices in a media production facility (such as an optical disc stamping facility) may receive encoded video data from a source device and produce an optical disc containing the encoded video data. Therefore, in various examples, computer-readable media can be understood to include one or more forms of computer-readable media.
[0221] The input interface of the destination device receives information from a computer-readable medium. The information in the computer-readable medium may include syntax information defined by the video encoder, which is also used by the video decoder. This syntax information includes characteristics and / or processed syntax elements of description blocks and other decoded units (e.g., groups of pictures (GOPs)). The display device displays the decoded video data to the user and may include any of a variety of display devices, such as a cathode ray tube (CRT), liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or another type of display device. Various aspects of this application have been described.
[0222] The specific details of the encoding device 104 and the decoding device 112 are respectively in Figure 16 and Figure 17 As shown in the image. Figure 16 This is a block diagram illustrating an example encoding device 104 that can implement one or more of the techniques described in this disclosure. For example, encoding device 104 can generate the syntax structures described herein (e.g., syntax structures of VPS, SPS, PPS, or other syntax elements). Encoding device 104 can perform intra-frame prediction decoding and inter-frame prediction decoding of video blocks within a video slice. As previously described, intra-frame decoding relies at least in part on spatial prediction to reduce or remove spatial redundancy within a given video frame or picture. Inter-frame decoding relies at least in part on temporal prediction to reduce or remove temporal redundancy within neighboring or surrounding frames of a video sequence. Intra-frame mode (I-mode) can refer to any of several spatial-based compression modes. Inter-frame mode (such as one-way prediction (P-mode) or two-way prediction (B-mode)) can refer to any of several temporal-based compression modes.
[0223] Encoding device 104 includes a segmentation unit 35, a prediction processing unit 41, a filter unit 63, an image memory 64, a summer 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 includes a motion estimation unit 42, a motion compensation unit 44, and an intra-frame prediction processing unit 46. For video block reconstruction, encoding device 104 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and a summer 62. Filter unit 63 is intended to represent one or more loop filters, such as deblocking filters, adaptive loop filters (ALF), and sample adaptive offset (SAO) filters. Although filter unit 63 is... Figure 16 The filter unit 63 is shown as an in-loop filter, but in other configurations, it may be implemented as a post-loop filter. The post-processing device 57 may perform additional processing on the encoded video data generated by the encoding device 104. In some instances, the techniques of this disclosure may be implemented by the encoding device 104. However, in other instances, one or more of the techniques of this disclosure may be implemented by the post-processing device 57.
[0224] like Figure 16 As shown, encoding device 104 receives video data, and segmentation unit 35 segments the data into video blocks. Segmentation may also include segmentation into slices, segments, tiles, or other larger units, as well as video block segmentation (e.g., according to a quadtree structure of LCUs and CUs). Encoding device 104 generally illustrates the components for encoding video blocks within a video slice to be encoded. Slices may be divided into multiple video blocks (and possibly into a set of video blocks referred to as tiles). Prediction processing unit 41 may select one of a variety of possible decoding modes for the current video block based on error results (e.g., decoding rate and distortion level, etc.), such as one or more of a variety of intra-frame prediction decoding modes or inter-frame prediction decoding modes. Prediction processing unit 41 may provide the resulting intra-frame decoded or inter-frame decoded blocks to summer 50 to generate residual block data and to summer 62 to reconstruct the encoded blocks for use as reference pictures.
[0225] The intra-prediction processing unit 46 within the prediction processing unit 41 can perform intra-prediction decoding of the current video block relative to one or more adjacent blocks in the same frame or slice as the current block to be decoded to provide spatial compression. The motion estimation unit 42 and motion compensation unit 44 within the prediction processing unit 41 perform inter-prediction decoding of the current video block relative to one or more predictive blocks in one or more reference pictures to provide temporal compression.
[0226] Motion estimation unit 42 can be configured to determine the inter-frame prediction mode of a video slice based on a predetermined mode of the video sequence. The predetermined mode can designate a video slice in the sequence as a P-slice, B-slice, or GPB-slice. Motion estimation unit 42 and motion compensation unit 44 can be highly integrated, but are illustrated separately for conceptual purposes. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of the video block. For example, the motion vectors can indicate the displacement of the prediction unit (PU) of a video block within the current video frame or picture relative to a predictive block within a reference picture.
[0227] A predictive block is a block found to closely match the PU of the video block to be decoded in terms of pixel differences, which can be determined by the sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some examples, encoding device 104 may compute values for sub-integer pixel positions of a reference image stored in image memory 64. For example, encoding device 104 may interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference image. Thus, motion estimation unit 42 may perform motion search relative to full-pixel positions and fractional pixel positions and output a motion vector with fractional pixel precision.
[0228] The motion estimation unit 42 calculates the motion vector of the PU by comparing the position of the PU in the video block in the inter-frame decoded slice with the position of the predictive block in the reference picture. The reference picture can be selected from a first reference picture list (list 0) or a second reference picture list (list 1), each of which identifies one or more reference pictures stored in the picture memory 64. The motion estimation unit 42 transmits the calculated motion vector to the entropy coding unit 56 and the motion compensation unit 44.
[0229] Motion compensation performed by motion compensation unit 44 may involve extracting or generating predictive blocks based on motion vectors determined by motion estimation, thereby potentially performing interpolation with subpixel precision. Upon receiving the motion vector of the PU for the current video block, motion compensation unit 44 can locate the predictive block pointed to by the motion vector in a list of reference images. Encoding device 104 forms residual video blocks by subtracting the pixel values of the predictive blocks from the pixel values of the current video block being decoded, thus forming pixel differences. The pixel differences form the residual data of the block and may include both luminance difference components and chrominance difference components. Summer 50 represents one or more components for which subtraction is performed. Motion compensation unit 44 may also generate syntax elements associated with video blocks and video slices for use by decoding device 112 when decoding video blocks of video slices.
[0230] Intra-prediction processing unit 46 can perform intra-prediction on the current block as an alternative to inter-prediction performed by motion estimation unit 42 and motion compensation unit 44, as described above. Specifically, intra-prediction processing unit 46 can determine the intra-prediction mode to be used for encoding the current block. In some examples, intra-prediction processing unit 46 can use various intra-prediction modes to encode the current block, for example, during a separate encoding process, and intra-prediction processing unit 46 can select an appropriate intra-prediction mode from tested modes for use. For example, intra-prediction processing unit 46 can use rate-distortion analysis for various tested intra-prediction modes to calculate rate-distortion values, and can select the intra-prediction mode with the best rate-distortion characteristics from the tested modes. Rate-distortion analysis typically determines the amount of distortion (or error) between the coded block and the original uncoded block encoded to produce the coded block, as well as the bit rate (i.e., the number of bits) used to produce the coded block. Intra-prediction processing unit 46 can calculate the ratio based on the distortion and rate of various coded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block.
[0231] In any case, after selecting an intra-prediction mode for a block, the intra-prediction processing unit 46 may provide information indicating the selected intra-prediction mode for the block to the entropy coding unit 56. The entropy coding unit 56 may encode the information indicating the selected intra-prediction mode. The encoding device 104 may include in the transmitted bitstream configuration data definitions for various blocks and indications of the most probable intra-prediction mode for each of the contexts, an intra-prediction mode index table, and a modified intra-prediction mode index table. The bitstream configuration data may include multiple intra-prediction mode index tables and multiple modified intra-prediction mode index tables (also referred to as codeword maps).
[0232] After prediction processing unit 41 generates a predictive block for the current video block via inter-frame prediction or intra-frame prediction, encoding device 104 forms a residual video block by subtracting the predictive block from the current video block. The residual video data in the residual block may be included in one or more TUs and applied to transform processing unit 52. Transform processing unit 52 uses a transform (such as Discrete Cosine Transform (DCT) or a conceptually similar transform) to transform the residual video data into residual transform coefficients. Transform processing unit 52 can transform the residual video data from the pixel domain to the transform domain, such as the frequency domain.
[0233] The transform processing unit 52 can transmit the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, the quantization unit 54 can then perform a scan of the matrix including the quantized transform coefficients. Alternatively, the entropy coding unit 56 can perform the scan.
[0234] After quantization, entropy coding unit 56 performs entropy coding on the quantized transform coefficients. For example, entropy coding unit 56 may perform context-adaptive variable-length decoding (CAVLC), context-adaptive binary arithmetic decoding (CABAC), syntax-based context-adaptive binary arithmetic decoding (SBAC), probabilistic interval partitioned entropy (PIPE) decoding, or another entropy coding technique. After entropy coding by entropy coding unit 56, the encoded bitstream can be sent to decoding device 112, or archived for later transmission or retrieval by decoding device 112. Entropy coding unit 56 may also perform entropy coding on the motion vectors and other syntax elements of the current video slice being decoded.
[0235] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct residual blocks in the pixel domain for later use as reference blocks in reference images. Motion compensation unit 44 can compute reference blocks by adding the residual blocks to a predictive block of a reference image from the list of reference images. Motion compensation unit 44 can also apply one or more interpolation filters to the reconstructed residual blocks to compute sub-integer pixel values for use in motion estimation. Summer 62 adds the reconstructed residual blocks to the motion-compensated predictive blocks generated by motion compensation unit 44 to produce reference blocks for storage in image memory 64. The reference blocks can be used by motion estimation unit 42 and motion compensation unit 44 for inter-frame prediction of blocks in subsequent video frames or images.
[0236] In this way, Figure 16 The encoding device 104 is configured to perform any of the techniques described herein (including those mentioned above). Figure 15 Examples of video encoders (described in the process). In some cases, some techniques of this disclosure can also be implemented by post-processing device 57.
[0237] Figure 17 This is a block diagram illustrating an example decoding device 112. Decoding device 112 includes an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, a summer 90, a filter unit 91, and an image memory 92. The prediction processing unit 81 includes a motion compensation unit 82 and an intra-frame prediction processing unit 84. In some examples, decoding device 112 may perform operations roughly related to... Figure 16 The encoding device 104 describes the encoding cycle as the inverse of the decoding cycle.
[0238] During the decoding process, decoding device 112 receives an encoded video bitstream representing video blocks and associated syntax elements of an encoded video slice transmitted by encoding device 104. In some aspects, decoding device 112 may receive the encoded video bitstream from encoding device 104. In some aspects, decoding device 112 may receive the encoded video bitstream from network entity 79 (such as a server, a media-aware network element (MANE), a video editor / stitcher, or other such device configured to implement one or more of the techniques described above). Network entity 79 may or may not include encoding device 104. Some of the techniques described in this disclosure may be implemented by network entity 79 before it sends the encoded video bitstream to decoding device 112. In some video decoding systems, network entity 79 and decoding device 112 may be part of a separate device, while in other instances, the functionality described with respect to network entity 79 may be performed by the same device including decoding device 112.
[0239] The entropy decoding unit 80 of the decoding device 112 performs entropy decoding on the bitstream to generate quantized coefficients, motion vectors, and other syntax elements. The entropy decoding unit 80 forwards the motion vectors and other syntax elements to the prediction processing unit 81. The decoding device 112 can receive syntax elements at the video slice level and / or the video block level. The entropy decoding unit 80 can process and parse both fixed-length and variable-length syntax elements from one or more parameter sets (such as VPS, SPS, and PPS).
[0240] When a video slice is decoded into an intra-frame decoded (I) slice, the intra-frame prediction processing unit 84 of the prediction processing unit 81 can generate prediction data for the video block of the current video slice based on the intra-frame prediction mode notified by a signal and data from previously decoded blocks from the current frame or picture. When a video frame is decoded into an inter-frame decoded (i.e., B, P, or GPB) slice, the motion compensation unit 82 of the prediction processing unit 81 generates predictive blocks for the video block of the current video slice based on motion vectors and other syntax elements received from the entropy decoding unit 80. Predictive blocks can be generated from one of the reference pictures in the reference picture list. The decoding device 112 can construct a reference frame list (list 0 and list 1) based on the reference pictures stored in the picture memory 92 using a default construction technique.
[0241] Motion compensation unit 82 determines prediction information for video blocks in the current video slice by parsing motion vectors and other syntax elements, and uses this prediction information to generate predictive blocks for the current video slice being decoded. For example, motion compensation unit 82 may use one or more syntax elements in the parameter set to determine the prediction mode (e.g., intra-frame prediction or inter-frame prediction) for decoding video blocks in the video slice, the inter-frame prediction slice type (e.g., B-slice, P-slice, or GPB-slice), the construction information of one or more reference picture lists for the slice, the motion vectors of each inter-frame coded video block in the slice, the inter-frame prediction state of each inter-frame decoded video block in the slice, and other information for decoding video blocks in the current video slice.
[0242] The motion compensation unit 82 can also perform interpolation based on an interpolation filter. The motion compensation unit 82 can use an interpolation filter, such as that used by the encoding device 104 during the encoding of a video block, to calculate the interpolated values of sub-integer pixels of a reference block. In this case, the motion compensation unit 82 can determine the interpolation filter used by the encoding device 104 from the received syntax elements and can use that interpolation filter to generate predictive blocks.
[0243] The inverse quantization unit 86 inverse-quantizes or dequantizes the quantized transform coefficients provided in the bitstream and decoded by the entropy decoding unit 80. The inverse quantization process may include determining the degree of quantization using quantization parameters calculated by the encoding device 104 for each video block in the video slice, and similarly determining the degree of inverse quantization to be applied. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT or other suitable inverse transform), inverse integer transform, or conceptually similar inverse transform process to the transform coefficients to produce residual blocks in the pixel domain.
[0244] After the motion compensation unit 82 generates a predictive block for the current video block based on motion vectors and other syntax elements, the decoding device 112 forms a decoded video block by adding the residual block from the inverse transform processing unit 88 to the corresponding predictive block generated by the motion compensation unit 82. The summer 90 represents one or more components for which the summation operation is performed. Loop filters (in or after the decoding loop) can also be used, if needed, to smooth pixel transitions or otherwise improve video quality. The filter unit 91 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although the filter unit 91 is in Figure 17 The image shown is an in-loop filter, but in other configurations, filter unit 91 can be implemented as a post-loop filter. The decoded video block from a given frame or image is then stored in image memory 92, which stores reference images for subsequent motion compensation. Image memory 92 also stores the decoded video for later display on a display device (such as...). Figure 1 The video is displayed on the destination device 122.
[0245] In this way, Figure 17 The decoding device 112 is configured to perform any of the techniques described herein (including those mentioned above). Figure 15 An example of a video decoder (describing the process).
[0246] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media include, but are not limited to, magnetic disks or magnetic tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, memory, or memory devices. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments may be coupled to other code segments or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, independent variables, parameters, data, etc., can be transmitted, forwarded, or sent through any suitable means, including memory sharing, message passing, token passing, network sending, etc.
[0247] In some examples, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and the signals themselves.
[0248] Specific details are provided in the foregoing description to provide a thorough understanding of the aspects and examples provided herein. However, those skilled in the art will understand that the examples can be practiced without these specific details. For clarity, in some instances, the technology may be presented as comprising individual functional blocks, including functional blocks containing devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes and other components may be shown as components in block diagram form so as not to obscure aspects of the present disclosure with unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures and techniques may be shown without unnecessary detail to avoid obscuring aspects of the present disclosure.
[0249] The examples above may be described as processes or methods, depicted as flowcharts, diagrams, data flow diagrams, structure diagrams, or block diagrams. While a flowchart may describe operations as a sequential process, many operations within an operation can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but a process may have additional steps not included in the accompanying diagrams. A process may correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, its termination may correspond to the function returning to its calling function or the main function.
[0250] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that configure, or otherwise configure, a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. The portion may be accessible via a network of the computer resources used. The computer-executable instructions may be, for example, binary, intermediate format instructions, such as assembly language, firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, information used, and / or information created during the methods according to the described examples include disks or optical discs, flash memory, USB devices with non-volatile memory, networked storage devices, etc.
[0251] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptop computers, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or interlocking cards. By further example, such functionality may also be implemented on circuit boards of different chips or different processes executed on a single device.
[0252] Instructions, media for delivering such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means of providing the functionality described in this disclosure.
[0253] In the foregoing description, aspects of this application have been described with reference to specific examples of this application; however, those skilled in the art will recognize that this application is not limited thereto. Therefore, while illustrative examples of this application have been described in detail herein, it should be understood that the inventive concepts can be implemented and employed in a variety of other ways, and the appended claims are intended to be construed as including these variations, unless limited by prior art. Various features and aspects of the above applications may be used individually or in combination. Furthermore, aspects of this disclosure may be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that, in alternative examples, the methods may be performed in a different order than that described.
[0254] Those skilled in the art will understand that, without departing from the scope of this specification, the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced by the less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively.
[0255] When a component is described as being “configured” to perform certain operations, such a configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.
[0256] The phrase “coupled to” means any component is physically connected directly or indirectly to another component, and / or any component communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0257] The claim language or other language in this disclosure that states "at least one of" and / or "one or more of" in a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, the claim language stating "at least one of A and B" means A, B, or A and B. In another example, the claim language stating "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language set "at least one of" and / or "one or more of" in a set does not limit the set to the items listed in the set. For example, the claim language stating "at least one of A and B" may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.
[0258] The various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the examples disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been broadly described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.
[0259] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices (mobile phones), or integrated circuit devices with multiple uses, including applications in wireless communication devices (mobile phones) and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagated signals or waves.
[0260] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternatives, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within dedicated software or hardware modules configured for encoding and decoding, or incorporated into a combined video encoder-decoder (CODEC).
[0261] The exemplary aspects of this disclosure include:
[0262] Aspect 1: An apparatus for processing video data, the apparatus comprising: at least one memory; and at least one processor coupled to the at least one memory, the at least one processor being configured to: obtain one or more first sets of juxtaposed motion vector data associated with a first block of video data included in a current frame of the video data; project each of the one or more first sets of juxtaposed motion vector data into a first projection motion field associated with a first buffer; obtain one or more second sets of juxtaposed motion vector data associated with a second block of video data included in the current frame of the video data; project each of the one or more second sets of juxtaposed motion vector data into the first projection motion field associated with the first buffer; and decode the first block of video data based on the first projection motion field associated with the first buffer, based on each of the one or more first sets of juxtaposed motion vector data and each of the one or more second sets of juxtaposed motion vector data projected into the first projection motion field associated with the first buffer.
[0263] Aspect 2: The apparatus according to aspect 1, wherein the at least one processor is further configured to: project each of the one or more first sets of juxtaposed motion vector data into a second projected motion field associated with a second buffer, wherein the second projected motion field is associated with decoding the second block of video data.
[0264] Aspect 3: The apparatus according to aspect 2, wherein the second block of video data is adjacent to the first block of video data.
[0265] Aspect 4: The apparatus according to any one of Aspects 2 to 3, wherein the at least one processor is further configured to: project each of the one or more second sets of juxtaposed motion vector data into a second projected motion field associated with the second buffer; and project each of the one or more second sets of juxtaposed motion vector data into a third projected motion field associated with a third buffer, wherein the third projected motion field is associated with decoding a third block of video data, the third block of video data being located adjacent to the second block of video data.
[0266] Aspect 5: The apparatus according to aspect 4, wherein the at least one processor is further configured to: obtain one or more third sets of juxtaposed motion vector data associated with the third block of video data included in the current frame of video data; project each of the one or more third sets of juxtaposed motion vector data into a second projected motion field associated with the second buffer; and decode the second block of video data based on the second projected motion field associated with the second buffer.
[0267] Aspect 6: The apparatus according to aspect 5, wherein the at least one processor is further configured to: project each of the one or more third sets of juxtaposed motion vector data into the third projected motion field associated with the third buffer; delete the first projected motion field from the first buffer based on successful decoding of the first block of video data; and project each of the one or more third sets of juxtaposed motion vector data into a fourth projected motion field, wherein the fourth projected motion field is written into the first buffer.
[0268] Aspect 7: The apparatus according to any one of Aspects 5 to 6, wherein the at least one processor is further configured to: remove the second projected motion field from the second buffer based on successful decoding of the second block of video data.
[0269] Aspect 8: The apparatus according to any one of Aspects 6 to 7, wherein the first buffer, the second buffer and the third buffer are included in a sliding buffer.
[0270] Aspect 9: The apparatus according to any one of Aspects 6 to 8, wherein the first buffer, the second buffer and the third buffer are included in a ring buffer or a ping-pong memory.
[0271] Aspect 10: The apparatus according to any one of Aspects 1 to 9, wherein: the one or more first sets of juxtaposed motion vector data comprise three first sets of juxtaposed motion vector data obtained from three corresponding reference frames associated with the current frame of the video data; and the one or more second sets of juxtaposed motion vector data comprise three second sets of juxtaposed motion vector data obtained from the three corresponding reference frames associated with the current frame of the video data.
[0272] Aspect 11: The apparatus according to aspect 10, wherein: the positions of the three first sets of juxtaposed motion vector data in the corresponding three reference frames are the same; and the positions of the three first sets of juxtaposed motion vector data in the corresponding three reference frames are the same as the positions of the first block of video data in the current frame of the video data.
[0273] Aspect 12: The apparatus according to any one of Aspects 10 and 11, wherein: the positions of the three second sets of juxtaposed motion vector data in the corresponding three reference frames are the same; and the positions of the three second sets of juxtaposed motion vector data in the corresponding three reference frames are the same as the positions of the second block of video data in the current frame of the video data.
[0274] Aspect 13: The apparatus according to any one of Aspects 1 to 12, wherein: the first block of video data is a 128×128 block of video data; and the one or more first sets of juxtaposed motion vector data have the same size as the first block of video data.
[0275] Aspect 14: The apparatus according to aspect 13, wherein, in order to project each of the one or more first sets of juxtaposed motion vector data, the at least one processor is configured to: split each of the one or more first sets of juxtaposed motion vector data into a first 64×128 set of juxtaposed motion vector data and a second 64×128 set of juxtaposed motion vector data; and project the first 64×128 set and the second 64×128 set in reverse N order.
[0276] Aspect 15: The apparatus according to aspect 14, wherein the first 64×128 set and the second 64×128 set are non-overlapping.
[0277] Aspect 16: A method for processing video data, the method comprising: obtaining one or more first sets of juxtaposed motion vector data associated with a first block of video data included in a current frame of the video data; projecting each of the one or more first sets of juxtaposed motion vector data into a first projection motion field associated with a first buffer; obtaining one or more second sets of juxtaposed motion vector data associated with a second block of video data included in the current frame of the video data; projecting each of the one or more second sets of juxtaposed motion vector data into the first projection motion field associated with the first buffer; and decoding the first block of video data based on the first projection motion field associated with the first buffer, based on each of the one or more first sets of juxtaposed motion vector data and each of the one or more second sets of juxtaposed motion vector data projected into the first projection motion field associated with the first buffer.
[0278] Aspect 17: The method according to aspect 16, the method further comprising: projecting each of the one or more first sets of juxtaposed motion vector data into a second projected motion field associated with a second buffer, wherein the second projected motion field is associated with decoding the second block of video data.
[0279] Aspect 18: According to the method of aspect 17, the second block of video data is adjacent to the first block of video data.
[0280] Aspect 19: The method according to any one of Aspects 17 to 18, the method further comprising: projecting each of the one or more second sets of juxtaposed motion vector data into a second projected motion field associated with the second buffer; and projecting each of the one or more second sets of juxtaposed motion vector data into a third projected motion field associated with a third buffer, wherein the third projected motion field is associated with decoding a third block of video data, the third block of video data being located adjacent to the second block of video data.
[0281] Aspect 20: The method according to aspect 19, the method further comprising: obtaining one or more third sets of juxtaposed motion vector data associated with the third block of video data included in the current frame of video data; projecting each of the one or more third sets of juxtaposed motion vector data into a second projected motion field associated with the second buffer; and decoding the second block of video data based on the second projected motion field associated with the second buffer.
[0282] Aspect 21: The method according to aspect 20, the method further comprising: projecting each of the one or more third sets of juxtaposed motion vector data into a third projected motion field associated with the third buffer; deleting the first projected motion field from the first buffer based on successful decoding of the first block of video data; and projecting each of the one or more third sets of juxtaposed motion vector data into a fourth projected motion field, wherein the fourth projected motion field is written into the first buffer.
[0283] Aspect 22: The method according to any one of Aspects 20 to 21, the method further comprising: removing the second projected motion field from the second buffer based on successful decoding of the second block of video data.
[0284] Aspect 23: The method according to any one of Aspects 21 to 22, wherein the first buffer, the second buffer and the third buffer are included in a sliding buffer.
[0285] Aspect 24: The method according to any one of aspects 21 to 23, wherein the first buffer, the second buffer and the third buffer are included in a ring buffer or a ping-pong memory.
[0286] Aspect 25: The method according to any one of Aspects 16 to 24, wherein: the one or more first sets of juxtaposed motion vector data comprise three first sets of juxtaposed motion vector data obtained from three corresponding reference frames associated with the current frame of the video data; and the one or more second sets of juxtaposed motion vector data comprise three second sets of juxtaposed motion vector data obtained from the three corresponding reference frames associated with the current frame of the video data.
[0287] Aspect 26: According to the method of aspect 25, wherein: the positions of the three first sets of juxtaposed motion vector data in the corresponding three reference frames are the same; and the positions of the three first sets of juxtaposed motion vector data in the corresponding three reference frames are the same as the positions of the first block of video data in the current frame of video data.
[0288] Aspect 27: The method according to any one of Aspects 25 to 26, wherein: the positions of the three second sets of juxtaposed motion vector data in the corresponding three reference frames are the same; and the positions of the three second sets of juxtaposed motion vector data in the corresponding three reference frames are the same as the positions of the second block of video data in the current frame of the video data.
[0289] Aspect 28: The method according to any one of Aspects 16 to 27, wherein: the first block of video data is a 128×128 block of video data; and the one or more first sets of juxtaposed motion vector data have the same size as the first block of video data.
[0290] Aspect 29: The method according to any one of aspects 16 to 28, wherein each of the one or more first sets of projected juxtaposed motion vector data comprises: splitting each of the one or more first sets of juxtaposed motion vector data into a first 64×128 set of juxtaposed motion vector data and a second 64×128 set of juxtaposed motion vector data; and projecting the first 64×128 set and the second 64×128 set in reverse N order.
[0291] Aspect 30: The method according to aspect 29, wherein the first 64×128 set and the second 64×128 set are non-overlapping.
[0292] Aspect 31: A non-transitory computer-readable storage medium having instructions stored thereon, the instructions, when executed by one or more processors, causing the one or more processors to perform any of the operations described in aspects 1 to 30.
[0293] Aspect 32: An apparatus comprising components for performing any of the operations described in aspects 1 to 30.
[0294] Aspect 33: A method for performing any of the operations described in aspects 1 to 15.
Claims
1. An apparatus for processing video data, the apparatus comprising: At least one memory; and At least one processor, coupled to the at least one memory, is configured to: Obtain one or more first sets of juxtaposed motion vector data associated with a first block of video data included in the current frame of the video data; Each of the one or more first sets of juxtaposed motion vector data is split into two or four parts of the juxtaposed motion vector data; Two or four portions of each of the one or more first sets of juxtaposed motion vector data are projected into a first projected motion field associated with a first buffer in reverse N order. Obtain one or more second sets of juxtaposed motion vector data associated with a second block of video data included in the current frame of the video data; Each of the one or more second sets of juxtaposed motion vector data is split into two or four parts of the juxtaposed motion vector data; Two or four portions of each of the one or more second sets of juxtaposed motion vector data are projected into the first projected motion field associated with the first buffer in reverse N order. as well as The first block of video data is decoded based on the first projected motion field associated with the first buffer, which projects each of the one or more first sets of juxtaposed motion vector data and each of the one or more second sets of juxtaposed motion vector data onto the first projected motion field associated with the first buffer.
2. The apparatus of claim 1, wherein the at least one processor is further configured to: Each of the one or more first sets of juxtaposed motion vector data is projected into a second projected motion field associated with a second buffer, wherein the second projected motion field is associated with decoding the second block of video data.
3. The apparatus of claim 2, wherein the second block of video data is adjacent to the first block of video data.
4. The apparatus of claim 2, wherein the at least one processor is further configured to: Each of the one or more second sets of juxtaposed motion vector data is projected into the second projected motion field associated with the second buffer.
5. The apparatus of claim 4, wherein the at least one processor is further configured to: Obtain one or more third sets of juxtaposed motion vector data associated with a third block of video data included in the current frame of the video data; Project each of the one or more third sets of juxtaposed motion vector data into the second projected motion field associated with the second buffer; and The second block of video data is decoded based on the second projected motion field associated with the second buffer.
6. The apparatus of claim 5, wherein the at least one processor is further configured to: The second projected motion field is removed from the second buffer based on the successful decoding of the second block of video data.
7. The apparatus of claim 5, wherein the first buffer and the second buffer are included in a sliding buffer.
8. The apparatus of claim 5, wherein the first buffer and the second buffer are included in a ring buffer or a ping-pong memory.
9. The apparatus according to claim 1, wherein: The one or more first sets of juxtaposed motion vector data comprise three first sets of juxtaposed motion vector data obtained from three corresponding reference frames associated with the current frame of the video data; and The one or more second sets of juxtaposed motion vector data include three second sets of juxtaposed motion vector data obtained from the corresponding three reference frames associated with the current frame of the video data.
10. The apparatus according to claim 9, wherein: The three first sets of juxtaposed motion vector data are positioned identically within the corresponding three reference frames; and The positions of the three first sets of juxtaposed motion vector data within the corresponding three reference frames are the same as the positions of the first block of video data within the current frame of the video data.
11. The apparatus according to claim 9, wherein: The three second sets of juxtaposed motion vector data are positioned identically within the corresponding three reference frames; and The positions of the three second sets of juxtaposed motion vector data within the corresponding three reference frames are the same as the positions of the second block of video data within the current frame of the video data.
12. The apparatus according to claim 1, wherein: The first block of video data is a 128×128 block of video data; and The first set of one or more juxtaposed motion vector data has the same size as the first block of video data.
13. The apparatus of claim 12, wherein, for projecting and juxtaposing each of the one or more first sets of motion vector data, the at least one processor is configured to: Each of the one or more first sets of juxtaposed motion vector data is split into a first 64×128 set of juxtaposed motion vector data and a second 64×128 set of juxtaposed motion vector data; and The first 64×128 set and the second 64×128 set are projected in reverse N order.
14. The apparatus of claim 13, wherein the first 64×128 set and the second 64×128 set are non-overlapping.
15. A method for processing video data, the method comprising: Obtain one or more first sets of juxtaposed motion vector data associated with a first block of video data included in the current frame of the video data; Each of the one or more first sets of juxtaposed motion vector data is split into two or four parts of the juxtaposed motion vector data; Two or four portions of each of the one or more first sets of juxtaposed motion vector data are projected into a first projected motion field associated with a first buffer in reverse N order. Obtain one or more second sets of juxtaposed motion vector data associated with a second block of video data included in the current frame of the video data; Each of the one or more second sets of juxtaposed motion vector data is split into two or four parts of the juxtaposed motion vector data; Two or four portions of each of the one or more second sets of juxtaposed motion vector data are projected into the first projected motion field associated with the first buffer in reverse N order. as well as The first block of video data is decoded based on the first projected motion field associated with the first buffer, which projects each of the one or more first sets of juxtaposed motion vector data and each of the one or more second sets of juxtaposed motion vector data onto the first projected motion field associated with the first buffer.
16. The method according to claim 15, further comprising: Each of the one or more first sets of juxtaposed motion vector data is projected into a second projected motion field associated with a second buffer, wherein the second projected motion field is associated with decoding the second block of video data.
17. The method of claim 16, wherein the second block of video data is adjacent to the first block of video data.
18. The method according to claim 16, further comprising: Each of the one or more second sets of juxtaposed motion vector data is projected into the second projected motion field associated with the second buffer.
19. The method according to claim 18, further comprising: Obtain one or more third sets of juxtaposed motion vector data associated with a third block of video data included in the current frame of the video data; Each of the one or more third sets of juxtaposed motion vector data is projected into the second projected motion field associated with the second buffer; as well as The second block of video data is decoded based on the second projected motion field associated with the second buffer.
20. The method of claim 19, further comprising: The second projected motion field is removed from the second buffer based on the successful decoding of the second block of video data.
21. The method of claim 19, wherein the first buffer and the second buffer are included in a sliding buffer.
22. The method of claim 19, wherein the first buffer and the second buffer are included in a ring buffer or a ping-pong memory.
23. The method according to claim 15, wherein: The one or more first sets of juxtaposed motion vector data comprise three first sets of juxtaposed motion vector data obtained from three corresponding reference frames associated with the current frame of the video data; and The one or more second sets of juxtaposed motion vector data include three second sets of juxtaposed motion vector data obtained from the corresponding three reference frames associated with the current frame of the video data.
24. The method according to claim 23, wherein: The three first sets of juxtaposed motion vector data are positioned identically within the corresponding three reference frames; and The positions of the three first sets of juxtaposed motion vector data within the corresponding three reference frames are the same as the positions of the first block of video data within the current frame of the video data.
25. The method according to claim 23, wherein: The three second sets of juxtaposed motion vector data are positioned identically within the corresponding three reference frames; and The positions of the three second sets of juxtaposed motion vector data within the corresponding three reference frames are the same as the positions of the second block of video data within the current frame of the video data.
26. The method according to claim 15, wherein: The first block of video data is a 128×128 block of video data; and The first set of one or more juxtaposed motion vector data has the same size as the first block of video data.
27. The method of claim 15, wherein each of the one or more first sets of projected and juxtaposed motion vector data comprises: Each set in the one or more first sets of juxtaposed motion vector data is split into a first 64×128 set of juxtaposed motion vector data and a second 64×128 set of juxtaposed motion vector data; as well as The first 64×128 set and the second 64×128 set are projected in reverse N order.
28. The method of claim 27, wherein the first 64×128 set and the second 64×128 set are non-overlapping.
29. A non-transitory computer-readable storage medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 15 to 28.
30. A computer program product comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 15 to 28.
Citation Information
Patent Citations
Method and apparatus for generating motion field motion vectors for blocks of current frame in on-the-fly manner
CN112954363A