Early notifications for low latency video decoders

By dividing the image data into multiple parts and parallelizing the process, the downstream equipment is notified in advance to complete the decoding, which solves the problem of excessively long image data decoding and transmission time in the prior art, and realizes low-latency video decoding, which is suitable for extended reality, vehicle and mobile applications.

CN120435863APending Publication Date: 2025-08-05QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380089509.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-01-05
Filing Date
2023-12-14
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing video decoding technology has bottlenecks in low-latency processing, resulting in the long time of decoding and sending image data, which cannot meet the real-time needs of high-fidelity and high-resolution video.

Method used

The image data is divided into multiple parts by early notification system, and the parallel processing technology is used to notify the downstream equipment in advance that the image part is decoding is completed, and the image data is sent in parallel to reduce the decoding delay.

Benefits of technology

It effectively reduces the total time for image data decoding and sending, improves the real-timeness of video processing, and is suitable for scenarios such as extended reality, vehicle applications and mobile applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120435863A_ABST
    Figure CN120435863A_ABST
Patent Text Reader

Abstract

Systems and techniques for processing video data are described herein. For example, a process may include determining a number of rows of pixels for one or more portions of an image. The process may also include obtaining first encoded data for a first portion of the image; decoding the first encoded data to generate first pixel data for the number of pixel rows of the first portion of the image; outputting first pixel data for a first portion of the image to a memory; and outputting an indication that the first portion of the image is available.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates generally to video processing. For example, aspects of the present application relate to improving video coding techniques (eg, encoding and / or decoding video) with respect to early notification for low-latency video decoders. Background Art

[0002] Wireless communication systems. Digital video capabilities can be incorporated into a variety of devices, including digital televisions, digital live broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radio phones (so-called "smart phones"), video teleconferencing equipment, video streaming equipment, and the like. Such devices allow processing and output of video data for consumption. Digital video data includes a large amount of data for satisfying the needs of consumers and video providers. For example, consumers of video data expect the highest quality video with high fidelity, high resolution, high frame rate, etc. As a result, the large amount of video data required to meet these demands has burdened the communication networks and devices that process and store video data.

[0003] Digital video devices can implement video decoding technology to compress video data. Video decoding is performed according to one or more video decoding standards or formats. For example, video coding standards or formats include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), MPEG-2 Part 2 coding (MPEG stands for Moving Picture Experts Group), Basic Video Coding (EVC), etc., and proprietary video encoder-decoder (codec) / format (such as AOMedia Video 1 (AV1) or Universal Bandwidth Compression (UBWC) developed by the Alliance for Open Media). One goal of video decoding technology is to compress video data into a form that uses a lower bit rate while avoiding or minimizing the degradation of video quality. As evolving video services become available, there is a need for technology to accelerate video decoding to allow video data to be decompressed so that it can be displayed in the shortest possible time. Summary of the Invention

[0004] Systems and techniques for improved video processing, such as video encoding and / or decoding, are described herein. For example, a system can provide advance notification to a low-latency video decoder to help parallelize decoding image data and sending the image data. According to at least one example, a device for processing video data is provided. The device includes a memory and a processor (e.g., configured in a circuit) coupled to the memory. The processor is configured to: determine a number of pixel rows for one or more portions of an image; obtain first encoded data for a first portion of the image; decode the first encoded data to generate first pixel data for the number of pixel rows for the first portion of the image; output the first pixel data for the first portion of the image to the memory; and output an indication that the first portion of the image is available.

[0005] As another illustrative example, a method for processing video data is provided. The method includes obtaining second encoded data for a second portion of an image, wherein the second portion of the image does not overlap with the first portion of the image; decoding the second encoded data to generate second pixel data for the second portion of the image for the number of pixel rows; outputting the second pixel data for the second portion of the image to the memory; and outputting an indication that the second portion of the image is available.

[0006] In another example, a non-transitory computer-readable medium is provided. The non-transitory computer-readable medium has instructions stored thereon that, when executed by a processor, cause the processor to: determine a number of pixel rows for one or more portions of an image; obtain first encoded data for a first portion of the image; decode the first encoded data to generate first pixel data for the number of pixel rows of the first portion of the image; output the first pixel data for the first portion of the image to a memory; and output an indication that the first portion of the image is available.

[0007] As another example, an apparatus for processing video data is provided, the apparatus comprising: means for obtaining second encoded data for a second portion of the image, wherein the second portion of the image does not overlap with the first portion of the image; means for decoding the second encoded data to generate second pixel data for the second portion of the image for the number of pixel rows; means for outputting the second pixel data for the second portion of the image to the memory; and means for outputting an indication that the second portion of the image is available.

[0008] In some aspects, one or more of the apparatus or devices described herein are part of and / or include a mobile device (e.g., a mobile phone or so-called "smartphone" or other mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a camera, a personal computer, a laptop computer, a server computer, a vehicle or a computing device or component of a vehicle, a robotic device or system, a television, or other device. In some aspects, the apparatus or device includes a camera or multiple cameras for capturing one or more pictures, images, or frames. In some aspects, the apparatus or device includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the apparatus or device may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyrometers, one or more accelerometers, any combination thereof, and / or other sensors).

[0009] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.

[0010] The foregoing and other features and aspects will become more fully apparent upon reference to the following description, claims and accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Illustrative examples of the present application are described in detail below with reference to the following drawings:

[0012] Figure 1 is a block diagram illustrating examples of encoding devices and decoding devices according to some aspects of the present disclosure;

[0013] Figure 2 is a block diagram illustrating an example architecture of a video coding hardware engine according to some aspects of the present disclosure;

[0014] Figure 3 is a block diagram illustrating an example architecture of a video coding system according to some aspects of the present disclosure;

[0015] Figure 4 is a block diagram illustrating mapping of image data in a grid to a bitstream according to aspects of the present disclosure;

[0016] Figure 5 is a block diagram illustrating an example architecture for an early notification technique for a low-latency video coding system in accordance with aspects of the present disclosure;

[0017] Figure 6 An example application of in-loop filtering to an image according to aspects of the present disclosure is illustrated;

[0018] Figure 7 is an example illustrating how a partial image for early notification interacts with image tiles of an image according to aspects of the present disclosure;

[0019] Figure 8 is a flow chart of a process for processing video data according to aspects of the present disclosure

[0020] Figure 9 is a block diagram illustrating an example video encoding device according to some aspects of the present disclosure; and

[0021] Figure 10 is a block diagram illustrating an example video decoding device according to some aspects of the present disclosure. DETAILED DESCRIPTION

[0022] Provide certain aspects of the present disclosure below. As will be apparent to those skilled in the art, some of these aspects can be applied independently, and some of them can be applied in combination. In the following description, for the purpose of explanation, specific details are set forth in order to provide a thorough understanding of various aspects of the application. However, it will be apparent that various aspects can be implemented without these specific details. The accompanying drawings and description are not intended to be restrictive.

[0023] The following description provides only example aspects and is not intended to limit the scope, applicability, or configuration of the present disclosure. Rather, the subsequent description of the example aspects will provide those skilled in the art with an enabling description for implementing the example aspects. It should be understood that various changes may be made to the function and arrangement of elements without departing from the spirit and scope of the present application as set forth in the appended claims.

[0024] Video decoding devices implement video compression techniques to efficiently encode and decode video data. Video compression techniques may include applying different prediction modes to reduce or eliminate redundancy inherent in video sequences, including spatial prediction (e.g., intra-frame prediction or intra-prediction), temporal prediction (e.g., inter-frame prediction or inter-prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques. A video encoder may divide each picture of an original video sequence into rectangular regions, referred to as video blocks or coding units (described in more detail below). These video blocks may be encoded using specific prediction modes.

[0025] A video block may be partitioned into one or more groups of smaller blocks in one or more ways. A block may include a coding tree block, a prediction block, a transform block, or other suitable blocks. Unless otherwise specified, general references to a "block" may refer to such a video block (e.g., a coding tree block, a coding block, a prediction block, a transform block, or other suitable blocks or sub-blocks, as will be understood by one of ordinary skill in the art). Furthermore, each of these blocks may also be referred to interchangeably herein as a "unit" (e.g., a coding tree unit (CTU), a coding unit, a prediction unit (PU), a transform unit (TU), etc.). In some cases, a unit may indicate a unit of decoding logic encoded in a bitstream, while a block may indicate a portion of a video frame buffer that is processed.

[0026] For inter-frame prediction mode, the video encoder can search for blocks similar to the block being encoded in a frame (or picture) located at another time position (which is called a reference frame or reference picture). The video encoder can limit the search to a certain spatial displacement from the block to be encoded. A two-dimensional (2D) motion vector including a horizontal displacement component and a vertical displacement component can be used to locate the best match. For intra-frame prediction mode, the video encoder can use spatial prediction techniques to form the predicted block based on data from previously encoded neighboring blocks in the same picture.

[0027] The video encoder can determine a prediction error. For example, the prediction can be determined as the difference between the pixel values in the block being encoded and the predicted block. The prediction error can also be referred to as a residual. The video encoder can also apply a transform to the prediction error (e.g., a discrete cosine transform (DCT) or other suitable transform) to generate transform coefficients. After the transform, the video encoder can quantize the transform coefficients. The quantized transform coefficients and motion vectors can be represented using syntax elements and, together with control information, form a decoded representation of the video sequence. In some cases, the video encoder can perform entropy decoding on the syntax elements to further reduce the number of bits required for their representation.

[0028] The video decoder can use the syntax elements and control information discussed above to construct predictive data (e.g., a predictive block) for decoding the current frame. For example, the video decoder can add the predicted block and the compressed prediction error. The video decoder can determine the compressed prediction error by weighting the transform basis function using the quantized coefficients. The difference between the reconstructed frame and the original frame is called the reconstruction error.

[0029] In some cases, after an image is decoded, a notification may be sent indicating that the decoded image is ready for display. As part of processing the notification, the image data may be sent to a downstream device. In some cases, sending all the image data may take a certain amount of time in addition to the time used to decode the image. In some cases, it may be desirable to reduce the total amount of time used to decode and send the image data.

[0030] Described herein are systems, apparatus, processes (also referred to as methods), and computer-readable media (collectively referred to herein as "systems and techniques") for providing early notification for a low-latency video decoder. In some aspects, early notification can help parallelize decoding image data and sending the image data. For example, an image can be divided into image portions, where an image portion can include a number of pixel rows. After decoding a first image portion, a notification can be sent indicating that image data for the first image portion is ready. The image data for the first image portion can then be sent to (e.g., accessed to / loaded to) a downstream device while the image data for the second image portion is being decoded. In some cases, the number of rows in a particular image portion can be determined based on a target latency. In some cases, the number of rows in an image portion can be determined additionally or alternatively based on the size of the decoded portion used to encode the image (e.g., tile size). In some cases, early notification can be adjusted for in-loop filtering.

[0031] The early notification systems and techniques described herein can reduce the latency associated with decoding video data. Such reduced latency can be useful, and in some cases necessary, for specific applications such as extended reality (XR) applications (e.g., virtual reality (VR), augmented reality (AR), mixed reality (MR), etc.), vehicular applications, robotics applications, mobile applications (e.g., media streaming for mobile devices), and the like.

[0032] Various aspects of the application are described herein with reference to the accompanying drawings.

[0033] The systems and techniques described herein may be applied to any of the existing video codecs, such as Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), Essential Video Coding (EVC), VP9, AV1 formats / codecs and / or other video coding standards, codecs, formats, etc., that are in development or to be developed.

[0034] Figure 1 1 is a block diagram illustrating an example of a system 100 including an encoding device 104 and a decoding device 112. The encoding device 104 can be part of a source device, and the decoding device 112 can be part of a receiving device. The source device and / or the receiving device can include an electronic device, such as a mobile or fixed-line telephone handset (e.g., a smartphone, a cellular phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the source device and the receiving device can include one or more wireless transceivers for wireless communication. The decoding techniques described herein can be applied to video decoding in various multimedia applications, including streaming video transmission (e.g., via the Internet), television broadcasting or transmission, encoding of digital video stored on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. As used herein, the term decoding can refer to encoding and / or decoding. In some examples, system 100 may support one-way or two-way video transmission to support applications such as video conferencing, video streaming, video playback, video broadcasting, gaming, and / or video telephony.

[0035] The encoding device 104 (or encoder) can be used to encode video data using a video coding standard, format, codec, or protocol to generate an encoded video bitstream. Examples of video coding standards and formats / codecs include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) (including its Scalable Video Coding (SVC) and Multi-View Video Coding (MVC) extensions), High Efficiency Video Coding (HEVC) or ITU-T H.265, and Versatile Video Coding (VVC) or ITU-T H.266. Various extensions to HEVC handle multi-layer video coding, including range and screen content coding extensions, 3D video coding (3D-HEVC), multi-view extensions (MV-HEVC), and scalable extensions (SHVC). HEVC and its extensions have been developed by the Joint Collaborative Team on Video Coding (JCT-VC), the Joint Collaborative Team on 3D Video Coding Extensions (JCT-3V) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG). VP9, AOMedia Video 1 (AV1) developed by the Alliance for Open Media (AOMedia), and Essential Video Coding (EVC) are other video coding standards for which the techniques described herein may be applied.

[0036] The techniques described herein can be applied to any existing video codec (e.g., High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), or other suitable existing video codecs), and / or can be efficient coding tools for any video coding standard being developed and / or future video coding standards (such as, for example, VVC and / or other video coding standards under development or to be developed). For example, the examples described herein can be performed using video codecs such as VVC, HEVC, AVC, and / or their extensions. However, the techniques and systems described herein can also be applicable to other coding standards, codecs, or formats, such as MPEG, JPEG (or other coding standards for still images), EVC, VP9, AV1, their extensions, or other suitable coding standards that are already available or not yet available or not yet developed. For example, in some examples, the encoding device 104 and / or the decoding device 112 can operate according to a proprietary video codec / format (such as AV1, an extension of AVI, and / or a subsequent version of AV1 (e.g., AV2), or other proprietary formats or industry standards). Thus, although the techniques and systems described herein may be described with reference to a particular video coding standard, one of ordinary skill in the art will understand that the description should not be construed as applying only to that particular standard.

[0037] Reference Figure 1 , video source 102 can provide video data to encoding device 104. Video source 102 can be part of a source device or can be part of a device other than a source device. Video source 102 can include a video capture device (e.g., a video camera, a camera phone, a video phone, etc.), a video archive containing stored video, a video server or content provider that provides video data, a video feed interface that receives video from a video server or content provider, a computer graphics system for generating computer graphics video data, a combination of such sources, or any other suitable video source.

[0038] Video data from video source 102 may include one or more input pictures or frames. A picture or frame is a still image that, in some cases, is part of a video. In some examples, the data from video source 102 may be a still image that is not part of a video. In HEVC, VVC, and other video coding specifications, a video sequence may include a series of pictures. A picture may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luma samples, SCb is a two-dimensional array of Cb chroma samples, and SCr is a two-dimensional array of Cr chroma samples. Chroma samples may also be referred to herein as "chroma" samples. A pixel may refer to all three components (luma samples and chroma samples) for a given position in the array of a picture. In other cases, a picture may be monochrome and may include only an array of luma samples, in which case the terms pixel and sample may be used interchangeably. While the example techniques described herein relate to individual samples for illustrative purposes, the same techniques may be applied to pixels (e.g., all three sample components for a given position in the array of a picture). With respect to the example techniques described herein involving pixels (eg, all three sample components for a given position in an array of a picture) for illustration purposes, the same techniques can be applied to individual samples.

[0039] The encoder engine 106 (or encoder) of the encoding device 104 encodes video data to generate an encoded video bitstream. In some examples, the encoded video bitstream (or "video bitstream" or "bitstream") is a series of one or more decoded video sequences. A decoded video sequence (CVS) includes a series of access units (AUs) starting with an AU having a random access point picture with specific properties in the base layer, up to but not including the next AU having a random access point picture with specific properties in the base layer. For example, the specific properties of the random access point picture that starts a CVS may include a RASL flag equal to 1 (e.g., NoRaslOutputFlag). Otherwise, the random access point picture (with a RASL flag equal to 0) does not start a CVS. An access unit (AU) includes one or more decoded pictures and control information corresponding to decoded pictures that share the same output time. Decoded slices of a picture are encapsulated at the bitstream level into data units called network abstraction layer (NAL) units. For example, an HEVC video bitstream may include one or more CVSs containing NAL units. Each NAL unit in the NAL unit has a NAL unit header. In one example, the header is one byte for H.264 / AVC (except for multi-layer extensions) and two bytes for HEVC. Syntax elements in the NAL unit header use designated bits and are therefore visible to all types of systems and transport layers (such as transport streams, real-time transport (RTP) protocols, file formats, etc.).

[0040] There are two types of NAL units in the HEVC standard, including video coding layer (VCL) NAL units and non-VCL NAL units. VCL NAL units include decoded picture data that form a decoded video bitstream. For example, the bit sequence that forms the decoded video bitstream is present in a VCL NAL unit. A VCL NAL unit may include a slice or slice segment (described below) of decoded picture data, and a non-VCL NAL unit includes control information related to one or more decoded pictures. In some cases, a NAL unit may be referred to as a packet. An HEVC AU includes a VCL NAL unit (which contains decoded picture data) and non-VCL NAL units corresponding to the decoded picture data (if any). In addition to other information, non-VCL NAL units may also contain parameter sets that have high-level information related to the encoded video bitstream. For example, a parameter set may include a video parameter set (VPS), a sequence parameter set (SPS), and a picture parameter set (PPS). In some cases, each slice or other portion of the bitstream may reference a single active PPS, SPS, and / or VPS to allow the decoding device 112 to access information that can be used to decode the slice or other portion of the bitstream.

[0041] A NAL unit may contain a sequence of bits that form a decoded representation of video data (e.g., a coded video bitstream, a CVS of a bitstream, etc.), such as a decoded representation of a picture in a video. The encoder engine 106 generates a decoded representation of a picture by dividing each picture into a plurality of slices. A slice is independent of other slices, so that information in the slice can be decoded without relying on data from other slices within the same picture. A slice includes one or more slice segments (including independent slice segments and one or more dependent slice segments that depend on previous slice segments, if any).

[0042] In HEVC, the slice is then divided into coding tree blocks (CTBs) of luma samples and chroma samples. The CTB of luma samples and one or more CTBs of chroma samples, along with the syntax of the samples, are called coding tree units (CTUs). A CTU can also be called a "tree block" or "largest coding unit" (LCU). The CTU is the basic processing unit for HEVC encoding. The CTU can be split into multiple coding units (CUs) of different sizes. A CU contains an array of luma and chroma samples, which is called a coding block (CB).

[0043] Luma CBs and chroma CBs can be further split into prediction blocks (PBs). A PB is a block of samples of a luma or chroma component that uses the same motion parameters for inter-frame prediction or intra-block copy (IBC) prediction (when available or enabled). The luma PB and one or more chroma PBs, along with associated syntax, form a prediction unit (PU). For inter-frame prediction, a set of motion parameters (e.g., one or more motion vectors, reference indices, etc.) is signaled in the bitstream for each PU and is used for inter-frame prediction of the luma PB and one or more chroma PBs. Motion parameters may also be referred to as motion information. CBs can also be divided into one or more transform blocks (TBs). A TB represents a square block of samples of a color component, to which a residual transform (e.g., in some cases, the same two-dimensional transform) is applied for decoding the prediction residual signal. A transform unit (TU) represents a TB of luma and chroma samples, as well as the corresponding syntax elements. Transform coding is described in more detail below.

[0044] The size of the CU corresponds to the size of the decoding mode and the shape can be square. For example, the size of the CU can be 8×8 samples, 16×16 samples, 32×32 samples, 64×64 samples, or any other appropriate size (up to the size of the corresponding CTU). The phrase "N×N" is used herein to refer to the pixel dimensions of a video block in terms of vertical and horizontal dimensions (e.g., 8 pixels×8 pixels). The pixels in a block can be arranged in rows and columns. In some implementations, a block may not have the same number of pixels in the horizontal direction as in the vertical direction. The syntax data associated with a CU can describe, for example, the division of the CU into one or more PUs. The division mode can differ between whether the CU is encoded in an intra-frame prediction mode or an inter-frame prediction mode. The PU can be divided into a non-square shape. The syntax data associated with a CU can also describe, for example, the division of the CU into one or more TUs according to a CTU. The shape of a TU can be square or non-square.

[0045] According to the HEVC standard, transforms can be performed using transform units (TUs). TUs may vary for different CUs. The size of a TU may be set based on the size of the PU within a given CU. The size of a TU may be the same as or smaller than the size of the PU. In some examples, a quadtree structure called a residual quadtree (RQT) may be used to subdivide the residual samples corresponding to a CU into smaller units. The leaf nodes of the RQT may correspond to TUs. The pixel difference values associated with the TU may be transformed to produce transform coefficients. The transform coefficients may then be quantized by the encoder engine 106.

[0046] Once a picture of video data is partitioned into CUs, the encoder engine 106 predicts each PU using a prediction mode. The prediction unit, or prediction block, is then subtracted from the original video data to obtain a residual (described below). For each CU, the prediction mode can be signaled within the bitstream using syntax data. Prediction modes can include intra-frame prediction (or intra-picture prediction) or inter-frame prediction (or inter-picture prediction). Intra-frame prediction exploits correlation between spatially adjacent samples within a picture. For example, using intra-frame prediction, each PU is predicted from adjacent image data in the same picture using, for example, DC prediction to find the average value for the PU, planar prediction to fit a planar surface to the PU, directional prediction to extrapolate from adjacent data, or any other suitable type of prediction. Inter-frame prediction exploits temporal correlation between pictures to derive motion-compensated predictions for blocks of image samples. For example, using inter-frame prediction, each PU is predicted using motion-compensated prediction from image data in one or more reference pictures (either preceding or following the current picture in output order). For example, the decision whether to use inter-picture prediction or intra-picture prediction for coding a picture region can be made at the CU level.

[0047] The encoder engine 106 and the decoder engine 116 (described in more detail below) can be configured to operate according to VVC. According to VVC, a video decoder (such as the encoder engine 106 and / or the decoder engine 116) divides a picture into multiple coding tree units (CTUs) (wherein a CTB for luma samples and one or more CTBs for chroma samples, together with syntax for the samples, is referred to as a CTU). The video decoder can divide the CTU according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure removes the concept of multiple partition types, such as the separation between CU, PU, and TU of HEVC. The QTBT structure includes two levels, including a first level divided according to quadtree partitioning and a second level divided according to binary tree partitioning. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary tree correspond to coding units (CUs).

[0048] In the MTT partitioning structure, blocks can be partitioned using quadtree partitioning, binary tree partitioning, and one or more types of ternary tree partitioning. Ternary tree partitioning is a partitioning in which a block is split into three sub-blocks. In some examples, ternary tree partitioning divides a block into three sub-blocks without partitioning the original block through the center. Partition types in MTT (e.g., quadtree, binary tree, and ternary tree) can be symmetric or asymmetric.

[0049] When operating according to the AV1 codec, the encoder engine 106 and the decoder engine 116 can be configured to decode video data in blocks. In AV1, the largest coded block that can be processed is called a super block. In AV1, a super block can be 128x128 luma samples or 64x64 luma samples. However, in subsequent video coding formats (e.g., AV2), super blocks can be defined by different (e.g., larger) luma sample sizes. In some examples, a super block is the top level of a block quadtree. The encoder engine 106 can further divide the super block into smaller coded blocks. The encoder engine 106 can use square or non-square partitioning to divide the super block and other coded blocks into smaller blocks. Non-square blocks can include N / 2xN, NxN / 2, N / 4xN, and NxN / 4 blocks. The encoder engine 106 and the decoder engine 116 can perform separate prediction and transform processing on each coded block in the coded block.

[0050] AV1 also defines tiles of video data. A tile is a rectangular array of superblocks that can be decoded independently of other tiles. That is, the encoder engine 106 and decoder engine 116 can encode and decode coded blocks within a tile, respectively, without using video data from other tiles. However, the encoder engine 106 and decoder engine 116 can perform filtering across tile boundaries. Tiles can be uniform or non-uniform in size. Tile-based decoding enables parallel processing and / or multi-threading for encoder and decoder implementations.

[0051] In some examples, the video coder may use a single QTBT or MTT structure to represent each of the luma and chroma components, while in other examples, the video coder may use two or more QTBT or MTT structures, such as one QTBT or MTT structure for the luma component and another QTBT or MTT structure for the two chroma components (or two QTBT and / or MTT structures for corresponding chroma components).

[0052] The video coder may be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, superblock partitioning, or other partitioning structures.

[0053] In some examples, a slice type is assigned to one or more slices of a picture. Slice types include intra-coded slices (I slices), inter-coded P slices, and inter-coded B slices. An I slice (Intra-coded Frame, Independently Decodable) is a slice of a picture that is coded only using intra prediction and is therefore independently decodable because an I slice only requires data within the frame to predict any prediction unit or prediction block of the slice. A P slice (Uni-directionally Predicted Frame) is a slice of a picture that can be coded using both intra prediction and uni-directional inter prediction. Each prediction unit or prediction block within a P slice is coded using either intra prediction or inter prediction. When inter prediction is applied, the prediction unit or prediction block is predicted using only one reference picture, and therefore the reference samples come from only one reference region of a frame. A B slice (Bi-directionally Predicted Frame) is a slice of a picture that can be coded using both intra prediction and inter prediction (e.g., bi-directional prediction or uni-directional prediction). A prediction unit or prediction block of a B slice can be predicted bidirectionally from two reference pictures, where each picture contributes one reference region and the sample sets of the two reference regions are weighted (e.g., with equal weights or with different weights) to produce a prediction signal for the bidirectionally predicted block. As explained above, slices of a picture are coded independently. In some cases, a picture may be coded as only one slice.

[0054] As described above, intra-picture prediction of a picture exploits the correlation between spatially adjacent samples within the picture. There are multiple intra-frame prediction modes (also referred to as "intra-frame modes"). In some examples, intra-frame prediction of a luminance block includes 35 modes, including a planar mode, a DC mode, and 33 angular modes (e.g., a diagonal intra-frame prediction mode and an angular mode adjacent to the diagonal intra-frame prediction mode). The 35 modes of intra-frame prediction are indexed as shown in Table 1 below. In other examples, more intra-modes may be defined, including prediction angles that may not yet be represented by the 33 angular modes. In other examples, the prediction angles associated with the angular modes may be different from those used in HEVC.

[0055] Intra prediction mode Association Name 0 INTRA_PLANAR 1 INTRA_DC 2..34 INTRA_ANGULAR2..INTRA_ANGULAR34

[0056] Table 1 – Specification of intra prediction modes and associated names

[0057] Inter-picture prediction uses temporal correlations between pictures to derive a motion-compensated prediction for the current block of image samples. Using a translational motion model, the position of a block in a previously decoded picture (reference picture) is indicated by a motion vector (Δx, Δy), where Δx specifies the horizontal displacement of the reference block relative to the position of the current block, and Δy specifies the vertical displacement of the reference block relative to the position of the current block. In some cases, the motion vector (Δx, Δy) can have integer sample accuracy (also referred to as integer accuracy), in which case the motion vector refers to the integer pixel grid (or integer pixel sampling grid) of the reference frame. In some cases, the motion vector (Δx, Δy) can have fractional sample accuracy (also referred to as fractional pixel accuracy or non-integer accuracy) to more accurately capture the movement of the underlying object without being restricted by the integer pixel grid of the reference frame. The accuracy of the motion vector can be represented by the quantization level of the motion vector. For example, the quantization level can be integer accuracy (e.g., 1-pixel) or fractional pixel accuracy (e.g., 1 / 4-pixel, 1 / 2-pixel, or other sub-pixel values). When the corresponding motion vector has fractional sample accuracy, interpolation is applied to the reference picture to derive the prediction signal. For example, the samples available at integer positions can be filtered (e.g., using one or more interpolation filters) to estimate the values at the fractional positions. The previously decoded reference picture is indicated by a reference index (refIdx) of the reference picture list. The motion vector and the reference index can be referred to as motion parameters. Two categories of inter-picture prediction can be performed, including unidirectional prediction and bidirectional prediction.

[0058] In the case of inter-frame prediction using bidirectional prediction (also known as bidirectional inter prediction), two sets of motion parameters (Δx0, y0, refIdx0 and Δx1, y1, refIdx1) are used to generate two motion-compensated predictions (based on the same reference picture or possibly different reference pictures). For example, in the case of bidirectional prediction, two motion-compensated prediction signals are used for each prediction block, and B prediction units are generated. The two motion-compensated predictions are then combined to obtain the final motion-compensated prediction. For example, the two motion-compensated predictions can be combined by averaging. In another example, weighted prediction can be used, in which case different weights can be applied to each motion-compensated prediction. The reference pictures that can be used in bidirectional prediction are stored in two separate lists (denoted as list 0 and list 1). The motion parameters can be derived at the encoder using a motion estimation process.

[0059] In the case of inter prediction using unidirectional prediction (also known as unidirectional inter prediction), a set of motion parameters (Δx0, y0, refIdx0) is used to generate a motion-compensated prediction based on the reference picture. For example, in the case of unidirectional prediction, each prediction block uses at most one motion-compensated prediction signal and generates P prediction units.

[0060] The PU may include data related to the prediction process (e.g., motion parameters, or other suitable data). For example, when the PU is encoded using intra prediction, the PU may include data describing the intra prediction mode for the PU. As another example, when the PU is encoded using inter prediction, the PU may include data defining a motion vector for the PU. The data defining the motion vector for the PU may describe, for example, the horizontal component of the motion vector (Δx), the vertical component of the motion vector (Δy), the resolution for the motion vector (e.g., integer precision, quarter-pixel precision, or eighth-pixel precision), the reference picture to which the motion vector points, a reference index, a reference picture list for the motion vector (e.g., List 0, List 1, or List C), or any combination thereof.

[0061] AV1 includes two general techniques for encoding and decoding coded blocks of video data. These two general techniques are intra prediction (e.g., intra frame prediction or spatial prediction) and inter prediction (e.g., inter frame prediction or temporal prediction). In the context of AV1, when an intra prediction mode is used to predict a block of a current frame of video data, the encoder engine 106 and the decoder engine 116 do not use video data from other frames of the video data. For most intra prediction modes, the encoding device 104 encodes the block of the current frame based on the difference between the sample values in the current block and the predicted values generated from reference samples in the same frame. The encoding device 104 determines the predicted values generated from the reference samples based on the intra prediction mode.

[0062] After performing prediction using intra prediction and / or inter prediction, the encoding device 104 may perform transformation and quantization. For example, after prediction, the encoder engine 106 may calculate a residual value corresponding to the PU. The residual value may include pixel difference values between the current pixel block (PU) being decoded and the prediction block used to predict the current block (e.g., a predicted version of the current block). For example, after generating a prediction block (e.g., using inter prediction or intra prediction), the encoder engine 106 may generate a residual block by subtracting the prediction block generated by the prediction unit from the current block. The residual block includes a set of pixel difference values that quantizes the difference between the pixel values of the current block and the pixel values of the prediction block. In some examples, the residual block may be represented in a two-dimensional block format (e.g., a two-dimensional matrix or an array of pixel values). In such examples, the residual block is a two-dimensional representation of the pixel values.

[0063] Any residual data that may remain after prediction is performed is transformed using a block transform, which may be based on a discrete cosine transform, a discrete sine transform, an integer transform, a wavelet transform, other suitable transform functions, or any combination thereof. In some cases, one or more block transforms (e.g., sizes 32x32, 16x16, 8x8, 4x4, or other suitable sizes) may be applied to the residual data in each CU. In some aspects, a TU may be used for the transform and quantization process implemented by the encoder engine 106. A given CU with one or more PUs may also include one or more TUs. As described in further detail below, the residual values may be transformed into transform coefficients using a block transform, and may then be quantized and scanned using the TUs to produce serialized transform coefficients for entropy coding.

[0064] In some aspects, after performing intra-frame prediction decoding or inter-frame prediction decoding using the PU of the CU, the encoder engine 106 may calculate residual data for the TU of the CU. The PU may include pixel data in the spatial domain (or pixel domain). The TU may include coefficients in the transform domain after applying the block transform. As previously described, the residual data may correspond to the pixel difference between the pixel of the unencoded picture and the predicted value corresponding to the PU. The encoder engine 106 may form a TU including the residual data for the CU, and may then transform the TU to generate transform coefficients for the CU.

[0065] The encoder engine 106 can perform quantization of the transform coefficients. Quantization provides further compression by quantizing the transform coefficients to reduce the amount of data used to represent the coefficients. For example, quantization can reduce the bit depth associated with some or all coefficients. In one example, a coefficient with an n-bit value can be rounded down to an m-bit value during quantization, where n is greater than m.

[0066] Once quantization is performed, the decoded video bitstream includes quantized transform coefficients, prediction information (e.g., prediction mode, motion vectors, block vectors, etc.), partition information, and any other suitable data (such as other syntax data). The different elements of the decoded video bitstream can then be entropy encoded by the encoder engine 106. In some examples, the encoder engine 106 can use a predefined scan order to scan the quantized transform coefficients to produce a serialized vector that can be entropy encoded. In some examples, the encoder engine 106 can perform adaptive scanning. After scanning the quantized transform coefficients to form a vector (e.g., a one-dimensional vector), the encoder engine 106 can entropy encode the vector. For example, the encoder engine 106 can use context-adaptive variable length coding, context-adaptive binary arithmetic coding, syntax-based context-adaptive binary arithmetic coding, probability interval partitioning entropy coding, or another suitable entropy coding technique.

[0067] The output 110 of the encoding device 104 can send the NAL units constituting the encoded video bitstream data to the decoding device 112 of the receiving device via a communication link 120. The input 114 of the decoding device 112 can receive the NAL units. The communication link 120 can include a channel provided by a wireless network, a wired network, or a combination of a wired network and a wireless network. The wireless network can include any wireless interface or a combination of wireless interfaces, and can include any suitable wireless network (e.g., the Internet or other wide area network, a packet-based network, WiFi™, radio frequency (RF), ultra-wideband (UWB), WiFi direct connection, cellular, long-term evolution (LTE), WiMax™, etc.). The wired network can include any wired interface (e.g., optical fiber, Ethernet, power line Ethernet, Ethernet on coaxial cable, digital signal line (DSL), etc.). Various devices such as base stations, routers, access points, bridges, gateways, switches, etc. can be used to implement wired and / or wireless networks. The encoded video bitstream data can be modulated according to a communication standard (e.g., a wireless communication protocol) and sent to the receiving device.

[0068] In some examples, encoding device 104 may store the encoded video bitstream data in storage 108. Output 110 may retrieve the encoded video bitstream data from encoder engine 106 or from storage 108. Storage 108 may include any of a variety of distributed or locally accessible data storage media. For example, storage 108 may include a hard drive, a storage disk, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. Storage 108 may also include a decoded picture buffer (DPB) for storing reference pictures used in inter-frame prediction. In another example, storage 108 may correspond to a file server or another intermediate storage device that can store encoded video generated by a source device. In such cases, a receiving device including decoding device 112 can access the stored video data from the storage device via streaming or downloading. The file server can be any type of server capable of storing encoded video data and transmitting the encoded video data to a receiving device. Example file servers include a network server (e.g., for a website), an FTP server, a network attached storage (NAS) device, or a local disk drive. The receiving device can access the encoded video data through any standard data connection, including an Internet connection. This can include a wireless channel suitable for accessing encoded video data stored on a file server (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of the two. The transmission of the encoded video data from storage 108 can be a streaming transmission, a download transmission, or a combination thereof.

[0069] The input 114 of the decoding device 112 receives the encoded video bitstream data and can provide the video bitstream data to the decoder engine 116 or to the storage 118 for later use by the decoder engine 116. For example, the storage 118 may include a DPB for storing reference pictures used in inter-frame prediction. A receiving device including the decoding device 112 can receive the encoded video data to be decoded via the storage 108. The encoded video data can be modulated according to a communication standard (such as a wireless communication protocol) and sent to the receiving device. The communication medium for sending the encoded video data can include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include a router, a switch, a base station, or any other device that can help facilitate communication from a source device to a receiving device.

[0070] The decoder engine 116 can decode the coded video bitstream data by entropy decoding (e.g., using an entropy decoder) and extracting the elements of one or more decoded video sequences that constitute the coded video data. The decoder engine 116 can then rescale the coded video bitstream data and perform an inverse transform on the coded video bitstream data. The residual data is then passed to the prediction stage of the decoder engine 116. The decoder engine 116 then predicts the current pixel block (e.g., PU). In some examples, the prediction is added to the output of the inverse transform (residual data).

[0071] The decoding device 112 may output the decoded video to a video destination device 122, which may include a display or other output device for displaying the decoded video data to a consumer of the content. In some aspects, the video destination device 122 may be part of a receiving device that includes the decoding device 112. In some aspects, the video destination device 122 may be part of a separate device other than the receiving device.

[0072] In some aspects, the encoding device 104 and / or the decoding device 112 can be integrated with the audio encoding device and the audio decoding device, respectively. The encoding device 104 and / or the decoding device 112 can also include other hardware or software necessary to implement the decoding techniques described above, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. The encoding device 104 and the decoding device 112 can be integrated as part of a combined encoder / decoder (codec) in the respective devices.

[0073] Figure 1 The example system shown is an illustrative example that can be used herein. The technology for processing video data using the technology described herein can be performed by any digital video encoding and / or decoding device. Although generally speaking, the technology of the present disclosure is performed by a video encoding device or a video decoding device, these technologies can also be performed by a combined video encoder-decoder commonly referred to as a "CODEC". In addition, the technology of the present disclosure can also be performed by a video preprocessor. The source device and the receiving device are merely examples of such decoding devices, wherein the source device generates decoded video data for transmission to the receiving device. In some examples, the source device and the receiving device can operate in a substantially symmetrical manner so that each of these devices includes a video encoding and decoding component. Therefore, the example system can support one-way or two-way video transmission between video devices, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0074] Extensions to the HEVC standard include a multi-view video coding extension called MV-HEVC and a scalable video coding extension called SHVC. MV-HEVC and SHVC extensions share the concept of layered coding, in which different layers are included in the coded video bitstream. Each layer in the coded video sequence is addressed by a unique layer identifier (ID). The layer ID may be present in the header of a NAL unit to identify the layer to which the NAL unit is associated. In MV-HEVC, different layers typically represent different views of the same scene in the video bitstream. In SHVC, different scalable layers are provided that represent the video bitstream at different spatial resolutions (or picture resolutions) or at different reconstruction fidelity. Scalable layers may include a base layer (with layer ID = 0) and one or more enhancement layers (with layer ID = 1, 2, and n). The base layer may conform to the profile of the first version of HEVC and represent the lowest available layer in the bitstream. Compared to the base layer, the enhancement layer has increased spatial resolution, temporal resolution or frame rate, and / or reconstruction fidelity (or quality). Enhancement layers are organized hierarchically and may or may not depend on lower layers. In some examples, a single standard codec may be used to code the different layers (e.g., all layers may be coded using HEVC, SHVC, or other coding standards). In some examples, multiple standard codecs may be used to code the different layers. For example, the base layer may be coded using AVC, while one or more enhancement layers may be coded using SHVC and / or MV-HEVC extensions to the HEVC standard.

[0075] In general, a layer comprises a set of VCL NAL units and a corresponding set of non-VCL NAL units. NAL units are assigned specific layer ID values. Layers can be hierarchical in the sense that layers can depend on lower layers. A layer set refers to a set of self-contained layers represented within a bitstream, where self-contained means that a layer within a layer set can depend on other layers in the layer set during decoding, but does not depend on any other layer for decoding. Thus, the layers in a layer set can form an independent bitstream that can represent video content. The set of layers in a layer set can be obtained from another bitstream through the operation of a sub-bitstream extraction process. A layer set can correspond to a set of layers to be decoded when a decoder wants to operate according to specific parameters.

[0076] As previously described, an HEVC bitstream includes a set of NAL units, including VCL NAL units and non-VCL NAL units. VCL NAL units contain decoded picture data that form the decoded video bitstream. For example, the bit sequence that forms the decoded video bitstream is present in a VCL NAL unit. In addition to other information, non-VCL NAL units may also contain parameter sets, which contain high-level information related to the encoded video bitstream. For example, parameter sets may include a video parameter set (VPS), a sequence parameter set (SPS), and a picture parameter set (PPS). Examples of parameter set objectives include bitrate efficiency, error resilience, and providing a system layer interface. Each slice references a single active PPS, SPS, and VPS to access information that the decoding device 112 can use to decode the slice. An identifier (ID) may be decoded for each parameter set, including a VPS ID, an SPS ID, and a PPS ID. The SPS includes an SPS ID and a VPS ID. The PPS includes a PPS ID and an SPS ID. Each slice header includes a PPS ID. Using the ID, the active parameter set can be identified for a given slice.

[0077] The PPS includes information applicable to all slices in a given picture. Therefore, all slices in a picture reference the same PPS. Slices in different pictures can also reference the same PPS. The SPS includes information applicable to all pictures in the same coded video sequence (CVS) or bitstream. As previously described, a coded video sequence is a series of access units (AUs) starting with a random access point picture (RAP) in the base layer with specific properties (as described above) (e.g., an instantaneous decoding reference (IDR) picture or a broken link access (BLA) picture, or other appropriate RAP picture), up to but not including the next AU (or the end of the bitstream) with a RAP picture in the base layer with specific properties. The information in the SPS can remain constant from picture to picture within the coded video sequence. Pictures in a coded video sequence can use the same SPS. The VPS includes information applicable to all layers within the coded video sequence or bitstream. The VPS includes a syntax structure with syntax elements that apply to the entire coded video sequence. In some aspects, the VPS, SPS, or PPS can be sent in-band with the coded bitstream. In some aspects, the VPS, SPS, or PPS may be sent out-of-band in a separate transmission compared to the NAL units containing the coded video data.

[0078] The present disclosure may generally refer to "signaling" certain information (such as syntax elements). The term "signaling" may generally refer to the transmission of values for syntax elements and / or other data used to decode encoded video data. For example, encoding device 104 may signal values for syntax elements in a bitstream. Generally, signaling refers to generating values in a bitstream. As described above, video device 102 may transmit the bitstream to video destination device 122 in substantially real time or in non-real time (such as may occur when storing syntax elements to storage 108 for later retrieval by video destination device 122).

[0079] The video bitstream may also include supplemental enhancement information (SEI) messages. For example, SEI NAL units may be part of the video bitstream. In some cases, SEI messages may contain information that is not required for the decoding process. For example, the information in the SEI message may not be necessary for the decoder to decode the video pictures of the bitstream, but the decoder may use the information to improve the display or processing of the pictures (e.g., the decoded output). The information in the SEI message may be embedded metadata. In one illustrative example, the information in the SEI message may be used by a decoder-side entity to improve the visibility of the content. In some cases, specific application standards may mandate the presence of such SEI messages in the bitstream so that quality improvements can be brought to all devices that comply with the application standards (e.g., the transport of frame packing SEI messages for frame-compatible planar stereoscopic 3DTV video formats (wherein SEI messages are carried for each frame of the video), the handling of recovery point SEI messages, and the use of pan scan rectangle SEI messages in DVB, among many other examples).

[0080] Figure 2 is a block diagram illustrating an example architecture 200 of a video decoding hardware engine. In some cases, the architecture 200 may be composed of Figure 1 In some examples, the architecture 200 can be implemented by the encoder engine 106 of the encoding device 104 or by the decoder engine 116 of the decoding device 112, as shown. Figure 1 shown.

[0081] In this example, architecture 200 of a video decoding hardware engine may include a control processor 210, an interface 222, a video stream processor (VSP) 212, processing pipelines 214-220 (also referred to as "pipelines"), a direct memory access (DMA) subsystem 230, and one or more buffers 232. In some examples, architecture 200 may include memory 240 for storing data such as frames, video, decoding information, output, etc. In other examples, memory 240 may be external memory on a decoding device implementing the video decoding hardware engine.

[0082] The interface 222 can transfer data between components of the video decoding hardware engine and / or the video decoding device through a communication system or system bus on the video decoding hardware engine and / or the decoding device implementing the video decoding hardware engine. For example, the interface 222 can connect the control processor 210, the VSP 212, the processing pipelines 214-220 (e.g., the video pixel processor (VPP)), the DMA subsystem 230, and / or one or more buffers 232 to the system bus on the video decoding hardware engine and / or the decoding device. In some examples, the interface 222 can include a network-based communication subsystem, such as a network on a chip (NoC).

[0083] DMA subsystem 230 can allow other components of the video decoding hardware engine (e.g., other components in architecture 200) to access memory on the video decoding hardware engine and / or a video decoding device implementing the video decoding hardware engine. For example, DMA subsystem 230 can provide access to memory 240 and / or one or more buffers 232. In some examples, DMA subsystem 230 can manage access to common memory locations and associated data traffic (e.g., tiles 202, blocks 204A-D, bitstream 236, entropy-decoded data 238, etc.).

[0084] The memory 240 may include one or more internal or external memory devices, such as, but not limited to, one or more random access memory (RAM) components, read-only memory (ROM) components, cache memory components, buffer components, and / or other memory devices. The memory 240 may store data used by the video decoding hardware engine and / or the video decoding device, such as frames, processing parameters, input data, output data, and / or any other type of data.

[0085] The control processor 210 may include one or more processors. The control processor 210 may control and / or program components of the video decoding hardware engine (e.g., other components in the architecture 200). In some examples, the control processor 210 may be coupled to Figure 2 For example, in some cases, the control processor 210 may interface with an application processor on a video decoding device.

[0086] The VSP 212 may perform bitstream analysis (e.g., separating the network abstraction layer, picture layer, and slice layer) and entropy decoding operations. In some examples, the VSP 212 may perform decoding functions such as variable length encoding or decoding. For example, the VSP 212 may implement a lossless compression / decompression algorithm to compress or decompress the bitstream 236. In some examples, the VSP 212 may perform arithmetic decoding, such as context-sensitive, adaptive binary arithmetic coding (CABAC), and / or any other decoding algorithm.

[0087] The processing pipelines 214-220 can perform video pixel operations such as motion estimation, motion compensation, transforms and quantization, image deblocking, and / or any other video pixel operations. In some cases, the processing pipelines 214-220 can perform video pixel operations based on the output of the VSP 212. In some cases, the output of a VSP 212 can be processed by multiple processing pipelines 214-220. The processing pipelines 214-220 (and / or each processing pipeline) can perform specific video pixel operations in parallel. For example, each processing pipeline can perform multiple operations (and / or process data) simultaneously and / or significantly in parallel. As another example, multiple processing pipelines can perform operations (and / or process data) simultaneously and / or significantly in parallel.

[0088] exist Figure 2 In some embodiments, the processing pipelines 214-220 can store and retrieve video pixel processing data (e.g., video pixel processing output, input, parameters, pixel data, processing synchronization data, etc.) to and from one or more buffers 232. In some cases, the one or more buffers 232 can include a single buffer. In other cases, the one or more buffers 232 can include multiple buffers. In some examples, the one or more buffers 232 can include global input / output line buffers and pipeline synchronization buffers. In some cases, the pipeline synchronization buffers can temporarily store data used to synchronize data and / or results from video pixel processing operations performed by the processing pipelines 214-220.

[0089] In some examples, the VSP 212 can decompress a bitstream 236 associated with a video or frame sequence and store entropy-decoded data 238 associated with the bitstream 236 for processing by the processing pipelines 214-220. In some cases, the entropy-decoded data 238 can be stored in a memory or buffer, and the memory or buffer can be part of or separate from the buffer 232. In some cases, the VSP 212 can retrieve the bitstream 236 and store the entropy-decoded data 238 to and from the memory using the DMA subsystem 230, which can manage access to the memory components and / or units as previously mentioned. In some cases, the VSP 212 can store the decoded data in a bitstream-based order. For example, where the bitstream organizes image information based on tiles, the decoded data can be grouped so that the decoded data for the tiles are stored together in the order in which the tiles were decoded (e.g., in tile order (also known as bitstream order)). The processing pipelines 214 - 220 may retrieve the entropy decoded data 238 (eg, via the DMA subsystem 230 ) and perform video pixel processing operations on the blocks 204A-D of the tile 202 associated with the bitstream 236 .

[0090] The processing pipelines 214-220 can perform video pixel processing operations in parallel, as previously described. The processing pipelines 214-220 can retrieve and store video pixel processing inputs and outputs from / in one or more buffers 232 (e.g., via the DMA subsystem 230). For example, a motion estimation algorithm implemented by the processing pipeline 214 can perform motion estimation on block 204A and store the motion estimation information calculated for block 204A in the one or more buffers 232. A motion compensation algorithm implemented by the processing pipeline 214 can retrieve the motion estimation information from the one or more buffers 232 and use the motion estimation information to perform motion compensation for block 204A. While the motion compensation algorithm is performing motion compensation, the motion estimation algorithm can perform motion estimation for the next block.

[0091] The motion compensation algorithm may store the motion compensation results in one or more buffers 232, which may be accessed and used by the transform, quantization, and deblocking algorithms to perform transform, quantization, and deblocking on block 204A. The motion compensation algorithm may perform motion compensation on the next block while the transform, quantization, and / or deblocking algorithms perform transform, quantization, and / or deblocking on block 204A. The transform, quantization, and deblocking algorithms may similarly perform the corresponding operations on block 204A and the next block in parallel. In some examples, the motion estimation, motion compensation, transform, quantization, and deblocking algorithms may perform the corresponding operations on different blocks in parallel.

[0092] The processing pipelines 214-220 can be implemented by hardware components and / or software components. For example, the processing pipelines 214-220 can be implemented by one or more pixel processors. In some examples, each processing pipeline can be implemented by one or more hardware components. In some cases, each processing pipeline can use different hardware units and / or components to implement different stages in the pipeline of the processing pipeline. After performing video pixel operations to generate output pixels for display, the output pixels for display can be output to a memory, such as memory 240 or one or more buffers 232 (such as display buffers). In some cases, the memory 204 can be system memory or a similar memory device, such as a double data rate (DDR) synchronous dynamic random access memory (SDRAM) or any other memory device. The memory 240 can store the output pixels to be displayed on a display device.

[0093] Figure 2 The number of processing pipelines shown in FIG is merely an example provided for explanation purposes. One of ordinary skill in the art will appreciate that architecture 200 may include more than Figure 2 For example, the number of processing pipelines implemented by architecture 200 may be increased or decreased to include more or fewer processing pipelines. Furthermore, although architecture 200 is shown as including specific components, one of ordinary skill in the art will appreciate that architecture 200 may include more than Figure 2 For example, in some cases, the architecture 200 may also include other memory devices (e.g., one or more random access memory (RAM) components, read-only memory (ROM) components, cache memory components, buffer components, database components, and / or other memory devices), processing devices (e.g., one or more CPUs, GPUs, and / or other processing devices), interfaces (e.g., internal buses, etc.), and / or Figure 2 Other components not shown.

[0094] Figure 3is a block diagram illustrating an example architecture 300 of a video decoding system. In architecture 300, application software 302 can direct video firmware 304 and video hardware 306 to decode a bitstream 308 to memory 310 for a downstream device 312 (e.g., a display device, a network device that sends decoded images to a display device, etc.). In some cases, application software 302 can be a driver, an operating system, advanced user software, etc. In some cases, application software 302 can be executing on a CPU or other general-purpose processor. Application software 302 can indicate the decoded bitstream 308 to video firmware 304. In some cases, video firmware 304 can be a control processor for video firmware 304, such as, Figure 2 The video hardware 306 may include video hardware components for processing video data, such as Figure 2 Components include VSP 212, processing pipelines 214-220, DMA subsystem 230, interface 222, etc.

[0095] Video firmware 304 can configure video hardware 306 to obtain and decode bitstream 308. In some cases, as video hardware 306 decodes bitstream 308 into portions of images, video hardware 306 can store one or more portions of images in memory 310. In some examples, memory 310 can be similar to Figure 2 310 . In some cases, after the image is decoded and ready for display, the image can be stored in memory 310 by video hardware 306. Video hardware 306 can also send an interrupt 320 to video firmware 304 indicating that the image is ready for display. Video firmware 304 sends an interrupt 322 to application software 302 indicating that the image is ready for display. Application software 302 can receive interrupt 322, and application software 302 can indicate 324 to downstream device 312 to obtain 326 the decoded image for display. In some cases, downstream device 312 can obtain (e.g., receive) 326 the decoded image from memory 310.

[0096] Figure 4 is a block diagram illustrating the mapping of image data to bitstreams in grid 400 according to aspects of the present disclosure. Figure 4, the image data associated with each pixel can be represented by each square of grid 400, which is arranged in grid 400 as the pixels are to be displayed. Thus, square 402 represents the image data for the upper left corner of the image. In some cases, the image data representing the pixels of the image can be sent to the downstream device in a raster scan order. In raster scan order, the image data for the pixels can be sent starting from the left-hand side of the top row and proceeding down the row until the last (e.g., the rightmost) pixel. Then, starting from the left again, the image data for the pixels in the next row can be sent. Thus, the image data for the pixels represented by square 402 can be sent first, followed by square 404, and then continuing to square 406. After sending the image data for square 406, the image data for square 408 is sent, and so on, until all the image data is sent. In some cases, sending all the image data may take a certain amount of time. In some cases, if image data is sent in response to an interrupt indicating that an image is ready, the latency for displaying the image will include the time to process the interrupt and the time to send all the image data, in addition to the time to render the image. In some cases, sending an interrupt before rendering the entire image can provide early notification to a low-latency video decoder. Early interrupts (e.g., notifications) can allow overlapping decoding / rendering of an image from a bitstream and sending of the image data, thereby reducing latency.

[0097] Figure 5 is a block diagram illustrating an example architecture 500 for early notification techniques for a low-latency video decoding system in accordance with aspects of the present disclosure. In the architecture 500, application software 502 can direct video firmware 504 and video hardware 506 to decode a bitstream 508 to a memory 510 for a downstream device 512. In some cases, the application software 502 can be a driver, an operating system, advanced user software, etc. In some cases, the application software 502 can be executing on a CPU or other general-purpose processor. The application software 502 can indicate the decoded bitstream 508 to the video firmware 504. In some cases, the video firmware 504 can be a control processor for the video firmware 504, such as, Figure 2 The video hardware 506 may include video hardware components for processing video data, such as Figure 2 Components include VSP 212, processing pipelines 214-220, DMA subsystem 230, interface 222, etc.

[0098] The application software 502 may send timing information 550 for decoding to the video firmware 504. For example, the application software 502 may send an indication of a target latency for a decoded image from the bitstream 508. For example, the application software 502 may indicate a target latency of 5 milliseconds (ms). Based on the target latency, the video firmware 504 may determine a portion of the image that may be updated for the target latency. For example, if an image is processed at 60 frames per second (FPS) or one image is processed every 16.6 milliseconds (ms) at a resolution of 1920x1080, the number of rows of the image may be 1080x 5 / 16.6=325.3. In some cases, the number of rows may be rounded down, for example, based on the size of a tile, CTU, or other portion of the image. For example, the number of rows may be rounded down to 320 rows. Based on the number of rows, the image may be segmented into a set of rows (e.g., a partial image). 1-1080. The video firmware 504 may also indicate to the video hardware 506 the number of lines to be processed before sending an interrupt 552. Although discussed in the context of interrupts, it will be appreciated that any hardware notification system may be used, such as a flag, memory bit, or the like. For example, rather than sending an interrupt to indicate that an image is ready to be displayed (e.g., output to a downstream device), a flag or memory bit may be set to indicate that the image is ready.

[0099] In some cases, the video hardware 506 can decode the bitstream 508 in raster order and store decoded portions of the image in memory 510 as the bitstream 508 is decoded. In some cases, a first interrupt 556 can be sent by the video hardware 506 to the video firmware 504 based on the indicated number of lines (such as the first partial image 554) being decoded and sent to memory 510 for storage. The video firmware 504 can send a second interrupt 558 to the application software 502 to indicate that the decoded first partial image 554 of the image is ready. In some cases, the video firmware 504 can continue decoding the bitstream 508. The application software 502 can receive the second interrupt 558 and can instruct 524 the downstream device 512 to obtain 526 the decoded first partial image 554 of the image for display.

[0100] In some cases, the downstream device 512 may obtain 526 the first partial image 554 in raster scan order. As described above, the video hardware 506 may continue decoding the second partial image 560 from the bitstream 508 while the downstream device 512 obtains 526 the first partial image 554. In some cases, before the downstream device 512 obtains 526 all of the first partial image 554, the video hardware 506 may decode lines of the second partial image 560 (e.g., lines 321-640) and store the decoded second partial image 560 in the memory 510. The video hardware 506 may send a third interrupt 562 to the video firmware 504 to indicate that the second partial image 560 is ready. The video firmware 504 may send a fourth interrupt 564 to the application software 502 to indicate that the decoded second partial image 560 is ready. The application software 502 may indicate 572 to the downstream device 512 that the decoded second partial image 560 is ready. The downstream device 512 may then obtain 566 the second partial image 560. The above-described technique may then be repeated with respect to third partial image 568 and fourth partial image 570 .

[0101] In some cases, in-loop filtering can be performed as the image is decoded. In-loop filtering can be filtering applied in the decoding loop. In some cases, in-loop filtering can be used to smooth pixel transitions or to improve video quality in other ways. Examples of in-loop filters can include deblocking filters, adaptive loop filters (ALFs), sample adaptive offset (SAO) filters, cross-component adaptive loop filters, etc. In some cases, after decoding the first row, in-loop filtering can be applied to the pixels of the first row based on the pixels in the second row below the first row in raster scan order.

[0102] Figure 6 An example application of in-loop filtering to an image 600 according to aspects of the present disclosure is illustrated. Image 600 includes at least CTU A 602, CTU B 604, CTU x 606, and CTU y 608. Each CTU includes a number N of rows. For simplicity, a single column 610 of pixels is shown with multiple rows. Operations performed on rows of this single column 610 of pixels are intended to represent operations performed across rows. In some cases, rows N-2, N-1, and N of CTU A 602 may be decoded in raster scan order. Rows 1, 2, and 3 of CTU B 604 may then be decoded and in-loop filtering applied along the edges of CTU B 604, for example, to smooth transitions across different CTUs. This in-loop filtering (such as a deblocking filter) may result in changes to the pixels of rows N-2, N-1, and N of CTU A 602.

[0103] As discussed above, for early notification, if the edge of a partial image is aligned with the edge of a CTU (such as CTU A), an interrupt indicating that the partial image is ready for display can be sent after row N of CTU A 602 is decoded and stored in memory. This interrupt can be sent before a portion of CTU B 604 (e.g., rows 1, 2, and 3) is decoded and in-loop filtering is applied. Thus, in some cases, early notification can result in an interrupt being sent before in-loop filtering can be applied. This can result in visual artifacts along the boundaries of a set of rows. Early notification can be adapted to accommodate in-loop filtering without turning it off.

[0104] In some cases, for early notification adapted to accommodate in-loop filtering, early notification (e.g., interruption) can be delayed until after in-loop filtering is complete across the rows of the partial image, compared to when the last row of the partial image is decoded. For example, assuming a first partial image includes rows N-2, N-1, and N of CTU A 602, and a second partial image includes rows 1, 2, and 3 of CTU B 604, rows N-2, N-1, and N of CTU A 602 can be decoded. Rows 1, 2, and 3 of CTU B 604 can also be decoded and in-loop filtering applied to the edges of CTU B 604 based on the in-loop filter being used. Any changes to the pixels in rows N-2, N-1, and N of CTU A 602 can be made based on the in-loop filtering. After in-loop filtering is completed, the decoded first partial image including line N-2, line N-1, and line N of CTU A 602 may be output to memory, and an interrupt may be sent indicating that the first partial image is ready for display.

[0105] As discussed above, in some cases, the number of rows for a portion of an image can be based on a tile size. In some cases, this tile size can be defined based on the decoding applied to the image. In some cases, for some codecs, a tile size can also be defined to help optimize reading / writing image data to memory and / or cache. Thus, a codec can have a fixed tile size for a particular resolution and / or bit rate. As an example, for a 1920x1080 resolution with 8-bit pixel depth, some non-codec tiles (e.g., a data tile format used to optimize image storage in memory) can have a resolution of 256x64, and a 1920x1080 image can include 17 such tiles.

[0106] Figure 7is an example illustrating how a partial image for early notification interacts with image tiles of image 700 according to aspects of the present disclosure. Image 700 may have a resolution of 1920x1080. In some cases, the application software may indicate a target latency that will result in two partial images. For an image with a height of 1080 pixels, the number of pixel rows for the two partial images will be 540 pixel rows. In some cases, the number of rows in the partial image may be based on the latency and the tile height or other partitioning height. For non-codec tiles, the tile height may be 64 pixel rows, and thus the number of pixel rows for the partial image may be a factor of the number of rows in the tile (e.g., 64). This helps avoid splitting tiles across multiple partial images.

[0107] In some cases, tile size can be selected to assist memory transactions and / or cache coherence. It can be useful to avoid tiles being in multiple partial images. In some cases, the number of pixel rows can be rounded down from 540 to the nearest factor of tile height. Thus, in this example, 540 rows can be rounded down to 512 rows (e.g., 64x8), and image 700 can be split into three partial images: partial image 1 702, partial image 2 704, and partial image 3 706. In some cases, rounding down the number of rows in the partial images helps ensure that the target latency is met. As shown, the partial images can have different numbers of pixel rows.

[0108] Figure 8 800 is a flow chart of a process 800 for processing video data according to aspects of the present disclosure. Process 800 can be performed by a computing device (or apparatus) or a component of a computing device (e.g., a chipset, a codec, etc.). The computing device can be a mobile device (e.g., a mobile phone), a networked wearable device (such as a watch), an extended reality (XR) device (such as a virtual reality (VR) device or an augmented reality (AR) device), a vehicle or a component or system of a vehicle, or other type of computing device. The operations of process 800 can be implemented as software components that execute and run on one or more processors.

[0109] At block 802, a computing device (or a component thereof) may determine a number of pixel rows for one or more portions of an image (e.g., a partial image). The computing device (or a component thereof) may obtain an indication of a target latency (e.g., from an application executing on the computing device). The computing device (or a component thereof) may determine the number of pixel rows for a first portion of the image based on the target latency. The computing device (or a component thereof) may further determine the number of pixel rows for the first portion of the image based on a partition height of the first encoded data. In some cases, the number of pixel rows for the first portion of the image is aligned with the partition height of the first encoded data.

[0110] At block 804 , a computing device (or a component thereof) may obtain first encoded data for a first portion of an image.

[0111] At block 806 , the computing device (or a component thereof) may decode the first encoded data to generate first pixel data for the number of pixel rows of the first portion of the image.

[0112] At block 808, the computing device (or its components) may output the first pixel data for the first portion of the image to a memory. The computing device (or its components) may obtain second encoded data for the second portion of the image. In some cases, the second portion of the image does not overlap with the first portion of the image. The computing device (or its components) may decode the second encoded data to generate second pixel data for the second portion of the image for the number of pixel rows. The computing device (or its components) may output the second pixel data for the second portion of the image to a memory. The computing device (or its components) may output an indication that the second portion of the image is available.

[0113] At block 810, the computing device (or a component thereof) may output an indication that the first portion of the image is available. In some cases, the indication includes at least one of an interrupt, a flag, or a memory bit. The computing device (or a component thereof) may output an indication that the first portion of the image is available after decoding the first encoded data and before decoding the second encoded data. The computing device (or a component thereof) may delay outputting the indication until at least a portion of the second encoded data is decoded. The computing device (or a component thereof) may apply an in-loop filter to update one or more pixel rows of the first portion of the image based on the second pixel data. In some cases, the delayed output is based on the amount of time used to apply the in-loop filter.

[0114] The processes (or methods) described herein may be used individually or in any combination. In some implementations, the processes (or methods) described herein may be performed by a computing device or apparatus (such as, Figure 1For example, the process may be performed by Figure 1 and Figure 10 , and / or performed by another client-side device (such as a player device, a display, or any other client-side device). In some cases, the computing device or apparatus may include: one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or one or more other components configured to perform the steps of one or more processes described herein.

[0115] In some examples, the computing device may include a mobile device, a desktop computer, a server computer and / or a server system, or other types of computing devices. The components of the computing device (e.g., one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components) may be implemented in circuits. For example, the components may include electronic circuits or other electronic hardware, and / or may be implemented using electronic circuits or other electronic hardware, the electronic circuits or other electronic hardware may include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or the components may include computer software, firmware, or a combination thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or a combination thereof for performing the various operations described herein. In some examples, the computing device or apparatus may include a camera configured to capture video data (e.g., a video sequence) comprising video frames. In some examples, the camera or other capture device that captures the video data is separate from the computing device, in which case the computing device receives or obtains the captured video data. The computing device may include a network interface configured to transmit the video data. The network interface may be configured to transmit Internet Protocol (IP) based data or other types of data.In some examples, the computing device or apparatus may include a display for displaying output video content (such as samples of pictures of a video bitstream).

[0116] The process may be described with respect to a logical flow diagram, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform specific functions or implement specific data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the process.

[0117] Furthermore, process 600 may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed collectively on one or more processors, implemented by hardware, or a combination thereof. As mentioned above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions that may be executed by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0118] The decoding techniques discussed herein may be implemented in an example video encoding and decoding system (e.g., system 100). In some examples, a system includes a source device that provides encoded video data to be decoded by a destination device at a later time. In particular, the source device provides the video data to the destination device via a computer-readable medium. The source device and the destination device may include any of a variety of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called "smart" phones, so-called "smart" tablets, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, and the like. In some cases, the source device and the destination device may be equipped for wireless communication.

[0119] The destination device can receive the encoded video data to be decoded via a computer-readable medium. The computer-readable medium may include any type of medium or device capable of moving the encoded video data from the source device to the destination device. In one example, the computer-readable medium may include a communication medium so that the source device can directly send the encoded video data to the destination device in real time. The encoded video data can be modulated according to a communication standard (such as a wireless communication protocol) and sent to the destination device. The communication medium may include any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, a switch, a base station, or any other device that can be used to facilitate communication from the source device to the destination device.

[0120] In some examples, the encoded data can be output from the output interface to a storage device. Similarly, the encoded data can be accessed from the storage device via the input interface. The storage device can include any of a variety of distributed or locally accessible data storage media, such as a hard drive, Blu-ray disc, DVD, CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. In another example, the storage device can correspond to a file server or another intermediate storage device that can store the encoded video generated by the source device. The destination device can access the stored video data from the storage device via streaming or downloading. The file server can be any type of server that can store the encoded video data and send the encoded video data to the destination device. Example file servers include network servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The destination device can access the encoded video data via any standard data connection (including an internet connection). This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.) suitable for accessing the encoded video data stored on the file server, or a combination of the two. The transmission of the encoded video data from the storage device may be a streaming transmission, a download transmission, or a combination thereof.

[0121] The technology of the present disclosure is not necessarily limited to wireless applications or settings. The technology can be applied to support video decoding of any of a variety of multimedia applications, such as: over-the-air television broadcasting, cable television transmission, satellite television transmission, Internet streaming video transmission (such as, dynamic adaptive streaming over HTTP (DASH)), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, the system can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting and / or video telephony.

[0122] In one example, a source device includes a video source, a video encoder, and an output interface. A destination device may include an input interface, a video decoder, and a display device. The video encoder of the source device may be configured to apply the technology disclosed herein. In other examples, the source device and the destination device may include other components or arrangements. For example, the source device may receive video data from an external video source such as an external camera. Similarly, the destination device may be connected to an external display device via an interface, rather than including an integrated display device.

[0123] The above example system is merely an example. The techniques for processing video data in parallel can be performed by any digital video encoding and / or decoding device. Although generally speaking, the techniques of the present disclosure are performed by a video encoding device, these techniques can also be performed by a video encoder / decoder, which is commonly referred to as a "CODEC." In addition, the techniques of the present disclosure can also be performed by a video preprocessor. The source device and the destination device are merely examples of such decoding devices, wherein the source device generates decoded video data for transmission to the destination device. In some examples, the source device and the destination device can operate in a substantially symmetrical manner, so that each of these devices includes a video encoding and decoding component. Therefore, the example system can support one-way or two-way video transmission between video devices, for example, for video streaming transmission, video playback, video broadcasting, or video telephony.

[0124] The source device may include a video capture device, such as a video camera, a video archive comprising previously captured video, and / or a video feed interface for receiving video from a video content provider. As another alternative, the video source may generate computer graphics-based data as the source video, or a combination of real-time video, archived video, and computer-generated video. In some cases, if the video source is a video camera, the source device and the destination device may form a so-called camera phone or video phone. However, as mentioned above, the technology described in this disclosure may generally be applicable to video decoding and may be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video may be encoded by a video encoder. The encoded video information may then be output to a computer-readable medium by an output interface.

[0125] As described above, computer-readable media may include transient media (such as wireless broadcast or wired network transmission), or storage media (i.e., non-transitory storage media) (such as a hard disk, flash drive, compact disc, digital video disc, Blu-ray disc, or other computer-readable media). In some examples, a network server (not shown) can receive encoded video data from a source device and provide the encoded video data to a destination device, for example, via a network transmission. Similarly, a computing device of a media production facility such as an optical disc stamping device can receive encoded video data from a source device and generate an optical disc containing the encoded video data. Thus, in various examples, computer-readable media may be understood to include one or more computer-readable media in various forms.

[0126] The input interface of the destination device receives information from the computer-readable medium. The information of the computer-readable medium may include syntax information defined by the video encoder, which is also used by the video decoder, and includes syntax elements that describe characteristics of blocks and other decoded units (e.g., groups of pictures (GOPs)) and / or processing of blocks and other decoded units. The display device displays the decoded video data to a user and may include any of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device. Various aspects of the present application have been described.

[0127] The details of the encoding device 104 and the decoding device 112 are respectively Figure 9 and Figure 10 Shown in. Figure 9is a block diagram illustrating an example encoding device 104 that can implement one or more of the techniques described in this disclosure. The encoding device 104 can, for example, generate syntax elements and / or structures described herein (e.g., syntax elements and / or structures of green metadata (such as complexity measures (CM)) or other syntax elements and / or structures). The encoding device 104 can perform intra-frame prediction and inter-frame prediction decoding on video blocks within video slices, tiles, sub-pictures, etc. As previously described, intra-frame decoding relies at least in part on spatial prediction to reduce or remove spatial redundancy within a given video frame or picture. Inter-frame decoding relies at least in part on temporal prediction to reduce or remove temporal redundancy within adjacent or surrounding frames of a video sequence. Intra-frame mode (I-mode) can refer to any of several spatial-based compression modes. Inter-frame modes such as unidirectional prediction (P-mode) or bidirectional prediction (B-mode) can refer to any of several temporal-based compression modes.

[0128] The encoding device 104 includes a partitioning unit 35, a prediction processing unit 41, a filter unit 63, a picture memory 64, a summer 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. The prediction processing unit 41 includes a motion estimation unit 42, a motion compensation unit 44, and an intra-frame prediction processing unit 46. For video block reconstruction, the encoding device 104 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and a summer 62. The filter unit 63 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although in Figure 6 Filter unit 63 is shown as an in-loop filter in FIG. 1 , but in other configurations, filter unit 63 may be implemented as a post-loop filter. Post-processing device 57 may perform additional processing on the encoded video data generated by encoding device 104. In some cases, the techniques of this disclosure may be implemented by encoding device 104. However, in other cases, one or more of the techniques of this disclosure may be implemented by post-processing device 57.

[0129] like Figure 9As shown in , encoding device 104 receives video data, and partitioning unit 35 partitions the data into video blocks. Partitioning may also include partitioning into slices, slice segments, tiles, or other larger units, as well as video block partitioning based on, for example, a quadtree structure of LCUs and CUs. Encoding device 104 generally illustrates components for encoding video blocks within a video slice to be encoded. A slice may be partitioned into multiple video blocks (and potentially into sets of video blocks called tiles). Prediction processing unit 41 may select one of multiple possible decoding modes for the current video block based on error results (e.g., decoding rate and distortion level, etc.), such as one of multiple intra-frame prediction decoding modes or one of multiple inter-frame prediction decoding modes. Prediction processing unit 41 may provide the resulting intra-frame decoded or inter-frame decoded block to summer 50 to generate residual block data and to summer 62 to reconstruct the encoded block for use as a reference picture.

[0130] Intra-prediction processing unit 46 within prediction processing unit 41 may perform intra-predictive coding of the current video block relative to one or more neighboring blocks in the same frame or slice as the current block to be coded to provide spatial compression. Motion estimation unit 42 and motion compensation unit 44 within prediction processing unit 41 may perform inter-predictive coding of the current video block relative to one or more predictive blocks in one or more reference pictures to provide temporal compression.

[0131] Motion estimation unit 42 may be configured to determine an inter-prediction mode for a video slice according to a predetermined pattern for a video sequence. The predetermined pattern may designate a video slice in the sequence as a P slice, a B slice, or a GPB slice. Motion estimation unit 42 and motion compensation unit 44 may be highly integrated but are described separately for conceptual purposes. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors, which estimate motion for a video block. For example, a motion vector may indicate the displacement of a prediction unit (PU) of a video block within a current video frame or picture relative to a predictive block within a reference picture.

[0132] A predictive block is a block that is found to closely match a PU of the video block to be coded in terms of pixel difference, which may be determined by sum of absolute difference (SAD), sum of squared difference (SSD), or other difference metrics. In some examples, encoding device 104 may calculate values for sub-integer pixel positions of a reference picture stored in picture memory 64. For example, encoding device 104 may interpolate values for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference picture. Thus, motion estimation unit 42 may perform motion searches relative to full pixel positions as well as fractional pixel positions and output motion vectors with fractional pixel precision.

[0133] Motion estimation unit 42 calculates a motion vector for a PU of a video block in a slice being inter-coded by comparing the position of the PU with the position of a predictive block of a reference picture. The reference picture may be selected from a first reference picture list (list 0) or a second reference picture list (list 1), each of which identifies one or more reference pictures stored in picture memory 64. Motion estimation unit 42 sends the calculated motion vector to entropy encoding unit 56 and motion compensation unit 44.

[0134] Motion compensation performed by motion compensation unit 44 may involve extracting or generating a predictive block based on the motion vector determined by motion estimation, possibly interpolating to sub-pixel precision. Upon receiving the motion vector for the PU of the current video block, motion compensation unit 44 may locate the predictive block to which the motion vector points in a reference picture list. Encoding device 104 forms a residual video block by subtracting the pixel values of the predictive block from the pixel values of the current video block being decoded to form pixel difference values. The pixel difference values form residual data for the block and may include both luma and chroma difference components. Summer 50 represents one or more components that perform this subtraction operation. Motion compensation unit 44 may also generate syntax elements associated with the video block and video slice for use by decoding device 112 when decoding the video block of the video slice.

[0135] As an alternative to the inter-frame prediction performed by motion estimation unit 42 and motion compensation unit 44 as described above, intra-frame prediction processing unit 46 can perform intra-frame prediction on the current block. In particular, intra-frame prediction processing unit 46 can determine an intra-frame prediction mode to be used to encode the current block. In some examples, intra-frame prediction processing unit 46 can encode the current block using various intra-frame prediction modes, for example during separate encoding passes, and intra-frame prediction processing unit 46 can select an appropriate intra-frame prediction mode to use from the tested modes. For example, intra-frame prediction processing unit 46 can calculate rate-distortion values using a rate-distortion analysis for the various tested intra-frame prediction modes and can select the intra-frame prediction mode with the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between the coded block and the original uncoded block that was coded to produce the coded block, as well as the bit rate (i.e., the number of bits) used to produce the coded block. Intra-prediction processing unit 46 may calculate ratios from the distortions and rates for the various encoded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block.

[0136] In any case, after selecting an intra-prediction mode for a block, the intra-prediction processing unit 46 may provide information indicating the selected intra-prediction mode for the block to the entropy coding unit 56. The entropy coding unit 56 may encode the information indicating the selected intra-prediction mode. The encoding device 104 may include in the transmitted bitstream configuration data defining the coding context for each block, as well as an indication of the most probable intra-prediction mode to be used for each of the contexts, an intra-prediction mode index table, and a modified intra-prediction mode index table. The bitstream configuration data may include multiple intra-prediction mode index tables and multiple modified intra-prediction mode index tables (also known as codeword mapping tables).

[0137] After the prediction processing unit 41 generates a predictive block for the current video block via inter-frame prediction or intra-frame prediction, the encoding device 104 forms a residual video block by subtracting the predictive block from the current video block. The residual video data in the residual block may be included in one or more TUs and applied to the transform processing unit 52. The transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform (such as a discrete cosine transform (DCT) or a conceptually similar transform). The transform processing unit 52 may convert the residual video data from the pixel domain to a transform domain (such as the frequency domain).

[0138] Transform processing unit 52 may send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all coefficients. The degree of quantization may be modified by adjusting a quantization parameter. In some examples, quantization unit 54 may then perform a scan of the matrix comprising the quantized transform coefficients. Alternatively, entropy coding unit 56 may perform the scan.

[0139] After quantization, entropy coding unit 56 performs entropy encoding on the quantized transform coefficients. For example, entropy coding unit 56 may perform context adaptive variable length coding (CAVLC), context adaptive binary arithmetic coding (CABAC), syntax-based context adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding, or another entropy coding technique. After entropy encoding by entropy coding unit 56, the encoded bitstream may be sent to decoding device 112 or archived for later transmission or retrieval by decoding device 112. Entropy coding unit 56 may also entropy encode motion vectors and other syntax elements for the current video slice being decoded.

[0140] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct the residual block in the pixel domain for later use as a reference block of a reference picture. Motion compensation unit 44 may calculate a reference block by adding the residual block to a predictive block of one of the reference pictures in the reference picture list. Motion compensation unit 44 may also apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer pixel values for motion estimation. Summer 62 adds the reconstructed residual block to the motion compensated prediction block produced by motion compensation unit 44 to produce a reference block for storage in picture memory 64. The reference block may be used by motion estimation unit 42 and motion compensation unit 44 as a reference block for inter-frame prediction of blocks in subsequent video frames or pictures.

[0141] In this way, Figure 9 The encoding device 104 represents an example of a video encoder configured to perform any of the techniques described herein. In some cases, some of the techniques of this disclosure may also be implemented by the post-processing device 57.

[0142] Figure 10 is a block diagram illustrating an example decoding device 112. The decoding device 112 includes an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, a summer 90, a filter unit 91, and a picture memory 92. The prediction processing unit 81 includes a motion compensation unit 82 and an intra-frame prediction processing unit 84. In some examples, the decoding device 112 may perform general processing related to the processing of the image from the decoded image. Figure 6 The encoding passes described by the encoding device 104 are reciprocal to the decoding passes.

[0143] During the decoding process, decoding device 112 receives an encoded video bitstream representing video blocks of an encoded video slice and associated syntax elements sent by encoding device 104. In some aspects, decoding device 112 may receive the encoded video bitstream from encoding device 104. In some aspects, decoding device 112 may receive the encoded video bitstream from network entity 79 (such as a server, a media-aware network element (MANE), a video editor / splicer, or other such device configured to implement one or more of the techniques described above). Network entity 79 may or may not include encoding device 104. Some of the techniques described in this disclosure may be implemented by network entity 79 before it sends the encoded video bitstream to decoding device 112. In some video decoding systems, network entity 79 and decoding device 112 may be parts of separate devices, while in other cases, the functionality described with respect to network entity 79 may be performed by the same device including decoding device 112.

[0144] The entropy decoding unit 80 of the decoding device 112 performs entropy decoding on the bitstream to generate quantized coefficients, motion vectors, and other syntax elements. The entropy decoding unit 80 forwards the motion vectors and other syntax elements to the prediction processing unit 81. The decoding device 112 can receive syntax elements at the video slice level and / or the video block level. The entropy decoding unit 80 can process and analyze both fixed-length syntax elements and variable-length syntax elements in one or more parameter sets (such as VPS, SPS, and PPS).

[0145] When the video slice is decoded as an intra-coded (I) slice, intra-prediction processing unit 84 of prediction processing unit 81 may generate prediction data for the video block of the current video slice based on the signaled intra-prediction mode and data from previously decoded blocks of the current frame or picture. When the video frame is decoded as an inter-coded (i.e., B, P, or GPB) slice, motion compensation unit 82 of prediction processing unit 81 generates a predictive block for the video block of the current video slice based on the motion vector and other syntax elements received from entropy decoding unit 80. The predictive block may be generated based on one of the reference pictures in the reference picture list. Decoding device 112 may construct reference frame lists: List 0 and List 1, based on the reference pictures stored in picture memory 92 using a default construction technique.

[0146] Motion compensation unit 82 determines prediction information for a video block of the current video slice by parsing motion vectors and other syntax elements, and uses the prediction information to generate a predictive block for the current video block being decoded. For example, motion compensation unit 82 may use one or more syntax elements in a parameter set to determine a prediction mode (e.g., intra-prediction or inter-prediction) used to code the video block of the video slice, an inter-prediction slice type (e.g., a B slice, a P slice, or a GPB slice), construction information for one or more reference picture lists for the slice, a motion vector for each inter-coded video block of the slice, an inter-prediction state for each inter-coded video block of the slice, and other information for decoding the video blocks in the current video slice.

[0147] Motion compensation unit 82 may also perform interpolation based on interpolation filters. Motion compensation unit 82 may use interpolation filters, as used by encoding device 104 during encoding of the video block, to calculate interpolated values for sub-integer pixels of a reference block. In this case, motion compensation unit 82 may determine the interpolation filters, as used by encoding device 104, from received syntax elements and may use the interpolation filters to produce a predictive block.

[0148] Inverse quantization unit 86 inverse quantizes or dequantizes the quantized transform coefficients provided in the bitstream and decoded by entropy decoding unit 80. The inverse quantization process may include using a quantization parameter calculated by encoding device 104 for each video block in a video slice to determine the degree of quantization and, likewise, the degree of inverse quantization that should be applied. Inverse transform processing unit 88 applies an inverse transform (e.g., an inverse DCT or other suitable inverse transform), an inverse integer transform, or a conceptually similar inverse transform process to the transform coefficients to produce a residual block in the pixel domain.

[0149] After the motion compensation unit 82 generates a predictive block for the current video block based on the motion vector and other syntax elements, the decoding device 112 forms a decoded video block by summing the residual block from the inverse transform processing unit 88 with the corresponding predictive block generated by the motion compensation unit 82. Summer 90 represents one or more components that perform this summing operation. If desired, loop filters (in the decoding loop or after the decoding loop) can also be used to smooth pixel transitions, or to improve video quality in other ways. Filter unit 91 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although in Figure 10 Filter unit 91 is shown as an in-loop filter in FIG, but in other configurations, filter unit 91 may be implemented as a post-loop filter. The decoded video blocks in a given frame or picture are then stored in picture memory 92, which stores reference pictures used for subsequent motion compensation. Picture memory 92 also stores the decoded video for later display on a display device such as a Figure 1 122).

[0150] In this way, Figure 10 The decoding device 112 represents an example of a video decoder configured to perform any of the techniques described herein.

[0151] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. A computer-readable medium may include a non-transitory medium in which data can be stored, which does not include a carrier wave and / or a temporary electronic signal that is transmitted wirelessly or over a wired connection. Examples of non-transitory media may include, but are not limited to, a disk or tape, an optical storage medium (such as a compact disc (CD) or a digital versatile disc (DVD)), a flash memory, a memory, or a storage device. A computer-readable medium may have code and / or machine-executable instructions stored thereon (which may represent any combination of a process, function, subroutine, program, routine, subroutine, module, software package, class, or instruction, data structure, or program statement). A code segment may be coupled to another code segment or a hardware circuit by transmitting and / or receiving information, data, parameters, or memory contents. Information, parameters, parameters, data, etc. may be transmitted, forwarded, or sent via any suitable means, including memory sharing, message passing, token passing, network transmission, etc.

[0152] In some aspects, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media expressly excludes media such as energy, carrier signals, electromagnetic waves, and signals themselves.

[0153] Specific details are provided in the above description to provide a thorough understanding of the aspects and examples provided herein. However, one of ordinary skill in the art will appreciate that aspects can be practiced without these specific details. For clarity of explanation, in some cases, the technology herein can be presented as comprising individual functional blocks comprising the following functional blocks, which comprise devices, device components, steps or routines in a method implemented with software, or a combination of hardware and software. In addition to the components shown in the accompanying drawings and / or described herein, additional components can also be used. For example, circuits, systems, networks, processes, and other components can be shown as components in block diagram form to avoid blurring aspects with unnecessary details. In other cases, well-known circuits, processes, algorithms, structures, and techniques can be shown without unnecessary details to avoid blurring aspects.

[0154] Various aspects may be described above as processes or methods depicted as flowcharts, flow diagrams, data flow diagrams, structure diagrams, or block diagrams. Although flowcharts may describe operations as sequential processes, many operations may be performed in parallel or simultaneously. Additionally, the order of these operations may be rearranged. When the operations of a process are completed, the process is terminated, but may have additional steps not included in the accompanying drawings. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, the termination of the process may correspond to the function returning to the calling function or main function.

[0155] The processes and methods according to the above examples can be implemented using computer-executable instructions stored or otherwise accessible from a computer-readable medium. For example, such instructions may include instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a specific function or group of functions. Portions of computer resources that can be accessed for use on a network. Computer-executable instructions can be, for example, binary files, intermediate format instructions, such as assembly language, firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, information used, and / or information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, networked storage devices, etc.

[0156] The equipment implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description language or any combination thereof, and may adopt any of a variety of form factors. When implemented in software, firmware, middleware or microcode, the program code or code segments (e.g., computer program products) for performing the necessary tasks may be stored in a computer-readable or machine-readable medium. The processor may perform the necessary tasks. Typical examples of form factors include personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc., of laptop computers, smart phones, mobile phones, tablet devices or other small form factors. The functions described herein may also be implemented in peripheral devices or add-on cards. By further example, such functions may also be implemented in the middle of different processes performed on different chips or in a single device on a circuit board.

[0157] Instructions, media for transmitting such instructions, computing resources for executing them, and other structure for supporting such computing resources are example means for providing the functionality described in this disclosure.

[0158] In the foregoing description, although various aspects of the present application have been described with reference to specific aspects of the present application, those skilled in the art will recognize that the present application is not limited thereto. Therefore, although the illustrative aspects of the present application have been described in detail herein, it should be understood that the inventive concept can be implemented and adopted differently in other ways, and the appended claims are intended to be interpreted as including such variations, except as limited by the prior art. The various features and aspects of the above-mentioned application can be used individually or in combination. In addition, the various aspects can be used in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the description and drawings should be considered to be illustrative rather than restrictive. For illustration purposes, the methods are described in a particular order. It should be understood that, in alternative aspects, these methods can be performed in an order different from the described order.

[0159] One of ordinary skill will recognize that the less than ("<") and greater than (">") symbols or terms used herein may be replaced with less than or equal to ("≤") and greater than or equal to ("≥") symbols, respectively, without departing from the scope of this specification.

[0160] When a component is described as being “configured to” perform a particular operation, such configuration may be achieved, for example, by designing electronic circuits or other hardware to perform the operation, programming a programmable electronic circuit (e.g., a microprocessor or other suitable electronic circuit) to perform the operation, or any combination thereof.

[0161] The phrase "coupled to" refers to any component that is directly or indirectly physically connected to another component, and / or any component that is directly or indirectly in communication with another component (e.g., connected to another component over a wired or wireless connection and / or other appropriate communication interface).

[0162] Claim language or other language in this disclosure that recites "at least one of" a set and / or "one or more items" in a set indicates that one member of the set or multiple members of the set (in any combination) satisfies the claim. For example, claim language that recites "at least one of A and B" means A, B, or A and B. In another example, claim language that recites "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one of" a set and / or "one or more items" in a set does not limit the set to the items listed in the set. For example, claim language that recites "at least one of A and B" means A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0163] The various illustrative logic blocks, modules, circuits and algorithmic steps described in conjunction with aspects disclosed herein can be implemented as electronic hardware, computer software, firmware or a combination thereof. In order to clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits and steps have been generally described in terms of their functionality. Whether this functionality is implemented as hardware or software depends on specific application and the design constraints imposed on the entire system. Technicians can implement the described functions in different ways for each specific application, but such implementation decision-making should not be interpreted as causing departure from the scope of the application.

[0164] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices, handheld devices, or integrated circuit devices with multiple uses, including applications in wireless communication devices, handheld devices, and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be implemented at least in part by a computer-readable data storage medium comprising program code that, when executed, includes instructions for performing one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage media, and the like. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.

[0165] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated logic circuits or discrete logic circuits. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor can be a microprocessor, but the processor can also be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein can refer to any of the aforementioned structures, any combination of the aforementioned structures, or any other structure or device suitable for implementing the technology described herein. In addition, in some aspects, the functions described herein can be provided within a dedicated software module and / or hardware module configured for encoding and decoding, or incorporated into a combined video codec (CODEC).

[0166] Illustrative aspects of the present disclosure include:

[0167] Aspect 1. A device for processing video data, comprising: a memory; and a processor coupled to the memory, the processor being configured to: determine the number of pixel rows for one or more portions of an image; obtain first encoded data for a first portion of the image; decode the first encoded data to generate first pixel data for the number of pixel rows for the first portion of the image; output the first pixel data for the first portion of the image to the memory; and output an indication that the first portion of the image is available.

[0168] Aspect 2. An apparatus according to aspect 1, wherein the processor is further configured to: obtain an indication of a target latency; and determine the number of pixel rows for the first portion of the image based on the target latency.

[0169] Aspect 3. The apparatus according to aspect 2, wherein the processor is configured to: determine the number of pixel rows for the first portion of the image further based on a partition height of the first encoded data.

[0170] Clause 4. The apparatus of clause 3, wherein the number of pixel rows for the first portion of the image is aligned with the partition height of the first encoded data.

[0171] Aspect 5. The apparatus of any one of aspects 1-4, wherein the indication comprises at least one of an interrupt, a flag, or a memory bit.

[0172] Aspect 6. An apparatus according to any one of Aspects 1-5, wherein the processor is further configured to: obtain second encoded data for a second part of the image, wherein the second part of the image does not overlap with the first part of the image; decode the second encoded data to generate second pixel data for the second part of the image for the number of pixel rows; output the second pixel data for the second part of the image to the memory; and output an indication that the second part of the image is available.

[0173] Aspect 7. An apparatus according to aspect 6, wherein the processor is configured to: output the indication that the first portion of the image is available after decoding the first encoded data and before decoding the second encoded data.

[0174] Aspect 8. The apparatus according to aspect 6, wherein the processor is further configured to: delay output of the indication until at least a portion of the second encoded data is decoded.

[0175] Aspect 9. An apparatus according to Aspect 8, wherein the processor is further configured to: apply an in-loop filter to update one or more pixel rows of the first portion of the image based on the second pixel data, and wherein the delayed output is based on the amount of time used to apply the in-loop filter.

[0176] Aspect 10. The apparatus according to any one of aspects 1 to 9, wherein the apparatus comprises a decoder device.

[0177] Aspect 11. A method for processing video data, comprising: determining a number of pixel rows for one or more portions of an image; obtaining first encoded data for a first portion of the image; decoding the first encoded data to generate first pixel data for the number of pixel rows for the first portion of the image; outputting the first pixel data for the first portion of the image to a memory; and outputting an indication that the first portion of the image is available.

[0178] Aspect 12. The method according to aspect 11, further comprising: obtaining an indication of a target latency; and determining the number of pixel rows for the first portion of the image based on the target latency.

[0179] Clause 13. The method of clause 12, wherein determining the number of pixel rows for the first portion of the image is further based on a partition height of the first encoded data.

[0180] Clause 14. The method of clause 13, wherein the number of pixel rows for the first portion of the image is aligned with the partition height of the first encoded data.

[0181] Aspect 15. The method according to any one of aspects 11-14, wherein the indication comprises at least one of an interrupt, a flag, or a memory bit.

[0182] Aspect 16. The method according to any one of Aspects 11-15 further includes: obtaining second encoded data for a second part of the image, wherein the second part of the image does not overlap with the first part of the image; decoding the second encoded data to generate second pixel data for the second part of the image for the number of pixel rows; outputting the second pixel data for the second part of the image to the memory; and outputting an indication that the second part of the image is available.

[0183] Clause 17. The method according to clause 16, further comprising: outputting the indication that the first portion of the image is available after decoding the first encoded data and before decoding the second encoded data.

[0184] Aspect 18. The method according to aspect 16, further comprising: delaying output of the indication until at least a portion of the second encoded data is decoded.

[0185] Aspect 19. The method according to Aspect 18 further includes: applying an in-loop filter to update one or more pixel rows of the first portion of the image based on the second pixel data, and wherein the delayed output is based on the amount of time used to apply the in-loop filter.

[0186] Aspect 20. A non-transitory computer-readable medium having instructions stored thereon, which, when executed by a processor, causes the processor to: determine a number of pixel rows for one or more portions of an image; obtain first encoded data for a first portion of the image; decode the first encoded data to generate first pixel data for the number of pixel rows for the first portion of the image; output the first pixel data for the first portion of the image to a memory; and output an indication that the first portion of the image is available.

[0187] Aspect 21. The non-transitory computer-readable medium of aspect 20, wherein the instructions further cause the processor to: obtain an indication of a target latency; and determine the number of pixel rows for the first portion of the image based on the target latency.

[0188] Aspect 22. The non-transitory computer-readable medium of aspect 21, wherein the instructions further cause the processor to determine the number of pixel rows for the first portion of the image based further on a partition height of the first encoded data.

[0189] Clause 23. The non-transitory computer-readable medium of clause 22, wherein the number of pixel rows for the first portion of the image is aligned with the partition height of the first encoded data.

[0190] Aspect 24. The non-transitory computer-readable medium of any one of aspects 20-23, wherein the indication comprises at least one of an interrupt, a flag, or a memory bit.

[0191] Aspect 25. A non-transitory computer-readable medium according to any one of Aspects 20-24, wherein the instructions further cause the processor to: obtain second encoded data for a second portion of the image, wherein the second portion of the image does not overlap with the first portion of the image; decode the second encoded data to generate second pixel data for the second portion of the image for the number of pixel rows; output the second pixel data for the second portion of the image to the memory; and output an indication that the second portion of the image is available.

[0192] Aspect 26. The non-transitory computer-readable medium of aspect 25, wherein the instructions further cause the processor to output the indication that the first portion of the image is available after decoding the first encoded data and before decoding the second encoded data.

[0193] Aspect 27. The non-transitory computer-readable medium of aspect 25, wherein the instructions further cause the processor to delay outputting the indication until at least a portion of the second encoded data is decoded.

[0194] Aspect 28. A non-transitory computer-readable medium according to Aspect 27, wherein the instructions further cause the processor to apply an in-loop filter to update one or more pixel rows of the first portion of the image based on the second pixel data, and wherein the delayed output is based on an amount of time used to apply the in-loop filter.

[0195] Aspect 29. A device for processing video data, comprising: a unit for determining the number of pixel rows for one or more portions of an image; a unit for obtaining first encoded data for a first portion of the image; a unit for decoding the first encoded data to generate first pixel data for the number of pixel rows for the first portion of the image; a unit for outputting the first pixel data for the first portion of the image to a memory; and a unit for outputting an indication that the first portion of the image is available.

[0196] Clause 30. The apparatus of clause 29, further comprising: means for obtaining an indication of a target latency; and means for determining the number of pixel rows for the first portion of the image based on the target latency.

[0197] Aspect 32: An apparatus for processing video data, comprising one or more means for performing the operations according to aspects 12 to 19.

Claims

1. A device for processing video data, comprising: Memory; as well as a processor coupled to the memory, the processor configured to: determining a number of pixel rows for one or more portions of an image; obtaining first encoded data for a first portion of the image; decoding the first encoded data to generate first pixel data for the number of pixel rows of the first portion of the image; outputting the first pixel data for the first portion of the image to the memory; as well as An indication is output that the first portion of the image is available.

2. The device according to claim 1, wherein The processor is further configured to: obtaining an indication of a target delay; and The number of pixel rows for the first portion of the image is determined based on the target latency.

3. The device according to claim 2, wherein The processor is configured to determine the number of pixel rows for the first portion of the image further based on a partition height of the first encoded data.

4. The device according to claim 3, wherein The number of pixel rows for the first portion of the image is aligned with the division height of the first encoded data.

5. The device according to claim 1, wherein The indication comprises at least one of an interrupt, a flag, or a memory bit.

6. The device according to claim 1, wherein The processor is further configured to: obtaining second encoded data for a second portion of the image, wherein the second portion of the image does not overlap with the first portion of the image; decoding the second encoded data to generate second pixel data for the number of pixel rows for the second portion of the image; outputting the second pixel data for the second portion of the image to the memory; and An indication is output that the second portion of the image is available.

7. The device according to claim 6, wherein The processor is configured to output the indication that the first portion of the image is available after decoding the first encoded data and before decoding the second encoded data.

8. The device according to claim 6, wherein The processor is further configured to delay outputting the indication until at least a portion of the second encoded data is decoded.

9. The device according to claim 8, wherein The processor is further configured to apply an in-loop filter to update one or more rows of pixels of the first portion of the image based on the second pixel data, and wherein the delayed output is based on an amount of time for applying the in-loop filter.

10. The device according to claim 1, wherein The apparatus comprises a decoder device.

11. A method for processing video data, comprising: determining a number of pixel rows for one or more portions of an image; obtaining first encoded data for a first portion of the image; decoding the first encoded data to generate first pixel data for the number of pixel rows of the first portion of the image; outputting the first pixel data for the first portion of the image to a memory; as well as An indication is output that the first portion of the image is available.

12. The method according to claim 11, further comprising: obtaining an indication of a target delay; as well as The number of pixel rows for the first portion of the image is determined based on the target latency.

13. The method according to claim 12, wherein: Determining the number of pixel rows for the first portion of the image is further based on a partition height of the first encoded data.

14. The method according to claim 13, wherein The number of pixel rows for the first portion of the image is aligned with the division height of the first encoded data.

15. The method according to claim 11, wherein The indication comprises at least one of an interrupt, a flag, or a memory bit.

16. The method according to claim 11, further comprising: obtaining second encoded data for a second portion of the image, wherein the second portion of the image does not overlap with the first portion of the image; decoding the second encoded data to generate second pixel data for the number of pixel rows for the second portion of the image; outputting the second pixel data for the second portion of the image to the memory; and An indication is output that the second portion of the image is available.

17. The method according to claim 16, further comprising: The indication that the first portion of the image is available is output after decoding the first encoded data and before decoding the second encoded data.

18. The method according to claim 16, further comprising: Outputting the indication is delayed until at least a portion of the second encoded data is decoded.

19. The method according to claim 18, further comprising: An in-loop filter is applied to update one or more rows of pixels of the first portion of the image based on the second pixel data, and wherein the delayed output is based on an amount of time used to apply the in-loop filter.

20. A non-transitory computer-readable medium having stored thereon instructions that, when executed by a processor, cause the processor to: determining a number of pixel rows for one or more portions of an image; obtaining first encoded data for a first portion of the image; decoding the first encoded data to generate first pixel data for the number of pixel rows of the first portion of the image; outputting the first pixel data for the first portion of the image to a memory; and outputting an indication that the first portion of the image is available.

21. The non-transitory computer readable medium of claim 20, wherein: The instructions further cause the processor to: obtaining an indication of a target delay; and The number of pixel rows for the first portion of the image is determined based on the target latency.

22. The non-transitory computer readable medium of claim 21, wherein: The instructions also cause the processor to determine the number of pixel rows for the first portion of the image further based on a partition height of the first encoded data.

23. The non-transitory computer readable medium of claim 22, wherein: The number of pixel rows for the first portion of the image is aligned with the division height of the first encoded data.

24. The non-transitory computer readable medium of claim 20, wherein: The indication comprises at least one of an interrupt, a flag, or a memory bit.

25. The non-transitory computer readable medium of claim 20, wherein: The instructions further cause the processor to: obtaining second encoded data for a second portion of the image, wherein the second portion of the image does not overlap with the first portion of the image; decoding the second encoded data to generate second pixel data for the number of pixel rows for the second portion of the image; outputting the second pixel data for the second portion of the image to the memory; and An indication is output that the second portion of the image is available.

26. The non-transitory computer readable medium of claim 25, wherein: The instructions further cause the processor to output the indication that the first portion of the image is available after decoding the first encoded data and before decoding the second encoded data.

27. The non-transitory computer-readable medium of claim 25, wherein: The instructions further cause the processor to delay outputting the indication until at least a portion of the second encoded data is decoded.

28. The non-transitory computer readable medium of claim 27, wherein: The instructions further cause the processor to apply an in-loop filter to update one or more rows of pixels of the first portion of the image based on the second pixel data, and wherein the delayed output is based on an amount of time for applying the in-loop filter.

29. An apparatus for processing video data, comprising: means for determining a number of pixel rows for one or more portions of an image; means for obtaining first encoded data for a first portion of the image; means for decoding the first encoded data to generate first pixel data for the number of pixel rows of the first portion of the image; means for outputting the first pixel data for the first portion of the image to a memory; as well as Means for outputting an indication that the first portion of the image is available.

30. The apparatus of claim 29, further comprising: means for obtaining an indication of a target delay; as well as Means for determining the number of pixel rows for the first portion of the image based on the target latency.