Enhanced video decoder using tile-to-grating reordering
By reordering the processing order from bitstream to raster scanning order in video encoding and decoding technology, the problems of initial delay and performance imbalance in the multi-processing pipeline are solved, and more efficient video data processing is achieved.
Patent Information
- Application Number
- CN202380084628.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-12-15
- Filing Date
- 2023-11-21
- Publication Date
- 2025-07-08
AI Technical Summary
Existing video encoding and decoding techniques have initial delay problems when processing video data, especially when multiple processing pipelines process independent decodeable tiles, resulting in inconsistent processing order and unbalanced performance.
By reordering the processing order from the bitstream sequence into a raster scan sequence across images, the tile data is processed in parallel using multiple processing pipelines, and the raster scan sequence is indicated by the control information to store and process intermediate data, reducing initial delay and improving processing efficiency.
Reduces the initial delay per frame, maintains processing equality and consistency, and improves the overall performance of the video decoder.
Smart Images

Figure CN120283401A_ABST
Abstract
Description
Technical Field
[0001] The present application generally relates to video processing. For example, aspects of the present application relate to improving video codec techniques (e.g., encoding and / or decoding video) for an enhanced video decoder that uses tile-to-raster reordering. Background Art
[0002] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radiotelephones, so-called "smart phones", video teleconferencing devices, video streaming devices, and the like. Such devices allow video data to be processed and output for consumption. Digital video data includes a large amount of data to meet the needs of consumers and video providers. For example, consumers of video data expect video with the highest quality, high fidelity, resolution, frame rate, etc. As a result, the large amount of video data required to meet these needs burdens the communication networks and devices that process and store the video data.
[0003] Digital video devices can implement video codec techniques to compress video data. Video codec is performed according to one or more video codec standards or formats. For example, video codec standards or formats include Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), MPEG-2 Part 2 coding (MPEG stands for Moving Picture Experts Group), Essential Video Coding (EVC), etc., as well as proprietary video encoder-decoder (codec) / formats, such as AOMedia Video 1 (AV1) developed by the Alliance for Open Media. Video codec generally employs prediction methods that utilize the redundancy present in a video image or sequence (e.g., inter-frame prediction, intra-frame prediction, etc.). The goal of video codec techniques is to compress video data into a form that uses a lower bit rate while avoiding or minimizing degradation of video quality. As evolving video services become available, there is a need for encoding and decoding techniques that can improve efficiency or reduce hardware costs. Summary of the Invention
[0004] Systems and techniques for processing video data are described herein. According to at least one example, an apparatus for processing video data. The apparatus includes at least one memory and at least one processor (e.g., configured in a circuit) coupled to the at least one memory. The at least one processor is configured to: obtain first encoded data of a first portion of an image, wherein the image is encoded in a plurality of independently decodable portions; generate first intermediate data of the first portion of the image; store the first intermediate data in the at least one memory in bitstream order; obtain second encoded data of a second portion of the image; generate second intermediate data of the second portion of the image; and store the second intermediate data in the at least one memory in bitstream order, wherein the first intermediate data is stored separately from the second intermediate data; the plurality of processing pipelines are configured to process the first intermediate data and the second intermediate data in a raster scan order across the first portion of the image and the second portion of the image to generate a portion of the image.
[0005] In another example, a method for processing video data is provided. The method includes: obtaining first encoded data of a first portion of an image, wherein the image is encoded in a plurality of independently decodable portions; generating first intermediate data of the first portion of the image; storing the first intermediate data in the at least one memory in bitstream order; obtaining second encoded data of a second portion of the image; generating second intermediate data of the second portion of the image; storing the second intermediate data in the at least one memory in bitstream order, wherein the first intermediate data is stored separately from the second intermediate data; and processing the first intermediate data and the second intermediate data in a raster scan order across the first portion of the image and the second portion of the image to generate a portion of the image.
[0006] As another example, a non-transitory computer-readable medium storing instructions is provided. The instructions, when executed by at least one processor, cause the at least one processor to: obtain first encoded data of a first portion of an image, wherein the image is encoded in a plurality of independently decodable portions; generate first intermediate data of the first portion of the image; store the first intermediate data in the at least one memory in bitstream order; obtain second encoded data of a second portion of the image; generate second intermediate data of the second portion of the image; and store the second intermediate data in the at least one memory in bitstream order, wherein the first intermediate data is stored separately from the second intermediate data; and wherein the instructions, when executed by a plurality of processing pipelines, further cause the plurality of processing pipelines to process the first intermediate data and the second intermediate data in a raster scan order across the first portion of the image and the second portion of the image to generate a portion of the image.
[0007] In another example, a device for processing video data is provided. The device includes: a component for obtaining first encoded data of a first portion of an image, wherein the image is encoded in a plurality of independently decodable portions; a component for generating first intermediate data of the first portion of the image; a component for storing the first intermediate data in at least one memory in bitstream order; a component for obtaining second encoded data of a second portion of the image; a component for generating second intermediate data of the second portion of the image; a component for storing the second intermediate data in at least one memory in bitstream order, wherein the first intermediate data and the second intermediate data are stored separately; and a component for processing the first intermediate data and the second intermediate data in a raster scan order across the first portion of the image and the second portion of the image to generate a portion of the image.
[0008] In some aspects, any one of the devices or apparatuses described above is a mobile device (e.g., a mobile phone or a so-called "smartphone" or other mobile device), a wearable device, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a camera, a personal computer, a laptop computer, a server computer, a vehicle or a computing device or component of a vehicle, a robotic device or system, a television, or other device, and is part of and / or includes the above items. In some aspects, the device or apparatus includes a camera or multiple cameras for capturing one or more pictures, images, or frames. In some aspects, the device or apparatus includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the device or apparatus can include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more accelerometers, any combination thereof, and / or other sensors).
[0009] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood by reference to the appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.
[0010] The foregoing and other features and embodiments will become more apparent by reference to the following specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The following describes illustrative examples of the present application in detail with reference to the following drawings:
[0012] Figure 1 is a block diagram showing examples of an encoding device and a decoding device according to some aspects of the present disclosure;
[0013] Figure 2is a block diagram showing an example architecture of a video codec hardware engine;
[0014] Figure 3 is a diagram showing an example of multiple processing pipelines configured to process video data across multiple tiles according to aspects of the present disclosure;
[0015] Figure 4A is a block diagram showing a part of a video codec hardware engine implementing an enhanced video decoder using tile-to-raster reordering according to aspects of the present disclosure;
[0016] Figure 4B is a diagram showing multiple processing pipelines configured to process video data across tiles according to aspects of the present disclosure;
[0017] Figure 5 is a flowchart showing a process for processing video data according to aspects of the present disclosure;
[0018] Figure 6 is a block diagram showing an example video encoding device according to some aspects of the present disclosure; and
[0019] Figure 7 is a block diagram showing an example video decoding device according to some aspects of the present disclosure. Detailed Description
[0020] Certain aspects and embodiments of the present disclosure are provided below. Some of these aspects and embodiments can be applied independently, and some of them can be applied in combination, which will be clear to those skilled in the art. In the following description, for the purpose of explanation, specific details are set forth to provide a thorough understanding of the embodiments of the present application. However, it is clear that the various embodiments can be practiced without these specific details. The drawings and the description are not intended to be restrictive.
[0021] The following description only provides exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. Instead, the following description of the exemplary embodiments will provide those skilled in the art with an enabling description for implementing the exemplary embodiments. It should be understood that various changes can be made to the functions and arrangements of the elements without departing from the spirit and scope of the present application set forth in the appended claims.
[0022] Video encoding and decoding devices implement video compression techniques to effectively encode and decode video data. Video compression techniques can include applying different prediction modes, including spatial prediction (e.g., intra prediction), temporal prediction (e.g., inter prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques for reducing or removing redundancy inherent in a video sequence. A video encoder is capable of splitting each picture of an original video sequence into rectangular regions called video blocks or coding units (described in more detail below). These video blocks can be encoded using a specific prediction mode.
[0023] Video blocks can be divided into one or more groups of smaller blocks in one or more ways. Blocks can include coding tree blocks, prediction blocks, transform blocks, or other suitable blocks. Unless otherwise specified, a reference to "block" generally refers to such video blocks (e.g., coding tree blocks, coding blocks, prediction blocks, transform blocks, or other appropriate blocks or sub-blocks, as would be understood by a person skilled in the art). Additionally, each of these blocks may also be interchangeably referred to herein as a "unit" (e.g., coding tree unit (CTU), coding unit, prediction unit (PU), transform unit (TU), etc.). In some cases, a unit may indicate a coding logic unit encoded in a bitstream, while a block may indicate a portion of a video frame buffer that the process is targeted at.
[0024] In some cases, an image can be encoded using independently decodable parts such as tiles. Since these tiles are independently decodable, the encoded and decoded data of a tile contains sufficient data for all tiles of the image to be decoded without reference to other tiles. The blocks of a tile can be encoded and stored in the raster scan order of the blocks of the tile. Thus, the bitstream of the encoded data can describe all blocks of a tile in raster scan order before describing the blocks of another tile.
[0025] In some cases, such a bitstream can be decoded in part by multiple processing pipelines configured to operate in parallel. In some cases, a processing pipeline can refer to multiple processing elements used to perform operations, where the processing elements are serially coupled such that data processed by the processing pipeline is passed from one processing element to the next processing element. As an example, a single processing pipeline can be composed of multiple processing elements for decoding blocks of an encoded video. However, since many codecs include blocks that depend on decoded information in other blocks, within a tile, the processing pipeline can be configured to decode one or more rows of blocks within the tile in a wave-like manner. For example, a first processing pipeline can process the blocks in the first row, while a second processing pipeline can process the blocks in the second row immediately below the first row. The first processing pipeline can process the blocks in the following columns of the first row, where the columns are y columns ahead of the blocks in the second row being processed by the second processing pipeline. Thus, when processing of a tile begins, there is a delay after the first processing pipeline starts processing the first row and before the second processing pipeline can start processing the second row. Similarly, there may be additional delays for other processing pipelines among the multiple processing pipelines. For each tile in an image, this initial delay can be incurred. In some cases, it may be useful to reduce the initial delay.
[0026] This disclosure describes systems, devices, electronic devices, methods (also referred to as processes), and computer-readable media (collectively referred to herein as "systems and techniques") for enhancing a decoder, where the enhanced decoder can reorder decoding from a bitstream order (e.g., tile order) to a raster order across independently decodable portions (e.g., tiles). In some cases, an encoded bitstream can be preprocessed such that intermediate data is generated and stored in tile order. For example, a portion of the bitstream of a first tile of an encoded image can be processed before a portion of the bitstream of a second tile of the encoded image is processed. Control data can be generated and stored along with the intermediate data. The control information can indicate a raster scan order for processing the intermediate data across the first and second tiles. In some cases, the processing pipeline can process the intermediate data in a raster scan order for the entire frame across the tiles. In some cases, the processing pipeline can process the intermediate data in a raster scan order based on the stored control data.
[0027] The systems and techniques described herein can provide various advantages. For example, by reordering the processing order from a bitstream order to a raster order across portions of a frame, the initial delay can be reduced from once per portion (e.g., tile) of the frame to once per frame. Additionally, reordering the processing order can help maintain consistent performance independent of the tile size / number. In some cases, load balancing across processing pipelines can be enhanced in situations where processing of the starting row of a portion is always performed by certain processing pipelines, since the processing of the starting row is per frame rather than per portion.
[0028] The systems and techniques described herein can be applied to any of the existing video codecs, such as Versatile Video Coding (VVC), High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), Essential Video Coding (EVC), VP9, AV1 format / codec, and / or other video coding standards, codecs, formats, etc. under development or to be developed.
[0029] Figure 1 FIG. 6 is a block diagram illustrating an example of a system 100 that includes an encoding device 104 and a decoding device 112. The encoding device 104 can be part of a source device, and the decoding device 112 can be part of a receiving device. The source device and / or the receiving device can include an electronic device, such as a mobile or fixed telephone handset (e.g., a smart phone, a cellular phone, etc.), a desktop computer, a laptop or notebook computer, a tablet computer, a set-top box, a television, a camera, a display device, a digital media player, a video game console, a video streaming device, an Internet Protocol (IP) camera, or any other suitable electronic device. In some examples, the source device and the receiving device can include one or more wireless transceivers for wireless communication. The coding and decoding techniques described herein are applicable to video coding and decoding in a variety of multimedia applications, including streaming video transmission (e.g., via the Internet), television broadcast or transmission, encoding of digital video for storage on a data storage medium, decoding of digital video stored on a data storage medium, or other applications. As used herein, the term coding and decoding can refer to encoding and / or decoding. In some examples, the system 100 can support unidirectional or bidirectional video transmission to support applications such as video conferencing, video streaming, video playback, video broadcast, gaming, and / or video telephony.
[0030] Encoding device 104 (or encoder) can be used to encode video data using a video coding standard, format, codec, or protocol to generate an encoded video bitstream. Examples of video coding standards and formats / codecs include ITU-T H.261, ISO / IEC MPEG-1 Visual, ITU-T H.262 or ISO / IEC MPEG-2 Visual, ITU-T H.263, ISO / IEC MPEG-4 Visual, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC), including its scalable video coding (SVC) and multi-view video coding (MVC) extensions, High Efficiency Video Coding (HEVC) or ITU-T H.265, and Versatile Video Coding (VVC) or ITU-T H.266. There are various extensions to HEVC for handling multi-layer video coding, including range and screen content coding extensions, 3D video coding (3D-HEVC) and multi-view extensions (MV-HEVC), and scalable extensions (SHVC). HEVC and its extensions have been developed by the Joint Collaborative Team on Video Coding (JCT-VC) of the ITU-T Video Coding Experts Group (VCEG) and the ISO / IEC Moving Picture Experts Group (MPEG), and the Joint Collaborative Team on 3D Video Coding Extension Development (JCT-3V). VP9, AOMedia Video 1 (AV1) developed by the Alliance for Open Media (AOMedia), and Essential Video Coding (EVC) are other video coding standards for which the techniques described herein can be applied.
[0031] The techniques described herein can be applied to any of the existing video codecs (e.g., High Efficiency Video Coding (HEVC), Advanced Video Coding (AVC), or other suitable existing video codecs), and / or can be efficient coding and decoding tools for any video coding standard being developed and / or future video coding standards (such as, for example, VVC and / or other video coding standards being developed or to be developed). For example, the examples described herein can be performed using video codecs such as VVC, HEVC, AVC, and / or their extensions. However, the techniques and systems described herein can also be applicable to other coding standards, codecs, or formats, such as MPEG, JPEG (or other coding standards for still images), EVC, VP9, AV1, their extensions, or other suitable coding standards that are already available or not yet available or being developed. For example, in some examples, the encoding device 104 and / or the decoding device 112 can operate according to a proprietary video codec / format (such as, an extension of AV1, AVI, and / or a successor to AV1 (e.g., AV2)) or other proprietary format or industry standard. Thus, while the techniques and systems described herein may be described with reference to a particular video coding standard, those skilled in the art will recognize that the description should not be construed as being applicable only to that particular standard.
[0032] Reference Figure 1 , the video source 102 can provide video data to the encoding device 104. The video source 102 can be part of the source device, or can be part of a device other than the source device. The video source 102 can include a video capture device (e.g., a video camera, a camera phone, a video phone, etc.), a video archive containing stored video, a video server or content provider that provides video data, a video feed interface that receives video from a video server or content provider, a computer graphics system for generating computer graphics video data, a combination of such sources, or any other suitable video source.
[0033] Video data from video source 102 may include one or more input pictures or frames. A picture or frame is a still image, which in some cases is part of a video. In some examples, data from video source 102 can be a still image that is not part of a video. In HEVC, VVC, and other video coding and decoding specifications, a video sequence can include a series of pictures. A picture can include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luminance samples, SCb is a two-dimensional array of Cb chrominance samples, and SCr is a two-dimensional array of Cr chrominance samples. Chrominance samples may also be referred to herein as "chroma" samples. A pixel can refer to all three components (luminance and chrominance samples) at a given position in the picture array. In other cases, a picture can be monochrome and can include only an array of luminance samples, in which case the terms pixel and sample can be used interchangeably. Regarding the example techniques described herein that refer to individual samples for illustrative purposes, the same techniques can be applied to pixels (e.g., all three sample components at a given position in the picture array). Regarding the example techniques described herein that refer to pixels (e.g., all three sample components at a given position in the picture array) for illustrative purposes, the same techniques can be applied to individual samples.
[0034] The encoder engine 106 (or encoder) of the encoding device 104 encodes the video data to generate an encoded video bitstream. In some examples, the encoded video bitstream (or "video bitstream" or "bitstream") is a series of one or more coded video sequences. A coded video sequence (CVS) includes a series of access units (AUs), which begins with an AU having a random access point picture in the base layer and having certain properties, until and excluding the next AU having a random access point picture in the base layer and having certain properties. For example, certain properties of the random access point picture that starts a CVS may include a RASL flag (e.g., NoRaslOutputFlag) equal to 1. Otherwise, a random access point picture (having a RASL flag equal to 0) does not start a CVS. An access unit (AU) includes one or more coded pictures and control information corresponding to the coded pictures sharing the same output time. The coded slices of a picture are encapsulated into data units called network abstraction layer (NAL) units at the bitstream level. For example, an HEVC video bitstream can include one or more CVSs including NAL units. Each of the NAL units has a NAL unit header. In one example, the header is one byte for H.264 / AVC (except for multi-layer extensions) and two bytes for HEVC. The syntax elements in the NAL unit header take specified bits and are thus visible to all kinds of system and transport layers, such as transport stream, real-time transport (RTP) protocol, file format, etc.
[0035] There are two types of NAL units in the HEVC standard, including Video Coding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units include coded picture data that forms the coded video bitstream. For example, the bit sequence that forms the coded video bitstream is present in the VCL NAL unit. A VCL NAL unit can include a slice or a slice segment of the coded picture data (described below), and a non-VCL NAL unit includes control information related to one or more coded pictures. In some cases, a NAL unit can be referred to as a packet. A HEVC AU includes a VCL NAL unit containing coded picture data and a non-VCL NAL unit (if any) corresponding to the coded picture data. Among other information, the non-VCL NAL unit can also contain parameter sets having high-level information related to the coded video bitstream. For example, the parameter sets can include a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), and a Picture Parameter Set (PPS). In some cases, each slice or other part of the bitstream can refer to a single active PPS, SPS, and / or VPS to allow the decoding device 112 to access information that can be used to decode the slice or other part of the bitstream.
[0036] A NAL unit can contain a bit sequence that forms a coded representation (such as a coded representation of a picture in a video) of video data (e.g., a coded video bitstream, the CVS of the bitstream, etc.). The encoder engine 106 generates a coded representation of a picture by dividing each picture into multiple slices. Slices are independent of other slices such that the information in a slice is coded and decoded without relying on data from other slices within the same picture. A slice includes one or more slice segments, which include independent slice segments and (if any) one or more dependent slice segments that depend on previous slice segments.
[0037] In HEVC, a slice is then divided into Coding Tree Blocks (CTBs) of luminance samples and chrominance samples. A CTB of luminance samples and one or more CTBs of chrominance samples along with the syntax of the samples are referred to as a Coding Tree Unit (CTU). A CTU can also be referred to as a "tree block" or a "Largest Coding Unit" (LCU). A CTU is the basic processing unit for HEVC coding. A CTU can be divided into multiple Coding Units (CUs) of varying sizes. A CU contains an array of luminance and chrominance samples called a Coding Block (CB).
[0038] The luminance CB and the chrominance CB can be further divided into prediction blocks (PBs). A PB is a block of samples of a luminance component or a chrominance component that uses the same motion parameters for inter - prediction or intra - block copy (IBC) prediction (when it is available or enabled for use). The luminance PB and one or more chrominance PBs together with the associated syntax form a prediction unit (PU). For inter - prediction, a set of motion parameters (e.g., one or more motion vectors, reference indices, etc.) is signaled in the bitstream for each PU, and the set of motion parameters is used for inter - prediction of the luminance PB and one or more chrominance PBs. The motion parameters can also be referred to as motion information. The CB can also be partitioned into one or more transform blocks (TBs). A TB represents a square block of samples of a color component to which a residual transform (e.g., the same 2D transform in some cases) is applied to encode - decode the prediction residual signal. A transform unit (TU) represents a TB of luminance and chrominance samples and the corresponding syntax elements. Transform coding is described in more detail below.
[0039] The size of a CU corresponds to the size of the coding mode and can be square - shaped. For example, the size of a CU can be 8×8 samples, 16×16 samples, 32×32 samples, 64×64 samples, or any other suitable size up to the size of the corresponding CTU. The phrase “N×N” is used herein to refer to the pixel dimensions of a video block in the vertical and horizontal dimensions (e.g., 8 pixels×8 pixels). The pixels in a block can be arranged in rows and columns. In some implementations, a block may not have the same number of pixels in the horizontal direction as in the vertical direction. The syntax data associated with a CU can describe, for example, the partitioning of the CU into one or more PUs. The partitioning pattern can differ between when the CU is encoded in an intra - prediction mode or an inter - prediction mode. A PU can be partitioned into a non - square shape. The syntax data associated with a CU can also describe, for example, the partitioning of the CU into one or more TUs according to the CTU. The shape of a TU can be square or non - square.
[0040] According to the HEVC standard, a transform unit (TU) can be used to perform a transform. The TU can vary for different CUs. The TU can be sized based on the size of the PU within a given CU. The TU can be the same size as the PU or smaller than the PU. In some examples, a quadtree structure called a residual quadtree (RQT) can be used to subdivide the residual samples corresponding to a CU into smaller units. The leaf nodes of the RQT can correspond to TUs. The pixel differences associated with a TU can be transformed to produce transform coefficients. The transform coefficients can then be quantized by the encoder engine 106.
[0041] Once the pictures of the video data are segmented into CUs, the encoder engine 106 uses prediction modes to predict each PU. A prediction unit or prediction block is then subtracted from the original video data to obtain a residual (described below). For each CU, the prediction mode can be signaled within the bitstream using syntax data. The prediction mode can include intra prediction (or intra-picture prediction) or inter prediction (or inter-picture prediction). Intra prediction exploits the correlation between spatially adjacent samples within a picture. For example, using intra prediction, each PU is predicted from adjacent picture data using, for example, DC prediction to find the average value of the PU, planar prediction to fit a planar surface to the PU, directional prediction to extrapolate from adjacent data, or any other suitable type of prediction. Inter prediction uses the temporal correlation between pictures to derive a motion-compensated prediction of a block of picture samples. For example, using inter prediction, each PU is predicted from the picture data in one or more reference pictures (before or after the current picture in output order). A decision as to whether to use inter-picture prediction or intra-picture prediction to encode and decode a picture region can be made, for example, at the CU level.
[0042] The encoder engine 106 and the decoder engine 116 (described in more detail below) can be configured to operate according to VVC. According to VVC, a video codec (such as the encoder engine 106 and / or the decoder engine 116) divides a picture into multiple coding tree units (CTUs) (wherein the CTB of the luminance samples and one or more CTBs of the chrominance samples together with the syntax of the samples are referred to as a CTU). The video codec is capable of dividing a CTU according to a tree structure such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure removes the concept of multiple partitioning types, such as the separation between the CU, PU, and TU of HEVC. The QTBT structure includes two levels, including a first level that is partitioned according to quadtree partitioning and a second level that is partitioned according to binary tree partitioning. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary tree correspond to coding units (CUs).
[0043] In the MTT partitioning structure, a block can be partitioned using quadtree partitioning, binary tree partitioning, and one or more types of ternary tree partitioning. Ternary tree partitioning is a partitioning in which a block is divided into three sub-blocks. In some examples, the ternary tree partitioning divides the block into three sub-blocks without dividing the original block through the center. The partitioning types in the MTT (e.g., quadtree, binary tree, and ternary tree) can be symmetric or asymmetric.
[0044] When operating according to the AV1 codec, the video encoder engine 106 and the video decoder engine 116 can be configured to encode and decode video data in blocks. In AV1, the largest codec block that can be processed is called a superblock. In AV1, a superblock can be 128×128 luma samples or 64×64 luma samples. However, in a successor video codec format (e.g., AV2), a superblock can be defined by a different (e.g., larger) luma sample size. In some examples, a superblock is the top level of a block quadtree. The video encoder engine 106 can further divide a superblock into smaller codec blocks. The video encoder engine 106 can divide a superblock and other codec blocks into smaller blocks using square or non-square partitioning. Non-square blocks can include N / 2×N, N×N / 2, N / 4×N, and N×N / 4 blocks. The video encoder engine 106 and the video decoder engine 116 can perform separate prediction and transform processes on each of the codec blocks.
[0045] AV1 also defines tiles of video data. A tile is a rectangular array of superblocks that can be encoded and decoded independently of other tiles. That is, the video encoder engine 106 and the video decoder engine 116 can encode and decode the codec blocks within a tile separately without using video data from other tiles. However, the video encoder engine 106 and the video decoder engine 116 can perform filtering across tile boundaries. The size of a tile can be uniform or non-uniform. Tile-based encoding and decoding can enable parallel processing and / or multithreading for encoder and decoder implementations.
[0046] In some examples, a video codec can use a single QTBT or MTT structure to represent each of the luma and chroma components, while in other examples, a video codec can use two or more than two QTBT or MTT structures, such as one QTBT or MTT structure for the luma component and another QTBT or MTT structure for the two chroma components (or two QTBTs and / or MTTs for the respective chroma components).
[0047] A video codec can be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, superblock partitioning, or other partitioning structures.
[0048] In some examples, one or more slices of a picture are assigned a slice type. The slice types include intra-coded slices (I slices), inter-coded P slices, and inter-coded B slices. An I slice (intra-coded frame, independently decodable) is a slice of a picture that is coded only by intra prediction and thus is independently decodable because an I slice requires only intra data to predict any prediction unit or prediction block of the slice. A P slice (unidirectional prediction frame) is a slice of a picture that can be coded using intra prediction and using unidirectional inter prediction. Each prediction unit or prediction block within a P slice is coded using intra prediction or inter prediction. When inter prediction is applied, the prediction unit or prediction block is predicted by only one reference picture, so the reference samples are from only one reference region of one frame. A B slice (bi-directional prediction frame) is a slice of a picture that can be coded using intra prediction and using inter prediction (e.g., bi-directional prediction or unidirectional prediction). The prediction units or prediction blocks of a B slice can be predicted bi-directionally from two reference pictures, where each picture contributes one reference region and the sample sets of the two reference regions are weighted (e.g., with equal weights or with different weights) to produce the prediction signal for the bi-directional prediction block. As explained above, the slices of a picture are coded independently. In some cases, a picture can be coded as only one slice.
[0049] As described above, intra-picture prediction of a picture exploits the correlation between spatially adjacent samples within the picture. There are multiple intra prediction modes (also referred to as “intra modes”). In some examples, the intra prediction of a luminance block includes 35 modes, including a planar mode, a DC mode, and 33 angular modes (e.g., diagonal intra prediction modes and angular modes adjacent to the diagonal intra prediction modes). The 35 intra prediction modes are indexed as shown in Table 1 below. In other examples, more intra modes can be defined that include prediction angles that may not be represented by the 33 angular modes. In other examples, the prediction angles associated with the angular modes can be different from those used in HEVC.
[0050]
[0051] Table 1 - Specification of Intra Prediction Modes and Associated Names
[0052] Inter-picture prediction uses the temporal correlation between pictures to derive a motion compensated prediction of the current block of image samples. Using a translational motion model, the position of a block in a previously decoded picture (reference picture) is indicated by a motion vector ( ) where specifies the horizontal displacement of the reference block relative to the current block and specifies the vertical displacement of the reference block relative to the current block. In some cases, the motion vector ( can be integer sample precision (also known as integer precision), in which case the motion vector points to the integer pixel grid (or integer pixel sampling grid) of the reference frame. In some cases, the motion vector ( can have fractional sample precision (also known as fractional pixel precision or non-integer precision) to more accurately capture the movement of the underlying object, not limited to the integer pixel grid of the reference frame. The precision of the motion vector can be represented by the quantization level of the motion vector. For example, the quantization level can be integer precision (e.g., 1 pixel) or fractional pixel precision (e.g., 1 / 4 pixel, 1 / 2 pixel, or other sub-pixel values). When the corresponding motion vector has fractional sample precision, interpolation is applied to the reference picture to derive the prediction signal. For example, samples available at integer positions can be filtered (e.g., using one or more interpolation filters) to estimate the values at fractional positions. The previously decoded reference pictures are indicated to the reference picture list by the reference index (refIdx). The motion vector and the reference index can be referred to as motion parameters. Two types of inter-picture prediction can be performed, including uni-directional prediction and bi-directional prediction.
[0053] With inter-frame prediction using bi-directional prediction (also known as bi-directional inter-frame prediction), two sets of motion parameters ( and are used to generate two motion-compensated predictions (from the same reference picture or possibly from different reference pictures). For example, with bi-directional prediction, each prediction block uses two motion-compensated prediction signals and generates B prediction units. Then the two motion-compensated predictions are combined to obtain the final motion-compensated prediction. For example, the two motion-compensated predictions can be combined by averaging. In another example, weighted prediction can be used, in which case different weights can be applied to each motion-compensated prediction. The reference pictures that can be used in bi-directional prediction are stored in two separate lists, denoted as list 0 and list 1. The motion parameters can be derived at the encoder using a motion estimation process.
[0054] In the case of inter-frame prediction using uni-directional prediction (also known as uni-directional inter-frame prediction), a set of motion parameters ( is used to generate a motion-compensated prediction from the reference picture. For example, in the case of uni-directional prediction, each prediction block uses at most one motion-compensated prediction signal and generates P prediction units.
[0055] A PU can include data related to the prediction process (e.g., motion parameters or other suitable data). For example, when encoding a PU using intra-frame prediction, the PU can include data describing the intra-frame prediction mode used for the PU. As another example, when encoding a PU using inter-frame prediction, the PU can include data defining the motion vector of the PU. The data defining the motion vector of the PU can describe, for example, the horizontal component of the motion vector ( ), the vertical component of the motion vector ( ), the resolution of the motion vector (e.g., integer precision, quarter-pixel precision, or eighth-pixel precision), the reference picture to which the motion vector points, the list of reference pictures of the motion vector (e.g., list 0, list 1, or list C), the reference index, or any combination thereof.
[0056] AV1 includes two general techniques for encoding and decoding coding blocks of video data. The two general techniques are intra prediction (e.g., intra-frame prediction or spatial prediction) and inter prediction (e.g., inter-frame prediction or temporal prediction). In the context of AV1, when using an intra prediction mode to predict a block of the current frame of video data, the video encoder engine 106 and the video decoder engine 116 do not use video data from other frames of the video data. For most intra prediction modes, the video coding device 104 encodes the block of the current frame based on the difference between the sample values in the current block and the predicted value generated from the reference samples in the same frame. The video coding device 104 determines the predicted value generated from the reference samples based on the intra prediction mode.
[0057] After performing prediction using intra and / or inter prediction, the coding device 104 is capable of performing transformation and quantization. For example, after prediction, the encoder engine 106 may calculate a residual value corresponding to the PU. The residual value may include the pixel difference between the current block (PU) of the positively coded and decoded pixels and the predicted block used to predict the current block (e.g., the predicted version of the current block). For example, after generating a predicted block (e.g., issuing an inter prediction or an intra prediction), the encoder engine 106 is capable of generating a residual block by subtracting the predicted block generated by the prediction unit from the current block. The residual block includes a set of pixel differences that quantize the difference between the pixel values of the current block and the pixel values of the predicted block. In some examples, the residual block may be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of pixel values). In such examples, the residual block is a two-dimensional representation of pixel values.
[0058] Use a block transform to transform any residual data that may remain after performing prediction, where the block transform may be based on a discrete cosine transform, a discrete sine transform, an integer transform, a wavelet transform, other suitable transform functions, or any combination thereof. In some cases, one or more block transforms (e.g., of size 32×32, 16×16, 8×8, 4×4, or other suitable size) may be applied to the residual data in each CU. In some embodiments, the TU may be used for the transformation and quantization processes implemented by the encoder engine 106. A given CU having one or more PUs may also include one or more TUs. As described in further detail below, the residual values can be transformed into transform coefficients using a block transform, and then the TUs can be used to quantize and scan the residual values to generate serialized transform coefficients for entropy coding and decoding.
[0059] In some embodiments, after performing intra prediction or inter prediction coding / decoding on a PU of a CU, the encoder engine 106 may compute residual data for a TU of the CU. The PU may include pixel data in the spatial domain (or pixel domain). The TU may include coefficients in the transform domain after applying a block transform. As previously described, the residual data may correspond to the pixel differences between the pixels of the uncoded picture and the predicted values corresponding to the PU. The encoder engine 106 may form a TU including the residual data of the CU, and then may transform the TU to produce transform coefficients of the CU.
[0060] The encoder engine 106 may perform quantization of the transform coefficients. Quantization provides further compression by quantizing the transform coefficients to reduce the amount of data used to represent the coefficients. For example, quantization may reduce the bit depth associated with some or all of the coefficients. In one example, a coefficient having an n-bit value may be rounded down to an m-bit value during quantization, where n is greater than m.
[0061] Once quantization is performed, the coded video bitstream includes the quantized transform coefficients, prediction information (e.g., prediction mode, motion vector, block vector, etc.), segmentation information, and any other suitable data, such as other syntax data. The different elements of the coded video bitstream may then be entropy coded by the encoder engine 106. In some examples, the encoder engine 106 may scan the quantized transform coefficients using a predefined scan order to produce a serialized vector that can be entropy coded. In some examples, the encoder engine 106 may perform adaptive scanning. After scanning the quantized transform coefficients to form a vector (e.g., a one-dimensional vector), the encoder engine 106 may entropy code the vector. For example, the encoder engine 106 may use context-adaptive variable length coding / decoding, context-adaptive binary arithmetic coding, syntax-based context-adaptive binary arithmetic coding, probability interval segmentation entropy coding, or another suitable entropy coding technique.
[0062] The output 110 of the encoding device 104 can send NAL units that constitute encoded video bitstream data to the decoding device 112 of the receiving device via the communication link 120. The input 114 of the decoding device 112 can receive the NAL units. The communication link 120 can include a channel provided by a wireless network, a wired network, or a combination of wired and wireless networks. The wireless network can include any wireless interface or combination of wireless interfaces and can include any suitable wireless network (e.g., the Internet or other wide area network, packet-based network, WiFiTM, radio frequency (RF), ultra-wideband (UWB), WiFi Direct, cellular, Long Term Evolution (LTE), WiMaxTM, etc.). The wired network can include any wired interface (e.g., fiber optic, Ethernet, powerline Ethernet, Ethernet over coaxial cable, digital subscriber line (DSL), etc.). The wired and / or wireless network can be implemented using various equipment such as base stations, routers, access points, bridges, gateways, switches, etc. The encoded video bitstream data can be modulated according to a communication standard (such as a wireless communication protocol) and sent to the receiving device.
[0063] In some examples, the encoding device 104 can store the encoded video bitstream data in the storage device 108. The output 110 can retrieve the encoded video bitstream data from the encoder engine 106 or from the storage device 108. The storage device 108 can include any one of various distributed or locally accessible data storage media. For example, the storage device 108 can include a hard disk drive, a storage optical disc, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. The storage device 108 can also include a decoded picture buffer (DPB) for storing reference pictures used in inter-frame prediction. In another example, the storage device 108 can correspond to a file server or another intermediate storage device that can store the encoded video generated by the source device. In such a case, the receiving device including the decoding device 112 can access the stored video data from the storage device via streaming or downloading. The file server can be any type of server capable of storing encoded video data and sending the encoded video data to the receiving device. Example file servers include web servers (e.g., for websites), FTP servers, network-attached storage (NAS) devices, or local disk drives. The receiving device can access the encoded video data through any standard data connection (including an Internet connection). This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device 108 can be a streaming transmission, a download transmission, or a combination thereof.
[0064] Input 114 of the decoding device 112 receives the encoded video bitstream data and can provide the video bitstream data to the decoder engine 116, or provide it to the storage device 118 for later use by the decoder engine 116. For example, the storage device 118 can include a DPB for storing reference pictures used in inter prediction. The receiving device including the decoding device 112 can receive the encoded video data to be decoded via the storage device 108. The encoded video data can be modulated according to a communication standard (such as a wireless communication protocol) and sent to the receiving device. The communication medium for sending the encoded video data can include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network (such as a local area network, a wide area network, or a global network such as the Internet). The communication medium can include routers, switches, base stations, or any other equipment that can be used to facilitate communication from the source device to the receiving device.
[0065] The decoder engine 116 can decode the encoded video bitstream data by entropy decoding (e.g., using an entropy decoder) and extracting elements of one or more coded video sequences that make up the encoded video data. The decoder engine 116 can then rescale the encoded video bitstream data and perform an inverse transform on the encoded video bitstream data. The residual data is then passed to the prediction stage of the decoder engine 116. The decoder engine 116 then predicts the current pixel block (e.g., PU). In some examples, the prediction is merged into the output of the inverse transform (residual data).
[0066] The video decoding device 112 can output the decoded video to the video destination device 122, where the video destination device 122 can include a display or other output device for displaying the decoded video data to the consumer of the content. In some aspects, the video destination device 122 can be part of the receiving device including the decoding device 112. In some aspects, the video destination device 122 can be part of a separate device other than the receiving device.
[0067] In some embodiments, the video encoding device 104 and / or the video decoding device 112 can be integrated with an audio encoding device and an audio decoding device, respectively. The video encoding device 104 and / or the video decoding device 112 can also include other hardware or software necessary to implement the encoding and decoding techniques described above, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. The video encoding device 104 and the video decoding device 112 can be integrated as part of a combined encoder / decoder (codec) in the respective devices.
[0068] Figure 1The example system shown is an illustrative example that can be used herein. The techniques for processing video data using the techniques described herein can be performed by any digital video encoding and / or decoding device. Although the techniques of the present disclosure are generally performed by a video encoding device or a video decoding device, the techniques can also be performed by a combined video encoder-decoder (commonly referred to as a "codec"). In addition, the techniques of the present disclosure can also be performed by a video preprocessor. The source device and the receiving device are only examples of such codec devices where the source device generates encoded and decoded video data for transmission to the receiving device. In some examples, the source device and the receiving device can operate in a substantially symmetric manner such that each of the devices includes video encoding and decoding components. Thus, the example system can support one-way or two-way video transmission between video devices, such as for video streaming, video playback, video broadcasting, or video telephony.
[0069] Extensions to the HEVC standard include multi-view video coding extensions (referred to as MV-HEVC) and scalable video coding extensions (referred to as SHVC). The MV-HEVC and SHVC extensions share the concept of hierarchical coding, where different layers are included in the encoded video bitstream. Each layer in the encoded video sequence is addressed by a unique layer identifier (ID). The layer ID can be present in the header of the NAL unit to identify the layer with which the NAL unit is associated. In MV-HEVC, different layers typically represent different views of the same scene in the video bitstream. In SHVC, different scalable layers are provided to represent the video bitstream at different spatial resolutions (or picture resolutions) or at different reconstruction fidelities. The scalable layers can include a base layer (with layer ID = 0) and one or more enhancement layers (with layer ID = 1, 2,... n). The base layer can conform to a profile of the first version of HEVC and represents the lowest available layer in the bitstream. Compared to the base layer, the enhancement layers have increased spatial resolution, temporal resolution, or frame rate and / or reconstruction fidelity (or quality). The enhancement layers are hierarchically organized and can (or may not) depend on the lower layers. In some examples, a single standard codec can be used to encode and decode different layers (e.g., using HEVC, SHVC, or other coding standards to encode all layers). In some examples, a multi-standard codec can be used to encode and decode different layers. For example, the base layer can be encoded and decoded using AVC, while one or more enhancement layers can be encoded and decoded using the SHVC and / or MV-HEVC extensions to the HEVC standard.
[0070] Generally, a layer includes a set of VCL NAL units and a corresponding set of non-VCL NAL units. The NAL units are assigned specific layer ID values. A layer can be hierarchical in the sense that a layer can depend on lower layers. A layer set refers to a self - contained set of layers represented within a bitstream, meaning that the layers within a layer set can depend on other layers within the layer set during the decoding process, but not on any other layers used for decoding. Thus, the layers within a layer set can form an independent bitstream capable of representing video content. The set of layers within a layer set can be obtained from another bitstream by the operation of a sub - bitstream extraction process. A layer set can correspond to the set of layers to be decoded when a decoder wants to operate according to certain parameters.
[0071] As previously described, the HEVC bitstream includes a set of NAL units, including VCL NAL units and non - VCL NAL units. The VCL NAL units include the coded picture data that forms the coded video bitstream. For example, the bit sequence that forms the coded video bitstream is present in the VCL NAL units. Among other information, the non - VCL NAL units can also contain parameter sets with high - level information related to the coded video bitstream. For example, the parameter sets can include a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), and a Picture Parameter Set (PPS). Examples of the objectives of the parameter sets include bit - rate efficiency, error - recovery capabilities, and providing a system - layer interface. Each slice references a single active PPS, SPS, and VPS to access the information that the decoding device 112 can use to decode the slice. An identifier (ID) can be coded for each parameter set, including the VPS ID, the SPS ID, and the PPS ID. The SPS includes the SPS ID and the VPS ID. The PPS includes the PPS ID and the SPS ID. Each slice header includes the PPS ID. Using the IDs, the active parameter sets can be identified for a given slice.
[0072] The PPS includes information that applies to all slices in a given picture. Thus, all slices in a picture refer to the same PPS. Slices in different pictures can also refer to the same PPS. The SPS includes information that applies to all pictures in the same coded video sequence (CVS) or bitstream. As previously described, a coded video sequence is a series of access units (AUs) that begins with a random access point picture in the base layer (e.g., an instantaneous decoding reference (IDR) picture or a broken link access (BLA) picture, or other suitable random access point picture) and has certain properties (described above) until and excluding the next AU (or the end of the bitstream) that has a random access point picture in the base layer and has certain properties. The information in the SPS may not change from picture to picture within the coded video sequence. Pictures in a coded video sequence may use the same SPS. The VPS includes information that applies to all layers within a coded video sequence or bitstream. The VPS includes a syntax structure having syntax elements that apply to the entire coded video sequence. In some embodiments, the VPS, SPS, or PPS may be sent in-band with the encoded bitstream. In some embodiments, the VPS, SPS, or PPS may be sent out-of-band in a transmission separate from the NAL unit containing the coded video data.
[0073] The present disclosure may generally refer to "signaling" certain information, such as syntax elements. The term "signaling" may generally refer to the communication of the values of syntax elements and / or other data used to decode the encoded video data. For example, the video encoding device 104 may signal the value of a syntax element in the bitstream. Generally speaking, signaling refers to generating a value in the bitstream. As described above, the video source 102 may transmit the bitstream to the video destination device 122 substantially in real time or non-real time. For example, this may occur when the syntax element is stored in the storage device 108 for later retrieval by the video destination device 122.
[0074] The video bitstream can also include Supplemental Enhancement Information (SEI) messages. For example, an SEI NAL unit can be part of the video bitstream. In some cases, the SEI message can contain information that is not required for the decoding process. For example, the information in the SEI message may not be necessary for the decoder to decode the video pictures of the bitstream, but the decoder can use the information to improve the display or processing of the pictures (e.g., decoded output). The information in the SEI message can be embedded metadata. In an illustrative example, the information in the SEI message can be used by decoder-side entities to improve the visibility of the content. In some cases, certain application standards may mandate the presence of such SEI messages in the bitstream, so that quality improvements can be brought to all devices compliant with the application standard (e.g., carrying frame-packing SEI messages for frame-compatible planar stereoscopic 3DTV video formats, where an SEI message is carried for each frame of the video, handling recovery point SEI messages, using pan-scan rectangle SEI messages in DVB, and many other examples).
[0075] In some cases, the video codec hardware can include multiple subsystems and processing pipelines, some of which are shown in Figure 2 In some examples, the processing pipeline can include various video processing operations such as motion estimation, motion compensation, transformation, and quantization. In some cases, the processing pipeline can perform processing operations in parallel.
[0076] Figure 2 is a block diagram showing an example architecture 200 of a video codec hardware engine. In some cases, the architecture 200 can be implemented by Figure 1 the encoding device 104 and / or the decoding device 112 shown in Figure 1 In some examples, the architecture 200 can be implemented by the encoder engine 106 of the encoding device 104 or by the decoder engine 116 of the decoding device 112, as shown in
[0077] In this example, the architecture 200 of the video codec hardware engine can include a control processor 210, an interface 222, a Video Stream Processor (VSP) 212, processing pipelines 214 - 220 (also referred to as "pipelines"), a Direct Memory Access (DMA) subsystem 230, and one or more buffers 232. In some examples, the architecture 200 can include a memory 240 for storing data such as frames, videos, codec information, outputs, etc. In other examples, the memory 240 can be an external memory on the codec device implementing the video codec hardware engine.
[0078] Interface 222 can transfer data between the video codec hardware engine and / or components of the video codec device through a communication system or system bus on the video codec hardware engine and / or a codec device implementing the video codec hardware engine. For example, interface 222 can connect the control processor 210, VSP 212, processing pipelines 214-220 (e.g., video pixel processor (VPP)), direct memory access (DMA) subsystem 230, and / or one or more buffers 232 to the video codec hardware engine and / or the system bus on the codec device. In some examples, interface 222 can include a network-based communication subsystem, such as a network-on-chip (NoC).
[0079] DMA subsystem 230 can allow other components of the video codec hardware engine (e.g., other components in architecture 200) to access the memory on the video codec hardware engine and / or the video codec device implementing the video codec hardware engine. For example, DMA subsystem 230 can provide access to memory 240 and / or one or more buffers 232. In some examples, DMA subsystem 230 can manage access to common memory units and associated data traffic (e.g., tile 202, blocks 204A-D, bitstream 236, decoded data 238, etc.).
[0080] Memory 240 can include one or more internal or external memory devices, such as, by way of example and not limitation, one or more random access memory (RAM) components, read-only memory (ROM) components, cache memory components, buffer components, and / or other memory devices. Memory 240 can store data used by the video codec hardware engine and / or the video codec device, such as frames, processing parameters, input data, output data, and / or any other type of data.
[0081] Control processor 210 can include one or more processors. Control processor 210 can control and / or program components of the video codec hardware engine (e.g., other components in architecture 200). In some examples, control processor 210 can interface with Figure 2 other drivers, applications, and / or components not shown. For example, in some cases, control processor 210 can interface with an application processor on the video codec device.
[0082] The VSP 212 is capable of performing bitstream parsing (e.g., separating the network abstraction layer, picture layer, and slice layer) and entropy coding / decoding operations. In some examples, the VSP 212 is capable of performing coding / decoding functions, such as variable length coding or decoding. For example, the VSP 212 can implement lossless compression / decompression algorithms to compress or decompress the bitstream 236. In some examples, the VSP 212 is capable of performing arithmetic coding / decoding, such as context-adaptive binary arithmetic coding (CABAC) and / or any other coding / decoding algorithm.
[0083] The processing pipelines 214 - 220 are capable of performing video pixel operations, such as motion estimation, motion compensation, transform and quantization, image deblocking, and / or any other video pixel operations. In some cases, the processing pipelines 214 - 220 can perform video pixel operations based on the output of the VSP 212. In some cases, the output of one VSP 212 can be processed by multiple processing pipelines 214 - 220. The processing pipelines 214 - 220 (and / or each individual processing pipeline) are capable of performing specific video pixel operations in parallel. For example, each processing pipeline can perform multiple operations (and / or process data) simultaneously and / or significantly in parallel. As another example, multiple processing pipelines can perform operations (and / or process data) simultaneously and / or significantly in parallel.
[0084] In Figure 2 it, the processing pipelines 214 - 220 can store video pixel processing data (e.g., video pixel processing output, input, parameters, pixel data, processing synchronization data, etc.) into one or more buffers 232 and retrieve video pixel processing data (e.g., video pixel processing output, input, parameters, pixel data, processing synchronization data, etc.) from one or more buffers 232. In some cases, one or more buffers 232 can include a single buffer. In other cases, one or more buffers 232 can include multiple buffers. In some examples, one or more buffers 232 can include a global input / output line buffer and a pipeline synchronization buffer. In some cases, the pipeline synchronization buffer can temporarily store data and / or results for synchronizing data used in video pixel processing operations performed by the processing pipelines 214 - 220.
[0085] In some examples, the VSP 212 is capable of compressing a bitstream 236 associated with a video or sequence of frames and storing entropy decoded data 238 associated with the bitstream 236 for processing by the processing pipeline 214 - 220. In some cases, the entropy decoded data 238 may be stored in a memory or buffer, and the memory or buffer may be part of buffer 232 or separate from buffer 232. In some cases, the VSP 212 is capable of using the DMA subsystem 230 to retrieve the bitstream 236 and store the entropy decoded data 238 to / from the memory, where the DMA subsystem 230 is capable of managing access to memory components and / or cells, as previously described. In some cases, the VSP 212 may store the decoded data in bitstream-based order. For example, in cases where the bitstream organizes image information based on tiles, the decoded data may be grouped such that the decoded data for a tile is stored together in the order in which the tile is decoded (e.g., in tile order, also known as bitstream order). The processing pipeline 214 - 220 is capable of retrieving the entropy decoded data 238 (e.g., via the DMA subsystem 230) and performing video pixel processing operations on blocks 204A - D of tile 202 associated with the bitstream 236.
[0086] As previously described, the processing pipeline 214 - 220 is capable of performing video pixel processing operations in parallel. The processing pipeline 214 - 220 is capable of retrieving and storing video pixel processing inputs and outputs from / into one or more buffers 232 (e.g., via the DMA subsystem 230). For example, the motion estimation algorithm implemented by the processing pipeline 214 is capable of performing motion estimation on block 204A and storing the motion estimation information calculated for block 204A in one or more buffers 232. The motion compensation algorithm implemented by the processing pipeline 214 is capable of retrieving the motion estimation information from one or more buffers 232 and using the motion estimation information to perform motion compensation for block 204A. While the motion compensation algorithm is performing motion compensation, the motion estimation algorithm is capable of performing motion estimation on the next block.
[0087] The motion compensation algorithm is capable of storing the motion compensation result in one or more buffers 232, which can be accessed and used by the transform, quantization, and deblocking algorithms to perform transform, quantization, and deblocking on block 204A. The motion compensation algorithm is capable of performing motion compensation on the next block while the transform, quantization, and / or deblocking algorithms perform transform, quantization, and / or deblocking on block 204A. The transform, quantization, and deblocking algorithms can similarly perform the corresponding operations on block 204A and the next block in parallel. In some examples, the motion estimation, motion compensation, transform, quantization, and deblocking algorithms can perform the corresponding operations on different blocks in parallel.
[0088] The processing pipelines 214-220 can be implemented by hardware and / or software components. For example, the processing pipelines 214-220 can be implemented by one or more pixel processors. In some examples, each processing pipeline can be implemented by one or more hardware components. In some cases, each processing pipeline can use different hardware units and / or components to implement different stages in the pipeline of the processing pipeline.
[0089] Figure 2 The number of processing pipelines shown in is merely an example provided for explanatory purposes. A person of ordinary skill in the art will recognize that the architecture 200 can include more or fewer processing pipelines than Figure 2 shown. For example, the number of processing pipelines implemented by the architecture 200 can be increased or decreased to include more or fewer processing pipelines. Additionally, although the architecture 200 is shown as including certain components, a person of ordinary skill in the art will understand that the architecture 200 can include more or fewer components than Figure 2 shown. For example, in some cases, the architecture 200 can also include other memory devices (e.g., one or more random access memory (RAM) components, read-only memory (ROM) components, cache memory components, buffer components, database components, and / or other memory devices), processing devices (e.g., one or more CPUs, GPUs, and / or other processing devices), interfaces (e.g., internal buses, etc.), and / or Figure 2 other components not shown in.
[0090] In some cases, the bitstream of the image to be decoded can be organized into one or more independently decodable parts, such as tiles. In some cases, the processing pipelines 214-220 can perform video pixel operations across one or more parts of the tile 202 in bitstream order and in a wave-like manner. For example, a frame can be composed of a grid of tiles, where each tile can have dimensions of h×w number of pixels. The image information of a particular tile can be encapsulated together in the bitstream before the image information of another tile. In some cases, a frame can be decoded one tile at a time. Thus, the image can be decoded tile by tile and in a raster scan mode within a particular tile. The raster scan mode order can be a mode for covering (e.g., processing, scanning, etc.) a rectangular grid area that covers each cell (e.g., pixel, block, etc.) of the grid row by row from left to right and top to bottom.
[0091] In some cases, a wave-like manner can be used to decode blocks within a tile because there may be dependencies across blocks. For example, as a result of processing a first block (such as block 204A), it can be used to process a second block (such as block 204B on the next line) (e.g., for inter-frame prediction, motion vector information, in-loop filtering, etc.). In some cases, the first processing pipeline 214 can perform pixel operations on the first row of tile 202 in a certain direction, such as in the raster scan order of tile 202 for the frame (e.g., from columns 0...7 in tile 202 from left to right), while the second processing pipeline can perform pixel operations on the second row of the block in the same direction, and so on for any additional processing pipelines. In some cases, a processing pipeline (such as the first processing pipeline 214) can process a block on a row above and at a higher column number (e.g., to the right when in raster scan order) than the block being processed by the next processing block (such as the second processing pipeline 216). For example, the first processing pipeline 214 can be processing block 204A, where block 204A is on a row above the block 204B being processed by the second processing pipeline 216. Additionally, the block 204A being processed by the first processing pipeline 214 is two columns earlier (e.g., column 7 vs. column 5) than the block 204B being processed by the second processing pipeline 216. This pattern can be repeated across the processing pipelines 214 - 220. When the first processing pipeline 214 reaches the end of the columns of tile 202, the first processing pipeline 214 can start processing the blocks in the next row (e.g., the row below the row being processed by pipeline N 220) and in the first column (e.g., column 0). In some cases, the first block (such as block 204A) can be several columns (e.g., Figure 2 two columns earlier as in
[0092] as indicated above, the processing pipelines 214 - 220 can perform pixel processing operations on various parts of the frame, such as blocks for a tile, and the processing pipelines 214 - 220 can operate in a wave-like manner. For example, each processing pipeline can start at the first column of the tile (e.g., column 0). The first processing pipeline of the processing pipelines 214 - 220 (such as the processing pipeline 214) can start processing the blocks in the first row of the tile, while the other pipelines can start processing the blocks on their other rows after an initial delay. In some cases, when starting to process another part of the frame (such as for a new tile), this initial delay that causes the processing pipelines 214 - 220 to ramp up may occur.
[0093] Figure 3FIG. 300 is an illustration showing an example of multiple processing pipelines for processing across multiple tiles according to aspects of the present disclosure. As indicated above, processing pipelines 302-308 may perform pixel processing operations on portions of a frame, such as for blocks of a tile, and processing pipelines 302-308 may operate in a wave-like manner. For example, each processing pipeline may start at the first column of a tile (e.g., column 0). Thus, the first processing pipeline among processing pipelines 302-308 (such as processing pipeline 302) may process blocks of a tile (such as tile 0), starting from a block 314 in the first column (such as column 0) of the first row 312 of the tile. Other pipelines may start processing their blocks on other rows after an initial delay. For example, the first processing pipeline 302 may start processing tile 0 by processing block 314. Then, the first processing pipeline 302 may process block 316 and block 318. When the first processing pipeline 302 starts processing block 318, the second processing pipeline 304 may start processing a block in row 320 and column 0. Similarly, when the second processing pipeline 304 starts processing a block in row 320 and column 2 and the first processing pipeline 302 starts processing a block in row 312 and column 4, the third processing pipeline 306 may start processing a block in row 322 and column 0. This pattern where the next processing pipeline starts after the previous processing pipeline has completed processing at least one block may be referred to as an initial delay or ramp-up processing pipeline, and this pattern may continue for each processing pipeline. In some cases, when starting to process another portion of the frame (such as for a new tile), this initial delay that causes the processing pipelines 214-220 to ramp up may occur.
[0094] In some cases, a particular processing pipeline (such as the first processing pipeline 302) may be used to start processing each new portion of the frame (e.g., tile), and this may further increase the initial delay. For example, processing of a new tile may always start from the first processing pipeline 302. Thus, if the first processing pipeline 302 is processing a block in the last row 310 of tile 0, the second processing pipeline 304 cannot start processing a block in the first row 312 of tile 1. Instead, as the first processing pipeline 302 finishes processing the block in the last row 310 of tile 0 and starts processing a block in the first row 312 of tile 1, there may be a delay. Then, after the initial delay, the second processing pipeline 304 may start processing a block in the second row 320 of tile 1. Generally, the initial delay is a relatively small portion of the processing time (e.g., because the ratio of the initial delay to the processing time is relatively low). However, in cases where the width of a portion of the frame (e.g., tile) is relatively narrow (e.g., relatively few block columns) and there are many portions (e.g., tiles) in the frame, as the ratio of the initial delay to the processing time increases, the resulting initial delay may become more problematic when accumulated. To help reduce the impact of the initial delay, the video decoder may be enhanced to use tile-to-raster reordering.
[0095] Figure 4A FIG. 400 is a block diagram showing a portion of a video codec hardware engine implementing enhanced video decoding using tile-to-raster reordering, in accordance with aspects of the present disclosure. It may be discussed in conjunction with Figure 4B to Figure 4A , Figure 4B FIG. 450 is a diagram showing multiple processing pipelines for processing across tiles, in accordance with aspects of the present disclosure. In some cases, a portion of the video codec hardware engine may be implemented by a decoding engine 116 of a decoding device 112 as shown in Figure 1 . In FIG. 400, a video stream processor (VSP) 402 may receive an encoded bitstream 404 for decoding. In some cases, the bitstream 404 may be generally similar to the bitstream 236 of Figure 2 . The VSP 402 may be similar to the VSP 212 in that the VSP 402 may parse and / or partially decode (e.g., process) the encoded bitstream 404 to generate intermediate data. For example, the VSP 212 may parse the encoded bitstream 404 to extract syntax information about blocks, such as the block type, motion vectors, coefficients, parameter sets, etc., of the blocks. This intermediate data may be stored in a bin buffer 406 in tile order for further decoding and / or processing by processing pipelines 410-416. Then, the processing pipelines 410-416 may process the intermediate data (e.g., perform prediction, apply motion vector information, etc.) to generate pixels of the decoded frame. In some cases, the bin buffer 406 may be a unified buffer or memory. In some cases, the bin buffer 406 may be a part of a larger buffer or memory.
[0096] In some video codec formats, the encoded image data of a frame may be decoded tile-by-tile in tile order and in raster scan order within each tile. For example, the encoded data may be partially decoded by the VSP 402 from the bitstream of blocks 454 to 456 of tile 0 452. After block 456, the VSP 402 may then process blocks 458 to 460 of tile 0 452. This pattern of processing the blocks of tile 0 is repeated until the blocks of tile 0 452 are processed. Then, starting from block 472, the VSP may start processing tile 1 470 row-by-row in a similar manner to tile 0 452. In some cases, the processed blocks may be written to the bin buffer 406 in an order that is generally the same as the order in which they are processed (e.g., in tile order).
[0097] To allow the processing pipeline 410-416 to process a frame in raster order of the frame and across tiles, rather than in tile order, the VSP 402 can be enhanced to output control information indicating the order for further processing the intermediate data to the bin buffer 406. The control information can indicate the access order of the intermediate data in the bin buffer 406 such that the intermediate data can be accessed in raster scan order of the frame (e.g., across tiles). For example, the VSP 402 can output control information that indicates that the output blocks of the first row of blocks can be processed from block 454 to 456 and then to block 472 to 474. Similarly, the control information can also indicate that the output blocks of the second row of blocks can be processed from block 458 to block 460 and then to block 476 to 478. The control information can be used by the processing pipeline 410-416 to perform pixel processing operations on the frame in raster scan order. In some cases, the control information can include pointers (e.g., storage addresses) to the storage locations from block to block. Since the size of a block may be different from the size of another block, the VSP 402 can track the storage locations of these blocks as some blocks are written to the bin buffer 406 to generate the control information. For example, the VSP 402 can track the storage addresses associated with the block at the start of a row (such as block 454), the block at the end of a row (such as block 456), and / or a pointer to the next row that should be processed by the processing pipeline 410-416 and write them to the control buffer 408. The control buffer 408 can be part of a larger buffer or memory. In some cases, the control buffer 408 can be a separate part of a larger buffer or memory that includes the bin buffer 406. In other cases, the control buffer 408 can be a buffer or memory separate from the bin buffer 406.
[0098] In some cases, load balancing across processing pipelines 410-416 can be enhanced, such as in the case where the processing of the starting line is performed by some of the processing pipelines, because the processing of the starting line is per frame rather than per section. For example, based on the control information in control buffer 408, processing pipelines 410-416 can then process the intermediate data in bin buffer 406 in raster scan order. For example, the first processing pipeline 410 can start processing the blocks of the first row from block 454 of tile 0 452 to block 456 of tile 0 452. Then, the first processing pipeline 410 can continue processing the blocks of the first row by processing blocks 472 of tile 1 470 until block 474 of tile 1 470. After a small delay (e.g., for two blocks) after the first processing pipeline 410 starts processing the blocks of the first row, the second processing pipeline 412 can start processing the blocks of the second row, which starts at block 458 of tile 0 450 and goes until block 478 of tile 1 470. The other processing pipelines in processing pipelines 410-416 can also process the blocks of other rows. By processing the rows of a portion (e.g., a tile) in raster order for the entire frame, the amount of initial delay caused by ramping up processing pipelines 214-220 can be reduced from once per portion of the frame (e.g., per tile) to once per frame.
[0099] Figure 5 FIG. 4 is a flowchart illustrating an example of a process 500 for processing video data in accordance with aspects of the present disclosure. Process 500 may be performed by a computing device (or apparatus) or a component of a computing device (e.g., an example architecture 200 of a video codec hardware engine, chipset, codec, etc.). The computing device may be a mobile device (e.g., a mobile phone), a network-connected wearable device (such as a watch), an extended reality (XR) device (such as a virtual reality (VR) device or an augmented reality (AR) device), a vehicle or a component or system of a vehicle, or other types of computing devices. In some cases, the computing device may be or may include a codec device, such as an encoding device 104, a decoding device 112, or a combined encoding device (or codec). The operations of process 500 may be implemented as software components executed and run on one or more processors.
[0100] At block 502, the computing device (or its component) may obtain first encoded data of a first portion of an image, where the image is encoded in a plurality of independently decodable portions. In some cases, a bitstream including the first encoded data and second encoded data is encoded tile-by-tile in raster scan order for each tile.
[0101] At block 504, the computing device (or its component) may generate first intermediate data of the first portion of the image. In some cases, the first intermediate data includes syntax information associated with the blocks of the first tile.
[0102] At block 506, the computing device (or its components) may store the first intermediate data in a memory in bitstream order.
[0103] At block 508, the computing device (or its components) may obtain second encoded data for a second portion of the image. In some cases, the first portion of the image includes a first tile of the image, and wherein the second portion of the image includes a second tile of the image. In some cases, the first portion of the image includes a first plurality of blocks of rows, and wherein the second portion of the image includes a second plurality of blocks of rows.
[0104] At block 510, the computing device (or its components) may generate second intermediate data for the second portion of the image. In some cases, the computing device (or its components) may generate control data indicative of a raster scan order across the first portion of the image and the second portion of the image. In some cases, the computing device (or its components) may store the control data. In some cases, a plurality of processing pipelines of the computing device (or its components) (e.g., processing pipelines 214 - 220, processing pipelines 302 - 308, and / or processing pipelines 410 - 416) are configured to process the first intermediate data and the second intermediate data based on the control data. In some cases, the control data includes storage addresses for the first intermediate data and the second intermediate data.
[0105] At block 512, the computing device (or its components) may store the second intermediate data in at least one memory in bitstream order. In some cases, the first intermediate data is stored separately from the second intermediate data. In some cases, the first intermediate data, the second intermediate data, and the control data are stored in separate portions of at least one memory, the control data indicative of a raster scan order across the first portion of the image and the second portion of the image.
[0106] At block 514, the computing device (or its components) may process the first intermediate data and the second intermediate data in a raster scan order across the first portion of the image and the second portion of the image to generate a portion of the image. In some cases, to process the first intermediate data and the second intermediate data in a raster scan order, the computing device (or its components) may use a first processing pipeline to process a first row of the first portion of the image and process a corresponding first row of the second portion of the image before the first processing pipeline begins processing a second row of the first portion of the image. In some cases, the computing device (or its components) may process a portion of a bitstream including the first encoded data of the first portion of the image before processing a portion of a bitstream including the second encoded data of the second portion of the image.
[0107] The processes (or methods) of this disclosure can be used alone or in any combination. In some implementations, the processes (or methods) described herein can be performed by a computing device or apparatus (such as Figure 1 the system 100 shown in Figure 1 ). For example, the process can be performed by Figure 6 the encoding device 104 shown in Figure 1 and Figure 7 the decoding device 112 shown in
[0108] and / or by another video source side device or video transmission device, and / or by
[0109] another client side device (such as a player device, a display, or any other client side device). In some cases, the computing device or apparatus can include one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other component(s) configured to perform the steps of one or more of the processes described herein.
[0108] In some examples, the computing device can include a mobile device, a desktop computer, a server computer, and / or a server system, or other types of computing devices. The components of the computing device (e.g., one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components) can be implemented in circuitry. For example, the components can include electronic circuits or other electronic hardware and / or can be implemented using electronic circuits or other electronic hardware, where the electronic circuits or other electronic hardware can include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuits), and / or can include computer software, firmware, or any combination thereof and / or be implemented using computer software, firmware, or any combination thereof to perform the various operations described herein. In some examples, the computing device or apparatus can include a camera configured to capture video data (e.g., a video sequence) including video frames. In some examples, the camera or other capture device that captures the video data is separate from the computing device, in which case the computing device receives or obtains the captured video data. The computing device can include a network interface configured to transmit the video data. The network interface can be configured to transmit data based on the Internet Protocol (IP) or other types of data. In some examples, the computing device or apparatus can include a display for displaying output video content (such as samples of pictures of a video bitstream).
[0109] A process can be described with respect to a logic flow diagram, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, an operation represents computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the operations. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the process.
[0110] Additionally, the process can be executed under the control of one or more computer systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, by hardware, or a combination thereof. As noted above, the code can be stored, for example, in the form of a computer program that includes multiple instructions executable by one or more processors, on a computer-readable or machine-readable storage medium. The computer-readable or machine-readable storage medium can be non-transitory.
[0111] The encoding and decoding techniques discussed herein can be implemented in an example video encoding and decoding system (e.g., system 100). In some examples, the system includes a source device that provides encoded video data to be decoded later by a destination device. Specifically, the source device provides the video data to the destination device via a computer-readable medium. The source device and the destination device can include any one of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, tablet computers, set-top boxes, telephone handsets such as so-called "smart" phones, so-called "smart" tablets, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, the source device and the destination device can be equipped for wireless communication.
[0112] The destination device may receive encoded video data to be decoded via a computer-readable medium. The computer-readable medium may include any type of medium or device capable of moving the encoded video data from the source device to the destination device. In one example, the computer-readable medium may include a communication medium to enable the source device to send the encoded video data directly to the destination device in real time. The encoded video data may be modulated according to a communication standard, such as a wireless communication protocol, and sent to the destination device. The communication medium may include any wireless or wired communication medium, such as the radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be used to facilitate communication from the source device to the destination device.
[0113] In some examples, the encoded data may be output from an output interface to a storage device. Similarly, the encoded data may be accessed from the storage device via an input interface. The storage device may include any one of a variety of distributed or locally accessible data storage media, such as a hard disk drive, a Blu-ray disc, a DVD, a CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. In another example, the storage device may correspond to another intermediate storage device or a file server that may store the encoded video generated by the source device. The destination device may access the stored video data from the storage device via streaming or downloading. The file server may be any type of server capable of storing the encoded video data and sending the encoded video data to the destination device. Example file servers include web servers (e.g., for websites), FTP servers, network-attached storage (NAS) devices, or local disk drives. The destination device may access the encoded video data through any standard data connection, including an Internet connection. This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both suitable for accessing the encoded video data stored on the file server. The transmission of the encoded video data from the storage device may be a streaming transmission, a download transmission, or a combination thereof.
[0114] The techniques of the present disclosure are not necessarily limited to wireless applications or settings. The techniques can be applied to video encoding and decoding to support any of a variety of multimedia applications, such as over-the-air television broadcasts, cable television transmissions, satellite television transmissions, Internet streaming video transmissions (such as HTTP Dynamic Adaptive Streaming over HTTP (DASH)), digital video encoded onto a data storage medium, decoding of digital video stored on a data storage medium, or other applications. In some examples, the system can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.
[0115] In one example, the source device includes a video source, a video encoder, and an output interface. The destination device can include an input interface, a video decoder, and a display device. The video encoder of the source device can be configured to apply the techniques disclosed herein. In other examples, the source device and the destination device can include other components or arrangements. For example, the source device can receive video data from an external video source (such as an external camera). Similarly, the destination device can interface with an external display device rather than include an integrated display device.
[0116] The above example systems are merely examples. The techniques for parallel processing of video data can be performed by any digital video encoding and / or decoding device. Although the techniques of the present disclosure are generally performed by a video encoding device, the techniques can also be performed by a video encoder / decoder (commonly referred to as a “codec”). Additionally, the techniques of the present disclosure can also be performed by a video preprocessor. The source device and the destination device are merely examples of such encoding and decoding devices where the source device generates encoded and decoded video data for transmission to the destination device. In some examples, the source device and the destination device can operate in a substantially symmetric manner such that each of the devices includes video encoding and decoding components. Thus, the example systems can support one-way or two-way video transmission between video devices, such as for video streaming, video playback, video broadcasting, or video telephony.
[0117] The video source can include a video capture device, such as a camera, a video archive containing previously captured video, and / or a video feed interface for receiving video from a video content provider. As another alternative, the video source can generate computer graphics-based data as the source video, or a combination of live video, archived video, and computer-generated video. In some cases, if the video source is a camera, the source device and the destination device can form a so-called camera phone or video phone. However, as mentioned above, the techniques described in the present disclosure generally apply to video encoding and decoding and can be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video can be encoded by a video encoder. The encoded video information can then be output by the output interface onto a computer-readable medium.
[0118] As described, the computer-readable medium can include a transient medium, such as a wireless broadcast or a wired network transmission, or a storage medium (i.e., a non-transient storage medium), such as a hard disk, a flash drive, a compact disc, a digital video disc, a Blu-ray disc, or other computer-readable media. In some examples, a network server (not shown) can receive encoded video data from a source device and provide the encoded video data to a destination device via a network transmission (e.g.). Similarly, a computing device of a media production facility (e.g., a disc stamping facility) can receive encoded video data from a source device and produce a disc containing the encoded video data. Thus, in various examples, the computer-readable medium can be understood to include one or more computer-readable media in various forms.
[0119] An input interface of the destination device receives information from the computer-readable medium. The information of the computer-readable medium can include syntax information defined by a video encoder, which is also used by the video decoder, and the syntax information includes syntax elements that describe the characteristics and / or processing of blocks and other coding units (e.g., groups of pictures (GOPs)). A display device displays the decoded video data to a user and can include any one of a variety of display devices, such as a cathode ray tube (CRT), a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device. Various embodiments of the present application have been described.
[0120] Specific details of the encoding device 104 and the decoding device 112 are shown respectively in Figure 6 and Figure 7 as shown. Figure 6 is a block diagram showing an example encoding device 104 that can implement one or more of the technologies described in the present disclosure. The encoding device 104 can, for example, generate the syntax elements and / or structures described herein (e.g., the syntax elements and / or structures of green metadata, such as a complexity metric (CM), or other syntax elements and / or structures). The encoding device 104 can perform intra prediction and inter prediction coding of video blocks within video slices, tiles, sub-pictures, etc. As previously described, intra coding depends at least in part on spatial prediction to reduce or remove spatial redundancy within a given video frame or picture. Inter coding depends at least in part on temporal prediction to reduce or remove temporal redundancy within adjacent or surrounding frames of a video sequence. An intra mode (I mode) can refer to any one of several spatial-based compression modes. Inter modes such as unidirectional prediction (P mode) or bidirectional prediction (B mode) can refer to any one of several time-based compression modes.
[0121] The encoding device 104 includes a splitting unit 35, a prediction processing unit 41, a filter unit 63, a picture memory 64, an adder 50, a transform processing unit 52, a quantization unit 54, and an entropy encoding unit 56. The prediction processing unit 41 includes a motion estimation unit 42, a motion compensation unit 44, and an intra-prediction processing unit 46. For video block reconstruction, the encoding device 104 further includes an inverse quantization unit 58, an inverse transform processing unit 60, and an adder 62. The filter unit 63 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although the filter unit 63 is shown as an in-loop filter in Figure 6 , in other configurations, the filter unit 63 can be implemented as a post-loop filter. The post-processing device 57 can perform additional processing on the encoded video data generated by the encoding device 104. The techniques of the present disclosure can be implemented by the encoding device 104 in some cases. However, in other cases, one or more of the techniques of the present disclosure can be implemented by the post-processing device 57.
[0122] As Figure 6 shown, the encoding device 104 receives video data, and the splitting unit 35 splits the data into video blocks. The splitting can also include splitting into slices, slice segments, tiles, or other larger units, and video block splitting according to, for example, the quadtree structure of the LCU and CU. The encoding device 104 generally shows components for encoding video blocks within a video slice to be encoded. A slice can be divided into multiple video blocks (and possibly divided into sets of video blocks called tiles). The prediction processing unit 41 can select one of multiple possible encoding and decoding modes for the current video block based on error results (such as encoding and decoding rate and distortion level, etc.), for example, one of multiple intra-prediction encoding and decoding modes or one of multiple inter-prediction encoding and decoding modes. The prediction processing unit 41 can provide the resulting intra- or inter-encoded and decoded block to the adder 50 to generate residual block data and to the adder 62 to reconstruct the encoded block for use as a reference picture.
[0123] The intra-prediction processing unit 46 within the prediction processing unit 41 can perform intra-prediction encoding and decoding of the current video block relative to one or more adjacent blocks in the same frame or slice as the current block to be encoded to provide spatial compression. The motion estimation unit 42 and the motion compensation unit 44 within the prediction processing unit 41 perform inter-prediction encoding and decoding of the current video block relative to one or more prediction blocks in one or more reference pictures to provide temporal compression.
[0124] The motion estimation unit 42 may be configured to determine an inter-frame prediction mode of a video slice according to a predetermined pattern for a video sequence. The predetermined pattern may designate video slices in the sequence as P slices, B slices, or GPB slices. The motion estimation unit 42 and the motion compensation unit 44 may be highly integrated but are illustrated separately for conceptual purposes. The motion estimation performed by the motion estimation unit 42 is a process of generating a motion vector, where the motion vector estimates the motion of a video block. For example, the motion vector may indicate the displacement of a prediction unit (PU) of a video block within a current video frame or picture relative to a prediction block within a reference picture.
[0125] The prediction block is a block of PUs that is found to closely match the video block to be encoded in terms of pixel difference, where the pixel difference may be determined by sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some examples, the encoding device 104 may calculate values at sub-integer pixel positions of a reference picture stored in the picture memory 64. For example, the encoding device 104 may interpolate values at quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference picture. Thus, the motion estimation unit 42 may perform a motion search relative to full pixel positions and fractional pixel positions and output a motion vector with fractional pixel accuracy.
[0126] The motion estimation unit 42 calculates a motion vector of a PU of a video block in an inter-frame encoded / decoded slice by comparing the position of the PU with the position of a prediction block of a reference picture. The reference picture may be selected from a first reference picture list (list 0) or a second reference picture list (list 1), where each of the lists identifies one or more reference pictures stored in the picture memory 64. The motion estimation unit 42 sends the calculated motion vector to the entropy encoding unit 56 and the motion compensation unit 44.
[0127] The motion compensation performed by the motion compensation unit 44 may involve extracting or generating a prediction block based on the motion vector determined by motion estimation (which may perform interpolation to sub-pixel accuracy). Upon receiving the motion vector of a PU of a current video block, the motion compensation unit 44 may locate the prediction block pointed to by the motion vector in the reference picture list. The encoding device 104 forms a residual video block by subtracting the pixel values of the prediction block from the pixel values of the current video block being encoded / decoded to form a pixel difference value. The pixel difference value forms the residual data of the block and may include both luminance and chrominance difference components. The summer 50 represents one or more components that perform this subtraction operation. The motion compensation unit 44 may also generate syntax elements associated with the video block and the video slice for use by the decoding device 112 in decoding the video block of the video slice.
[0128] As an alternative to the inter prediction performed by the motion estimation unit 42 and the motion compensation unit 44 as described above, the intra prediction processing unit 46 may perform intra prediction on the current block. Specifically, the intra prediction processing unit 46 may determine an intra prediction mode for encoding the current block. In some examples, the intra prediction processing unit 46 may, for example, encode the current block using various intra prediction modes during a separate encoding pass, and the intra prediction processing unit 46 may select an appropriate intra prediction mode from the tested modes for use. For example, the intra prediction processing unit 46 may calculate rate-distortion values using rate-distortion analysis of various tested intra prediction modes and may select an intra prediction mode having the best rate-distortion characteristics among the tested modes. Rate-distortion analysis generally determines the amount of distortion (or error) between an encoded block and the original unencoded block that was encoded to produce the encoded block, as well as the bit rate (i.e., the number of bits) used to produce the encoded block. The intra prediction processing unit 46 may calculate ratios of the distortion and rate for various encoded blocks to determine which intra prediction mode exhibits the best rate-distortion value for the block.
[0129] In any case, after selecting the intra prediction mode for the block, the intra prediction processing unit 46 may provide information indicating the selected intra prediction mode for the block to the entropy encoding unit 56. The entropy encoding unit 56 may encode the information indicating the selected intra prediction mode. The encoding device 104 may include in the transmitted bitstream configuration data a definition of the coding context for various blocks and an indication of the most likely intra prediction mode, the intra prediction mode index table, and the modified intra prediction mode index table for each in the context. The bitstream configuration data may include a plurality of intra prediction mode index tables and a plurality of modified intra prediction mode index tables (also referred to as codeword mapping tables).
[0130] After the prediction processing unit 41 generates a predicted block for the current video block via inter prediction or intra prediction, the encoding device 104 forms a residual video block by subtracting the predicted block from the current video block. The residual video data in the residual block may be included in one or more TUs and applied to the transform processing unit 52. The transform processing unit 52 transforms the residual video data into residual transform coefficients using a transform such as a discrete cosine transform (DCT) or a conceptually similar transform. The transform processing unit 52 may convert the residual video data from the pixel domain to the transform domain, such as the frequency domain.
[0131] The transform processing unit 52 may send the resulting transform coefficients to the quantization unit 54. The quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process may reduce the bit depth associated with some or all of the coefficients. The degree of quantization may be modified by adjusting the quantization parameter. In some examples, the quantization unit 54 may then perform a scan of the matrix including the quantized transform coefficients. Alternatively, the entropy coding unit 56 may perform the scan.
[0132] After quantization, the entropy coding unit 56 performs entropy coding on the quantized transform coefficients. For example, the entropy coding unit 56 may perform context-adaptive variable length coding / decoding (CAVLC), context-adaptive binary arithmetic coding / decoding (CABAC), syntax-based context-adaptive binary arithmetic coding (SBAC), probability interval partitioning entropy (PIPE) coding / decoding, or another entropy coding technique. After entropy coding by the entropy coding unit 56, the encoded bitstream may be sent to the decoding device 112 or archived for later transmission or retrieval by the decoding device 112. The entropy coding unit 56 may also entropy code the motion vectors and other syntax elements of the current video slice being coded / decoded.
[0133] The inverse quantization unit 58 and the inverse transform processing unit 60 respectively apply inverse quantization and inverse transform to reconstruct the residual block in the pixel domain for later use as a reference block of a reference picture. The motion compensation unit 44 may calculate the reference block by combining (e.g., adding, summing) the residual block to a predicted block of one of the reference pictures within the reference picture list. The motion compensation unit 44 may also apply one or more interpolation filters to the reconstructed residual block to calculate sub-integer pixel values for motion estimation. The summer 62 combines the reconstructed residual block to the motion compensation prediction block generated by the motion compensation unit 44 to generate a reference block for storage in the picture memory 64. The reference block may be used as a reference block by the motion estimation unit 42 and the motion compensation unit 44 for inter-frame prediction of blocks in subsequent video frames or pictures.
[0134] In this way, Figure 6 the encoding device 104 represents an example of a video encoder configured to perform any of the techniques described herein. In some cases, some of the techniques of the present disclosure may also be implemented by a post-processing device 57.
[0135] Figure 7 is a block diagram showing an example decoding device 112. The decoding device 112 includes an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, a summer 90, a filter unit 91, and a picture memory 92. The prediction processing unit 81 includes a motion compensation unit 82 and an intra-prediction processing unit 84. In some examples, the decoding device 112 may perform substantially the same as that with respect to from Figure 6The decoding path that is reciprocal to the encoding path described for the encoding device 104.
[0136] During the decoding process, the decoding device 112 receives an encoded video bitstream representing video blocks and associated syntax elements of an encoded video slice transmitted by the encoding device 104. In some embodiments, the decoding device 112 may receive the encoded video bitstream from the encoding device 104. In some embodiments, the decoding device 112 may receive the encoded video bitstream from a network entity 79 such as a server, a media-aware network element (MANE), a video editor / splicer, or other such device configured to implement one or more of the techniques described above. The network entity 79 may or may not include the encoding device 104. Some of the techniques described in this disclosure may be implemented by the network entity 79 before the network entity 79 sends the encoded video bitstream to the decoding device 112. In some video decoding systems, the network entity 79 and the decoding device 112 may be parts of separate devices, while in other cases, the functionality described with respect to the network entity 79 may be performed by the same device that includes the decoding device 112.
[0137] The entropy decoding unit 80 of the decoding device 112 performs entropy decoding on the bitstream to generate quantization coefficients, motion vectors, and other syntax elements. The entropy decoding unit 80 forwards the motion vectors and other syntax elements to the prediction processing unit 81. The decoding device 112 may receive syntax elements at the video slice level and / or the video block level. The entropy decoding unit 80 may process and parse both fixed-length syntax elements and variable-length syntax elements in one or more parameter sets such as VPS, SPS, and PPS.
[0138] When a video slice is coded as an intra-coded (I) slice, the intra prediction processing unit 84 of the prediction processing unit 81 may generate prediction data for video blocks of the current video slice based on data from previously decoded blocks of the current frame or picture and the intra prediction mode signaled. When a video frame is coded as an inter-coded (i.e., B, P, or GPB) slice, the motion compensation unit 82 of the prediction processing unit 81 generates a prediction block for video blocks of the current video slice based on other syntax elements and motion vectors received from the entropy decoding unit 80. The prediction block may be generated from one of the reference pictures within a reference picture list. The decoding device 112 may construct reference frame lists, i.e., list 0 and list 1, using default construction techniques based on the reference pictures stored in the picture memory 92.
[0139] The motion compensation unit 82 determines prediction information for video blocks of the current video slice by parsing motion vectors and other syntax elements, and uses the prediction information to generate a prediction block for the current video block being decoded. For example, the motion compensation unit 82 may use one or more syntax elements in the parameter set to determine a prediction mode (e.g., intra or inter prediction) for encoding and decoding video blocks of the video slice, an inter prediction slice type (e.g., B slice, P slice, or GPB slice), construction information for one or more reference picture lists of the slice, a motion vector for each inter-coded video block of the slice, an inter prediction state for each inter-encoded and decoded video block of the slice, and other information for decoding video blocks in the current video slice.
[0140] The motion compensation unit 82 may also perform interpolation based on an interpolation filter. The motion compensation unit 82 may use an interpolation filter such as that used by the encoding device 104 during the encoding of the video block to calculate the interpolated values of sub-integer pixels of the reference block. In this case, the motion compensation unit 82 may determine the interpolation filter used by the encoding device 104 from the received syntax elements and may use the interpolation filter to generate the prediction block.
[0141] The inverse quantization unit 86 inverse quantizes or de-quantizes the quantized transform coefficients provided in the bitstream and decoded by the entropy decoding unit 80. The inverse quantization process may include using the quantization parameter calculated by the encoding device 104 for each video block in the video slice to determine the degree of quantization that should be applied and the same degree of inverse quantization. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT or other suitable inverse transform), an inverse integer transform, or a conceptually similar inverse transform process to the transform coefficients in order to generate a residual block in the pixel domain.
[0142] After the motion compensation unit 82 generates a prediction block for the current video block based on motion vectors and other syntax elements, the decoding device 112 forms a decoded video block by summing the residual block from the inverse transform processing unit 88 and the corresponding prediction block generated by the motion compensation unit 82. The summer 90 represents one or more components that perform this summation operation. If needed, a loop filter (either in or after the encoding / decoding loop) may also be used to smooth the pixel transitions or otherwise improve the video quality. The filter unit 91 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although the filter unit 91 is Figure 7is shown as a loop-in filter in, but in other configurations, filter unit 91 can be implemented as a post-loop filter. A decoded video block in a given frame or picture is then stored in picture memory 92, where picture memory 92 stores reference pictures for subsequent motion compensation. Picture memory 92 also stores the decoded video for later presentation on a display device (such as, Figure 1 the video destination device 122 shown in).
[0143] In this way, Figure 7 the decoding device 112 represents an example of a video decoder configured to perform any of the techniques described herein.
[0144] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions (or sets thereof) and / or data. A computer-readable medium can include non-transitory media in which data can be stored and that exclude carrier waves and / or transient electronic signals propagated wirelessly or by wire. Examples of non-transitory media can include, but are not limited to, magnetic disks or tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, memory or memory devices. A computer-readable medium can have code and / or machine-executable instructions stored thereon, where the code and / or machine-executable instructions can represent a process, function, subroutine, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program statements. A code segment can be coupled to another code segment or hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. can be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.
[0145] In some embodiments, a computer-readable storage device, medium, and memory can include a cable or wireless signal containing a bitstream, etc. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals themselves.
[0146] Specific details are provided in the above description to provide a thorough understanding of the embodiments and examples provided herein. However, one of ordinary skill in the art will understand that the embodiments may be practiced without these specific details. For clarity, in some instances, the present technology may be presented as including separate functional blocks that include functional blocks containing devices, device components, steps or routines in a method embodied in software or a combination of hardware and software. In addition to those components shown and / or described herein, additional components may be used. For example, circuits, systems, networks, processes, and other components may be shown in block diagram form as components so as not to obscure the embodiments with unnecessary details. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary details so as not to obscure the embodiments.
[0147] The above may describe various embodiments as processes or methods depicted as flowcharts, flow diagrams, data flow diagrams, structure diagrams, or block diagrams. Although a flowchart may describe operations as a sequential process, many operations can be performed in parallel or concurrently. In addition, the order of the operations can be rearranged. A process terminates when its operations are completed, but can have additional steps not included in the figures. A process can correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, its termination can correspond to the function returning to the calling function or the main function.
[0148] The processes and methods according to the above examples can be implemented using computer-executable instructions stored in or otherwise accessible from a computer-readable medium. Such instructions can include, for example, instructions and data that cause or otherwise configure a general-purpose computer, a special-purpose computer, or a processing device to perform a particular function or a group of functions. Portions of the computer resources used can be accessible via a network. The computer-executable instructions can be, for example, binary files, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that can be used to store instructions, the information used, and / or the information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, networked storage devices, etc.
[0149] Devices implementing the disclosed processes and methods can include hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof, and can be in any of a variety of form factors. When program code or code segments for performing the necessary tasks (e.g., a computer program product) are implemented in software, firmware, middleware, or microcode, the program code or code segments can be stored in a computer-readable or machine-readable medium. A processor (or processors) can perform the necessary tasks. Typical examples of form factors include laptop computers, smart phones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functions described herein can also be embodied in a peripheral device or an add-on card. As another example, such functions can also be implemented between different chips or different processes executed in a single device on a circuit board.
[0150] Instructions, a medium for transmitting such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functions described in this disclosure.
[0151] In the foregoing description, aspects of the present application have been described with reference to their specific embodiments, but those skilled in the art will recognize that the present application is not limited thereto. Thus, while illustrative embodiments of the present application have been described in detail herein, it should be understood that the inventive concept can be implemented and employed differently in other ways, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. The various features and aspects of the application described above can be used singly or in combination. Additionally, embodiments can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. For purposes of illustration, the methods are described in a particular order. It should be recognized that in alternative embodiments, the methods can be performed in an order different from that described.
[0152] One of ordinary skill in the art will recognize that, without departing from the scope of this specification, the less than (“<”) and greater than (“>”) symbols or terms used herein can be replaced with the less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively.
[0153] In cases where a component is described as “configured to” perform certain operations, such configuration can be achieved, for example, by designing electronic circuitry or other hardware to perform the operations, by programming a programmable electronic circuit (e.g., a microprocessor or other suitable electronic circuit) to perform the operations, or any combination thereof.
[0154] The phrase "coupled to" refers to any component that is physically connected to another component either directly or indirectly, and / or any component that communicates with another component either directly or indirectly (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0155] The claim language or other language that recites "at least one" in a set and / or "one or more" in a set in this disclosure indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, the claim language that recites "at least one of A and B" means A, B, or A and B. In another example, the claim language that recites "at least one of A, B, and C" means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language "at least one" in a set and / or "one or more" in a set does not limit the set to the items listed in the set. For example, the claim language that recites "at least one of A and B" can mean A, B, or A and B, and can additionally include items not listed in the set of A and B.
[0156] The various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of this application.
[0157] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general purpose computer, a wireless communication device handset, or an integrated circuit device, which have a variety of uses including applications in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be at least partially realized by a computer-readable data storage medium including program code including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging material. The computer-readable medium may include a memory or data storage medium, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, and the like. Additionally or alternatively, the techniques may be at least partially realized by a computer-readable communication medium, where the computer-readable communication medium carries or transmits program code in the form of instructions or data structures and is accessible, readable, and / or executable by a computer, such as a propagated signal or wave.
[0158] The program code may be executed by a processor, where the processor may include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors may be configured to perform any of the techniques described in this disclosure. A general purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Thus, as used herein, the term “processor” may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within a special purpose software module or hardware module configured for encoding and decoding, or incorporated into a combined video encoder-decoder (CODEC).
[0159] Exemplary aspects of the present disclosure include:
[0160] Aspect 1. An apparatus for processing video data, comprising: at least one memory; at least one processor coupled to the at least one memory, the at least one processor being configured to: obtain first encoded data of a first portion of an image, wherein the image is encoded in a plurality of independently decodable portions; generate first intermediate data of the first portion of the image; store the first intermediate data in the at least one memory in bitstream order; obtain second encoded data of a second portion of the image; generate second intermediate data of the second portion of the image; and store the second intermediate data in the at least one memory in bitstream order, wherein the first intermediate data and the second intermediate data are stored separately; and a plurality of processing pipelines coupled to the at least one processor and the at least one memory, the plurality of processing pipelines being configured to process the first intermediate data and the second intermediate data in a raster scan order across the first portion of the image and the second portion of the image to generate a portion of the image.
[0161] Aspect 2. The apparatus according to aspect 1, wherein: the at least one processor is further configured to: generate control data indicating a raster scan order across the first portion of the image and the second portion of the image; and store the control data; and the plurality of processing pipelines are configured to process the first intermediate data and the second intermediate data based on the control data.
[0162] Aspect 3. The apparatus according to any one of aspects 1-2, wherein the bitstream including the first encoded data and the second encoded data is encoded tile-by-tile in a raster scan order for each tile.
[0163] Aspect 4. The apparatus according to aspect 3, wherein the first portion of the image includes a first tile of the image, and wherein the second portion of the image includes a second tile of the image.
[0164] Aspect 5. The apparatus according to aspect 4, wherein the first intermediate data includes syntax information associated with blocks of the first tile.
[0165] Aspect 6. The apparatus according to any one of aspects 1-5, wherein the first intermediate data, the second intermediate data, and the control data are stored in separate portions of the at least one memory, and the control data indicates a raster scan order across the first portion of the image and the second portion of the image.
[0166] Aspect 7. The apparatus according to aspect 6, wherein the control data includes storage addresses for the first intermediate data and the second intermediate data.
[0167] Aspect 8. The apparatus according to any one of aspects 1-7, wherein the first portion of the image includes a first plurality of blocks of rows, and wherein the second portion of the image includes a second plurality of blocks of rows.
[0168] Aspect 9. The apparatus according to any one of aspects 1 - 8, wherein, in order to process the first intermediate data and the second intermediate data in a raster scan order, a first processing pipeline among the plurality of processing pipelines is configured to process the first row of the first portion of the image and the corresponding first row of the second portion of the image before the first processing pipeline starts processing the second row of the first portion of the image.
[0169] Aspect 10. The apparatus according to any one of aspects 1 to 9, wherein at least one processor is further configured to process a portion of the bitstream including the first encoded data of the first portion of the image before processing a portion of the bitstream including the second encoded data of the second portion of the image.
[0170] Aspect 11. The apparatus according to any one of aspects 1 - 10, wherein the apparatus includes a decoder.
[0171] Aspect 12. The apparatus according to any one of aspects 1 - 11, further comprising a display configured to display the generated portion of the image.
[0172] Aspect 13. The apparatus according to any one of aspects 1 - 12, wherein the apparatus is a mobile device.
[0173] Aspect 14. A method for processing video, comprising: obtaining first encoded data of a first portion of an image, wherein the image is encoded in a plurality of independently decodable portions; generating first intermediate data of the first portion of the image; storing the first intermediate data in at least one memory in a bitstream order; obtaining second encoded data of a second portion of the image; generating second intermediate data of the second portion of the image; storing the second intermediate data in at least one memory in a bitstream order, wherein the first intermediate data and the second intermediate data are stored separately; and processing the first intermediate data and the second intermediate data in a raster scan order across the first portion of the image and the second portion of the image to generate a portion of the image.
[0174] Aspect 15. The method according to aspect 14, further comprising: generating control data indicating a raster scan order across the first portion of the image and the second portion of the image; storing the control data; and processing the first intermediate data and the second intermediate data based on the control data.
[0175] Aspect 16. The method according to any one of aspects 14 - 15, wherein the bitstreams including the first encoded data and the second encoded data are encoded tile - by - tile in a raster scan order for each tile.
[0176] Aspect 17. The method according to aspect 16, wherein the first portion of the image includes a first tile of the image, and wherein the second portion of the image includes a second tile of the image.
[0177] Aspect 18. The method according to aspect 17, wherein the first intermediate data includes syntax information associated with the blocks of the first tile.
[0178] Aspect 19. The method according to any one of aspects 14-18, wherein the first intermediate data, the second intermediate data, and the control data are stored in separate portions of at least one memory, and the control data indicates a raster scan order across a first portion of the image and a second portion of the image.
[0179] Aspect 20. The method according to aspect 19, wherein the control data includes storage addresses for the first intermediate data and the second intermediate data.
[0180] Aspect 21. The method according to any one of aspects 14-20, wherein the first portion of the image includes blocks of a first plurality of rows, and wherein the second portion of the image includes blocks of a second plurality of rows.
[0181] Aspect 22. The method according to any one of aspects 14-21, wherein processing the first intermediate data and the second intermediate data in raster scan order includes: processing a first row of a first portion of the image in a first processing pipeline; and processing a corresponding first row of a second portion of the image before the first processing pipeline begins processing a second row of the first portion of the image.
[0182] Aspect 23. The method according to any one of aspects 14-22, further comprising: processing a portion of a bitstream including first encoded data of a first portion of the image before processing a portion of a bitstream including second encoded data of a second portion of the image.
[0183] Aspect 24. A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one processor, cause the at least one processor to: obtain first encoded data of a first portion of an image, wherein the image is encoded in a plurality of independently decodable portions; generate first intermediate data of the first portion of the image; store the first intermediate data in at least one memory in bitstream order; obtain second encoded data of a second portion of the image; generate second intermediate data of the second portion of the image; and store the second intermediate data in at least one memory in bitstream order, wherein the first intermediate data is stored separately from the second intermediate data; and wherein the instructions, when executed by a plurality of processing pipelines, further cause the plurality of processing pipelines to process the first intermediate data and the second intermediate data in a raster scan order across the first portion of the image and the second portion of the image to generate a portion of the image.
[0184] Aspect 25. The non-transitory computer-readable medium according to aspect 24, wherein the bitstreams including the first encoded data and the second encoded data are encoded tile by tile in raster scan order for each tile.
[0185] Aspect 26. The non-transitory computer-readable medium according to aspect 25, wherein the first portion of the image includes a first tile of the image, and wherein the second portion of the image includes a second tile of the image.
[0186] Aspect 27. The non-transitory computer-readable medium according to aspect 26, wherein the first intermediate data includes syntax information associated with the block of the first tile.
[0187] Aspect 28. The non-transitory computer-readable medium according to any one of aspects 24-27, wherein the first intermediate data, the second intermediate data, and the control data are stored in separate portions of at least one memory, and the control data indicates a raster scan order across the first portion of the image and the second portion of the image.
[0188] Aspect 29. The non-transitory computer-readable medium according to aspect 28, wherein the control data includes storage addresses for the first intermediate data and the second intermediate data.
[0189] Aspect 30. The non-transitory computer-readable medium according to any one of aspects 24-29, wherein the first portion of the image includes blocks of a first plurality of rows, and wherein the second portion of the image includes blocks of a second plurality of rows.
[0190] Aspect 31. The non-transitory computer-readable medium according to any one of aspects 24-30, wherein, in order to process the first intermediate data and the second intermediate data in a raster scan order, the instructions cause a first processing pipeline among a plurality of processing pipelines to process a first row of the first portion of the image and a corresponding first row of the second portion of the image before the first processing pipeline begins processing a second row of the first portion of the image.
[0191] Aspect 32. The non-transitory computer-readable medium according to any one of aspects 24 to 31, wherein the instructions further cause at least one processor to process a portion of a bitstream of the first encoded data including the first portion of the image before processing a portion of a bitstream of the second encoded data including the second portion of the image.
[0192] Aspect 33. An apparatus for processing video data, comprising one or more components for performing operations according to aspects 1 to 32 or any combination thereof.
Claims
1. An apparatus for processing video data, comprising: at least one memory; at least one processor, coupled to the at least one memory, the at least one processor being configured to: obtain first encoded data of a first portion of an image, wherein the image is encoded in a plurality of independently decodable portions; generate first intermediate data of the first portion of the image; store the first intermediate data in the at least one memory in bitstream order; obtain second encoded data of a second portion of the image; generate second intermediate data of the second portion of the image; and store the second intermediate data in the at least one memory in the bitstream order, wherein the first intermediate data is stored separately from the second intermediate data; and a plurality of processing pipelines, coupled to the at least one processor and the at least one memory, the plurality of processing pipelines being configured to process the first intermediate data and the second intermediate data in a raster scan order across the first portion of the image and the second portion of the image to generate a portion of the image.
2. The apparatus according to claim 1, wherein: the at least one processor is further configured to: generate control data indicating a raster scan order across the first portion of the image and the second portion of the image; and store the control data; and the plurality of processing pipelines are configured to process the first intermediate data and the second intermediate data based on the control data.
3. The device according to claim 1, wherein The bitstream including the first encoded data and the second encoded data is encoded tile by tile, in the raster scan order of each tile.
4. The device according to claim 3, wherein, The first portion of the image includes a first tile of the image, and wherein the second portion of the image includes a second tile of the image.
5. The device according to claim 4, wherein, The first intermediate data includes syntax information associated with blocks of the first tile.
6. The apparatus according to claim 1, wherein The first intermediate data, the second intermediate data, and the control data are stored in separate portions of the at least one memory, the control data indicating a raster scan order across the first portion of the image and the second portion of the image.
7. The apparatus according to claim 6, wherein, The control data includes storage addresses for the first intermediate data and the second intermediate data.
8. The device according to claim 1, wherein, The first portion of the image includes a first plurality of rows of blocks, and wherein the second portion of the image includes a second plurality of rows of blocks.
9. The device according to claim 1, wherein To process the first intermediate data and the second intermediate data in the raster scan order, a first processing pipeline of the plurality of processing pipelines is configured to process a first row of the first portion of the image and a corresponding first row of the second portion of the image before the first processing pipeline begins processing a second row of the first portion of the image.
10. The device according to claim 1, wherein, The at least one processor is further configured to process a portion of the bitstream including the first encoded data of the first portion of the image before processing a portion of the bitstream including the second encoded data of the second portion of the image.
11. The device according to claim 1, wherein The apparatus includes a decoder.
12. The apparatus according to claim 1, further comprising a display, configured to display the generated portion of the image.
13. The device according to claim 1, wherein, The apparatus is a mobile device.
14. A method for processing video, comprising: obtaining first encoded data of a first portion of an image, wherein the image is encoded in a plurality of independently decodable portions; generating first intermediate data of the first portion of the image; storing the first intermediate data in the at least one memory in bitstream order; obtaining second encoded data of a second portion of the image; generating second intermediate data of the second portion of the image; storing the second intermediate data in the at least one memory in the bitstream order, wherein the first intermediate data is stored separately from the second intermediate data; and processing the first intermediate data and the second intermediate data in a raster scan order across the first portion of the image and the second portion of the image to generate a portion of the image.
15. The method according to claim 14, further comprising: generating control data indicating a raster scan order across the first portion of the image and the second portion of the image; storing the control data; and processing the first intermediate data and the second intermediate data based on the control data.
16. The method according to claim 14, wherein, The bitstream including the first encoded data and the second encoded data is encoded tile by tile in a raster scan order for each tile.
17. The method according to claim 16, wherein, The first portion of the image includes a first tile of the image, and wherein the second portion of the image includes a second tile of the image.
18. The method according to claim 17, wherein, The first intermediate data includes syntax information associated with blocks of the first tile.
19. The method according to claim 14, wherein, The first intermediate data, the second intermediate data, and the control data are stored in separate portions of the at least one memory, and the control data indicates a raster scan order across the first portion of the image and the second portion of the image.
20. The method according to claim 19, wherein, The control data includes storage addresses for the first intermediate data and the second intermediate data.
21. The method according to claim 14, wherein The first portion of the image includes blocks of a first plurality of rows, and wherein the second portion of the image includes blocks of a second plurality of rows.
22. The method according to claim 14, wherein, Processing the first intermediate data and the second intermediate data in the raster scan order includes: processing a first row of the first portion of the image in a first processing pipeline; and processing a corresponding first row of the second portion of the image before the first processing pipeline starts processing a second row of the first portion of the image.
23. The method according to claim 14 further comprises: Processing a portion of the bitstream including the first encoded data of the first portion of the image before processing a portion of the bitstream including the second encoded data of the second portion of the image.