Green metadata signaling

CN117769835BActive Publication Date: 2026-09-04QUALCOMM INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280054157.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-07-29
Filing Date
2022-08-01
Publication Date
2026-09-04
Estimated Expiration
2042-08-01

AI Technical Summary

Technical Problem

结果,为了满足这些需求所需要的大量视频数据为处理和存储视频数据的通信网络和设备带来了负担

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117769835B_ABST
    Figure CN117769835B_ABST
Patent Text Reader

Abstract

Systems, methods, apparatuses, and computer readable media for processing video data are disclosed. For example, an apparatus for processing video data can include at least one memory and at least one processor coupled to the at least one memory, the at least one processor configured to obtain a bitstream, retrieve a granularity type syntax element associated with the bitstream, the granularity type syntax element specifying a granularity type of one or more pictures to which a complexity metric (CM) associated with the bitstream applies, retrieve a periodicity type syntax element associated with the bitstream, the periodicity type syntax element indicating a forthcoming temporal period or set of pictures to which the CM applies, and decode a portion of the bitstream based on the granularity type syntax element and the periodicity type syntax element.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to video processing in general. For example, aspects of this application relate to improving video decoding techniques (e.g., video encoding and / or decoding) relative to green metadata. Background Technology

[0002] Digital video capabilities can be integrated into a wide variety of devices, including digital televisions, digital live broadcast systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, tablets, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radio phones (so-called "smartphones"), video conferencing equipment, video streaming devices, and more. These devices allow video data to be processed and output for consumption. Digital video data comprises vast amounts of data to meet the needs of both consumers and video providers. For example, consumers of video data expect the highest quality video, with high fidelity, high resolution, and high frame rates. As a result, the large amounts of video data required to meet these needs place a burden on the communication networks and equipment that process and store the video data.

[0003] Digital video devices implement video decoding techniques to compress video data. Video decoding is performed according to one or more video decoding standards or formats. Examples of video decoding standards or formats include Universal Video Decoding (VVC), High-Efficiency Video Decoding (HEVC), Advanced Video Decoding (AVC), MPEG-2 Part 2 Decoding (MPEG stands for Moving Picture Experts Group), and proprietary video decoder-decoder (codec) / formats such as AOMedia Video 1 (AV1) developed by the Alliance for Open Media. Video decoding typically utilizes prediction methods (e.g., inter-frame prediction, intra-frame prediction, etc.) that take advantage of redundancy present in video images or sequences. The goal of video decoding technology is to compress video data into a form using a lower bitrate while avoiding or minimizing video quality degradation. As more and more video services become available, there is a need for decoding technologies with better decoding efficiency. Summary of the Invention

[0004] This paper describes systems and techniques for processing video data. Based on at least one example, a method for processing video is provided, comprising: obtaining a bitstream; and retrieving a granularity type syntax element associated with the bitstream, the granularity type syntax element specifying the granularity type to which a complexity metric (CM) associated with the bitstream applies.

[0005] Systems, methods, apparatuses, and computer-readable media for processing video data are disclosed. In one exemplary example, an apparatus for processing video data is provided. The apparatus includes: at least one memory; and at least one processor (e.g., implemented in a circuit) coupled to the at least one memory, the at least one processor being configured to: obtain a bitstream; retrieve a granularity type syntax element associated with the bitstream, the granularity type syntax element specifying the granularity type of one or more pictures to which a complexity metric (CM) associated with the bitstream applies; retrieve a periodicity type syntax element associated with the bitstream, the periodicity type syntax element indicating an upcoming time period or set of pictures to which the CM applies; and decode a portion of the bitstream based on the granularity type syntax element and the periodicity type syntax element.

[0006] For example, a method for processing video data is provided. The method includes: obtaining a bitstream; retrieving a granularity type syntax element associated with the bitstream, the granularity type syntax element specifying the granularity type of one or more images to which a complexity metric (CM) associated with the bitstream applies; retrieving a periodicity type syntax element associated with the bitstream, the periodicity type syntax element indicating an upcoming time period or set of images to which the CM applies; and decoding a portion of the bitstream based on the granularity type syntax element and the periodicity type syntax element.

[0007] In another example, a non-transitory computer-readable medium is provided with instructions that, when executed by one or more processors, cause the one or more processors to: obtain a bitstream; retrieve a granularity type syntax element associated with the bitstream, the granularity type syntax element specifying the granularity type of one or more pictures to which a complexity metric (CM) associated with the bitstream applies; retrieve a periodicity type syntax element associated with the bitstream, the periodicity type syntax element indicating an upcoming time period or set of pictures to which the CM applies; and decode a portion of the bitstream based on the granularity type syntax element and the periodicity type syntax element.

[0008] For example, an apparatus for processing video data is provided. The apparatus includes: components for acquiring a bitstream; components for retrieving a granularity type syntax element associated with the bitstream, the granularity type syntax element specifying the granularity type of one or more images to which a complexity metric (CM) associated with the bitstream applies; components for retrieving a periodicity type syntax element associated with the bitstream, the periodicity type syntax element indicating an upcoming time period or set of images to which the CM applies; and components for decoding a portion of the bitstream based on the granularity type syntax element and the periodicity type syntax element.

[0009] For example, an apparatus for processing video data is provided. The apparatus includes: at least one memory; and at least one processor (e.g., implemented in a circuit) coupled to the at least one memory, the at least one processor being configured to: acquire video data; generate a granularity type syntax element for a bitstream, the granularity type syntax element specifying the granularity type of one or more pictures to which a complexity metric (CM) associated with the bitstream applies; generate a periodicity type syntax element for the bitstream associated with the bitstream, the periodicity type syntax element indicating an upcoming time period or set of pictures to which the CM applies; generate the bitstream associated with the video data, the bitstream including the granularity type syntax element and the periodicity type syntax element; and output the generated bitstream.

[0010] For example, a method for processing video data is provided. The method includes: acquiring video data; generating a granularity type syntax element for a bitstream, the granularity type syntax element specifying the granularity type of one or more images to which a complexity metric (CM) associated with the bitstream applies; generating a periodicity type syntax element for the bitstream associated with the bitstream, the periodicity type syntax element indicating an upcoming time period or set of images to which the CM applies; generating the bitstream associated with the video data, the bitstream including the granularity type syntax element and the periodicity type syntax element; and outputting the generated bitstream.

[0011] In another example, a non-transitory computer-readable medium is provided with instructions that, when executed by one or more processors, cause the one or more processors to: acquire video data; generate a granularity-type syntax element for a bitstream, the granularity-type syntax element specifying the granularity type of one or more pictures to which a complexity metric (CM) associated with the bitstream applies; generate a periodicity-type syntax element for the bitstream associated with the bitstream, the periodicity-type syntax element indicating an upcoming time period or set of pictures to which the CM applies; generate the bitstream associated with the video data, the bitstream including the granularity-type syntax element and the periodicity-type syntax element; and output the generated bitstream.

[0012] For example, an apparatus for processing video data is provided. The apparatus includes: components for acquiring video data; components for generating granularity type syntax elements for a bitstream, the granularity type syntax elements specifying the granularity type of one or more images to which a complexity metric (CM) associated with the bitstream applies; components for generating periodicity type syntax elements associated with the bitstream, the periodicity type syntax elements indicating an upcoming time period or set of images to which the CM applies; components for generating the bitstream associated with the video data, the bitstream including the granularity type syntax elements and the periodicity type syntax elements; and components for outputting the generated bitstream.

[0013] According to at least one other example, an apparatus for processing video data is provided, including at least one memory and one or more processors coupled to the memory (e.g., implemented in a circuit). The one or more processors are configured to: acquire a bitstream; and retrieve a granularity type syntax element associated with the bitstream, the granularity type syntax element specifying the granularity type to which a complexity metric (CM) associated with the bitstream applies.

[0014] According to at least one other example, a non-transitory computer-readable medium is provided that includes instructions which, when executed by one or more processors, cause the one or more processors to: obtain a bitstream; and retrieve a granularity type syntax element associated with the bitstream, the granularity type syntax element specifying the granularity type to which a complexity metric (CM) associated with the bitstream applies.

[0015] According to at least one other example, an apparatus for processing video data is provided, comprising: a component for acquiring a bitstream; and a component for retrieving a granularity type syntax element associated with the bitstream, the granularity type syntax element specifying the granularity type to which a complexity metric (CM) associated with the bitstream applies.

[0016] In some respects, the value of the granularity type syntax element specifies the image for which the CM applies to the bitstream.

[0017] In some respects, the value of the granularity type syntax element specifies the slice of the bitstream to which the CM applies.

[0018] In some respects, the value of the granularity type syntax element specifies the tiles for which the CM applies to the bitstream.

[0019] In some respects, the value of the granularity type syntax element specifies the sub-image for which the CM applies in the bitstream.

[0020] In some respects, the value of the granularity type syntax element specifies the scalable layer to which the CM applies to the bitstream.

[0021] In some respects, the value of a granularity type syntax element specifies the line of the decode tree unit (CTU) to which the CM applies for the bitstream.

[0022] In some aspects, the methods, apparatus, and nontransitory computer-readable media described above may include retrieving a period type syntax element associated with the bitstream, the period type syntax element specifying the type of the upcoming period to which the CM applies.

[0023] In some aspects, the methods, apparatuses, and nontransitory computer-readable media described above may include retrieving a picture-level CM syntax structure associated with the bitstream, the picture-level CM syntax structure specifying a complexity measure for one or more pictures within a period.

[0024] In some aspects, the methods, apparatus, and non-transitory computer-readable media described above may include retrieving a granular-level CM syntax structure associated with the bitstream, the granular-level CM syntax structure specifying a granular-level complexity measure for one or more entities within a period. In some aspects, the one or more entities include at least one of slices, tiles, sub-pictures, and layers.

[0025] In some aspects, the methods, apparatus, and nontransitory computer-readable media described above may include retrieving subpicture syntax elements associated with the bitstream, which, when the period spans multiple pictures, indicate the signaling of a subpicture identifier (ID) by the CM.

[0026] In some aspects, the methods, apparatus, and nontransitory computer-readable media described above may include retrieving a decoded tree block (CTB) quantity syntax element associated with the bitstream, wherein when the granularity type is equal to a slice or tile and the period spans multiple pictures, the CTB quantity syntax element indicates the total number of decoded tree luminance blocks that can be signaled by CM within the period.

[0027] In some aspects, the methods, apparatus, and nontransitory computer-readable media described above may include retrieving an average decoded tree block (CTB) count syntax element associated with the bitstream, the average CTB count syntax element indicating the average number of CTBs or 4×4 blocks per picture per granularity.

[0028] In some respects, when there are blocks available for intra-frame decoding in at least a portion of the bitstream, the block statistics of intra-frame decoding are signaled in association with at least that portion of the bitstream.

[0029] In some respects, when there are blocks of inter-frame decoding available in at least a portion of the bitstream, the block statistics of the inter-frame decoding are signaled in association with at least that portion of the bitstream.

[0030] In some aspects, the methods, apparatus, and non-transitory computer-readable media described above may include retrieving one or more quality recovery metrics associated with one or more granular segments of the bitstream. In some aspects, the one or more granular segments of the bitstream include at least one of slices, tiles, and sub-pictures.

[0031] In some aspects, the methods, apparatus, and nontransitory computer-readable media described above may include: receiving a Supplemental Enhancement Information (SEI) message; and retrieving the granularity type syntax element from the SEI message.

[0032] In some respects, the methods, apparatus, and non-transitory computer-readable media described above may include determining the operating frequency of the apparatus based on the CM associated with the bitstream.

[0033] In some respects, the device includes a decoder.

[0034] According to at least one other example, a method for processing video is provided, comprising: acquiring video data; generating a bitstream associated with the video data; and generating a granularity type syntax element for the bitstream, the granularity type syntax element specifying the granularity type to which a complexity metric (CM) associated with the bitstream applies.

[0035] According to at least one other example, an apparatus for processing video data is provided, including at least one memory and one or more processors coupled to the memory (e.g., implemented in a circuit). The one or more processors are configured to: acquire video data; generate a bitstream associated with the video data; and generate granularity type syntax elements for the bitstream, the granularity type syntax elements specifying the granularity type to which a complexity metric (CM) associated with the bitstream applies.

[0036] According to at least one other example, a non-transitory computer-readable medium is provided that includes instructions which, when executed by one or more processors, cause the one or more processors to: acquire video data; generate a bitstream associated with the video data; and generate a granularity type syntax element for the bitstream, the granularity type syntax element specifying the granularity type to which a complexity metric (CM) associated with the bitstream applies.

[0037] According to at least one other example, an apparatus for processing video data is provided, comprising: components for acquiring video data; components for generating a bitstream associated with the video data; and components for generating granularity type syntax elements for the bitstream, the granularity type syntax elements specifying the granularity type to which a complexity metric (CM) associated with the bitstream applies.

[0038] In some respects, the value of the granularity type syntax element specifies the image for which the CM applies to the bitstream.

[0039] In some respects, the value of the granularity type syntax element specifies the slice of the bitstream to which the CM applies.

[0040] In some respects, the value of the granularity type syntax element specifies the tiles for which the CM applies to the bitstream.

[0041] In some respects, the value of the granularity type syntax element specifies the sub-image for which the CM applies in the bitstream.

[0042] In some respects, the value of the granularity type syntax element specifies the scalable layer to which the CM applies to the bitstream.

[0043] In some respects, the value of a granularity type syntax element specifies the line of the decode tree unit (CTU) to which the CM applies for the bitstream.

[0044] In some respects, the methods, apparatus, and nontransitory computer-readable media described above may include generating periodic type syntax elements for the bitstream, which specify the type of the upcoming period to which the CM is applicable.

[0045] In some respects, the methods, apparatuses and nontransitory computer-readable media described above may include generating a picture-level CM syntax structure for the bitstream, the picture-level CM syntax structure specifying a complexity measure for one or more pictures within a period.

[0046] In some aspects, the methods, apparatus, and non-transitory computer-readable media described above may include generating a granular-level CM syntax structure for the bitstream, the granular-level CM syntax structure specifying a granular-level complexity measure for one or more entities within a period. In some aspects, the one or more entities include at least one of slices, tiles, sub-pictures, and layers.

[0047] In some respects, the methods, apparatus, and non-transitory computer-readable media described above may include generating subpicture syntax elements for the bitstream, which, when the period covers multiple pictures, indicate the signaling of a subpicture identifier (ID) by the CM.

[0048] In some aspects, the methods, apparatus, and nontransitory computer-readable media described above may include a decode tree block (CTB) quantity syntax element generated for the bitstream, which indicates the total number of decode tree blocks that can be signaled in the period according to CM when the granularity type is equal to a slice or tile and the period covers multiple pictures.

[0049] In some respects, the methods, apparatus, and nontransitory computer-readable media described above may include a syntax element for generating an average decode tree block (CTB) number for the bitstream, the average CTB number syntax element indicating the average number of CTBs or 4×4 blocks per image per granularity.

[0050] In some respects, when there are blocks available for intra-frame decoding in at least a portion of the bitstream, the block statistics of intra-frame decoding are signaled in association with at least that portion of the bitstream.

[0051] In some respects, when there are blocks of inter-frame decoding available in at least a portion of the bitstream, the block statistics of the inter-frame decoding are signaled in association with at least that portion of the bitstream.

[0052] In some aspects, the methods, apparatus, and nontransitory computer-readable media described above may include generating one or more quality recovery metrics associated with one or more granular segments of the bitstream.

[0053] In some respects, the one or more granular segments of the bitstream include at least one of slices, tiles, and sub-pictures.

[0054] In some aspects, the methods, apparatus, and nontransitory computer-readable media described above may include: generating a Supplemental Enhancement Information (SEI) message; and including the granularity type syntax element in the SEI message.

[0055] In some respects, the methods, apparatus, and non-transitory computer-readable media described above may include storing the bit stream.

[0056] In some respects, the methods, apparatus, and non-transitory computer-readable media described above may include the transmission of the bit stream.

[0057] In some respects, the device includes an encoder.

[0058] In some aspects, the device is, is part of, and / or includes: mobile devices (e.g., mobile phones or so-called "smartphones" or other mobile devices), wearable devices, extended reality devices (e.g., virtual reality (VR) devices, augmented reality (AR) devices, or mixed reality (MR) devices), cameras, personal computers, laptop computers, server computers, vehicles or computing devices or components of vehicles, robotic devices or systems, televisions, or other devices. In some aspects, the device includes one or more cameras for capturing one or more images. In some aspects, the device includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the device may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more accelerometers, any combination thereof, and / or other sensors).

[0059] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to define the scope of the claimed subject matter. This summary should be understood with reference to the appropriate portions of the entire specification, any or all drawings, and each claim.

[0060] The foregoing, as well as other features and embodiments, will become more apparent from the following description, claims and accompanying drawings. Attached Figure Description

[0061] The following description, with reference to the accompanying drawings, details exemplary examples of this application:

[0062] Figure 1 This is a block diagram illustrating examples of encoding and decoding devices according to some examples of this disclosure;

[0063] Figure 2 This is an illustration showing an example of how granularity-level complexity metrics are used for video images;

[0064] Figure 3 This is a flowchart illustrating techniques for decoding encoded video according to various aspects of this disclosure;

[0065] Figure 4 This is a flowchart illustrating techniques for encoding video according to various aspects of this disclosure;

[0066] Figure 5 This is a block diagram illustrating an exemplary video decoding device according to some examples of this disclosure; and

[0067] Figure 6 This is a block diagram illustrating an exemplary video encoding device according to some examples of this disclosure. Detailed Implementation

[0068] Certain aspects and embodiments of this disclosure are provided below. Some of these aspects and embodiments may be applied independently, and some may be combined, as will be apparent to those skilled in the art. In the following description, specific details are set forth for purposes of explanation to provide a thorough understanding of the various embodiments of this application. However, it will be apparent, however, that the various embodiments may be practiced without these specific details. The accompanying drawings and descriptions are not intended to be limiting.

[0069] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of the example embodiments will provide those skilled in the art with enabling descriptions for implementing the example embodiments. It should be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of this application as set forth in the appended claims.

[0070] Video decoding devices implement video compression techniques to efficiently encode and decode video data. Video compression techniques may include applying different prediction modes, including spatial prediction (e.g., intra-frame prediction or intra-prediction), temporal prediction (e.g., inter-frame prediction or inter-prediction), inter-layer prediction (across different layers of video data), and / or other prediction techniques to reduce or remove inherent redundancy in video sequences. A video encoder can divide each frame of the original video sequence into multiple rectangular regions, which are called video blocks or decoding units (described in more detail below). These video blocks may be encoded using specific prediction modes.

[0071] Video blocks can be divided into one or more smaller blocks in one or more ways. Blocks may include decode tree blocks, prediction blocks, transform blocks, or other suitable blocks. Unless otherwise specified, the reference to “block” generally refers to such a video block (e.g., decode tree block, decode block, prediction block, transform block, or other suitable block or sub-block, as will be understood by one of ordinary skill in the art). Furthermore, each of these blocks may also be interchangeably referred to herein as a “unit” (e.g., decode tree unit (CTU), decode unit, prediction unit (PU), transform unit (TU), etc.). In some cases, a unit may refer to a decoded logic unit encoded in the bitstream, while a block may refer to a portion of the process targeted in the video frame buffer.

[0072] For inter-frame prediction mode, the video encoder searches for blocks similar to the block to be encoded in a frame (or picture) located at another time position (called a reference frame or reference picture). The video encoder can restrict the search to a certain spatial displacement from the block to be encoded. The best match can be located using two-dimensional (2D) motion vectors that include horizontal and vertical displacement components. For intra-frame prediction mode, the video encoder can use spatial prediction techniques to form the prediction block based on data from previously encoded neighboring blocks within the same picture.

[0073] A video encoder can determine prediction error. For example, a prediction can be determined as the difference between the pixel values ​​in the block being encoded and the predicted block. Prediction error can also be called a residual. The video encoder can also apply a transform to the prediction error (e.g., a Discrete Cosine Transform (DCT) or other suitable transform) to generate transform coefficients. After the transform, the video encoder can quantize the transform coefficients. The quantized transform coefficients and motion vectors can be represented using syntax elements and, together with control information, form the decoded representation of the video sequence. In some cases, the video encoder can perform entropy decoding on the syntax elements, further reducing the number of bits required for its representation.

[0074] The video decoder can use the syntax elements and control information discussed above to construct prediction data (e.g., prediction blocks) for decoding the current frame. For example, the video decoder can add the prediction blocks to the compressed prediction error. The video decoder can determine the compressed prediction error by weighting the transform basis function using quantized coefficients. The difference between the reconstructed frame and the original frame is called the reconstruction error.

[0075] The Energy-Saving Media Consumption (Green Metadata or Green MPEG) standard, with the international standard number ISO / IEC 23001-11 (which is incorporated herein by reference in its entirety and for all purposes), specifies green metadata to facilitate reduced energy consumption during media consumption. Green metadata for energy-efficient decoding specifies two sets of information: Complexity Measurement (CM) metadata and Reduced Decoding Operation Request (DOR-Req) metadata. For example, a decoder can use CM metadata to change the processor's operating frequency and thus reduce decoder power consumption. In an exemplary example, in a peer-to-peer video conferencing application, a remote encoder (which generates the encoded bitstream) can receive DOR-Req metadata and use it to modify the decoding complexity of the bitstream, thus reducing the local decoder's power consumption. By signaling the decoding complexity of the bitstream, the local decoder may be able to estimate the amount of power required to decode the bitstream and potentially adjust the bitstream based on, for example, the remaining battery power before requesting less (or more) complex bitstreams. In some cases, supplemental enhancement information (SEI) messages can be used to signal green metadata in a bitstream (e.g., AVC, HEVC, VVC, AV1, or other streams).

[0076] Green metadata for AVC and HEVC is specified in ISO / IEC 23001-11, version 2. The draft of Green MPEG, version 3 (MPEG MDS20584_WG03_N00330), proposes new green metadata supporting VVC decoder-decoder (codec) and specifies CMs at various granularities. The syntax structure can be improved to support more granular types across various cycle types. Furthermore, using a single type to signal CMs for slice, tile, subpicture, or layer granularity can be problematic. In some cases, such as for VVC, the encoder may divide the picture (e.g., frame) of the video being encoded into one or more parts, such as slices, tiles, subpictures, layers, etc. For example, a picture may be divided into one or more tiles, and each tile may be divided into one or more blocks. A slice may include multiple tiles or multiple blocks within a tile. A subpicture may be one or more complete rectangular slices, each rectangular slice covering a rectangular area of ​​the picture. Subpictures may be decoded independently of or not independently of other subpictures of the same picture.

[0077] Currently, decoders such as those used for AVC / HEVC can use the number of slices and tiles to identify whether the computational granularity (CM) is calculated for slices or tiles. For example, identifying the CM granularity is complex when the number of slices equals the number of tiles. Furthermore, AVC and HEVC do not support sub-image granularity. Defining different types of slice and tile granularity, as well as defining sub-image and layer granularity, would be beneficial.

[0078] VVC allows subpicks to be replaced with different subpicks within a decoded layer video sequence (CLVS). The decoded video sequence (CVS) can be a layer-by-layer set of CLVSs. In some cases, signaling is necessary to map a CM to a specific subpick using a subpick identifier (ID).

[0079] VVC also allows for resolution variations within CLVS. In some cases, parsing each slice header to derive the total number of decoded blocks within a cycle to decode normalized coding statistics for each slice or tile can be complex. Slice headers can be included along with the slices and convey information about the associated slices. Information about all slices applied to the picture can be conveyed in the picture header. Syntax elements indicating the total number of CTBs will help simplify derivation.

[0080] Currently, when performing intra-frame decoding on all blocks, the CM provides block statistics for intra-frame decoding. It's possible that P and B slices may have more intra-frame decoded blocks than inter-frame decoded blocks, or that P or B images may have more intra-frame decoded blocks than inter-frame decoded blocks. Therefore, the CM may not accurately represent complexity. Intra-frame decoded blocks refer to blocks predicted based on another block within the same image, while inter-frame decoded blocks refer to blocks predicted based on another block from a different image. An I slice is a slice that includes intra-frame decoded blocks but does not include inter-frame decoded blocks. P and B slices may include both intra-frame decoded and inter-frame decoded blocks.

[0081] Furthermore, instead of applying quality metrics to the entire image, in VVC, quality metrics can be applied separately to each sub-image.

[0082] This disclosure describes systems, apparatus, methods, and computer-readable media (collectively, the “Systems and Technologies”) for providing enhanced green metadata signaling, such as signaling for improving complexity metrics (CM). For example, in some cases, granularity type indicator identifiers (e.g., granularity type syntax elements, such as granularity_type) are provided to support various granularities, such as slices, tiles, subpictures, scalable layers, and / or other granularities. In some examples, the semantics of periodic type syntax elements (e.g., period_type) are modified.

[0083] In some cases, these systems and techniques provide improved complexity metric (CM) signaling. For example, the systems and techniques described herein provide video codecs (e.g., video encoders, video encoders, or combined video encoder-decoders) with the ability to specify CM values ​​for portions of a video (such as slices, tiles, sub-pictures, and / or layers) for multiple pictures. For example, as previously described, sub-pictures can be defined for the encoded video. A sub-picture includes a portion of the image, such as the upper right corner of the image. A CM can be specified for a sub-picture, where the CM differs from at least one other CM specified for a slice (or other portion) of the image. CM values ​​associated with a sub-picture can be defined for multiple pictures at once (e.g., for 30 pictures from a first picture). CMs can be provided as part of the metadata included with the encoded video. Allowing a single CM value to be specified for a sub-picture across multiple frames helps reduce the size of the encoded video's metadata while allowing increased flexibility and granularity in defining CMs for portions of the image.

[0084] In some aspects, CM signaling changes associated with resolution changes are provided. In some cases, CM signaling changes regarding block statistics of intra-frame decoding are provided. In some aspects, sub-picture quality metrics are provided.

[0085] The systems and techniques described herein can be applied to any of the existing video codecs under development or yet to be developed, such as Universal Video Decoder (VVC), High Efficiency Video Decoder (HEVC), Advanced Video Decoder (AVC), VP9, ​​AV1 format / codec and / or other video decoding standards, codecs, formats, etc.

[0086] Figure 1This is a block diagram illustrating an example of a system 100 including an encoding device 104 and a decoding device 112. The encoding device 104 may be part of a source device, and the decoding device 112 may be part of a receiving device. The source device and / or receiving device may include electronic devices such as mobile or landline handsets (e.g., smartphones, cellular phones, etc.), desktop computers, laptops or notebook computers, tablet computers, set-top boxes, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, Internet Protocol (IP) cameras, or any other suitable electronic devices. In some examples, the source device and receiving device may include one or more wireless transceivers for wireless communication. The decoding techniques described herein are applicable to video decoding in a variety of multimedia applications, including streaming video (e.g., via the Internet), television broadcasting or transmission, encoding digital video for storage on a data storage medium, decoding digital video stored on a data storage medium, or other applications. As used herein, the term decoding may refer to encoding and / or decoding. In some examples, system 100 may support one-way or two-way video transmission to support applications such as video conferencing, video streaming, video playback, video broadcasting, gaming, and / or video telephony.

[0087] Encoding device 104 (or encoder) can be used to encode video data using video decoding standards, formats, codecs, or protocols to generate an encoded video bitstream. Examples of video decoding standards and formats, and codecs include ITU-T H.261, ISO / IEC MPEG-1 video, ITU-T H.262, or ISO / IEC MPEG-2 video, ITU-T H.263, ISO / IEC MPEG-4 video, ITU-T H.264 (also known as ISO / IEC MPEG-4 AVC) (including its Scalable Video Decoding (SVC) and Multi-View Video Decoding (MVC) extensions), High-Efficiency Video Decoding (HEVC) or ITU-T H.265, and Universal Video Decoding (VVC) or ITU-T H.266. Various extensions to HEVC are available for multi-layer video decoding, including range and screen content decoding extensions, 3D video decoding (3D-HEVC), multi-view extensions (MV-HEVC), and scalable extensions (SHVC). HEVC and its extensions have been developed by the Joint Collaborative Team for Video Decoding (JCT-VC) and the Joint Collaborative Team for the Development of 3D Video Decoding Extensions of the ITU-T Video Decoding Experts Group (VCEG) and the ISO / IEC Animation Experts Group (MPEG). VP9, ​​AOMedia Video 1 (AV1) developed by the Open Media Consortium, the Open Media Consortium (AOMedia), and Elementary Video Decoding (EVC) are other video coding standards to which the technologies described herein can be applied.

[0088] The techniques described herein are applicable to any of existing video codecs (e.g., High Efficiency Video Decoding (HEVC), Advanced Video Decoding (AVC), or other suitable existing video codecs), and / or can be efficient decoding tools for any video decoding standard under development and / or future video decoding standards (e.g., VVC and / or other video decoding standards under development or to be developed). For example, many of the examples described herein can be performed using video codecs such as VVC, HEVC, AVC, and / or their extensions. However, the techniques and systems described herein are also applicable to other decoding standards, codecs, or formats such as MPEG, JPEG (or other decoding standards for still images), VP9, ​​AV1, their extensions, or other suitable decoding standards that are already available or not yet available or under development. For example, in some examples, encoding device 104 and / or decoding device 112 may operate according to proprietary video codecs / formats such as AV1, extensions of AVI, and / or successors to AV1 (e.g., AV2), or other proprietary formats or industry standards. Therefore, although the techniques and systems described herein may be described with reference to specific video decoding standards, it will be understood by those skilled in the art that the description should not be construed as applicable only to that particular standard.

[0089] refer to Figure 1 Video source 102 can provide video data to encoding device 104. Video source 102 can be part of a source device, or it can be part of a device other than a source device. Video source 102 can include video capture devices (e.g., cameras, camera phones, video phones, etc.), video archives containing stored video, video servers or content providers that provide video data, video feed interfaces that receive video from video servers or content providers, computer graphics systems for generating computer graphics video data, combinations of such sources, or any other suitable video source.

[0090] Video data from video source 102 may include one or more input pictures or frames. A picture or frame is a still image, which in some cases is part of the video. In some examples, the data from video source 102 may be a still image that is not part of the video. In HEVC, VVC, and other video decoding specifications, a video sequence may include a series of pictures. A picture may include three sample arrays, denoted as SL, SCb, and SCr. SL is a two-dimensional array of luminance samples, SCb is a two-dimensional array of Cb chrominance samples, and SCr is a two-dimensional array of Cr chrominance samples. Chroma samples may also be referred to herein as “chroma” samples. A pixel may refer to all three components (luminance and chrominance samples) at a given location in the array of pictures. In other cases, a picture may be monochrome and may include only an array of luminance samples, in which case the terms pixel and sample are used interchangeably. The exemplary techniques described herein with reference to the various samples for illustrative purposes can be applied to pixels (e.g., all three sample components at a given location in the array of pictures). The example techniques described herein, which refer to pixels (e.g., all three sample components at a given location in an array of images) for illustrative purposes, can be applied to individual samples.

[0091] The encoder engine 106 (or encoder) of the encoding device 104 encodes the video data to generate an encoded video bitstream. In some examples, the encoded video bitstream (or “video bitstream” or “bitstream”) is a series of one or more decoded video sequences. The decoded video sequence (CVS) includes a series of access units (AUs) that begin with an AU that has a random access point picture in the base layer and has certain attributes, and continue until the next AU that has a random access point picture in the base layer and has certain attributes, but does not include that next AU. For example, certain attributes of the random access point picture that starts the CVS may include a RASL flag equal to 1 (e.g., NoRaslOutputFlag). Otherwise, a random access point picture (with a RASL flag equal to 0) does not start the CVS. An access unit (AU) includes one or more decoded pictures and control information corresponding to decoded pictures that share the same output time. Decoded slices of pictures are encapsulated as data units at the bitstream level, which are called Network Abstraction Layer (NAL) units. For example, an HEVC video bitstream may include one or more CVSs, which include NAL units. Each NAL unit has a NAL unit header. In one example, this header is one byte for H.264 / AVC (except for multi-layer extensions) and two bytes for HEVC. The syntax elements in the NAL unit header use specified bits and are therefore visible to all kinds of systems and transport layers, such as transport streams, Real-Time Transport (RTP) protocols, file formats, and others.

[0092] The HEVC standard contains two types of NAL units: Video Decoding Layer (VCL) NAL units and non-VCL NAL units. VCL NAL units contain decoded picture data that forms the decoded video bitstream. For example, the bit sequence forming the decoded video bitstream is present in a VCL NAL unit. A VCL NAL unit may include a slice or fragment of the decoded picture data (described below), and non-VCL NAL units contain control information relating to one or more decoded pictures. In some cases, NAL units may be referred to as packets. A HEVC AU includes VCL NAL units containing the decoded picture data and non-VCL NAL units (if any) corresponding to the decoded picture data. Among other information, non-VCL NAL units may also contain a set of parameters with high-level information relating to the encoded video bitstream. For example, the parameter set may include a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), and a Picture Parameter Set (PPS). In some cases, each slice or other portion of the bitstream may reference a single valid PPS, SPS, and / or VPS to allow the decoding device 112 to access information that can be used to decode the slice or other portion of the bitstream.

[0093] NAL units can include bit sequences that form a decoded representation of video data (e.g., an encoded video bitstream, a CVS of a bitstream, etc.), such as a decoded representation of images in a video. Encoder engine 106 generates a decoded representation of an image by dividing each image into multiple slices. Each slice is independent of other slices, allowing information in that slice to be decoded without depending on data from other slices within the same image. A slice includes one or more segments, comprising independent segments and (if present) one or more dependent segments that depend on previous segments.

[0094] In HEVC, the slice is then divided into decoder tree blocks (CTBs) for luma and chroma samples. One or more CTBs for luma samples and one CTB for chroma samples, along with the syntax used for the samples, are called decoder tree units (CTUs). CTUs are also referred to as "tree blocks" or "maximum decoder units" (LCUs). A CTU is the basic processing unit used for HEVC encoding. A CTU can be subdivided into multiple decoder units (CUs) of different sizes. A CU contains an array of luma and chroma samples called a decoder block (CB).

[0095] Luminance and chrominance CBs can be further subdivided into prediction blocks (PBs). A PB is a sample block of either the luminance or chrominance component that uses the same motion parameters for inter-frame prediction or intra-block copy (IBC) prediction (when available or enabled). A luminance PB and one or more chrominance PBs, along with their associated syntax, form a prediction unit (PU). For inter-frame prediction, the set of motion parameters (e.g., one or more motion vectors, reference indices, etc.) is signaled in the bitstream for each PU, along with inter-frame prediction for the luminance PB and one or more chrominance PBs. Motion parameters can also be referred to as motion information. CBs can also be subdivided into one or more transform blocks (TBs). A TB represents a square block of samples of the chrominance component, to which a residual transform (e.g., in some cases, the same two-dimensional transform) is applied to decode the prediction residual signal. A transform unit (TU) represents a TB of luminance and chrominance samples along with the corresponding syntax elements. Transform decoding is described in more detail below.

[0096] The size of a CU corresponds to the size of the decoding mode and can be square. For example, the size of a CU can be 8×8 samples, 16×16 samples, 32×32 samples, 64×64 samples, or any other suitable size up to the corresponding CTU size. The phrase "N×N" is used herein to refer to the pixel size of a video block in both the vertical and horizontal dimensions (e.g., 8 pixels × 8 pixels). Pixels in a block can be arranged in rows and columns. In some specific implementations, a block may not have the same number of pixels in the horizontal direction as it does in the vertical direction. The syntax data associated with a CU can describe, for example, the segmentation of the CU into one or more PUs. The segmentation mode can differ between intra-frame predictive mode coding and inter-frame predictive mode coding of the CU. PUs can be segmented into non-square shapes. The syntax data associated with a CU can also describe, for example, the segmentation of the CU into one or more TUs according to the CTU. TUs can be square or non-square shapes.

[0097] According to the HEVC standard, transform units (TUs) can be used to perform transforms. TUs can vary for different core cells (CUs). The size of the TU can be set based on the size of the functional unit (PU) within a given CU. A TU can have the same size as or smaller than the PU. In some examples, a quadtree structure called a residual quadtree (RQT) can be used to subdivide the residual samples corresponding to the CU into smaller units. The leaf nodes of the RQT can correspond to TUs. The pixel differences associated with the TU can be transformed to produce transform coefficients. The transform coefficients can then be quantized by the encoder engine 106.

[0098] Once the video data is segmented into Units (CUs), the encoder engine 106 uses a prediction mode to predict each Processing Unit (PU). The prediction unit or block is then subtracted from the original video data to obtain the residual (described below). For each CU, a prediction mode can be signaled within the bitstream using syntax data. Prediction modes can include intra-frame prediction (or intra-picture prediction) or inter-frame prediction (or inter-picture prediction). Intra-frame prediction utilizes the correlation between spatially adjacent samples within a picture. For example, using intra-frame prediction, each PU is predicted from neighboring image data in the same picture using, for example, DC prediction to find the average value for the PU, planar prediction to adapt a planar surface to the PU, orientation prediction to infer from neighboring data, or any other suitable prediction type. Inter-frame prediction uses temporal correlations between pictures to derive motion-compensated predictions for blocks of image samples. For example, using inter-frame prediction, each PU is predicted from image data in one or more reference pictures (in the output order before or after the current picture) using motion-compensated prediction. For example, a decision can be made at the CU level whether to use inter-picture prediction or intra-picture prediction to decode a picture region.

[0099] Encoder engine 106 and decoder engine 116 (described in more detail below) can be configured to operate according to VVC. According to VVC, the video decoder (such as encoder engine 106 and / or decoder engine 116) segments the image into multiple decoder tree units (CTUs) (where one or more CTBs of luminance samples and chrominance samples, together with the syntax used for the samples, are referred to as CTUs). The video decoder can segment CTUs according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure removes the concept of multiple partition types, such as the separation between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels: a first level segmented according to quadtree segmentation and a second level segmented according to binary tree segmentation. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to decoder units (CUs).

[0100] In the MTT partitioning structure, blocks can be partitioned using quadtree partitioning, binary tree partitioning, and one or more types of ternary tree partitioning. A ternary tree partition is a partition in which a block is divided into three sub-blocks. In some examples, a ternary tree partition divides a block into three sub-blocks without partitioning the original block through a center. The partitioning type in MTT (e.g., quadtree, binary tree, and ternary tree) can be symmetric or asymmetric.

[0101] When operating according to the AV1 codec, the video encoder 200 and video decoder 300 can be configured to decode video data block by block. In AV1, the largest decoded block that can be processed is called a superblock. In AV1, a superblock can be 128×128 luma samples or 64×64 luma samples. However, in subsequent video decoding formats (e.g., AV2), superblocks can be defined by different (e.g., larger) luma sample sizes. In some examples, the superblock is the top level of a block quadtree. The video encoder 200 can further divide the superblock into smaller decoded blocks. The video encoder 200 can divide the superblock and other decoded blocks into smaller blocks using square or non-square partitions. Non-square blocks can include N / 2×N blocks, N×N / 2 blocks, N / 4×N blocks, and N×N / 4 blocks. The video encoder 200 and video decoder 300 can perform separate prediction and transformation processes for each of the decoded blocks.

[0102] AV1 also defines video data tiles. A tile is a rectangular array of superblocks that can be decoded independently of other tiles. That is, the video encoder 200 and video decoder 300 can encode and decode the decoding blocks within a tile separately without using video data from other tiles. However, the video encoder 200 and video decoder 300 can perform filtering across tile boundaries. Tiles can be uniform or non-uniform in size. Tile-based decoding enables parallel processing and / or multithreading of the encoder and decoder.

[0103] In some examples, the video decoder may use a single QTBT or MTT structure to represent each of the luminance and chrominance components, while in other examples, the video decoder may use two or more QTBT or MTT structures, such as one QTBT or MTT structure for the luminance component and another QTBT or MTT structure for the two chrominance components (or two QTBT and / or MTT structures for the respective chrominance components).

[0104] The video decoder can be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, superblock partitioning, or other partitioning structures.

[0105] In some examples, a slice type is assigned to one or more slices of an image. Slice types include intra-frame decoded slices (I-slices), inter-frame decoded P-slices, and inter-frame decoded B-slices. An I-slice (an intra-frame decoded frame, independently decodeable) is a slice of an image that is decoded only by intra-frame prediction and is therefore independently decodeable because an I-slice only requires intra-frame data to predict any prediction unit or prediction block of the slice. A P-slice (a one-way prediction frame) is a slice of an image that can be decoded using both intra-frame prediction and one-way inter-frame prediction. Each prediction unit or prediction block within a P-slice is decoded using either intra-frame or inter-frame prediction. When inter-frame prediction is applied, the prediction unit or prediction block is predicted by only one reference image, and therefore the reference sample comes from only one reference region within a frame. A B-slice (a two-way prediction frame) is a slice of an image that can be decoded using both intra-frame and inter-frame prediction (e.g., either two-way or one-way prediction). Prediction units or blocks of a B-slice can be bidirectionally predicted from two reference images, where each image contributes to a reference region and the sample sets of the two reference regions are weighted (e.g., with equal weights or with different weights) to generate the prediction signal for the bidirectional prediction block. As described above, slices of an image are decoded independently. In some cases, an image can be decoded into only one slice.

[0106] As mentioned above, intra-image prediction utilizes the correlation between spatially adjacent samples within an image. Several intra-prediction modes exist (also referred to as "intra-frame modes"). In some examples, intra-frame prediction for luma blocks includes 35 modes, comprising planar modes, DC modes, and 33 angular modes (e.g., diagonal intra-prediction modes and adjacent angular modes). The 35 intra-prediction modes are indexed as shown in Table 1 below. In other examples, more intra-frame modes can be defined, including prediction angles that may not yet be represented by the 33 angular modes. In other examples, the prediction angles associated with angular modes may differ from those used in HEVC.

[0107] 0 INTRA_PLANAR 1 INTRA_DC 2..34 INTRA_ANGULAR2..INTRA_ANGULAR34

[0108] Table 1 - Specification of Intra-Frame Prediction Modes and Associated Names

[0109] Inter-image prediction uses temporal correlations between images to derive motion-compensated predictions for image sample blocks. Using a translational motion model, the position of a block in a previously decoded image (reference image) is indicated by a motion vector (Δx, Δy), where Δx specifies the horizontal displacement of the reference block relative to the current block's position, and Δy specifies the vertical displacement of the reference block relative to the current block's position. In some cases, the motion vector (Δx, Δy) may have integer sampling precision (also known as integer precision), in which case the motion vector points to an integer pixel grid (or integer pixel sample grid) of the reference frame. In other cases, the motion vector (Δx, Δy) may have fractional sampling precision (also known as fractional pixel precision or non-integer precision) to more accurately capture the movement of underlying objects, without being limited to an integer pixel grid of the reference frame. The precision of the motion vector can be represented by the quantization level of the motion vector. For example, the quantization level can be integer precision (e.g., 1 pixel) or fractional pixel precision (e.g., 1 / 4 pixel, 1 / 2 pixel, or other sub-pixel values). When the corresponding motion vector has fractional sampling precision, interpolation is applied to the reference image to derive the predicted signal. For example, samples available at integer positions can be filtered (e.g., using one or more interpolation filters) to estimate values ​​at fractional positions. Previously decoded reference images are indicated by reference indices (refIdx) in a list of reference images. Motion vectors and reference indices can be referred to as motion parameters. Two types of inter-image prediction can be performed, including one-way and two-way prediction.

[0110] In the case of inter-frame prediction using bidirectional prediction (also known as bidirectional inter-frame prediction), two sets of motion parameters (Δx0, y0, refIdx0 and Δx1, y1, refIdx1) are used to generate two motion-compensated predictions (from the same reference image or possibly from different reference images). For example, in the case of bidirectional prediction, each prediction block uses two motion-compensated prediction signals and generates B prediction units. The two motion-compensated predictions are then combined to obtain the final motion-compensated prediction. For example, the two motion-compensated predictions can be combined by averaging. In another example, weighted prediction can be used, in which case different weights can be applied to each motion-compensated prediction. The reference images that can be used in bidirectional prediction are stored in two separate lists, denoted as List 0 and List 1, respectively. The motion parameters can be derived at the encoder using a motion estimation process.

[0111] In the case of inter-frame prediction using unidirectional prediction (also known as one-way inter-frame prediction), motion-compensated predictions are generated from a reference image using a set of motion parameters (Δx0, y0, refIdx0). For example, in the case of unidirectional prediction, each prediction block uses at most one motion-compensated prediction signal, and P prediction units are generated.

[0112] The prediction unit (PU) may include data related to the prediction process (e.g., motion parameters or other suitable data). For example, when the PU is encoded using intra-frame prediction, the PU may include data describing the intra-frame prediction mode used for the PU. As another example, when the PU is encoded using inter-frame prediction, the PU may include data defining the motion vectors used for the PU. The data defining the motion vectors used for the PU may describe, for example, the horizontal component (Δx) of the motion vector, the vertical component (Δy) of the motion vector, the resolution used for the motion vector (e.g., integer precision, quarter-pixel precision, or eighth-pixel precision), the reference picture to which the motion vector points, the reference index, a list of reference pictures used for the motion vector (e.g., list 0, list 1, or list C), or any combination thereof.

[0113] AV1 includes two common techniques for encoding and decoding blocks of video data. These two common techniques are intra-frame prediction (e.g., intra-frame prediction or spatial prediction) and inter-frame prediction (e.g., inter-frame prediction or temporal prediction). In the context of AV1, when using intra-frame prediction modes to predict blocks of video data for the current frame, the video encoder 200 and video decoder 300 do not use video data from other frames of the video data. For most intra-frame prediction modes, the video encoding device 104 encodes blocks of the current frame based on the difference between sample values ​​in the current block and predicted values ​​generated from reference samples in the same frame. The video encoding device 104 determines the predicted values ​​generated from the reference samples based on the intra-frame prediction mode.

[0114] After performing prediction using intra-frame prediction and / or inter-frame prediction, encoding device 104 can perform transform and quantization. For example, after prediction, encoder engine 106 can compute a residual value corresponding to the PU. The residual value can include the pixel difference between the current pixel block (PU) being decoded and the prediction block used to predict the current block (e.g., a predicted version of the current block). For example, after generating a prediction block (e.g., using inter-frame prediction or intra-frame prediction), encoder engine 106 can generate a residual block by subtracting the prediction block generated by the prediction unit from the current block. The residual block includes a set of pixel differences that quantize the differences between the pixel values ​​of the current block and the pixel values ​​of the prediction block. In some examples, the residual block can be represented in a two-dimensional block format (e.g., a two-dimensional matrix or array of pixel values). In such examples, the residual block is a two-dimensional representation of the pixel values.

[0115] Block transforms are used to transform any residual data that may remain after prediction is performed. These block transforms can be based on discrete cosine transform, discrete sine transform, integer transform, wavelet transform, other suitable transform functions, or any combination thereof. In some cases, one or more block transforms (e.g., sizes of 32×32, 16×16, 8×8, 4×4, or other suitable sizes) can be applied to the residual data in each CU. In some implementations, TUs can be used for transform and quantization processes implemented by encoder engine 106. A given CU with one or more PUs may also include one or more TUs. As described further in detail below, residual values ​​can be transformed into transform coefficients using block transforms, and then quantized and scanned using TUs to produce serialized transform coefficients for entropy decoding.

[0116] In some implementations, after intra-frame prediction or inter-frame prediction decoding using the PU of the CU, the encoder engine 106 can compute residual data for the TU of the CU. The PU may include pixel data in the spatial domain (or pixel domain). The TU may include coefficients in the transform domain after applying the block transform. As previously described, the residual data may correspond to the pixel difference between a pixel in the unencoded image and the predicted value corresponding to the PU. The encoder engine 106 may form a TU including the residual data for the CU, and may then transform the TU to produce transform coefficients for the CU.

[0117] The encoder engine 106 can perform quantization on the transform coefficients. Quantization provides further compression by reducing the amount of data used to represent the coefficients. For example, quantization can reduce the bit depth associated with some or all of the coefficients. In one example, a coefficient with an n-bit value can be rounded down to an m-bit value during quantization, where n is greater than m.

[0118] Once quantization is performed, the decoded video bitstream includes the quantized transform coefficients, prediction information (e.g., prediction modes, motion vectors, block vectors, etc.), segmentation information, and any other suitable data (such as other syntax data). The distinct elements of the decoded video bitstream can then be entropy-coded by encoder engine 106. In some examples, encoder engine 106 may scan the quantized transform coefficients using a predefined scan order to produce a serialized vector that can be entropy-coded. In some examples, encoder engine 106 may perform an adaptive scan. After scanning the quantized transform coefficients to form a vector (e.g., a one-dimensional vector), encoder engine 106 may entropy-code that vector. For example, encoder engine 106 may use context-adaptive variable-length decoding, context-adaptive binary arithmetic decoding, syntax-based context-adaptive binary arithmetic decoding, probabilistic interval segmentation entropy decoding, or another suitable entropy coding technique.

[0119] The output 110 of the encoding device 104 can send the NAL units constituting the encoded video bitstream data to the decoding device 112 of the receiving device via the communication link 120. The input 114 of the decoding device 112 can receive the NAL units. The communication link 120 may include a channel provided by a wireless network, a wired network, or a combination of a wired and wireless network. The wireless network may include any wireless interface or combination of wireless interfaces, and may include any suitable wireless network (e.g., the Internet or other wide area networks, packet-based networks, WiFi). TM Radio frequency (RF), UWB, WiFi Direct, Cellular, 5G New Radio (NR), Long Term Evolution (LTE), WiMax TM Wired networks can include any wired interface (e.g., fiber optic, Ethernet, powerline Ethernet, coaxial cable-based Ethernet, digital signal line (DSL), etc.). Wired and / or wireless networks can be implemented using various equipment (such as base stations, routers, access points, bridges, gateways, switches, etc.). The encoded video bitstream data can be modulated according to communication standards (such as wireless communication protocols) and transmitted to the receiving device.

[0120] In some examples, encoding device 104 may store the encoded video bitstream data in storage device 108. Output 110 may retrieve the encoded video bitstream data from encoder engine 106 or from storage device 108. Storage device 108 may include any of a variety of distributed or locally accessed data storage media. For example, storage device 108 may include hard disk drives, storage disks, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data. Storage device 108 may also include a decoded picture buffer (DPB) for storing reference pictures used in inter-frame prediction. In further examples, storage device 108 may correspond to a file server or another intermediate storage device that may store the encoded video generated by the source device. In such a case, receiving device including decoding device 112 may access the stored video data from the storage device via streaming or downloading. The file server may be any type of server capable of storing the encoded video data and sending it to the receiving device. Example file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The receiving device can access the encoded video data via any standard data connection, including an internet connection. This can include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., DSL, cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on a file server. Transmission of the encoded video data from storage device 108 can be streaming, downloading, or a combination thereof.

[0121] Input 114 of decoding device 112 receives encoded video bitstream data and can provide the video bitstream data to decoder engine 116 or to storage device 118 for later use by decoder engine 116. For example, storage device 118 may include a DPB for storing reference pictures used in inter-frame prediction. A receiving device including decoding device 112 can receive the encoded video data to be decoded via storage device 108. The encoded video data may be modulated according to a communication standard (such as a wireless communication protocol) and transmitted to the receiving device. The communication medium used to transmit the encoded video data may include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include a router, a switch, a base station, or any other equipment that may be useful for facilitating communication from the source device to the receiving device.

[0122] Decoder engine 116 can decode the encoded video bitstream data by entropy decoding (e.g., using an entropy decoder) and extracting elements of one or more decoded video sequences that constitute the encoded video data. Decoder engine 116 can then rescale the encoded video bitstream data and perform an inverse transform on it. The residual data is then passed to the prediction stage of decoder engine 116. Decoder engine 116 then predicts pixel blocks (e.g., PUs). In some examples, the prediction is added to the output of the inverse transform (the residual data).

[0123] Video decoding device 112 can output the decoded video to video target device 122, which may include a display or other output device for displaying the decoded video data to a consumer of the content. In some aspects, video target device 122 may be part of a receiving device that includes decoding device 112. In some aspects, video target device 122 may be part of a separate device, distinct from the receiving device.

[0124] In some implementations, video encoding device 104 and / or video decoding device 112 may be integrated with audio encoding device and audio decoding device, respectively. Video encoding device 104 and / or video decoding device 112 may also include other hardware or software necessary for implementing the decoding techniques described above, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. Video encoding device 104 and video decoding device 112 may be integrated as part of a combined encoder / decoder (codec) in the respective device.

[0125] exist Figure 1 The example system shown is an illustrative example that can be used herein. The techniques used to process video data using the techniques described herein can be implemented by any digital video encoding and / or decoding device. Although the techniques disclosed herein are generally implemented by video encoding or video decoding devices, these techniques can also be implemented by a combined video encoder-decoder (commonly referred to as a "CODEC"). Furthermore, the techniques disclosed herein can also be implemented by a video preprocessor. The source device and receiving device are merely examples of such decoding devices, wherein the source device generates decoded video data for transmission to the receiving device. In some examples, the source device and receiving device can operate in a substantially symmetrical manner, such that each of these devices includes video encoding and decoding components. Therefore, the example system can support one-way or two-way video transmission between video devices, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0126] Extensions to the HEVC standard include the Multi-View Video Decoding Extension (MV-HEVC) and the Scalable Video Decoding Extension (SHVC). MV-HEVC and SHVC extensions share the concept of layered decoding, where different layers are included in the encoded video bitstream. Each layer in the decoded video sequence is addressed by a unique layer identifier (ID). The layer ID may exist in the header of a NAL unit to identify the layer associated with that NAL unit. In MV-HEVC, different layers typically represent different views of the same scene in the video bitstream. In SHVC, different scalable layers are provided to represent the video bitstream at different spatial resolutions (or picture resolutions) or different reconstruction fidelities. A scalable layer may include a base layer (with layer ID = 0) and one or more enhancement layers (with layer IDs = 1, 2, ..., n). The base layer may conform to the profile of the first version of HEVC and represents the lowest available layer in the bitstream. Compared to the base layer, enhancement layers have increased spatial resolution, temporal resolution, or frame rate and / or reconstruction fidelity (or quality). Enhancement layers are organized hierarchically and may (or may not) depend on lower layers. In some examples, a single standard codec can be used to decode different layers (e.g., using HEVC, SHVC, or other decoding standards to encode all layers). In other examples, multiple standard codecs can be used to decode different layers. For example, AVC can be used to decode the base layer, while SHVC and / or MV-HEVC extensions to the HEVC standard can be used to decode one or more enhancement layers.

[0127] Typically, a layer consists of a set of VCL NAL units and a corresponding set of non-VCL NAL units. Specific layer ID values ​​are assigned to the NAL units. Layers can be layered in the sense that they can depend on lower layers. A layer set refers to a self-contained set of layers represented within a bitstream, meaning that layers within a layer set can depend on other layers in that set during decoding, but not on any other layers to be decoded. Therefore, the layers in a layer set can form independent bitstreams that represent video content. The set of layers in a layer set can be obtained from another bitstream through sub-bitstream extraction operations. A layer set can correspond to the set of layers to be decoded when the decoder wants to operate according to certain parameters.

[0128] As previously described, the HEVC bitstream includes a group of NAL units, including VCL NAL units and non-VCL NAL units. VCL NAL units contain decoded picture data that forms the decoded video bitstream. For example, the bit sequence forming the decoded video bitstream resides in the VCL NAL units. Among other information, non-VCL NAL units may also contain parameter sets with high-level information relating to the encoded video bitstream. For example, parameter sets may include a Video Parameter Set (VPS), a Sequence Parameter Set (SPS), and a Picture Parameter Set (PPS). Examples of objectives for parameter sets include bit rate efficiency, fault tolerance, and providing a system-level interface. Each slice references a single active PPS, SPS, and VPS to access information available to the decoding device 112 for decoding the slice. An identifier (ID) can be decoded for each parameter set, including a VPS ID, an SPS ID, and a PPS ID. An SPS includes an SPS ID and a VPS ID. A PPS includes a PPS ID and an SPS ID. Each slice header includes a PPS ID. Using the ID, a valid parameter set can be identified for a given slice.

[0129] The PPS includes information applicable to all slices in a given picture. Therefore, all slices in a picture reference the same PPS. Slices in different pictures can also reference the same PPS. The SPS includes information applicable to all pictures in the same decoded video sequence (CVS) or bitstream. As previously described, the decoded video sequence is a series of access units (AUs) starting from a random access point picture (e.g., an instantaneous decode reference (IDR) picture or a broken link access (BLA) picture, or other suitable random access point picture) in the base layer and having certain properties (described above), and ending at the next AU (or the end of the bitstream) that has a random access point picture in the base layer and has certain properties, and does not include that next AU. The information in the SPS may remain unchanged within the decoded video sequence across pictures. Pictures in the decoded video sequence can use the same SPS. The VPS includes information applicable to all layers within the decoded video sequence or bitstream. The VPS includes a syntax structure with syntax elements applicable to the entire decoded video sequence. In some implementations, the VPS, SPS, or PPS may be transmitted in-band along with the encoded bitstream. In some implementations, the VPS, SPS, or PPS may be transmitted out-of-band in a separate transmission from the NAL unit containing the decoded video data.

[0130] This disclosure may generally refer to "signaling" certain information, such as syntax elements. The term "signaling" can generally refer to the transmission of values ​​for syntax elements and / or other data for decoding encoded video data. For example, video encoding device 104 may signal values ​​for syntax elements in a bitstream. Generally, "signaling" means generating a value in the bitstream. As described above, video device 102 may transmit the bitstream to video target device 122 substantially in real time or not in real time (e.g., it may occur while storing syntax elements to storage device 108 for later retrieval by video target device 122).

[0131] Video bitstreams may also include Supplemental Enhancement Information (SEI) messages. For example, an SEI NAL unit can be part of the video bitstream. In some cases, SEI messages may contain information not required by the decoding process. For example, the information in an SEI message may not be necessary for the decoder to decode the video images in the bitstream, but the decoder can use the information to improve the display or processing of the images (e.g., the decoded output). The information in an SEI message can be embedded metadata. In an exemplary example, the information in an SEI message can be used by a decoder-side entity to improve the visibility of the content. In some instances, certain application standards may indicate the presence of such SEI messages in the bitstream, enabling quality improvements for all devices conforming to the application standard (e.g., carrying frame encapsulation SEI messages for frame-compatible planar stereoscopic 3DTV video formats, where an SEI message is carried for each frame of the video; disposing of recovery point SEI messages; using pull-scan rectangle SEI messages in DVB; and many other examples).

[0132] As mentioned above, the energy-efficient media consumption standard (ISO / IEC 23001-11) specifies green metadata to facilitate reduced energy use during media consumption. Green metadata includes Complexity Measurement (CM) metadata and Reduced Decoding Operation Request (DOR-Req) metadata. Decoders can use CM metadata to help adjust the operating frequency of the processor performing decoding to help reduce power consumption. As previously described, this document describes systems and techniques for improving green metadata (e.g., CM signaling). For example, using a single type to signal CMs for slice, tile, and / or subframe granularity across multiple images can be problematic. In some aspects, the systems and techniques described herein improve the syntax structure of green metadata to support more granularity types (e.g., slice granularity, tile granularity, etc.) and, in some cases, various cycle types. In some aspects, the systems and techniques described herein provide signaling for mapping one or more CMs to a specific sub-image using a sub-image identifier (ID). In some aspects, the systems and techniques described herein provide signaling (e.g., syntax elements) indicating the total number of blocks (e.g., CTBs or other blocks). In some cases, such signaling (indicating the total number of blocks) simplifies the derivation of the total number of decoded blocks within a cycle to decode normalized coding statistics for each slice or tile. This signaling can be included in the CM metadata included with the encoded video. In some aspects, the systems and techniques described herein provide signaling for block statistics for intra-frame decoding. In some cases, such intra-frame decoding block statistics signaling addresses problems that arise when the CM provides block statistics for intra-frame decoding when all blocks are intra-frame decoded (e.g., when P and B slices have more intra-frame decoded blocks than inter-frame decoded blocks, when P or B pictures have more intra-frame decoded blocks than inter-frame decoded blocks, etc.). In some aspects, the systems and techniques described herein provide mechanisms for, for example, applying quality metrics separately to individual parts of a picture in a VVC (e.g., to individual sub-pictures), rather than applying quality metrics to the entire picture.

[0133] The various aspects of the aforementioned Complexity Measurement (CM) signaling will now be described. For example, in some aspects, granularity type indicator identifiers (e.g., granularity type syntax elements, such as granularity_type) are provided to support various granularities (e.g., granularity segments), such as slices, tiles, subpictures, scalable layers, and / or other granularities. For example, encoding device 104 may signal the granularity type indicator identifier in or with the bitstream. The granularity type indicator identifier may be used in combination with period type syntax elements to support granularity CM signaling applied to multiple pictures. In some examples, the semantics of the period type syntax element (e.g., period_type) are modified. In an illustrative example, CM signaling for VVC green metadata is provided in Table 2 below (where additions to ISO / IEC 23001-11 are shown between <> (e.g., <added statement>):

[0134] Table 2 - Syntax of VVC CM

[0135]

[0136] The period_type syntax element (e.g., a variable) specifies the type of the upcoming period to which the complexity measure applies, and the values ​​of the period_type syntax element can be defined in Table 3 below (as illustrative examples):

[0137] Table 3 - Specification for period_type in VVC

[0138]

[0139] The granularity_type syntax element specifies the type of granularity to which the complexity measure applies, and the values ​​for the granularity_type syntax element can be defined in Table 4 below (as illustrative examples):

[0140] Table 4 - Specification for granularity_type for VVC

[0141] 0x00 Image granularity, where CM is applicable to images 0x01 Slice grain size, where CM is applicable to slices 0x02 Tile granularity, where CM is applicable to tiles 0x03 Sub-image granularity, where CM is applicable to sub-images 0x04 Scalable layer granularity, where CM is suitable for scalable layers 0x05 CTU line granularity, where CM is applicable to CTU lines 0x07-0xFF Reserved

[0142] The `picture_level_CMs` syntax structure specifies a measure of the complexity of a particular image within a given period. In this paper, the `picture_level_CMs` syntax structure may be referred to as the image-level CM syntax structure.

[0143] The granularity_level_CMs syntax structure specifies the granularity-level complexity measure for each entity (such as a slice, tile, sub-image, or layer) within a cycle. In this paper, the granularity_level_CMs syntax structure may be referred to as the granularity-level CM syntax structure.

[0144] Figure 2 This is an illustration of an exemplary use of a granular-level CM for a video picture 200 (also referred to as a frame or image) according to various aspects of this disclosure. The video picture 200 includes a view of a cyclist 202 in motion while riding a bicycle across the video picture 200, and the cyclist 202 appears in a set of pictures 204 of the video picture 200. Each picture of the video picture 200 may (e.g., by an encoder such as encoding device 104) be divided into one or more portions, such as slices, tiles, sub-pictures, layers, etc. Picture 206 of the picture set 204 is shown divided into sixteen slices 208, four tiles 210, and one sub-picture 212, wherein each tile 210 includes four slices 208, and the sub-picture 212 includes two tiles 210 on the lower portion of the picture.

[0145] In some cases, being able to specify both the period type and granularity type for a decoding device (e.g., decoding device 112) allows for greater flexibility and reduced signaling by enabling the definition of granularity-level CMs for slices, tiles, subpictures, or layers of multiple pictures at once. For example, an encoding device (e.g., encoding device 104) can apply a single granularity-level CM to subpictures of all pictures at specified time intervals, instead of having to define a granularity-level CM for each subpicture of a picture. In video picture 200, the area where cyclist 202 appears as cyclist 202 moves in the video can be encoded / decoded in a more complex way compared to other areas of picture set 204 (which have less or no motion), and different CMs can be specified for those areas using granularity-level CMs. For example, a granularity-level CM (e.g., granularity_type = 3) can be specified for a subpicture 212 region of multiple pictures (e.g., six pictures of picture set 204, e.g., num_pictures = 6) at once. By allowing granularity-level CMs to be set for specific time intervals (e.g., a set of multiple images, a time period, etc.), a single granularity-level CM can be used in the metadata of the first image corresponding to image set 204, and this granularity-level CM can be applied to all images in image set 204 based on the specified time interval. After image set 204, the granularity-level CM of sub-image 212 can be adjusted because cyclist 202 is no longer in the area covered by sub-image 212, and that area can now be encoded / decoded more easily. Similarly, multiple (potentially different) granularity-level CMs can be specified for any number of slices, tiles, sub-images, or layers in an image, where each granularity-level CM can be applied in a different upcoming period (e.g., a single image, all images within a specified time interval, multiple images, all images until the image containing the next slice, etc.).

[0146] In some aspects, the encoding device (e.g., encoding device 104) can specify CM values ​​for portions of the images (e.g., slices, tiles, sub-images, and / or layers) applicable to multiple images of a video. For example, according to some aspects, the encoding device can generate and signal sub-image CM signaling. The sub-image CM signaling indicates which sub-images the CM applies to for one or more images. In one example, for sub-image granularity, a syntax element (e.g., referred to as a sub-image syntax element) indicates that the sub-image ID is signaled by CM metadata when the cycle spans multiple images (e.g., in the case where granularity-level CM applies to multiple images). An example is shown in Table 5 below:

[0147] Table 5 Sub-image CM

[0148]

[0149] subpic_id[i] specifies the subpic ID of the associated complexity metric (CM).

[0150] subpic_CM is the complexity metric structure for the i-th subpicture.

[0151] In some cases, `subpic_id` and / or `subpic_CM(i)` can be replaced by one or more syntax elements that reference a segment address. In some cases, a segment can be a slice, tile, or subpicture, and the segment address can identify, for example, a specific slice, tile, and / or subpicture of a picture. For example, a segment address [t] can indicate the address of the t-th segment. Therefore, when the granularity type specifies the subpicture granularity, the segment address [t] can indicate the subpicture ID of the t-th subpicture.

[0152] In some cases, aspects are associated with resolution changes. For example, in VVC, resolution changes within a decoded layer video sequence (CLVS) apply to picture, slice, and tile granularities, but not to subpicture granularities. Depending on some aspects, such as when the granularity type is equal to a slice (e.g., 0x01 from Table 4) or a tile (e.g., 0x02 from Table 4) and the period type spans multiple pictures, a syntax element indicating the total number of decoded tree luminance blocks within a period (e.g., called the decoded tree block (CTB quantity syntax element)) can be signaled in the green metadata (e.g., in a CM syntax table as one or more syntax elements, such as num_ctbs_minus1 in Table 6 below). An example is shown in Table 6 below:

[0153] Table 6. Syntax for Complexity Measurement

[0154]

[0155]

[0156] num_ctbs_minus1 specifies the total number of code tree blocks associated with the complexity metric within a period.

[0157] In some respects, alternative syntax elements (e.g., avg_number_ctbs_minus1) can indicate the average number of CTBs per image per granularity or per 4×4 block (or other block size) rather than the total number of CTBs over a period to reduce overhead. Such syntax elements may be referred to as average CTB number syntax elements.

[0158] In some cases, aspects are associated with the block statistics of intra-frame decoding. For example, when all blocks are intra-frame decoded, the current green metadata CM syntax only signals the block statistics of intra-frame decoding (e.g., portion_intra_predicted_blocks_area == 255). Table 7 below shows the proposed CM signaling changes, where additions are shown between <> (e.g., <added statement>), and deletions are shown as text with strikethrough (e.g., deleted statement). When a block of intra-frame decoding is available, the block statistics of intra-frame decoding are signaled. When a block of inter-frame decoding is available, the block statistics of inter-frame decoding are signaled.

[0159] Table 7 presents the proposed CM syntax structure.

[0160]

[0161]

[0162] The following are examples of definitions for various syntax elements from Table 7 for VVC:

[0163] `portion_intra_predicted_blocks_area` indicates the portion of the image covered by intra-frame predicted blocks in a specified period using a 4-sample granularity, and is defined as follows:

[0164]

[0165] NumIntraPredictedBlocks is the number of intra-predicted blocks using a 4-sample granularity in a specified period. On the encoder side, it is calculated as follows:

[0166]

[0167] Where NumIntraPredictedBlocks_X is the number of blocks predicted using intra-frame prediction in the specified period, where the number of samples comes from X = 4, 8, 16, 32, 64, 128, 256, 512, 1024, 2048, 4096.

[0168] In the decoder, NumIntraPredictedBlocks are derived from portion_intra_predicted_blocks_area and TotalNum4BlocksInPeriod.

[0169] `portion_planar_blocks_in_intra_area` indicates the portion of the intra-planar prediction block region within the intra-prediction region during the specified period, and is defined as follows:

[0170]

[0171] When it does not exist, it equals 0.

[0172] NumPlanarPredictedBlocks is the number of intra-planar prediction blocks with a 4-sample granularity used in a specified period. On the encoder side, it is calculated as follows:

[0173]

[0174] Where NumIntraPlanarBlocks_X is the number of blocks predicted using intraplane prediction in a specified period, with the number of samples coming from X = 4, 8, 16, 32, 64, 128, 256, 512, 1024, 2048, 4096.

[0175] In the decoder, NumPlanarPredictedBlocks are derived from portion_planar_blocks_in_intra_area and NumIntraPredictedBlocks.

[0176] The `portion_dc_blocks_in_intra_area` indicates the portion of the intra-DC prediction block region within the intra-prediction region during the specified period, and is defined as follows:

[0177]

[0178] When it does not exist, it equals 0.

[0179] NumDcPredictedBlocks is the number of intra-frame DC prediction blocks with a 4-sample granularity used in a specified period. On the encoder side, it is calculated as follows:

[0180]

[0181] Where NumIntraDcBlocks_X is the number of blocks predicted using intra-frame DC in a specified period, with the number of samples coming from X = 4, 8, 16, 32, 64, 128, 256, 512, 1024, 2048, 4096.

[0182] In the decoder, NumDcPredictedBlocks are derived from portion_dc_blocks_in_intra_area and NumIntraPredictedBlocks.

[0183] `portion_angular_hv_blocks_in_intra_area` (also known as `portion_hv_blocks_in_intra_area`) indicates the portion of the intra-horizontal and vertical prediction block regions within the intra-prediction region of a specified period, and is defined as follows:

[0184]

[0185] When it does not exist, it equals 0.

[0186] NumHvPredictedBlocks is the number of intra-frame horizontal and vertical prediction blocks using a 4-sample granularity within a specified period. On the encoder side, it is calculated as follows:

[0187]

[0188] Where NumIntraHvBlocks_X is the number of blocks predicted in the horizontal and vertical directions within a specified period, with the number of samples coming from X = 4, 8, 16, 32, 64, 128, 256, 512, 1024, 2048, 4096.

[0189] In the decoder, NumHvPredictedBlocks are derived from portion_hv_blocks_in_intra_area and NumIntraPredictedBlocks.

[0190] The `portion_mip_blocks_in_intra_area` indicates the portion of the intra-MIP prediction block region within the intra-prediction region during the specified period, and is defined as follows:

[0191]

[0192] When it does not exist, it equals 0.

[0193] NumMipPredictedBlocks is the number of intra-frame MIP prediction blocks using a 4-sample granularity in a specified period. On the encoder side, it is calculated as follows:

[0194]

[0195] Where NumIntraMipBlocks_X is the number of blocks predicted using intra-frame MIP in a specified period, with the number of samples coming from X = 4, 8, 16, 32, 64, 128, 256, 512, 1024, 2048, 4096.

[0196] In the decoder, NumMipPredictedBlocks are derived from portion_mip_blocks_in_intra_area and NumIntraPredictedBlocks.

[0197] `portion_bi_and_gpm_predicted_blocks_area` indicates the region covered by inter-frame bidirectional prediction or GPM prediction blocks in a specified period of image using 4-sample granularity, and is defined as follows:

[0198]

[0199] NumBiAndGpmPredictedBlocks is the number of inter-frame bidirectional prediction and GPM prediction blocks using a 4-sample granularity in a specified period. On the encoder side, it is calculated as follows:

[0200]

[0201] NumBiPredictedXBlocks is the number of blocks predicted using inter-frame bidirectional prediction or GPM prediction in a specified period, with the number of samples coming from X = 4, 8, 16, 32, 64, 128, 256, 512, 1024, 2048, 4096.

[0202] In the decoder, NumBiPredictedXBlocks are derived from portion_bi_and_gpm_predicted_blocks_area and TotalNum4BlocksInPeriod.

[0203] The portion_deblocking_instances indicates the portion of the deblocking filter instance (as defined in the terminology and definitions of this document) within a specified period, and is defined as follows:

[0204]

[0205] NumDeblockingInstances is the number of deblocking filter instances in a specified period. In the decoder, it is derived from portion_deblocking_instances and MaxNumDeblockingInstances.

[0206] `portion_sao_filtered_blocks` indicates the portion of blocks filtered by SAO with a 4-sample granularity used in a specified period. On the encoder side, it is calculated as follows:

[0207]

[0208] NumSaoFilteredBlocks is the number of portions of blocks that are SAO-filtered using a 4-sample granularity in a specified period. In the decoder, it is derived from portion_sao_filtered_blocks, TotalNum4BlocksInPeriod.

[0209] `portion_alf_filtered_blocks` indicates the portion of blocks filtered by ALF with a 4-sample granularity used in a specified period. On the encoder side, it is calculated as follows:

[0210]

[0211] NumAlfFilteredBlocks is the number of ALF-filtered block portions used in a specified period with a 4-sample granularity. In the decoder, it is derived from portion_alf_filtered_blocks, TotalNum4BlocksInPeriod.

[0212] In some cases, aspects are associated with sub-picture quality metrics. For example, a quality restoration metric can be applied to each granular segment. In some cases, a segment can be a slice, tile, or sub-picture. Table 8 provides examples of sub-picture-based metrics for quality restoration proposed for Green MPEG, where additions are shown between <> (e.g., <added statement>).

[0213] Table 8 Quality Recovery Metrics for Green Metadata

[0214]

[0215] xsd_subpic_number_minus1 specifies the number of subpicks available in the associated image. When xsd_subpic_number_minus1 equals 0, the quality restoration metric is applied to the entire image.

[0216] xsd_metric_type[i] indicates the type of objective quality metric for the i-th objective quality metric.

[0217] Xsd_metric_value[i][j] contains the value of the i-th objective quality metric associated with the j-th sub-image.

[0218] The current quality metric describes the quality of the final image for each segment. The aspects described herein allow SEI messages to carry quality metrics describing the quality of the associated images. For example, encoding devices (e.g., Figure 1 and Figure 4 The encoding device 104) can add quality metrics to the SEI message.

[0219] Figure 3 This is a flowchart illustrating a process 300 for decoding encoded video according to various aspects of this disclosure. At operation 302, process 300 may include obtaining a bitstream. At operation 304, process 300 may include retrieving a granularity type syntax element associated with the bitstream, the granularity type syntax element specifying the granularity type to which a complexity metric (CM) associated with the bitstream applies to one or more pictures. In some cases, the value of the granularity type syntax element specifies that the CM applies to a picture or a portion of a picture in the bitstream, the portion of which is smaller than the whole picture. In some cases, the value of the granularity type syntax element specifies that the CM applies to at least one of a slice, tile, subpicture, scalable layer, or decoder tree unit (CTU) row of the one or more pictures in the bitstream.

[0220] At operation 306, process 300 may include retrieving a period type syntax element associated with the bitstream, the period type syntax element indicating the upcoming time period or set of pictures to which the CM applies. In some cases, the period type syntax element indicates at least one of the following: a specified time interval for the upcoming period, the number of pictures for the upcoming period, the upcoming period including all pictures up to the point where the next slice is included, or the upcoming period including a single picture. In some cases, process 300 may also include retrieving a granularity-level CM syntax structure associated with the bitstream, the granularity-level CM syntax structure specifying a granularity-level complexity measure for one or more granular segments of the bitstream within the upcoming period. In some cases, process 300 may also include retrieving an additional period type syntax element associated with the bitstream, the additional period type syntax element being associated with the granularity type syntax element, wherein the additional period type syntax element is different from the period type syntax element; and decoding a portion of the bitstream based on the granularity type syntax element and the additional period type syntax element.

[0221] In some cases, process 300 may also include retrieving at least one of the following: a subpicture syntax element associated with the bitstream, which, when CM is applied to multiple pictures, indicates the signaling of a subpicture identifier (ID); a decode tree block (CTB) quantity syntax element associated with the bitstream, which, when the granularity type is equal to a slice or tile and the upcoming period covers multiple pictures, indicates the total number of decode tree blocks that can be signaled by CM in the upcoming period; or an average decode tree block (CTB) quantity syntax element associated with the bitstream, which indicates the average number of CTBs or 4×4 blocks per picture per granularity.

[0222] In some cases, for process 300, when a block of intra-frame decoding is available in at least a portion of the bitstream, block statistics of intra-frame decoding are signaled in association with at least that portion of the bitstream. In some cases, when a block of inter-frame decoding is available in at least a portion of the bitstream, block statistics of inter-frame decoding are signaled in association with at least that portion of the bitstream. In some cases, process 300 may also include displaying at least that portion of the bitstream on a display. In some cases, process 300 may also include determining the operating frequency of the device based on the CM associated with the bitstream.

[0223] At operation 308, process 300 may include decoding a portion of the bitstream based on granularity-type syntax elements and periodic-type syntax elements. In some cases, process 300 may be performed by one of a mobile device, wearable device, extended reality device, camera, personal computer, vehicle, robotic device, television, or computing device.

[0224] Figure 4 This is a flowchart 400 illustrating techniques for encoding video according to various aspects of this disclosure. At operation 402, process 400 may include acquiring video data. At operation 404, process 400 may include generating a granularity type syntax element for a bitstream, the granularity type syntax element specifying the granularity type to which a complexity metric (CM) associated with the bitstream applies to one or more pictures. In some cases, the value of the granularity type syntax element specifies that the CM applies to a picture or a portion of a picture in the bitstream, said portion being smaller than the whole picture. In some cases, the value of the granularity type syntax element specifies that the CM applies to at least one of a slice, tile, subpicture, scalable layer, or decoder tree unit (CTU) row of the one or more pictures in the bitstream.

[0225] At operation 406, process 400 may include generating a period type syntax element associated with the bitstream, the period type syntax element indicating the upcoming time period or set of pictures to which the CM applies. In some cases, the period type syntax element indicates at least one of the following: a specified time interval for the upcoming period, the number of pictures for the upcoming period, the upcoming period including all pictures up to the point where the next slice is included, or the upcoming period including a single picture. In some cases, the period type syntax element indicates at least one of the following: a specified time interval for the upcoming period, the number of pictures for the upcoming period, the upcoming period including all pictures up to the point where the next slice is included, or the upcoming period including a single picture.

[0226] In some cases, process 400 may also include generating a granular-level CM syntax structure for the bitstream, which specifies a granular-level complexity measure for one or more entities in an upcoming period. In some cases, process 400 may also include generating an additional periodic-type syntax element associated with the granular-type syntax element for the bitstream, wherein the additional periodic-type syntax element is different from the periodic-type syntax element, and wherein the additional periodic-type syntax element is used to decode a portion of the bitstream having the granular-type syntax element. In some cases, process 400 may also include generating at least one of the following for the bitstream: a subpicture syntax element associated with the bitstream, which, when CM is applied to multiple pictures, indicates the signaling of a subpicture identifier (ID); a decode tree block (CTB) quantity syntax element associated with the bitstream, which, when the granularity type is equal to a slice or tile and the upcoming period covers multiple pictures, indicates the total number of decode tree blocks that can be signaled by CM in the upcoming period; or an average decode tree block (CTB) quantity syntax element associated with the bitstream, which indicates the average number of CTBs or 4×4 blocks per picture per granularity.

[0227] In some cases, for process 400, when a block of intra-frame decoding is available in at least a portion of the bitstream, block statistics of intra-frame decoding are signaled in association with at least that portion of the bitstream. In some cases, when a block of inter-frame decoding is available in at least a portion of the bitstream, block statistics of inter-frame decoding are signaled in association with at least that portion of the bitstream. In some cases, process 400 may be performed by one of a mobile device, wearable device, extended reality device, camera, personal computer, vehicle, robotic device, television set, or computing device.

[0228] In some implementations, the processes (or methods) described herein may be performed by a computing device or apparatus (such as...) Figure 1 The system 100 shown is executed. For example, it can be executed by... Figure 1 and Figure 5 The encoding device 104 shown is derived from another video source-side device or video transmission device, or from... Figure 1 and Figure 6 The decoding device 112 shown may perform these processes by another client-side device, such as a player device, a display, or any other client-side device. In some cases, the computing device or apparatus may include one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components configured to implement the steps of one or more processes described herein.

[0229] In some examples, a computing device may include a mobile device, a desktop computer, a server computer and / or a server system, or other types of computing devices. Components of the computing device (e.g., one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, and / or other components) may be implemented in circuitry. For example, components may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or components may include computer software, firmware, or combinations thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or combinations thereof for performing the various operations described herein. In some examples, a computing device or apparatus may include a camera configured to capture video data (e.g., a video sequence) comprising video frames. In some examples, the camera or other capturing device that captures the video data is separate from the computing device, in which case the computing device receives or acquires the captured video data. The computing device may also include a network interface configured to transmit video data. The network interface can be configured to transmit Internet Protocol (IP) based data or other types of data. In some examples, the computing device or apparatus may include a display for showing samples of output video content, such as images of a video bitstream.

[0230] These processes can be described relative to logic flowcharts, whose operations represent sequences of operations that can be implemented by hardware, computer instructions, or combinations thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which operations are described is not intended to be construed as limiting, and any number of described operations can be combined in any order and / or in parallel to implement a process.

[0231] Furthermore, these processes can be executed under the control of one or more computer systems configured with executable instructions, and can be implemented in hardware as code (e.g., executable instructions, one or more computer programs, or one or more applications) or a combination thereof that executes jointly on one or more processors. As mentioned above, the code can be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium can be non-transitory.

[0232] The decoding techniques discussed herein can be implemented in exemplary video encoding and decoding systems (e.g., system 100). In some examples, a system includes a source device that provides encoded video data that will later be decoded by a target device. Specifically, the source device provides the video data to the target device via a computer-readable medium. The source and target devices can include any of a wide variety of devices, including desktop computers, laptop computers, tablet computers, set-top boxes, mobile phones such as so-called "smart" phones, so-called "smart" tablets, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, etc. In some cases, the source and target devices may be equipped for wireless communication.

[0233] The target device can receive encoded video data to be decoded via a computer-readable medium. The computer-readable medium can include any type of medium or device capable of moving encoded video data from a source device to the target device. In one example, the computer-readable medium can include a communication medium enabling the source device to transmit encoded video data directly to the target device in real time. The encoded video data can be modulated according to communication standards (such as wireless communication protocols) and transmitted to the target device. The communication medium can include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network, such as a local area network (LAN), a wide area network (WAN), or a global network like the Internet. The communication medium can include routers, switches, base stations, or any other equipment that may be useful in facilitating communication from the source device to the target device.

[0234] In some examples, the encoded data can be output from an output interface to a storage device. Similarly, the encoded data can be accessed from a storage device via an input interface. The storage device can include any data storage medium of various distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing the encoded video data. In another example, the storage device may correspond to a file server or another intermediate storage device that can store the encoded video generated by the source device. The target device can access the stored video data from the storage device via streaming or downloading. The file server can be any type of server capable of storing the encoded video data and sending it to the target device. Example file servers include web servers (e.g., for websites), FTP servers, network attached storage (NAS) devices, or local disk drives. The target device can access the encoded video data via any standard data connection, including an internet connection. This can include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., DSL, cable modems, etc.), or a combination of both, suitable for accessing the encoded video data stored on a file server. The transmission of encoded video data from the storage device can be streaming, downloading, or a combination thereof.

[0235] The technologies disclosed herein are not necessarily limited to wireless applications or setups. These technologies can be applied to video decoding to support any multimedia application in a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission (such as HTTP-based Dynamic Adaptive Streaming (DASH)), digital video encoded onto data storage media, decoding of digital video stored on data storage media, or other applications. In some examples, the system can be configured to support one-way or two-way video transmission to support applications such as video streaming, video playback, video broadcasting, and / or video telephony.

[0236] In one example, the source device includes a video source, a video encoder, and an output interface. The target device may include an input interface, a video decoder, and a display device. The video encoder of the source device may be configured to apply the techniques disclosed herein. In other examples, the source and destination devices may include other components or arrangements. For example, the source device may receive video data from an external video source such as an external camera. Similarly, the target device may interface with an external display device, rather than including an integrated display device.

[0237] The exemplary system described above is merely one example. Techniques for simultaneously processing video data can be implemented by any digital video encoding and / or decoding device. While the techniques disclosed herein are generally implemented by video encoding devices, they can also be implemented by video encoders / decoders (commonly referred to as "CODECs"). Furthermore, the techniques disclosed herein can also be implemented by video preprocessors. The source and target devices are merely examples of such decoding devices, where the source device generates decoded video data for transmission to the target device. In some examples, the source and target devices can operate in a substantially symmetrical manner, such that each of these devices includes both video encoding and decoding components. Therefore, the example system can support unidirectional or bidirectional video transmission between video devices, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0238] Video sources may include video capture devices, such as cameras, video archiving units including previously captured video, and / or video feed interfaces for receiving video from video content providers. Alternatively, a video source may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In some cases, if the video source is a camera, the source and target devices may form a so-called camera phone or video phone. However, as described above, the techniques described in this disclosure are generally applicable to video decoding and can be applied to wireless and / or wired applications. In each case, the captured, pre-captured, or computer-generated video may be encoded by a video encoder. The encoded video information may then be output to a computer-readable medium via an output interface.

[0239] As described above, computer-readable media may include transient media, such as wireless broadcasting or wired network transmissions, or storage media (i.e., non-transitory storage media), such as hard disks, flash drives, compressed optical discs, digital video optical discs, Blu-ray discs, or other computer-readable media. In some examples, a network server (not shown) may receive encoded video data from a source device and, for example, provide the encoded video data to a target device via network transmission. Similarly, a computing device in a media production facility (e.g., an optical disc stamping facility) may receive encoded video data from a source device and produce an optical disc containing the encoded video data. Therefore, in various examples, computer-readable media can be understood to include one or more forms of computer-readable media.

[0240] The target device's input interface receives information from a computer-readable medium. The information in the computer-readable medium may include syntax information defined by the video encoder, which is also used by the video decoder. This syntax information includes characteristics of descriptive blocks and other decoded units (e.g., groups of pictures (GOPs)) and / or processed syntax elements. The display device displays the decoded video data to the user and may include any of a variety of display devices, such as a cathode ray tube (CRT), liquid crystal display (LCD), plasma display, organic light-emitting diode (OLED) display, or another type of display device. Various embodiments of this application have been described.

[0241] The specific details of the encoding device 104 and the decoding device 112 are respectively in Figure 5 and Figure 6 As shown in the image. Figure 5 This is a block diagram illustrating an exemplary encoding device 104 capable of implementing one or more of the techniques described in this disclosure. Encoding device 104 may, for example, generate the syntax elements and / or structures described herein (e.g., syntax elements and / or structures of green metadata, such as complexity metrics (CM), or other syntax elements and / or structures). Encoding device 104 may perform intra-frame prediction and inter-frame prediction decoding of video blocks within video slices, tiles, sub-pictures, etc. As previously described, intra-frame decoding relies at least in part on spatial prediction to reduce or remove spatial redundancy within a given video frame or picture. Inter-frame decoding relies at least in part on temporal prediction to reduce or remove temporal redundancy within neighboring or surrounding frames of a video sequence. Intra-frame mode (I-mode) may refer to any of several spatially based compression modes. Inter-frame mode (e.g., one-way prediction (P-mode) or two-way prediction (B-mode)) may refer to any of several temporally based compression modes.

[0242] Encoding device 104 includes a segmentation unit 35, a prediction processing unit 41, a filter unit 63, an image memory 64, a summer 50, a transform processing unit 52, a quantization unit 54, and an entropy coding unit 56. Prediction processing unit 41 includes a motion estimation unit 42, a motion compensation unit 44, and an intra-frame prediction processing unit 46. For video block reconstruction, encoding device 104 also includes an inverse quantization unit 58, an inverse transform processing unit 60, and a summer 62. Filter unit 63 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although filter unit 63 is... Figure 5The filter unit 63 is shown as an in-loop filter, but in other configurations, it may be implemented as a post-loop filter. The post-processing device 57 may perform additional processing on the encoded video data generated by the encoding device 104. In some examples, the techniques of this disclosure may be implemented by the encoding device 104. However, in other cases, one or more of the techniques of this disclosure may be implemented by the post-processing device 57.

[0243] like Figure 5 As shown, encoding device 104 receives video data, and segmentation unit 35 segments the data into video blocks. Segmentation may also include segmentation into slices, segments, tiles, or other larger units, as well as video block segmentation (e.g., according to a quadtree structure of LCUs and CUs). Encoding device 104 generally illustrates the components for encoding video blocks within a video slice to be encoded. Slices may be divided into multiple video blocks (and possibly into a set of video blocks referred to as tiles). Prediction processing unit 41 may select one of a variety of possible decoding modes for the current video block based on error results (e.g., decoding rate and distortion level, etc.), such as one or more of a variety of intra-frame prediction decoding modes or inter-frame prediction decoding modes. Prediction processing unit 41 may provide the resulting intra-frame decoded or inter-frame decoded blocks to summer 50 to generate residual block data and to summer 62 to reconstruct the encoded blocks for use as reference pictures.

[0244] The intra-prediction processing unit 46 within the prediction processing unit 41 can perform intra-prediction decoding of the current video block relative to one or more adjacent blocks in the same frame or slice as the current block to be decoded, to provide spatial compression. The motion estimation unit 42 and motion compensation unit 44 within the prediction processing unit 41 perform inter-prediction decoding of the current video block relative to one or more prediction blocks in one or more reference pictures, to provide temporal compression.

[0245] Motion estimation unit 42 can be configured to determine the inter-frame prediction mode of video slices based on a predetermined mode of the video sequence. The predetermined mode can designate video slices in the sequence as P-slices, B-slices, or GPB-slices. Motion estimation unit 42 and motion compensation unit 44 can be highly integrated, but are shown separately for conceptual purposes. Motion estimation performed by motion estimation unit 42 is the process of generating motion vectors that estimate the motion of video blocks. For example, the motion vectors can indicate the displacement of the prediction unit (PU) of a video block within the current video frame or picture relative to a prediction block within a reference picture.

[0246] A predicted block is a block found to closely match the PU of the video block to be decoded in terms of pixel differences, which can be determined by the sum of absolute differences (SAD), sum of squared differences (SSD), or other difference metrics. In some examples, encoding device 104 may compute values ​​for sub-integer pixel positions of a reference image stored in image memory 64. For example, encoding device 104 may interpolate values ​​for quarter-pixel positions, eighth-pixel positions, or other fractional pixel positions of the reference image. Thus, motion estimation unit 42 may perform motion search relative to integer pixel positions and fractional pixel positions and output motion vectors with fractional pixel precision.

[0247] The motion estimation unit 42 calculates the motion vector of the PU by comparing the position of the PU in the video block of the inter-frame decoded slice with the position of the predicted block in the reference picture. The reference picture can be selected from a first reference picture list (list 0) or a second reference picture list (list 1), where each identifies one or more reference pictures stored in the picture memory 64. The motion estimation unit 42 sends the calculated motion vector to the entropy coding unit 56 and the motion compensation unit 44.

[0248] Motion compensation performed by motion compensation unit 44 may involve extracting or generating prediction blocks based on motion vectors determined by motion estimation, thereby potentially performing interpolation with subpixel precision. Upon receiving the motion vector of the PU for the current video block, motion compensation unit 44 can locate the prediction block pointed to by the motion vector in a list of reference images. Encoding device 104 forms residual video blocks by subtracting the pixel values ​​of the prediction blocks from the pixel values ​​of the current video block being decoded, thus forming pixel differences. The pixel differences form the residual data of the block and may include both luma and chroma difference components. Summer 50 represents one or more components performing this subtraction operation. Motion compensation unit 44 may also generate syntax elements associated with video blocks and video slices for use by decoding device 112 when decoding video blocks of video slices.

[0249] Intra-prediction processing unit 46 can perform intra-prediction on the current block as an alternative to inter-prediction performed by motion estimation unit 42 and motion compensation unit 44, as described above. Specifically, intra-prediction processing unit 46 can determine an intra-prediction mode for encoding the current block. In some examples, intra-prediction processing unit 46 can use various intra-prediction modes to encode the current block, for example, during individual encoding passes, and intra-prediction processing unit 46 can select an appropriate intra-prediction mode from tested modes for use. For example, intra-prediction processing unit 46 can use rate-distortion analysis for various tested intra-prediction modes to calculate rate-distortion values, and can select the intra-prediction mode with the best rate-distortion characteristics from the tested modes. Rate-distortion analysis typically determines the amount of distortion (or error) between the encoded block and the original uncoded block encoded to produce the encoded block, as well as the bit rate (i.e., the number of bits) used to produce the encoded block. Intra-prediction processing unit 46 can calculate the ratio based on the distortion and rate of various encoded blocks to determine which intra-prediction mode exhibits the best rate-distortion value for the block.

[0250] In any case, after selecting an intra-prediction mode for a block, the intra-prediction processing unit 46 may provide information indicating the selected intra-prediction mode for the block to the entropy coding unit 56. The entropy coding unit 56 may encode the information indicating the selected intra-prediction mode. The encoding device 104 may include in the transmitted bitstream configuration data definitions for various blocks, as well as indications of the intra-prediction mode most likely to be used for each of the contexts, an intra-prediction mode index table, and a modified intra-prediction mode index table. The bitstream configuration data may include multiple intra-prediction mode index tables and multiple modified intra-prediction mode index tables (also referred to as codeword mapping tables).

[0251] After prediction processing unit 41 generates a prediction block for the current video block via inter-frame prediction or intra-frame prediction, encoding device 104 forms a residual video block by subtracting the prediction block from the current video block. The residual video data in the residual block may be included in one or more TUs and applied to transform processing unit 52. Transform processing unit 52 uses a transform (such as Discrete Cosine Transform (DCT) or a conceptually similar transform) to transform the residual video data into residual transform coefficients. Transform processing unit 52 can transform the residual video data from the pixel domain to the transform domain, such as the frequency domain.

[0252] Transform processing unit 52 can send the resulting transform coefficients to quantization unit 54. Quantization unit 54 quantizes the transform coefficients to further reduce the bit rate. The quantization process can reduce the bit depth associated with some or all of the coefficients. The degree of quantization can be modified by adjusting the quantization parameters. In some examples, quantization unit 54 can then perform a scan of a matrix including the quantized transform coefficients. Alternatively, entropy coding unit 56 can perform the scan.

[0253] After quantization, entropy coding unit 56 performs entropy coding on the quantized transform coefficients. For example, entropy coding unit 56 may perform context-adaptive variable-length decoding (CAVLC), context-adaptive binary arithmetic decoding (CABAC), syntax-based context-adaptive binary arithmetic decoding (SBAC), probabilistic interval partitioned entropy (PIPE) decoding, or another entropy coding technique. After entropy coding by entropy coding unit 56, the encoded bitstream can be sent to decoding device 112, or archived for later transmission or retrieval by decoding device 112. Entropy coding unit 56 may also perform entropy coding on the motion vectors and other syntax elements of the current video slice being decoded.

[0254] Inverse quantization unit 58 and inverse transform processing unit 60 apply inverse quantization and inverse transform, respectively, to reconstruct residual blocks in the pixel domain for later use as reference blocks in reference images. Motion compensation unit 44 can compute reference blocks by adding the residual blocks to the prediction blocks of one of the reference images in the reference image list. Motion compensation unit 44 can also apply one or more interpolation filters to the reconstructed residual blocks to compute sub-integer pixel values ​​for use in motion estimation. Summer 62 adds the reconstructed residual blocks to the motion-compensated prediction blocks generated by motion compensation unit 44 to produce reference blocks for storage in image memory 64. The reference blocks can be used by motion estimation unit 42 and motion compensation unit 44 as reference blocks for inter-frame prediction of blocks in subsequent video frames or images.

[0255] In this way, Figure 5 The encoding device 104 represents an example of a video encoder configured to perform any of the techniques described herein. In some cases, some of the techniques of this disclosure may also be implemented by the post-processing device 57.

[0256] Figure 6 This is a block diagram illustrating an exemplary decoding device 112. The decoding device 112 includes an entropy decoding unit 80, a prediction processing unit 81, an inverse quantization unit 86, an inverse transform processing unit 88, a summer 90, a filter unit 91, and an image memory 92. The prediction processing unit 81 includes a motion compensation unit 82 and an intra-frame prediction processing unit 84. In some examples, the decoding device 112 may perform operations generally related to... Figure 5The encoding device 104 describes the encoding cycle as the inverse of the decoding cycle.

[0257] During the decoding process, decoding device 112 receives an encoded video bitstream representing video blocks and associated syntax elements of the encoded video slice sent by encoding device 104. In some embodiments, decoding device 112 may receive the encoded video bitstream from encoding device 104. In some embodiments, decoding device 112 may receive the encoded video bitstream from network entity 79, such as a server, a media-aware network element (MANE), a video editor / stitcher, or other such device configured to implement one or more of the techniques described above. Network entity 79 may or may not include encoding device 104. Some of the techniques described in this disclosure may be implemented by network entity 79 before it sends the encoded video bitstream to decoding device 112. In some video decoding systems, network entity 79 and decoding device 112 may be part of a separate device, while in other examples, the functionality described with respect to network entity 79 may be performed by the same device including decoding device 112.

[0258] The entropy decoding unit 80 of the decoding device 112 performs entropy decoding on the bitstream to generate quantization coefficients, motion vectors, and other syntax elements. The entropy decoding unit 80 forwards the motion vectors and other syntax elements to the prediction processing unit 81. The decoding device 112 can receive syntax elements at the video slice level and / or video block level. The entropy decoding unit 80 can process and parse both fixed-length and variable-length syntax elements from one or more parameter sets (such as VPS, SPS, and PPS).

[0259] When a video slice is decoded into an intra-frame decoded (I) slice, the intra-frame prediction processing unit 84 of the prediction processing unit 81 can generate prediction data for the video block of the current video slice based on the intra-frame prediction mode emitted by the signal and data from the previously decoded block from the current frame or picture. When a video frame is decoded into an inter-frame decoded (i.e., B, P, or GPB) slice, the motion compensation unit 82 of the prediction processing unit 81 generates a prediction block for the video block of the current video slice based on the motion vectors and other syntax elements received from the entropy decoding unit 80. A prediction block can be generated from one of the reference pictures in the reference picture list. The decoding device 112 can construct a reference frame list (list 0 and list 1) based on the reference pictures stored in the picture memory 92 using a default construction technique.

[0260] Motion compensation unit 82 determines prediction information for video blocks used in the current video slice by parsing motion vectors and other syntax elements, and uses this prediction information to generate prediction blocks for the current video slice being decoded. For example, motion compensation unit 82 may use one or more syntax elements in the parameter set to determine the prediction mode (e.g., intra-frame or inter-frame prediction), the inter-frame prediction slice type (e.g., B-slice, P-slice, or GPB-slice), the construction information of one or more reference picture lists for the slice, the motion vectors of each inter-frame encoded video block of the slice, the inter-frame prediction state of each inter-frame decoded video block of the slice, and other information for decoding video blocks in the current video slice.

[0261] The motion compensation unit 82 can also perform interpolation based on an interpolation filter. The motion compensation unit 82 can use an interpolation filter, such as that used by the encoding device 104 during the encoding of a video block, to calculate the interpolated values ​​of sub-integer pixels of the reference block. In this case, the motion compensation unit 82 can determine the interpolation filter used by the encoding device 104 from the received syntax elements and can use that interpolation filter to generate a prediction block.

[0262] The inverse quantization unit 86 inverse-quantizes or dequantizes the quantized transform coefficients provided in the bitstream and decoded by the entropy decoding unit 80. The inverse quantization process may include determining the degree of quantization using quantization parameters calculated by the encoding device 104 for each video block in the video slice, and similarly determining the degree of inverse quantization to be applied. The inverse transform processing unit 88 applies an inverse transform (e.g., inverse DCT or other suitable inverse transform), inverse integer transform, or conceptually similar inverse transform process to the transform coefficients to produce residual blocks in the pixel domain.

[0263] After the motion compensation unit 82 generates a prediction block for the current video block based on motion vectors and other syntax elements, the decoding device 112 forms the decoded video block by adding the residual block from the inverse transform processing unit 88 to the corresponding prediction block generated by the motion compensation unit 82. The summer 90 represents one or more components performing this summation operation. Loop filters (in or after the decoding loop) can also be used, if needed, to smooth pixel transitions or otherwise improve video quality. The filter unit 91 is intended to represent one or more loop filters, such as a deblocking filter, an adaptive loop filter (ALF), and a sample adaptive offset (SAO) filter. Although the filter unit 91 is in Figure 6The image shown is an in-loop filter, but in other configurations, filter unit 91 may be implemented as a post-loop filter. The decoded video block in a given frame or image is then stored in image memory 92, which stores reference images for subsequent motion compensation. Image memory 92 also stores the decoded video for later display on a display device (such as...). Figure 1 The video is displayed on the target device 122 shown.

[0264] In this way, Figure 6 The decoding device 112 represents an example of a video decoder configured to perform any of the techniques described herein.

[0265] As used herein, the term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored and which do not include carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or magnetic tapes, optical storage media such as CDs or DVDs, flash memory, memory, or memory devices. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments may be coupled to other code segments or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted via any suitable means, including memory sharing, messaging, token passing, network transmission, etc.

[0266] In some implementations, computer-readable storage devices, media, and memories may include wired or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0267] Specific details are provided in the foregoing description to provide a thorough understanding of the embodiments and examples provided herein. However, those skilled in the art will understand that these embodiments can be practiced without these specific details. For clarity, in some cases, the technology may be presented as comprising individual functional blocks, including functional blocks containing devices, device components, steps or routines in methods embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes and other components may be shown as components in block diagram form to avoid obscuring these embodiments in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures and techniques may be shown without the need for unnecessary detail to avoid confusing the embodiments.

[0268] Individual implementations can be described above as processes or methods depicted as flowcharts, flow diagrams, data flow diagrams, structure diagrams, or block diagrams. Although flowcharts can describe operations as sequential processes, many operations within an operation can be performed in parallel or simultaneously. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but a process may have additional steps not included in the accompanying drawings. A process can correspond to a method, function, program, subroutine, subroutines, etc. When a process corresponds to a function, its termination can correspond to the function returning to the calling function or the main function.

[0269] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise available from a computer-readable medium. These instructions may include, for example, instructions and data that cause or otherwise configure a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. The portion may be accessible via a network of computer resources used. The computer-executable instructions may be, for example, binary, intermediate format instructions, such as assembly language, firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during the methods according to the described examples include disks or optical discs, flash memory, USB devices with non-volatile memory, networked storage devices, etc.

[0270] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. A processor may perform the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or interlocking cards. By further example, such functionality may also be implemented on circuit boards of different chips or different processes executed on a single device.

[0271] Instructions, the medium for conveying such instructions, the computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functionality described in this disclosure.

[0272] In the foregoing description, various aspects of this application have been described with reference to specific embodiments thereof; however, those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative embodiments of this application have been described in detail herein, it is to be understood that the inventive concept can be implemented and employed in a variety of other ways, and the appended claims are not intended to be construed as including these variations unless limited by prior art. Various features and aspects of the above applications can be used individually or in combination. Furthermore, without departing from the broader spirit and scope of this specification, the embodiments can be used in any number of environments and applications beyond those described herein. Therefore, the specification and drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that in alternative embodiments, the methods may be performed in a different order than described.

[0273] Those skilled in the art will understand that the less than (“<”) and greater than (“>”) symbols or terms used herein may be replaced by the less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively, without departing from the scope of this specification.

[0274] When a component is described as being “configured” to perform certain operations, such a configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0275] The phrase “coupled to” means any component is physically connected directly or indirectly to another component, and / or any component communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0276] The language used in this disclosure to refer to “at least one of” and / or “one or more of” the set indicates that one or more members of the set (in any combination) satisfy the claim. For example, the claim language stating “at least one of A and B” means A, B, or A and B. In another example, the claim language stating “at least one of A, B, and C” means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language “at least one of” and / or “one or more of” the set does not limit the set to the items listed in the set. For example, the claim language stating “at least one of A and B” may mean A, B, or A and B, and may additionally include items not listed in the set of A and B.

[0277] The various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with the embodiments disclosed herein can be implemented as electronic hardware, computer software, firmware, or a combination thereof. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described in general terms of their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.

[0278] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication handheld devices, or integrated circuit devices with multiple uses, including applications in wireless communication handheld devices and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, these techniques can be implemented at least in part by a computer-readable data storage medium comprising program code, including instructions that, when executed, perform one or more of the methods described above. The computer-readable data storage medium can form part of a computer program product and may include packaging material. The computer-readable medium can include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Alternatively or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or transmits program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagating signals or waves.

[0279] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such a processor can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternatives, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, the term "processor" as used herein may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or apparatus suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein may be provided within a dedicated software or hardware module configured for encoding and decoding, or incorporated into a combined video encoder-decoder (codec).

[0280] The illustrative aspects of this disclosure include:

[0281] Aspect 1. An apparatus for processing video data, comprising: at least one memory; and at least one processor coupled to said at least one memory, said at least one processor being configured to: acquire a bitstream; retrieve a granularity type syntax element associated with said bitstream, said granularity type syntax element specifying the granularity type of one or more pictures to which a complexity metric (CM) associated with said bitstream applies; retrieve a periodicity type syntax element associated with said bitstream, said periodicity type syntax element indicating an upcoming time period or set of pictures to which said CM applies; and decode a portion of said bitstream based on said granularity type syntax element and said periodicity type syntax element.

[0282] Aspect 2. The apparatus of claim 1, wherein the value of the granularity type syntax element specifies that the CM applies to an image or a portion of the image of the bitstream, the portion of the image being smaller than the whole image.

[0283] Aspect 3. The apparatus of claim 1, wherein the value of the granularity type syntax element specifies that the CM is applicable to at least one of the slices, tiles, sub-pictures, scalable layers, or decoder tree unit (CTU) rows of the one or more pictures of the bitstream.

[0284] Aspect 4. The apparatus of claim 1, wherein the period type syntax element indicates at least one of the following: a specified time interval for the upcoming period, a number of pictures for the upcoming period, the upcoming period comprising all pictures up to the point where the next slice is included, or the upcoming period comprising a single picture.

[0285] Aspect 5. The apparatus of claim 1, wherein the at least one processor is configured to retrieve a granular-level CM syntax structure associated with the bitstream, the granular-level CM syntax structure specifying a granular-level complexity measure for one or more granular segments of the bitstream in the upcoming period.

[0286] Aspect 6. The apparatus of claim 1, wherein the at least one processor is configured to: retrieve an additional periodic type syntax element associated with the bitstream, the additional periodic type syntax element being associated with the granular type syntax element, wherein the additional periodic type syntax element is different from the periodic type syntax element; and decode the portion of the bitstream based on the granular type syntax element and the additional periodic type syntax element.

[0287] Aspect 7. The apparatus of claim 1, wherein the at least one processor is configured to retrieve at least one of the following: a sub-picture syntax element associated with the bitstream, wherein when the CM is applied to multiple pictures, the sub-picture syntax element indicates the signaling of a sub-picture identifier (ID); a decode tree block (CTB) quantity syntax element associated with the bitstream, wherein when the granularity type is equal to a slice or tile and the upcoming period covers multiple pictures, the CTB quantity syntax element indicates the total number of decode tree blocks that can be signaled by the CM in the upcoming period; or an average decode tree block (CTB) quantity syntax element associated with the bitstream, wherein the average CTB quantity syntax element indicates the average number of CTBs or 4×4 blocks per picture per granularity.

[0288] Aspect 8. The apparatus of claim 1, wherein when there are available blocks for intra-frame decoding in at least a portion of the bitstream, block statistics for intra-frame decoding are signaled in association with at least said portion of the bitstream.

[0289] Aspect 9. The apparatus of claim 1, wherein when there are blocks of inter-frame decoding available in at least a portion of the bitstream, block statistics of inter-frame decoding are signaled in association with at least said portion of the bitstream.

[0290] Aspect 10. The apparatus of claim 1, wherein the at least one processor is configured to determine the operating frequency of the apparatus based on the CM associated with the bit stream.

[0291] Aspect 11. The apparatus of claim 1, further comprising a display configured to display at least said portion of the bitstream.

[0292] Aspect 12. The apparatus of claim 1, wherein the apparatus is one of a mobile device, a wearable device, an extended reality device, a camera, a personal computer, a vehicle, a robotic device, a television set, or a computing device.

[0293] Aspect 13. A method for processing video data, comprising: obtaining a bitstream; retrieving a granularity type syntax element associated with the bitstream, the granularity type syntax element specifying a granularity type to which a complexity metric (CM) associated with the bitstream applies to one or more pictures; retrieving a periodicity type syntax element associated with the bitstream, the periodicity type syntax element indicating an upcoming time period or set of pictures to which the CM applies; and decoding a portion of the bitstream based on the granularity type syntax element and the periodicity type syntax element.

[0294] Aspect 14. The method of claim 13, wherein the value of the granularity type syntax element specifies that the CM applies to an image or a portion of the image of the bitstream, the portion of the image being smaller than the whole image.

[0295] Aspect 15. The method of claim 13, wherein the value of the granularity type syntax element specifies that the CM is applicable to at least one of the slices, tiles, sub-pictures, scalable layers, or decoder tree unit (CTU) rows of the one or more pictures of the bitstream.

[0296] Aspect 16. The method of claim 13, wherein the period type syntax element indicates at least one of the following: a specified time interval for the upcoming period, a number of pictures for the upcoming period, the upcoming period comprising all pictures up to the point where the next slice is included, or the upcoming period comprising a single picture.

[0297] Aspect 17. The method of claim 13 further includes retrieving a granular-level CM syntax structure associated with the bitstream, the granular-level CM syntax structure specifying a granular-level complexity measure for one or more granular segments of the bitstream in the upcoming period.

[0298] Aspect 18. The method of claim 13, further comprising: retrieving an additional periodic type syntax element associated with the bitstream, the additional periodic type syntax element being associated with the granularity type syntax element, wherein the additional periodic type syntax element is different from the periodic type syntax element; and decoding a portion of the bitstream based on the granularity type syntax element and the additional periodic type syntax element.

[0299] Aspect 19. The method of claim 13, further comprising retrieving at least one of the following: a sub-picture syntax element associated with the bitstream, wherein when the CM is applied to multiple pictures, the sub-picture syntax element indicates the signaling of a sub-picture identifier (ID); a decode tree block (CTB) quantity syntax element associated with the bitstream, wherein when the granularity type is equal to a slice or tile and the upcoming period covers multiple pictures, the CTB quantity syntax element indicates the total number of decode tree blocks that can be signaled by the CM in the upcoming period; or an average decode tree block (CTB) quantity syntax element associated with the bitstream, wherein the average CTB quantity syntax element indicates the average number of CTBs or 4×4 blocks per picture per granularity.

[0300] Aspect 20. The method of claim 13, wherein when there are available blocks for intra-frame decoding in at least a portion of the bitstream, block statistics for intra-frame decoding are signaled in association with at least said portion of the bitstream.

[0301] Aspect 21. The method of claim 13, wherein when there are blocks of inter-frame decoding available in at least a portion of the bitstream, block statistics of inter-frame decoding are signaled in association with at least said portion of the bitstream.

[0302] Aspect 22. The method of claim 13, further comprising: displaying at least said portion of the bit stream on a display.

[0303] Aspect 23. The method of claim 13, further comprising determining the operating frequency of the device based on the CM associated with the bit stream.

[0304] Aspect 24. An apparatus for processing video data, comprising: at least one memory; and at least one processor coupled to said at least one memory, said at least one processor being configured to: acquire video data; generate a granularity type syntax element for a bitstream, the granularity type syntax element specifying a granularity type to which a complexity metric (CM) associated with the bitstream applies to one or more pictures; generate a periodicity type syntax element for the bitstream associated with the bitstream, the periodicity type syntax element indicating an upcoming time period or set of pictures to which the CM applies; generate the bitstream associated with the video data, the bitstream including the granularity type syntax element and the periodicity type syntax element; and output the generated bitstream.

[0305] Aspect 25. The apparatus of claim 24, wherein the value of the granularity type syntax element specifies that the CM applies to an image or a portion of the image of the bitstream, the portion of the image being smaller than the whole image.

[0306] Aspect 26. The apparatus of claim 24, wherein the value of the granularity type syntax element specifies that the CM is applicable to at least one of the slices, tiles, sub-pictures, scalable layers, or decoder tree unit (CTU) rows of the one or more pictures of the bitstream.

[0307] Aspect 27. The apparatus of claim 24, wherein the period type syntax element indicates at least one of the following: a specified time interval for the upcoming period, a number of pictures for the upcoming period, the upcoming period comprising all pictures up to the point where the next slice is included, or the upcoming period comprising a single picture.

[0308] Aspect 28. The apparatus of claim 24, wherein the one or more processors are configured to generate a granular CM syntax structure for the bitstream, the granular CM syntax structure specifying a granular complexity metric for one or more entities within the upcoming period.

[0309] Aspect 29. The apparatus of claim 24, wherein the period type syntax element indicates at least one of the following: a specified time interval for the upcoming period, a number of pictures for the upcoming period, the upcoming period comprising all pictures up to the point where the next slice is contained, or the upcoming period comprising a single picture.

[0310] Aspect 30. The apparatus of claim 24, wherein the at least one processor is configured to generate associated additional periodic type syntax elements for the bitstream, the additional periodic type syntax elements being associated with the granular type syntax element, wherein the additional periodic type syntax elements are different from the periodic type syntax element, and wherein the additional periodic type syntax elements are used to decode a portion of the bitstream having the granular type syntax element.

[0311] Aspect 31. The apparatus of claim 24, wherein the at least one processor is configured to generate for the bitstream at least one of the following: a sub-picture syntax element associated with the bitstream, wherein when the CM is applied to multiple pictures, the sub-picture syntax element indicates the signaling of a sub-picture identifier (ID); a decode tree block (CTB) quantity syntax element associated with the bitstream, wherein when the granularity type is equal to a slice or tile and the upcoming period covers multiple pictures, the CTB quantity syntax element indicates the total number of decode tree blocks that can be signaled by the CM in the upcoming period; or an average decode tree block (CTB) quantity syntax element associated with the bitstream, wherein the average CTB quantity syntax element indicates the average number of CTBs or 4×4 blocks per picture per granularity.

[0312] Aspect 32. The apparatus of claim 24, wherein when there are available blocks of intra-frame decoding in at least a portion of the bitstream, block statistics of intra-frame decoding are signaled in association with at least said portion of the bitstream.

[0313] Aspect 33. The apparatus of claim 24, wherein when there are available blocks of inter-frame decoding in at least a portion of the bitstream, block statistics of inter-frame decoding are signaled in association with at least said portion of the bitstream.

[0314] Aspect 34. The apparatus of claim 24 further includes a camera configured to capture the video data.

[0315] Aspect 35. The apparatus of claim 24, wherein the apparatus is one of a mobile device, a wearable device, an extended reality device, a camera, a personal computer, a vehicle, a robotic device, a television set, or a computing device.

[0316] Aspect 36. The apparatus of claim 1, wherein the at least one processor is configured to retrieve one or more quality recovery metrics associated with one or more granular segments of the bitstream.

[0317] Aspect 37. The apparatus of claim 1, wherein the at least one processor is configured to: receive a Supplemental Enhancement Information (SEI) message; and retrieve the granularity type syntax element from the SEI message.

[0318] Aspect 38. The apparatus of claim 1, wherein the apparatus includes a decoder.

[0319] Aspect 39. The apparatus of claim 1, wherein the apparatus includes a camera configured to capture one or more images.

[0320] Aspect 40. The method of claim 13, wherein the at least one processor is configured to retrieve one or more quality recovery metrics associated with one or more granular segments of the bitstream.

[0321] Aspect 41. The method of claim 13, wherein the at least one processor is configured to: receive a Supplemental Enhancement Information (SEI) message; and retrieve the granularity type syntax element from the SEI message.

[0322] Aspect 42. The method of claim 13, wherein the apparatus includes a decoder.

[0323] Aspect 43. The method of claim 13, wherein the apparatus includes a camera configured to capture one or more images.

[0324] Aspect 44. The apparatus of claim 24, wherein the at least one processor is configured to encode one or more quality recovery metrics associated with one or more granular segments of the bitstream.

[0325] Aspect 45. The apparatus of claim 24, wherein the at least one processor is configured to encode a Supplemental Enhancement Information (SEI) message using the granularity type syntax element.

[0326] Aspect 46. The apparatus of claim 24, wherein the apparatus includes a decoder.

[0327] Aspect 47. The apparatus of claim 24, wherein the apparatus includes a camera configured to capture one or more images of the video data.

[0328] Aspect 48: A method for processing video data, comprising one or more operations according to any one of aspects 24 to 47.

[0329] Aspect 49: A non-transitory computer-readable medium having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform one or more operations according to any one of aspects 13 to 23 and aspect 48.

[0330] Aspect 50: An apparatus for processing video data, comprising components for performing one or more operations of any of the operations described in aspects 13 to 23 and aspect 48.

[0331] Aspect 1A: An apparatus for processing video data, comprising: at least one memory; and one or more processors coupled to said at least one memory, said one or more processors being configured to: acquire a bitstream; and retrieve a granularity type syntax element associated with said bitstream, said granularity type syntax element specifying the granularity type to which a complexity metric (CM) associated with said bitstream applies.

[0332] Aspect 2A: According to the apparatus of aspect 1A, wherein the value of the granularity type syntax element specifies the CM applicable to the picture of the bitstream.

[0333] Aspect 3A: According to the apparatus of aspect 2A, wherein the value of the granularity type syntax element specifies the slice of the bitstream to which the CM applies.

[0334] Aspect 4A: According to the apparatus of aspect 2A, wherein the value of the granularity type syntax element specifies the CM as a tile applicable to the bitstream.

[0335] Aspect 5A: According to the apparatus of aspect 2A, wherein the value of the granularity type syntax element specifies the CM applicable to the sub-picture of the bitstream.

[0336] Aspect 6A: According to the apparatus of aspect 2A, wherein the value of the granularity type syntax element specifies the scalable layer of the CM applicable to the bitstream.

[0337] Aspect 7A: The apparatus according to aspect 2A, wherein the value of the granularity type syntax element specifies the CM applicable to the decode tree unit (CTU) line of the bitstream.

[0338] Aspect 8A: An apparatus according to any one of aspects 1A to 7A, wherein one or more processors are configured to retrieve a period type syntax element associated with the bit stream, the period type syntax element specifying the type of the upcoming period to which the CM applies.

[0339] Aspect 9A: An apparatus according to any one of aspects 1A to 8A, wherein the one or more processors are configured to retrieve a picture-level CM syntax structure associated with the bitstream, the picture-level CM syntax structure specifying a complexity measure for one or more pictures within a period.

[0340] Aspect 10A: An apparatus according to any one of aspects 1A to 9A, wherein the one or more processors are configured to retrieve a granular CM syntax structure associated with the bitstream, the granular CM syntax structure specifying a granular complexity metric for one or more entities within a period.

[0341] Aspect 11A: The apparatus according to aspect 10A, wherein the one or more entities include at least one of slices, tiles, sub-pictures, and layers.

[0342] Aspect 12A: An apparatus according to any one of aspects 1A to 11A, wherein the one or more processors are configured to: retrieve sub-picture syntax elements associated with the bitstream, wherein, when the period covers multiple pictures, the sub-picture syntax elements indicate the signaling of a sub-picture identifier (ID) according to the CM.

[0343] Aspect 13A: An apparatus according to any one of aspects 1A to 12A, wherein the one or more processors are configured to retrieve a decode tree block (CTB) quantity syntax element associated with the bit stream, wherein the CTB quantity syntax element indicates the total number of decode tree luminance blocks that can be signaled in the period according to CM when the granularity type is equal to a slice or tile and the period covers multiple pictures.

[0344] Aspect 14A: An apparatus according to any one of aspects 1A to 13A, wherein the one or more processors are configured to retrieve an average decode tree block (CTB) count syntax element associated with the bit stream, the average CTB count syntax element indicating the average number of CTBs or 4×4 blocks per picture per granularity.

[0345] Aspect 15A: An apparatus according to any one of aspects 1A to 14A, wherein when there are available blocks of intra-frame decoding in at least a portion of the bitstream, block statistics of intra-frame decoding are signaled in association with at least said portion of the bitstream.

[0346] Aspect 16A: An apparatus according to any one of aspects 1A to 15A, wherein when there are blocks of inter-frame decoding available in at least a portion of the bitstream, block statistics of inter-frame decoding are signaled in association with at least said portion of the bitstream.

[0347] Aspect 17A: An apparatus according to any one of aspects 1A to 16A, wherein the one or more processors are configured to retrieve one or more quality recovery metrics associated with one or more granular segments of the bitstream.

[0348] Aspect 18A: The apparatus according to aspect 17A, wherein the one or more granular segments of the bitstream include at least one of slices, tiles, and sub-pictures.

[0349] Aspect 19A: An apparatus according to any one of aspects 1A to 18A, wherein the one or more processors are configured to: receive a Supplemental Enhancement Information (SEI) message; and retrieve the granularity type syntax element from the SEI message.

[0350] Aspect 20A: An apparatus according to any one of aspects 1A to 19A, wherein one or more processors are configured to determine the operating frequency of the apparatus based on the CM associated with the bit stream.

[0351] Aspect 21A: The apparatus according to any one of aspects 1A to 20A, wherein the apparatus includes a decoder.

[0352] Aspect 22A: The apparatus according to any one of aspects 1A to 21A further includes a display configured to display one or more output images.

[0353] Aspect 23A: The apparatus according to any one of aspects 1A to 22A further includes a camera configured to capture one or more images.

[0354] Aspect 24A: The apparatus according to any one of aspects 1A to 23A, wherein the apparatus is a mobile device.

[0355] Aspect 25A: A method for processing video data, comprising one or more operations of any one of aspects 1A to 24A.

[0356] Aspect 26A: A non-transitory computer-readable medium having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform one or more operations according to any one of aspects 1A to 24A.

[0357] Aspect 27A: An apparatus for processing video data, comprising components for performing one or more operations of any one of aspects 1A to 24A.

[0358] Aspect 28A: An apparatus for processing video data, comprising: a memory; and one or more processors coupled to the memory, the one or more processors being configured to: acquire video data; generate a bitstream associated with the video data; and generate granularity type syntax elements for the bitstream, the granularity type syntax elements specifying the granularity type to which a complexity metric (CM) associated with the bitstream applies.

[0359] Aspect 29A: The apparatus according to aspect 28A, wherein the value of the granularity type syntax element specifies the CM applicable to the picture of the bitstream.

[0360] Aspect 30A: An apparatus according to aspect 29A, wherein the value of the granularity type syntax element specifies the slice of the bitstream to which the CM applies.

[0361] Aspect 31A: The apparatus according to aspect 29A, wherein the value of the granularity type syntax element specifies the CM as a tile applicable to the bitstream.

[0362] Aspect 32A: The apparatus according to aspect 29A, wherein the value of the granularity type syntax element specifies the CM applicable to the sub-picture of the bitstream.

[0363] Aspect 33A: The apparatus according to aspect 29A, wherein the value of the granularity type syntax element specifies the CM as a scalable layer applicable to the bitstream.

[0364] Aspect 34A: The apparatus according to aspect 29A, wherein the value of the granularity type syntax element specifies the CM for the decode tree unit (CTU) row of the bitstream.

[0365] Aspect 35A: An apparatus according to any one of aspects 28A to 34A, wherein one or more processors are configured to generate a period type syntax element for the bit stream, the period type syntax element specifying the type of the upcoming period to which the CM is applicable.

[0366] Aspect 36A: An apparatus according to any one of aspects 28A to 35A, wherein the one or more processors are configured to generate a picture-level CM syntax structure for the bitstream, the picture-level CM syntax structure specifying a complexity measure for one or more pictures within a period.

[0367] Aspect 37A: An apparatus according to any one of aspects 28A to 36A, wherein the one or more processors are configured to generate a granular CM syntax structure for the bitstream, the granular CM syntax structure specifying a granular complexity metric for one or more entities within a period.

[0368] Aspect 38A: The apparatus according to aspect 37A, wherein the one or more entities include at least one of slices, tiles, sub-pictures and layers.

[0369] Aspect 39A: The apparatus according to any one of aspects 28A to 38A, wherein the one or more processors are configured to generate sub-picture syntax elements for the bitstream, wherein the sub-picture syntax elements indicate the issuance of a sub-picture identifier (ID) by signaling according to the CM as the period covers multiple pictures.

[0370] Aspect 40A: An apparatus according to any one of aspects 28A to 39A, wherein the one or more processors are configured to generate a decode tree block (CTB) quantity syntax element for the bit stream, wherein the CTB quantity syntax element indicates the total number of decode tree luminance blocks within the period when the granularity type is equal to a slice or tile and the period covers multiple pictures.

[0371] Aspect 41A: An apparatus according to any one of aspects 28A to 40A, wherein the one or more processors are configured to generate an average code tree block (CTB) number syntax element for the bit stream, the CTB number syntax element indicating the average number of CTBs or 4×4 blocks per image per granularity.

[0372] Aspect 42A: The apparatus according to any one of aspects 28A to 41A, wherein when there are available blocks of intra-frame decoding in at least a portion of the bitstream, block statistics of intra-frame decoding are signaled in association with at least said portion of the bitstream.

[0373] Aspect 43A: The apparatus according to any one of aspects 28A to 42A, wherein when there are blocks of inter-frame decoding available in at least a portion of the bitstream, block statistics of inter-frame decoding are signaled in association with at least said portion of the bitstream.

[0374] Aspect 44A: An apparatus according to any one of aspects 28A to 43A, wherein the one or more processors are configured to generate one or more quality recovery metrics associated with one or more granular segments of the bitstream.

[0375] Aspect 45A: The apparatus according to aspect 44A, wherein the one or more granular segments of the bitstream include at least one of slices, tiles, and sub-pictures.

[0376] Aspect 46A: The apparatus according to any one of aspects 28A to 45A, wherein the one or more processors are configured to: generate supplementary enhancement information (SEI) messages; and include the granularity type syntax element in the SEI messages.

[0377] Aspect 47A: An apparatus according to any one of aspects 28A to 46A, wherein one or more processors are configured to store the bit stream.

[0378] Aspect 48A: The apparatus according to any one of aspects 28A to 47A, wherein the one or more processors are configured to transmit the bit stream.

[0379] Aspect 49A: The apparatus according to any one of aspects 28A to 48A, wherein the apparatus includes an encoder.

[0380] Aspect 50A: The apparatus according to any one of aspects 28A to 49A further includes a display configured to display one or more output images.

[0381] Aspect 51A: The apparatus according to any one of aspects 28A to 50A further includes a camera configured to capture one or more images.

[0382] Aspect 52A: The apparatus according to any one of aspects 28A to 51A, wherein the apparatus is a mobile device.

[0383] Aspect 53A: A method for processing video data, comprising one or more operations according to any one of aspects 28A to 52A.

[0384] Aspect 54A: A non-transitory computer-readable medium having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to perform one or more operations of any one of the operations described in aspects 28A to 52A.

[0385] Aspect 55A: An apparatus for processing video data, comprising components for performing one or more operations of any of aspects 28A to 52A.

Claims

1. An apparatus for processing video data, comprising: At least one memory including instructions; and At least one processor, the at least one processor being configured to execute the instructions to cause the device to: Obtain the bitstream; Retrieve a granularity type syntax element associated with the bitstream, the granularity type syntax element specifying the granularity type of one or more images to which a complexity metric (CM) associated with the bitstream applies, wherein the CM is used to determine the operating frequency of the device, and wherein the value of the granularity type syntax element specifies that the CM applies to an image or a portion of the image of the bitstream, the portion of the image being smaller than the whole image; Retrieve the period type syntax element associated with the bitstream, the period type syntax element indicating the upcoming time period or set of images to which the CM applies; as well as A portion of the bitstream is decoded based on the granularity type syntax elements and the periodic type syntax elements.

2. The apparatus of claim 1, wherein the period type syntax element indicates at least one of the following: a specified time interval for the upcoming period, a number of pictures for the upcoming period, the upcoming period comprising all pictures up to the point where the next slice is included, or the upcoming period comprising a single picture.

3. The apparatus of claim 1, wherein the at least one processor is configured to cause the apparatus to retrieve a granular-level CM syntax structure associated with the bitstream, the granular-level CM syntax structure specifying a granular-level complexity measure for one or more granular segments of the bitstream in the upcoming period.

4. The apparatus of claim 1, wherein the at least one processor is configured such that the apparatus: Retrieve an additional periodic type syntax element associated with the bitstream, the additional periodic type syntax element being associated with the granularity type syntax element, wherein the additional periodic type syntax element is different from the periodic type syntax element; and The portion of the bitstream is decoded based on the granularity type syntax element and the additional periodic type syntax element.

5. The apparatus of claim 1, wherein the at least one processor is configured such that the apparatus retrieves at least one of the following: The sub-picture syntax element associated with the bitstream, when the CM is applied to multiple pictures, indicates that a sub-picture identifier (ID) is emitted by signaling. The CTB (Cracked Tree Blocks) quantity syntax element associated with the bitstream, when the granularity type is equal to a slice or tile and the upcoming period spans multiple images, indicates the total number of cracked tree blocks that can be signaled in the upcoming period according to CM; or The average CTB count syntax element associated with the bitstream indicates the average number of CTBs per granularity or 4×4 blocks per image.

6. The apparatus of claim 1, wherein when there are available blocks for intra-frame decoding in at least a portion of the bitstream, block statistics for intra-frame decoding are signaled in association with at least said portion of the bitstream.

7. The apparatus of claim 1, wherein when there are available blocks for inter-frame decoding in at least a portion of the bitstream, block statistics for inter-frame decoding are signaled in association with at least said portion of the bitstream.

8. The apparatus of claim 1, further comprising a display configured to display at least said portion of the bitstream.

9. The apparatus of claim 1, wherein the apparatus is one of a mobile device, a wearable device, an augmented reality device, a camera, a personal computer, a vehicle, a robotic device, a television set, or a computing device.

10. A method for processing video data performed by an apparatus, comprising: Obtain the bitstream; Retrieve a granularity type syntax element associated with the bitstream, the granularity type syntax element specifying the granularity type of one or more images to which a complexity metric (CM) associated with the bitstream applies, wherein the CM is used to determine the operating frequency of the device, and wherein the value of the granularity type syntax element specifies that the CM applies to an image or a portion of the image of the bitstream, the portion of the image being smaller than the whole image; Retrieve the period type syntax element associated with the bitstream, the period type syntax element indicating the upcoming time period or set of images to which the CM applies; as well as A portion of the bitstream is decoded based on the granularity type syntax elements and the periodic type syntax elements.

11. The method of claim 10, wherein the period type syntax element indicates at least one of the following: a specified time interval for the upcoming period, a number of pictures for the upcoming period, the upcoming period comprising all pictures up to the point where the next slice is included, or the upcoming period comprising a single picture.

12. The method of claim 10, further comprising retrieving a granular-level CM syntax structure associated with the bitstream, the granular-level CM syntax structure specifying a granular-level complexity measure for one or more granular segments of the bitstream in the upcoming period.

13. The method of claim 10, further comprising: Retrieve an additional periodic type syntax element associated with the bitstream, the additional periodic type syntax element being associated with the granularity type syntax element, wherein the additional periodic type syntax element is different from the periodic type syntax element; as well as A portion of the bitstream is decoded based on the granularity type syntax element and the additional periodic type syntax element.

14. The method of claim 10, further comprising retrieving at least one of the following: The sub-picture syntax element associated with the bitstream, when the CM is applied to multiple pictures, indicates that a sub-picture identifier (ID) is emitted by signaling. The CTB (Cracked Tree Blocks) quantity syntax element associated with the bitstream, when the granularity type is equal to a slice or tile and the upcoming period spans multiple images, indicates the total number of cracked tree blocks that can be signaled in the upcoming period according to CM; or The average CTB count syntax element associated with the bitstream indicates the average number of CTBs per granularity or 4×4 blocks per image.

15. The method of claim 10, wherein when there are available blocks for intra-frame decoding in at least a portion of the bitstream, block statistics for intra-frame decoding are signaled in association with at least said portion of the bitstream.

16. The method of claim 10, wherein when there are available blocks for inter-frame decoding in at least a portion of the bitstream, block statistics for inter-frame decoding are signaled in association with at least said portion of the bitstream.

17. The method of claim 10, further comprising displaying at least said portion of the bitstream on a display.

18. An apparatus for processing video data, comprising: At least one memory including instructions; and at least one processor, the at least one processor being configured to execute the instructions to cause the device to: Acquire video data; A granularity type syntax element is generated for a bitstream, wherein the granularity type syntax element specifies the granularity type of one or more images to which a complexity metric (CM) associated with the bitstream applies, wherein the CM is used to determine the operating frequency of the decoder, and wherein the value of the granularity type syntax element specifies that the CM applies to an image or a portion of the image of the bitstream, wherein the portion of the image is smaller than the whole image; Generate a periodic type syntax element associated with the bitstream, the periodic type syntax element indicating the upcoming time period or set of images to which the CM applies; Generate the bitstream associated with the video data, the bitstream including the granularity type syntax element and the periodic type syntax element; as well as Output the generated bitstream.

19. The apparatus of claim 18, wherein the period type syntax element indicates at least one of the following: a specified time interval for the upcoming period, a number of pictures for the upcoming period, the upcoming period comprising all pictures up to the point where the next slice is included, or the upcoming period comprising a single picture.

20. The apparatus of claim 18, wherein the at least one processor is configured to cause the apparatus to generate a granular CM syntax structure for the bitstream, the granular CM syntax structure specifying a granular complexity metric for one or more entities within the upcoming period.

21. The apparatus of claim 18, wherein the period type syntax element indicates at least one of the following: a specified time interval for the upcoming period, a number of pictures for the upcoming period, the upcoming period comprising all pictures up to the point where the next slice is contained, or the upcoming period comprising a single picture.

22. The apparatus of claim 18, wherein the at least one processor is configured to generate associated additional periodic type syntax elements for the bitstream, the additional periodic type syntax elements being associated with the granular type syntax element, wherein the additional periodic type syntax elements are different from the periodic type syntax element, and wherein the additional periodic type syntax elements are used to decode a portion of the bitstream having the granular type syntax element.

23. The apparatus of claim 18, wherein the at least one processor is configured such that the apparatus generates at least one of the following for the bitstream: The sub-picture syntax element associated with the bitstream, when the CM is applied to multiple pictures, indicates that a sub-picture identifier (ID) is emitted by signaling. The CTB (Cracked Tree Blocks) quantity syntax element associated with the bitstream, when the granularity type is equal to a slice or tile and the upcoming period spans multiple images, indicates the total number of cracked tree blocks that can be signaled in the upcoming period according to CM; or The average CTB count syntax element associated with the bitstream indicates the average number of CTBs per granularity or 4×4 blocks per image.

24. The apparatus of claim 18, wherein when there are available blocks for intra-frame decoding in at least a portion of the bitstream, block statistics for intra-frame decoding are signaled in association with at least said portion of the bitstream.

25. The apparatus of claim 18, wherein when there are available blocks for inter-frame decoding in at least a portion of the bitstream, block statistics for inter-frame decoding are signaled in association with at least said portion of the bitstream.

26. The apparatus of claim 18, further comprising a camera configured to capture the video data.

27. The apparatus of claim 18, wherein the apparatus is one of a mobile device, a wearable device, an extended reality device, a camera, a personal computer, a vehicle, a robotic device, a television set, or a computing device.

28. A method for processing video data, comprising: Acquire video data; A granularity type syntax element is generated for a bitstream, wherein the granularity type syntax element specifies the granularity type of one or more images to which a complexity metric (CM) associated with the bitstream applies, wherein the CM is used to determine the operating frequency of the decoder, and wherein the value of the granularity type syntax element specifies that the CM applies to an image or a portion of the image of the bitstream, wherein the portion of the image is smaller than the whole image; Generate a periodic type syntax element associated with the bitstream, the periodic type syntax element indicating the upcoming time period or set of images to which the CM applies; Generate the bitstream associated with the video data, the bitstream including the granularity type syntax element and the periodic type syntax element; as well as Output the generated bitstream.

29. An apparatus for processing video data, the apparatus comprising components for performing the method of any one of claims 10 to 17.

30. A computer-readable medium having program code recorded thereon, wherein the program code is executable by one or more processors of a device to cause the processor to perform the method of any one of claims 10 to 17.

31. A computer program product comprising computer-readable instructions, which, when executed by a processor, cause the processor to perform the method of any one of claims 10 to 17.

32. An apparatus for processing video data, the apparatus comprising components for performing the method of claim 28.

33. A computer-readable medium having program code recorded thereon, wherein the program code is executable by one or more processors of the device to cause the processors to perform the method of claim 28.

34. A computer program product comprising computer-readable instructions, which, when executed by a processor, cause the processor to perform the method of claim 28.

Citation Information

Patent Citations

  • Storage and carriage of green metadata for display adaptation

    CN106464962A

  • An encoder, a decoder with support of sub-layer picture rates

    WO2021061024A1