Picture orientation and quality metrics supplemental enhancement information message for video coding
By using SEI messages with orientation and quality syntax elements, video decoding technologies address the issue of improper image handling, ensuring correct orientation and quality-based processing for enhanced video display and prediction.
Patent Information
- Application Number
- TW111108836
- Authority / Receiving Office
- TW · TW
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-03-08
- Filing Date
- 2022-03-10
- Publication Date
- 2026-07-11
- Estimated Expiration
- 2042-03-09
AI Technical Summary
Existing video decoding technologies do not effectively handle orientation and quality metadata in video data, leading to improper display and suboptimal processing of video images.
Incorporating Supplemental Enhancement Information (SEI) messages with syntax elements that indicate image orientation and quality metrics to guide proper transformation and processing of video data, such as rotation and mirroring, enhancing image display and quality-based decisions.
Enables accurate orientation and quality-based processing of video images, improving display quality and inter-frame prediction efficiency.
Smart Images

Figure IMG-2_DRAW_111108836-A0304-14-0001-1 
Figure IMG-2_DRAW_111108836-A0304-14-0002-2 
Figure IMG-2_DRAW_111108836-A0304-14-0003-3
Abstract
Description
Technical Field
[0001] This application claims the rights to U.S. Provisional Application No. 63 / 170,267, filed April 2, 2021, and U.S. Provisional Application No. 63 / 214,378, filed June 24, 2021, the entire contents of each of which are incorporated herein by reference.
[0002] This disclosure relates to video encoding and video decoding. Prior Technology
[0003] Digital video capabilities can be incorporated into a wide range of devices, including digital televisions, digital live broadcasting systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, tablets, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite wireless phones, so-called "smartphones," video conferencing equipment, video streaming devices, and so on. Digital video devices implement video decoding technologies (such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 Part 10 (Advanced Video Decoding (AVC)), ITU-T H.265 / High Efficiency Video Decoding (HEVC), ITU-T H.266 / Various Video Decoding (VVC), and extensions to such standards) as well as proprietary video codecs / formats (such as AOMedia Video 1 (AV1) developed by the Open Media Consortium). By implementing such video decoding technology, video devices can more efficiently send, receive, encode, decode, and / or store digital video information.
[0004] Video decoding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in video sequences. For block-based video decoding, video slices (e.g., video pictures or portions of video pictures) can be segmented into video blocks, which may also be referred to as decoding tree units (CTUs), decoding units (CUs), and / or decoding nodes. Video blocks in a slice of a picture that has been intra-decoded (I) are encoded using spatial prediction relative to reference samples in adjacent blocks within the same picture. Video blocks in a slice of a picture that has been inter-decoded (P or B) can use spatial prediction relative to reference samples in adjacent blocks within the same picture or temporal prediction relative to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. Summary of the Invention
[0005] In summary, this disclosure describes techniques for decoding video data. Specifically, it describes techniques for encoding and decoding messages (e.g., Supplemental Enhancement Information (SEI) messages and / or other segmented structures) that include metadata that aids in processing (e.g., decoding, displaying, etc.) the video data. The messages in this disclosure may include syntax elements indicating the orientation of an image and / or transformations to be applied to the decoded image, which may be used to rotate and / or mirror the decoded image to a desired orientation. Syntax elements may indicate transformations for the entire image or constituent images (e.g., stereoscopic images of left and right views) used for display. In another example, the message may include syntax elements indicating image quality metrics. Image quality metrics may indicate the encoded quality of the image, such as quality-based view switching and quality-based metric measurements.
[0006] A video decoder or other device can decode the message and process the video data into images based on the message. Image orientation information can be used to provide the video decoder with instructions regarding recommended orientation transformations to be applied to the decoded images. In this way, the decoded images can be displayed in a more appropriate orientation. The video decoder can use quality metrics in post-processing of the decoded images, and / or can use quality metrics to select higher-quality images for inter-frame prediction.
[0007] In one example, this disclosure describes a method for processing video data, the method comprising: receiving an image; and decoding image location information including a transformation type syntax element, wherein the transformation type syntax element indicates a transformation to be applied to the image among a plurality of transformations.
[0008] In another example, this disclosure describes an apparatus configured to process video data, the apparatus comprising: a memory configured to store images; and one or more processors implemented in a circuit and communicating with the memory, the one or more processors being configured to: receive images; and decode image orientation information including a transformation type syntax element, wherein the transformation type syntax element indicates a transformation from among a plurality of transformations to be applied to the image.
[0009] In another example, this disclosure describes an apparatus configured to process video data, the apparatus comprising: a component for receiving an image; and a component for decoding image orientation information including a transformation type syntax element, wherein the transformation type syntax element indicates a transformation to be applied to the image from among a plurality of transformations.
[0010] In another example, this disclosure describes a non-transitory computer-readable storage medium for storing instructions that, when executed, cause one or more processors of a device configured to process video data to: receive an image; and decode an image orientation message including a transformation type syntax element, wherein the transformation type syntax element indicates a transformation from among a plurality of transformations to be applied to the image.
[0011] In another example, this disclosure describes a method for processing video data, the method comprising: receiving an image; and decoding a quality metric message including a quality metric syntax element, wherein the quality metric syntax element indicates a value of a quality metric associated with the image.
[0012] In another example, this disclosure describes an apparatus configured to process video data, the apparatus comprising: a memory configured to store images; and one or more processors implemented in a circuit and communicating with the memory, the one or more processors being configured to: receive images; and decode quality metric messages including quality metric syntax elements, wherein the quality metric syntax elements indicate values of quality metrics associated with the images.
[0013] In another example, this disclosure describes an apparatus configured to process video data, the apparatus comprising: a component for receiving an image; and a component for decoding a quality metric message including a quality metric syntax element, wherein the quality metric syntax element indicates a value of a quality metric associated with the image.
[0014] In another example, this disclosure describes a non-transitory computer-readable storage medium for storing instructions that, when executed, cause one or more processors of a device configured to process video data to: receive an image; and decode a quality metric message including a quality metric syntax element, wherein the quality metric syntax element indicates a value of a quality metric associated with the image.
[0015] Details of one or more examples are set forth in the drawings and the description below. Other features, objects, and advantages will be apparent from the specification, drawings, and claims. Simple Explanation of the Diagram
[0016] Figure 1 is a block diagram illustrating an example video encoding and decoding system that can perform the techniques of this disclosure.
[0017] Figure 2 is a conceptual diagram showing an example of image rotation.
[0018] Figure 3 is a conceptual diagram illustrating an example conversion type.
[0019] Figure 4 is a flowchart illustrating an example program for decoding image orientation enhancement information.
[0020] Figure 5 is a flowchart illustrating an example program for decoding supplementary information messages for quality metrics.
[0021] Figure 6 is a block diagram illustrating an example video encoder that can perform the techniques described in this disclosure.
[0022] Figure 7 is a block diagram illustrating an example video decoder that can perform the techniques of this disclosure.
[0023] Figure 8 is a flowchart illustrating an example method for encoding the current block according to the technology of this disclosure.
[0024] Figure 9 is a flowchart illustrating an example method for decoding the current block according to the technology of this disclosure. Implementation
[0025] This disclosure describes techniques for encoding and decoding messages (e.g., Supplemental Enhancement Information (SEI) messages and / or other segmented structures) including metadata for video data that aids in processing (e.g., decoding, displaying, etc.). The messages of this disclosure may include syntax elements indicating the orientation of an image and / or transformations to be applied to the decoded image, which may be used to rotate and / or mirror the decoded image to a desired orientation. Syntax elements may indicate transformations for the entire image or constituent images (e.g., stereoscopic images of left and right views) used for display. In another example, the message may include syntax elements indicating image quality metrics. Image quality metrics may indicate the encoded quality of the image, such as quality-based view switching and quality-based metric measurements.
[0026] A video decoder or other device can decode the message and process the video data into images based on the message. Image orientation information can be used to provide the video decoder with instructions regarding recommended orientation transformations to be applied to the decoded images. In this way, the decoded images can be displayed in a more suitable orientation. The video decoder can use quality metrics in post-processing of the decoded images, and / or can use quality metrics to select higher-quality images for inter-frame prediction.
[0027] Figure 1 is a block diagram illustrating an example video encoding and decoding system 100 capable of performing the techniques of this disclosure. In summary, the techniques of this disclosure are directed to decoding (encoding and / or decoding) video data. Typically, video data includes any data used for processing video. Therefore, video data can include raw, unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata (such as signaling data).
[0028] As shown in Figure 1, in this example, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116. Specifically, the source device 102 provides the video data to the destination device 116 via computer-readable media 110. The source device 102 and the destination device 116 can include any of a wide variety of devices, including desktop computers, notebook computers (i.e., laptop computers), mobile devices, tablet computers, set-top boxes, mobile phones (such as smartphones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, broadcast receivers, etc. In some cases, the source device 102 and the destination device 116 may be equipped for wireless communication and therefore may be referred to as wireless communication devices.
[0029] In the example of Figure 1, source device 102 includes a video source 104, memory 106, a video encoder 200, and an output interface 108. Destination device 116 includes an input interface 122, a video decoder 300, memory 120, and a display device 118. According to this disclosure, the video encoder 200 of source device 102 and the video decoder 300 of destination device 116 can be configured to apply techniques for SEI message decoding. Therefore, source device 102 represents an example of a video encoding device, and destination device 116 represents an example of a video decoding device. In other examples, the source device and destination device may include other components or arrangements. For example, source device 102 may receive video data from an external video source such as an external camera. Similarly, destination device 116 may interface with an external display device, rather than including an integrated display device.
[0030] System 100, as shown in Figure 1, is merely an example. Typically, any digital video encoding and / or decoding device can perform techniques for SEI message decoding. Source device 102 and destination device 116 are merely examples of such decoding devices, wherein source device 102 generates decoded video data for transmission to destination device 116. In this disclosure, "decoding" device refers to a device that performs the decoding (e.g., encoding and / or decoding) of data. Thus, video encoder 200 and video decoder 300 represent examples of decoding devices (specifically, video encoder and video decoder). In some examples, source device 102 and destination device 116 may operate in a substantially symmetrical manner, such that each of source device 102 and destination device 116 includes video encoding and decoding components. Therefore, system 100 can support one-way or two-way video transmission between source device 102 and destination device 116, for example, for video streaming, video playback, video broadcasting, or video telephony.
[0031] Typically, video source 104 represents the source of video data (i.e., raw, unencoded video data) and provides a series of consecutive pictures (also referred to as “frames”) of the video data to video encoder 200, which encodes the data for the pictures. Video source 104 of source device 102 may include video capturing devices such as cameras, video archives containing previously captured raw video, and / or video feed interfaces for receiving video from video content providers. Alternatively, video source 104 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In each case, video encoder 200 may encode the captured, pre-captured, or computer-generated video data. Video encoder 200 may rearrange the pictures from the received order (sometimes referred to as “display order”) to a decoding order for decoding. Video encoder 200 may generate a bitstream comprising the encoded video data. Then, the source device 102 can output encoded video data to a computer-readable medium 110 via the output interface 108 for reception and / or retrieval by, for example, the input interface 122 of the destination device 116.
[0032] The memory 106 of the source device 102 and the memory 120 of the destination device 116 represent general-purpose memory. In some examples, the memories 106 and 120 may store raw video data, such as raw video from the video source 104 and raw decoded video data from the video decoder 300. Alternatively, the memories 106 and 120 may store software instructions executable by, for example, the video encoder 200 and the video decoder 300. Although the memories 106 and 120 are shown separately from the video encoder 200 and the video decoder 300 in this example, it should be understood that the video encoder 200 and the video decoder 300 may also include internal memory for functionally similar or equivalent purposes. Furthermore, the memories 106 and 120 may store, for example, encoded video data output from the video encoder 200 and input to the video decoder 300. In some examples, portions of memory 106, 120 may be allocated as one or more video buffers, for example, to store raw, decoded and / or encoded video data.
[0033] Computer-readable medium 110 can represent any type of media or device capable of transmitting encoded video data from source device 102 to destination device 116. In one example, computer-readable medium 110 represents communication media that enables source device 102 to transmit encoded video data directly to destination device 116 in real time, for example via a radio frequency network or a computer-based network. According to communication standards such as wireless communication protocols, output interface 108 can modulate the transmitted signal including the encoded video data, and input interface 122 can demodulate the received transmitted signal. Communication media can include any wireless or wired communication media, such as radio frequency (RF) spectrum or one or more physical transmission lines. Communication media can form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. Communication media can include routers, switches, base stations, or any other devices that can facilitate communication from source device 102 to destination device 116.
[0034] In some examples, source device 102 can output encoded data from output interface 108 to storage device 112. Similarly, destination device 116 can access encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a wide variety of distributed or locally accessible data storage media, such as hard disk drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage media for storing encoded video data.
[0035] In some examples, source device 102 may output encoded video data to file server 114 or to another intermediate storage device that may store the encoded video data generated by source device 102. Destination device 116 may access the stored video data from file server 114 via streaming or downloading.
[0036] File server 114 can be any type of server device capable of storing encoded video data and sending the encoded video data to destination device 116. File server 114 can represent a web server (e.g., for a website), a server configured to provide file transfer protocol services (such as File Transfer Protocol (FTP) or FLUTE-based file delivery protocol), a Content Delivery Network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Service (MBMS) or Enhanced MBMS (eMBMS) server, and / or a Network Attached Storage (NAS) device. File server 114 may additionally or alternatively implement one or more HTTP streaming protocols, such as HTTP-based Dynamic Adaptive Streaming (DASH), HTTP Real-Time Streaming (HLS), Real-Time Streaming Protocol (RTSP), HTTP Dynamic Streaming, etc.
[0037] Destination device 116 can access encoded video data from file server 114 via any standard data connection (including an Internet connection). This can include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., Digital Subscriber Line (DSL), cable modems, etc.), or combinations thereof, suitable for accessing encoded video data stored on file server 114. Input interface 122 can be configured to operate according to any one or more of the various protocols discussed above for retrieving or receiving media data from file server 114, or other such protocols for retrieving media data.
[0038] Output interface 108 and input interface 122 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 may be configured to transmit data (such as encoded video data) according to cellular communication standards (such as 4G, 4G-LTE (Long Term Evolution), improved LTE, 5G, etc.). In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 may be configured to transmit data (such as encoded video data) according to other wireless standards (such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee™), the Bluetooth™ standard, etc.). In some examples, source device 102 and / or destination device 116 may include corresponding system-on-a-chip (SoC) devices. For example, source device 102 may include a SoC device for performing functions attributed to video encoder 200 and / or output interface 108, and destination device 116 may include a SoC device for performing functions attributed to video decoder 300 and / or input interface 122.
[0039] The technology disclosed herein can be applied to video decoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission (such as HTTP-based Dynamic Adaptive Streaming (DASH)), digital video encoded onto data storage media, decoding of digital video stored on data storage media, or other applications.
[0040] The input interface 122 of the destination device 116 receives an encoded video bitstream from computer-readable media 110 (e.g., communication media, storage device 112, file server 114, etc.). The encoded video bitstream may include signaling information defined by the video encoder 200, such as syntax elements (which are also used by the video decoder 300), which have values describing the characteristics and / or processing of video blocks or other decoding units (e.g., slices, pictures, picture groups, sequences, etc.). The display device 118 displays a decoded image of the decoded video data to the user. The display device 118 may represent any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.
[0041] Although not shown in Figure 1, in some examples, the video encoder 200 and the video decoder 300 may each be integrated with the audio encoder and / or audio decoder, and may include appropriate MUX-DEMUX units or other hardware and / or software to process multiplexed streams that include both audio and video in a public data stream.
[0042] The video encoder 200 and video decoder 300 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is partially implemented in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the technology of this disclosure. Each of the video encoder 200 and video decoder 300 may be included in one or more encoders or decoders, either of which may be integrated as part of a combined encoder / decoder (CODEC) in the respective device. Devices including the video encoder 200 and / or video decoder 300 may include integrated circuits, microprocessors, and / or wireless communication devices (such as cellular phones).
[0043] The video encoder 200 and video decoder 300 may operate according to video decoding standards such as ITU-T H.265 (also known as High Efficiency Video Decoding (HEVC)) or extensions thereof such as MultiView and / or Scalable Video Decoding Extensions. Alternatively, the video encoder 200 and video decoder 300 may operate according to other proprietary or industry standards such as the ITU-T H.266 standard (also known as Universal Video Decoding (VVC)). In other examples, the video encoder 200 and video decoder 300 may operate according to proprietary video codecs / formats such as AOMedia video 1 (AV1), extensions to AV1, and / or subsequent versions of AV1 (e.g., AV2)). In other examples, the video encoder 200 and video decoder 300 may operate according to other proprietary formats or industry standards. However, the technology of this disclosure is not limited to any particular decoding standard or format. Typically, the video encoder 200 and the video decoder 300 can be configured to perform the techniques of this disclosure in combination with any video decoding technique that uses SEI messages to determine image orientation and / or image quality metrics.
[0044] Typically, video encoder 200 and video decoder 300 can perform block-based decoding of images. The term "block" generally refers to a structure comprising data to be processed (e.g., encoded, decoded, or otherwise used in encoding and / or decoding processes). For example, a block may comprise a two-dimensional matrix of samples of luminance and / or chrominance data. Typically, video encoder 200 and video decoder 300 can decode video data represented in YUV (e.g., Y, Cb, Cr) format. That is, video encoder 200 and video decoder 300 can decode luminance and chrominance components, rather than decoding red, green, and blue (RGB) data samples for an image, where the chrominance components may include both red and blue hue chrominance components. In some examples, video encoder 200 converts the received RGB-formatted data to YUV representation before encoding, and video decoder 300 converts the YUV representation to RGB format. Alternatively, preprocessing and post-processing units (not shown) can perform these conversions.
[0045] In summary, this disclosure may relate to the decoding (e.g., encoding and decoding) of images to include procedures for encoding or decoding data of an image. Similarly, this disclosure may relate to the decoding of blocks of an image to include procedures for encoding or decoding data used for the blocks (e.g., prediction and / or residual decoding). Encoded video bitstreams typically include a series of values for representing decoding decisions (e.g., decoding modes) and syntax elements that segment the image into blocks. Therefore, references to decoding an image or block should generally be understood as decoding the values of the syntax elements used to form the image or block.
[0046] HEVC defines various blocks, including decoding units (CUs), prediction units (PUs), and transformation units (TUs). According to HEVC, a video decoder (such as a video encoder 200) partitions the decoding tree unit (CTU) into CUs based on a quadtree structure. That is, the video decoder partitions the CTU and CU into four equal, non-overlapping squares, and each node of the quadtree has zero or four child nodes. Nodes without child nodes can be called "leaf nodes," and the CU of such leaf nodes can include one or more PUs and / or one or more TUs. The video decoder can further partition the PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents a partition of the TU. In HEVC, PUs represent inter-frame prediction data, while TUs represent residual data. Intra-predicted CUs include intra-frame prediction information, such as intra-frame mode indication.
[0047] As another example, video encoder 200 and video decoder 300 can be configured to operate according to VVC. According to VVC, the video decoder (such as video encoder 200) segments the image into multiple decoding tree units (CTUs). Video encoder 200 can segment CTUs according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure eliminates the concept of multiple segmentation types, such as the separation between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels: a first level of segmentation based on quadtree segmentation and a second level of segmentation based on binary tree segmentation. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to decoding units (CUs).
[0048] In the MTT partitioning structure, blocks can be partitioned using quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) partitioning (also known as triplet tree (TT)) partitioning. A ternary tree or triplet tree partitioning is a partition in which a block is divided into three sub-blocks. In some examples, a ternary tree or triplet tree partitioning divides a block into three sub-blocks without partitioning the original block through a center. The partitioning types in MTT (e.g., QT, BT, and TT) can be symmetric or asymmetric.
[0049] When operating according to the AV1 codec, the video encoder 200 and video decoder 300 can be configured to decode video data in blocks. In AV1, the largest decoded block that can be processed is called a superblock. In AV1, a superblock can be 128x128 luma samples or 64x64 luma samples. However, in subsequent video decoding formats (e.g., AV2), superblocks can be defined by different (e.g., larger) luma sample sizes. In some examples, the superblock is the top level of a block quadtree. The video encoder 200 can further subdivide the superblock into smaller decoded blocks. The video encoder 200 can use square or non-square partitioning to divide the superblock and other decoded blocks into smaller blocks. Non-square blocks can include N / 2xN, NxN / 2, N / 4xN, and NxN / 4 blocks. The video encoder 200 and video decoder 300 can perform separate prediction and conversion procedures for each of the decoded blocks.
[0050] AV1 also defines video data tiles. A tile is a rectangular array of superblocks that can be decoded independently of other tiles. That is, the video encoder 200 and video decoder 300 can encode and decode the decoding blocks within a tile separately, without using video data from other tiles. However, the video encoder 200 and video decoder 300 can perform filtering across tile boundaries. Tiles can be uniform or non-uniform in size. Tile-based decoding enables parallel processing and / or multithreading for encoder and decoder implementations.
[0051] In some examples, the video encoder 200 and the video decoder 300 may use a single QTBT or MTT structure to represent each of the luma and chroma components, while in other examples, the video encoder 200 and the video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luma component and another QTBT / MTT structure for the two chroma components (or two QTBT / MTT structures for the respective chroma components).
[0052] The video encoder 200 and the video decoder 300 can be configured to use quadtree segmentation, QTBT segmentation, MTT segmentation, superblock segmentation or other segmentation structures.
[0053] In some examples, a CTU includes a decoded tree block (CTB) of luminance samples for an image with three sample arrays, two corresponding CTBs of chrominance samples, or a CTB of samples for a monochrome image or an image decoded using three separate color planes, and a syntax structure for decoding the samples. A CTB can be an NxN sample block for some value of N, such that dividing the components into CTBs is a partition. Components are arrays or single samples from one of the three arrays (luminance and two chrominance) constituting the image in a 4:2:0, 4:2:2, or 4:4:4 color format, or an array constituting the image in monochrome format or a single sample of that array. In some examples, a decoded block is an MxN sample block for some values of M and N, such that dividing the CTB into decoded blocks is a partition.
[0054] Blocks (e.g., CTUs or CUs) can be grouped in various ways within an image. As an example, a brick can refer to a rectangular area of a row of CTUs within a specific tile in an image. A tile can be a rectangular area of a CTU within a specific tile column or row in an image. A tile row refers to a rectangular area of a CTU with a height equal to the height of the image and a width specified by syntax elements (e.g., as in an image parameter set). A tile column refers to a rectangular area of a CTU with a height specified by syntax elements (e.g., as in an image parameter set) and a width equal to the width of the image.
[0055] In some examples, a tile can be divided into multiple brick-shaped regions, each of which can include one or more CTU columns within the tile. Tiles not divided into multiple brick-shaped regions can still be referred to as brick-shaped regions. However, brick-shaped regions that are a true subset of tiles may not be referred to as tiles. Brick-shaped regions in an image can also be arranged in slices. A slice can be an integer number of brick-shaped regions in the image, which can be uniquely contained within a single Network Abstraction Layer (NAL) unit. In some examples, a slice includes several complete tiles or a continuous sequence of complete brick-shaped regions of a single tile.
[0056] This disclosure uses “NxN” and “N by N” interchangeably to refer to the sample size of a block (such as a CU or other video block) in the vertical and horizontal dimensions, for example, 16x16 samples or 16 by 16 samples. Typically, a 16x16 CU will have 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an NxN CU typically has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. Samples in a CU can be arranged in rows and columns. Furthermore, a CU does not necessarily need to have the same number of samples in the horizontal direction as it does in the vertical direction. For example, a CU can include NxM samples, where M is not necessarily equal to N.
[0057] The video encoder 200 encodes video data for use in predicting and / or residual information, as well as other information, for the CU. Prediction information indicates how the CU will be predicted to form a prediction block for the CU. Residual information typically represents the sample-by-sample difference between a sample of the CU before encoding and the prediction block.
[0058] To predict the Cubic Frame (CU), the video encoder 200 typically forms prediction blocks for the CU through inter-frame prediction or intra-frame prediction. Inter-frame prediction typically refers to predicting the CU based on data from previously decoded images, while intra-frame prediction typically refers to predicting the CU based on data from previously decoded images of the same frame. To perform inter-frame prediction, the video encoder 200 can use one or more motion vectors to generate prediction blocks. The video encoder 200 can typically perform motion search to identify, for example, reference blocks that closely match the CU in terms of the difference between the CU and a reference block. The video encoder 200 can use sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to calculate a difference metric to determine whether the reference block closely matches the current CU. In some examples, the video encoder 200 can use unidirectional or bidirectional prediction to predict the current CU.
[0059] Some examples of VVC also provide an affine motion compensation mode, which can be considered an inter-frame prediction mode. In affine motion compensation mode, the video encoder 200 can determine two or more motion vectors representing non-translational motion (such as zooming in or out, rotation, perspective motion, or other irregular types of motion).
[0060] To perform intra-frame prediction, the video encoder 200 can select an intra-frame prediction mode to generate prediction blocks. Some examples of VVC provide sixty-seven intra-frame prediction modes, including various directional modes, as well as planar and DC modes. Typically, the video encoder 200 selects an intra-frame prediction mode that describes the samples adjacent to the current block (e.g., the block of the CU), and predicts samples for the current block based on these adjacent samples. Assuming the video encoder 200 decodes the CTU and CU in raster scan order (left to right, top to bottom), such samples can typically be located above, to the upper left, or to the left of the current block within the same image.
[0061] The video encoder 200 encodes data representing the prediction mode used for the current block. For example, for inter-frame prediction modes, the video encoder 200 may encode data indicating which of the various available inter-frame prediction modes is used, as well as motion information for the corresponding mode. For unidirectional or bidirectional inter-frame prediction, for example, the video encoder 200 may use Advanced Motion Vector Prediction (AMVP) or merging modes to encode motion vectors. The video encoder 200 may use similar modes to encode motion vectors used for affine motion compensation modes.
[0062] AV1 includes two common techniques for encoding and decoding decoded blocks of video data. These two common techniques are intra-frame prediction (e.g., intra-frame prediction or spatial prediction) and inter-frame prediction (e.g., inter-frame prediction or temporal prediction). In the context of AV1, when using intra-frame prediction modes to predict blocks of the current frame of video data, the video encoder 200 and video decoder 300 do not use video data from other frames of the video data. For most intra-frame prediction modes, the video encoder 200 encodes the blocks of the current frame based on the difference between sample values in the current block and predicted values generated based on reference samples in the same frame. The video encoder 200 determines the predicted values generated based on reference samples using the intra-frame prediction mode.
[0063] Following a prediction, such as intra-frame or inter-frame prediction of a block, the video encoder 200 can compute residual data for that block. The residual data (such as a residual block) represents the sample-by-sample difference between the block and the prediction block used to form that block using the corresponding prediction mode. The video encoder 200 can apply one or more transformations to the residual block to produce transformed data in the transform domain rather than the sample domain. For example, the video encoder 200 can apply Discrete Cosine Transform (DCT), integer transformation, wavelet transformation, or conceptually similar transformations to the residual video data. Additionally, the video encoder 200 can apply a secondary transformation after the first transformation, such as Mode-dependent Inseparable Secondary Transform (MDNSST), Signal-dependent Transform, Karhunen-Loeve Transform (KLT), etc. The video encoder 200 generates transformation coefficients after applying one or more transformations.
[0064] As described above, after any transformation used to generate the conversion coefficients, the video encoder 200 can perform quantization on the conversion coefficients. Quantization generally refers to a procedure in which the conversion coefficients are quantized to reduce the amount of data used to represent them as much as possible, thereby providing further compression. By performing the quantization procedure, the video encoder 200 can reduce the bit depth associated with some or all of the conversion coefficients. For example, the video encoder 200 can round an n-bit value down to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder 200 can perform a bitwise right shift of the value to be quantized.
[0065] After quantization, the video encoder 200 can scan the transform coefficients to generate a one-dimensional vector from a two-dimensional matrix including the quantized transform coefficients. The scan can be designed to place higher-energy (and therefore lower-frequency) transform coefficients before the vector and lower-energy (and therefore higher-frequency) transform coefficients after the vector. In some examples, the video encoder 200 can utilize a predefined scan order to scan the quantized transform coefficients to produce a serialized vector, and then entropy-encode the quantized transform coefficients of the vector. In other examples, the video encoder 200 can perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 200 can entropy-encode the one-dimensional vector, for example, according to context-adaptive binary arithmetic decoding (CABAC). The video encoder 200 can also entropy-encode the values of syntax elements used to describe metadata associated with the encoded video data for use by the video decoder 300 when decoding the video data.
[0066] To perform CABAC, the video encoder 200 can assign context within a context model to the symbols to be transmitted. Context may involve, for example, whether the neighboring values of a symbol are zero. Probability decisions can be based on the context assigned to the symbol.
[0067] The video encoder 200 can also generate, for example, syntax data (such as block-based syntax data, image-based syntax data, and sequence-based syntax data) or other syntax data (such as sequence parameter sets (SPS), image parameter sets (PPS), or video parameter sets (VPS)) destined for the video decoder 300 in image headers, block headers, and slice headers. Similarly, the video decoder 300 can decode such syntax data to determine how to decode the corresponding video data.
[0068] In this way, the video encoder 200 can generate a bitstream that includes encoded video data, such as syntax elements describing the segmentation of an image into blocks (e.g., CUs) and prediction and / or residual information for those blocks. Finally, the video decoder 300 can receive the bitstream and decode the encoded video data.
[0069] Typically, the video decoder 300 executes a program that is the inverse of the program executed by the video encoder 200 to decode the encoded video data of the bitstream. For example, the video decoder 300 can use CABAC to decode the values of syntax elements for the bitstream in a manner substantially similar to, but inverse of, the CABAC encoding program of the video encoder 200. Syntax elements can define segmentation information for segmenting a picture into CTUs and for segmenting each CTU according to a corresponding segmentation structure (such as a QTBT structure) to define the CUs of the CTU. Syntax elements can also define prediction and residual information for blocks (e.g., CUs) of the video data.
[0070] The residual information can be represented, for example, by quantized transformation coefficients. The video decoder 300 can inversely quantize and inversely transform the quantized transformation coefficients of a block to regenerate a residual block for that block. The video decoder 300 uses a signal-informed prediction mode (intra-frame prediction or inter-frame prediction) and associated prediction information (e.g., motion information for inter-frame prediction) to form a prediction block for that block. The video decoder 300 can then combine the prediction block and the residual block (on a sample-by-sample basis) to regenerate the original block. The video decoder 300 can perform additional processing, such as performing a deblocking procedure to reduce visual artifacts along the block boundaries.
[0071] In summary, this disclosure may involve "signaling" certain information (such as syntax elements). The term "signaling" can generally refer to the transmission of values for syntax elements and / or other data for decoding encoded video data. That is, the video encoder 200 can signal values for syntax elements in the bitstream. Generally, signaling refers to the generation of values in the bitstream. As described above, the source device 102 can transmit the bitstream to the destination device 116 substantially immediately or not immediately (such as when syntax elements are stored in the storage device 112 for later retrieval by the destination device 116).
[0072] In summary, this disclosure describes techniques for decoding video data. Specifically, it describes techniques for decoding SEI messages. The SEI message of this disclosure may include syntax elements indicating the orientation of an image. In another example, the SEI message may include syntax elements indicating a measure of image quality. A video decoder or other device may decode the SEI message and process images from the video data based on the SEI message.
[0073] The General Supplemental Enhancement Information (VSEI) standards (e.g., ITU-T H.274 and ISO / IEC 23002-7) specify some SEI messages in the Video Availability Information (VUI) message and the SEI messages used with the VVC bitstream. SEI messages enable the video encoder 200 to include metadata in the bitstream that is not necessary for the correct decoding of sample values of the output image but can be used for various other purposes. The video encoder 200 can be configured to include any number of SEI Network Abstraction Layer (NAL) units in its access unit, and each SEI NAL unit can include one or more SEI messages. Specifications and systems using VVC can specify the encoder to generate specific SEI messages or define specific processing for specific types of received SEI messages.
[0074] The following document specifies the display of the orientation SEI message to inform the decoder (e.g., video decoder 300) of the conversion recommended to be applied to the cropped decoded image before display: ISO / IEC JTC 1 / SC 29 / WG 11 N 18277, “Information technology — High efficiency coding and media delivery in heterogeneous environments — Part 2: High Efficiency Video Coding”, 2019 (“HEVC”). The syntax structure for displaying the orientation SEI message in HEVC is shown in Table 1 below.
[0075] Table 1 shows the syntax of the SEI (Search Engine Information) message. display_orientation( payloadSize ) { [Descriptor] [display_orientation_cancel_flag] u(1) if( !display_orientation_cancel_flag ) { [hor_flip] u(1) [ver_flip] u(1) [anticlockwise_rotation] u(16) [display_orientation_persistence_flag] u(1) } }
[0076] As can be seen in Table 1, the HEVC display orientation SEI message allows indications for horizontal flip (hor_flip), vertical flip (ver_flip), and anticlockwise rotation (anticlockwise_rotation) conversions.
[0077] 3GPP specifies Coordination of Video Orientation (CVO) in the following document: Technical Specification (TS) 26.114, “IP Multimedia Subsystem (IMS); Multimedia telephony; Media handling and interaction”, 2021. CVO signals the current orientation of the image captured at the transmitting end (e.g., at source device 102) to the receiver (e.g., destination device 116) for appropriate rendering and display. CVO information for lower rotation granularity is carried in bytes in the following format to support horizontal flipping and 90-degree rotation: Number of bits: 7 6 5 4 3 2 1 0 (LSB) Define 0 0 0 0 CF R1 R0 LSB represents the least significant bit.
[0078] CVO information for higher rotational granularity is carried in bytes in the following format: Number of bits: 7 6 5 4 3 2 1 0 (LSB) Define R5 R4 R3 R2 CF R1 R0
[0079] Some current examples of the VSEI standard do not support any orientation metadata. HEVC displays orientation SEI messages without considering frame encapsulation, where rotation should be applied to each component image rather than the entire image. Figure 2 shows an example of display rotation for a frame-encapsulated image, where each component image should be rotated separately. As shown in Figure 2, image 150 includes two component images (e.g., a left view image and a right view image for stereoscopic video). Using the techniques of this disclosure, video encoder 200 can send a code including a transformation type syntax element and an SEI message that instructs video decoder 300 to perform a rotation transformation on each component image to achieve the transformed image 152.
[0080] The example VSEI Region-Wrapped (RWP) SEI message provides information for remapping color samples from a cropped decoded image onto the projected image. However, the RWP SEI message is used when omnidirectional video projection is instructed to be applied to an image. An RWP SEI message with an rwp_cancel_flag equal to 0 should not exist in the decoded layer video sequence (CLVS) applied to the image.
[0081] Image quality metrics are used to evaluate image quality and decoding performance. The following document specifies the carrying of timed metadata metrics for media (such as Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), Video Quality Metrics (VQM), and Mean Opinion Score (MOS) in ISO BMFF (ISO / IEC Basic Media File Format): ISO / IEC 23001-10, “Information technology — MPEG systems technologies — Part 10: Carriage of Timed Metadata Metrics of Media in ISO Base Media File Format”, 2015. The following documents also specify image quality-related sorting to facilitate quality-related viewpoint switching and immersive media metrics: OMAF, ISO / IEC JTC1 / SC29 / WG11 N19042, “Text of ISO / IEC DIS 23090-2 2nd edition OMAF”, 2020; and Immersive Media Metrics (IMM), ISO / IEC JTC1 / SC29 / WG3 N0073, “IS of ISO / IEC 23090-6 Immersive Media Metrics”, 2020. Some image quality metrics (such as PSNR and SSIM) can be obtained only on the encoder side. SEI messages carrying this information can provide relevant information to system applications.
[0082] Image location SEI message
[0083] According to one example of this disclosure, video encoder 200 is configured to generate and signal a Picture Orientation SEI message, the Picture Orientation SEI message including one or more syntax elements shown in Table 2 below. Specifically, video encoder 200 may be configured to generate and encode a transformation type syntax element (e.g., por_transform_type), wherein the transformation type syntax element indicates a transformation from among multiple transformations to be applied to the picture. Video encoder 200 may also be configured to generate and encode one or more of the other syntax elements and flags listed in Table 2. Video decoder 300 may be configured to receive the Picture Orientation SEI message and may process and / or display the picture according to the syntax elements contained therein. For example, video decoder 300 may be configured to apply the transformation indicated by the transformation type syntax element to the decoded picture.
[0084] Table 2. Image Orientation SEI Message Syntax picture_orientation( payloadSize ) { [Descriptor] [por_cancel_flag] u(1) if (!por_cancel_flag) { [por_persistence_flag] u(1) [por] [_] [constituent_picture_matching_flag] u(1) [por_transform_type] u(5) } }
[0085] Typically, the Picture Orientation (POR) SEI message provides information to the video decoder 300 about the transformations recommended to be applied to the decoded image before display. In some examples, the decoded image may be a cropped image.
[0086] The syntax element `por_cancel_flag` having a value of 1 indicates that the current SEI message cancels the persistence of any previous POR SEI messages in the order they were output. The syntax element `por_cancel_flag` having a value of 0 indicates that a POR message follows.
[0087] The value of the syntax element por_persistence_flag specifies the persistence of the POR SEI message used for the current layer.
[0088] The value of the syntax element por_persistence_flag being equal to 0 indicates that the POR SEI message applies only to the currently decoded image.
[0089] The syntax element `por_persistence_flag` having a value of 1 specifies that the POR SEI message applies to the currently decoded image and continues for all subsequent images of the current layer in output order until one or more of the following conditions are true: – A new CLV begins for the current layer. – End of bitstream. - Outputs the image in the current layer of the access unit (AU) associated with the POR SEI message, following the current image in the output order.
[0090] A value of 1 for the syntax element `por_constituent_picture_matching_flag` indicates that the SEI message applies individually to each component picture, and the stereo frame encapsulation format is indicated by the frame encapsulation layout SEI message. A value of 0 for the syntax element `por_component_picture_matching_flag` indicates that the SEI message applies to cropped decoded pictures.
[0091] The value of the syntax element por_constituent_picture_matching_flag should be equal to 0 if any of the following conditions are true: – StereoFlag equals 0. – StereoFlag equals 1 and fp_arrangement_type equals 5.
[0092] A StereoFlag value of 0 indicates that there is no frame encapsulation arrangement SEI message with fp_arrangement_cancel_flag equal to 0 applicable to the image. A StereoFlag value of 1 indicates that the associated image is a frame-encapsulated image.
[0093] The value of the syntax element fp_arrangement_type equal to 5 indicates that the component planes of the cropped decoded images output in the output order form alternating first and second component frames with temporal interleaving.
[0094] The value of the syntax element `por_transform_type` specifies the transformations that can be applied to the image (e.g., rotation, mirroring, or a combination of rotation and mirroring). Note that in some examples, mirroring can be referred to as flipping. When the transformation indicated by `por_transform_type` specifies both rotation and mirroring, the video decoder 300 can be configured to apply the rotation transformation before applying the mirroring, or vice versa. Example values for `por_transform_type` are specified in Table 3 below. In one example, values of `por_transform_type` from 8 to 31 are reserved for future use by ITU-T|ISO / IEC.
[0095] Table 3 por_transform_type values [value] [describe] 0 No conversion 1 Horizontal mirror 2 Rotate 180 degrees (counter-clockwise) 3 Rotate 180 degrees (counter-clockwise) before mirroring horizontally. 4 Rotate 90 degrees (counter-clockwise) before mirroring horizontally. 5 Rotate 90 degrees (counterclockwise) 6 Rotate 270 degrees (counter-clockwise) before mirroring horizontally. 7 Rotate 270 degrees (counterclockwise) 8..31 Reserved
[0096] The specific values in Table 3 are merely examples. In other examples, more or fewer transformation types may be specified. Furthermore, transformations may be specified in an order different from that shown in Table 3.
[0097] Figure 3 is a conceptual diagram illustrating example transformation types. In the examples in Table 3, when the transformation syntax element has a value of 0, the video decoder 300 may not apply a transformation. Figure 3 shows the original image 160 with no transformation applied. Other transformation types in Figure 3 will be shown with reference to the original image 160. When the transformation syntax element has a value of 1, the video decoder 300 may apply a horizontal mirror transformation to the original image 160 to obtain image 162. Horizontal mirroring can also be referred to as horizontal flipping. When the transformation syntax element has a value of 2, the video decoder 300 may apply a 180-degree counterclockwise rotation transformation to the original image 160 to obtain image 164. When the transformation syntax element has a value of 3, the video decoder 300 may apply a 180-degree counterclockwise rotation transformation to the original image 160, followed by a horizontal mirror transformation, to obtain image 166.
[0098] When the transformation syntax element has a value of 4, the video decoder 300 can apply a 90-degree counterclockwise rotation transformation to the original image 160, followed by a horizontal mirror transformation, to obtain image 168. When the transformation syntax element has a value of 5, the video decoder 300 can apply a 90-degree counterclockwise rotation transformation to the original image 160 to obtain image 170. When the transformation syntax element has a value of 6, the video decoder 300 can apply a 270-degree counterclockwise rotation transformation to the original image 160, followed by a horizontal mirror transformation, to obtain image 172. When the transformation syntax element has a value of 7, the video decoder 300 can apply a 270-degree counterclockwise rotation transformation to the original image 160 to obtain image 174.
[0099] Figure 4 is a flowchart illustrating an example procedure for decoding image orientation enhancement information. Figure 4 shows both the encoding and decoding procedures of this disclosure. As shown in Figure 4, the encoding process can be performed by a source device 102 including a video encoder 200. The decoding process can be performed by a destination device 116 including a video decoder 300.
[0100] In one example of this disclosure, source device 102 may be configured to receive an image (400). Source device 102 may also be configured (e.g., using video encoder 200) to encode the image and send the encoded video bitstream to destination device 116. Source device 102 may also be configured to determine a recommended conversion type for the image (402). The recommended conversion type may be one of a plurality of conversion types. Source device 102 may also be configured to encode image orientation information including a conversion type syntax element, wherein the conversion type syntax element indicates one of a plurality of conversions to be applied to the image (404).
[0101] Destination device 116 can be configured to receive an image (410). Destination device 116 can also be configured to decode the image (e.g., using video decoder 300). Destination device 116 can also decode image orientation information including a transformation type syntax element, wherein the transformation type syntax element indicates a transformation (412) from among multiple transformations to be applied to the image. Destination device 116 can also be configured to apply a transformation to the image according to the transformation type syntax element to form a transformed image (414), and display the transformed image (416).
[0102] In one example of this disclosure, the image orientation information includes an image orientation SEI message. In another example, the image orientation information includes an image orientation open bitstream unit (OBU).
[0103] As described above, multiple transformations include two or more of rotation transformations, mirror transformations, or combinations of rotation and mirror transformations. In a more specific example, multiple transformations include: a first transformation that includes a horizontal mirror transformation; a second transformation that includes a 180-degree counterclockwise rotation transformation; a third transformation that includes a 180-degree counterclockwise transformation and a subsequent horizontal mirror transformation; a fourth transformation that includes a 90-degree counterclockwise transformation and a subsequent horizontal mirror transformation; a fifth transformation that includes a 90-degree counterclockwise transformation; a sixth transformation that includes a 270-degree counterclockwise transformation and a subsequent horizontal mirror transformation; and a seventh transformation that includes a 270-degree counterclockwise transformation. In another example, the transformation type syntax element also includes a value indicating that no transformation will be applied.
[0104] In some examples, the video encoder 200 can signal high-granularity rotation in the SEI message. High-granularity rotation and mirroring can be applied to each component image. Typically, high-granularity rotation can indicate the degree of rotation at relatively small intervals. For example, high-granularity rotation can include rotating the image at angles less than 90 degrees. Where each component image can be rotated differently, the video encoder 200 can specify a separate orientation transformation type or high-granularity rotation in the SEI message or other metadata type, each applied to a component image.
[0105] In some examples, the video encoder 200 can specify a component image matching flag in the CVO signaling. The component image matching flag can indicate the rotation granularity applied to each component image.
[0106] Image Quality Measurement (SEI) Message
[0107] According to another example of this disclosure, video encoder 200 is configured to generate a picture quality metric message (e.g., SEI message and / or other encapsulated structure) including one or more syntax elements shown in Table 4 below and to signal it. Video decoder 300 is configured to receive the picture quality metric SEI message and can process and / or display the picture according to the syntax elements contained therein. For example, destination device 116 and / or video decoder 300 may be configured to apply one or more post-processing techniques to the decoded picture according to the quality metric indicated in the picture quality metric SEI message. Example post-processing techniques may include upscaling the decoded picture based on picture quality. In other examples, video decoder 300 may be configured to use the quality metric to select certain pictures for inter-frame prediction. For example, when multiple versions of the same picture are available, video decoder 300 may be configured to select the picture with the highest quality metric (e.g., lowest signal-to-noise ratio) as a reference picture in inter-frame prediction.
[0108] Table 4 shows an example Image Quality Indicator (SEI) message. The Image Quality Indicator (SEI) message provides a quality metric for each color component of the currently decoded image.
[0109] Table 4 Image Quality Measurement SEI Message Syntax Picture_quality_metrics( payloadSize ) { [Descriptor] [pqm_metric_type] u(7) [pqm_single_component_flag] u(1) for( cIdx = 0; cIdx < ( dph_sei_single_component_flag ? 1 : 3 ); cIdx++ ){ if( pqm_sei_type == 0 ) [pqm] [_] [psnr][ cIdx] u(16) else if ( pqm_sei_type == 1 ) [pqm_ssim][ cIdx ] u(8) else if ( pqm_sei_type == 2 ) [pqm_msssim][ cIdx ] u(8) else if ( pqm_sei_type == 3 ) [pqm_vqm][ cIdx ] u(8) } }
[0110] The value of the syntax element pqm_metric_type indicates the type of quality metric associated with the component specified in Table 5. Values of pqm_metric_type from 4 to 127 are reserved for future use by ITU-T | ISO / IEC and should not exist in the payload data conforming to this version of the specification.
[0111] Table 5 Explanation of pqm_metric_type [pqm_metric_type] [measure] 0 PSNR 1 SSIM 2 MS-SSIM 3 VQM
[0112] PSNR quality metric is Peak Signal-to-Noise Ratio. SSIM quality metric is Structural Similarity Index. MS-SSIM quality metric is Multi-Scale Structural Similarity Index. VQM quality metric is Video Quality Metric.
[0113] The value of the syntax element `pqm_single_component_flag` equal to 1 indicates that the image associated with the Image Quality Indicator (SEI) message contains a single color component. The value of the syntax element `pqm_single_component_flag` equal to 0 indicates that the image associated with the SEI message contains three color components. The value of `pqm_single_component_flag` should be equal to (ChromaFormatIdc == 0).
[0114] The value of the syntax element pqm_psnr[ cIdx ] specifies the PSNR value. The PSNR corresponding to the color component cIdx of the decoded image is derived as follows (in floating-point): PSNR = pqm_psnr[cIdx] / 100; except when pqm_psnr[cIdx] equals 0, PSNR = infinity.
[0115] The value of the syntax element pqm_ssim[ cIdx ] specifies the SSIM value. The corresponding SSIM for the color component cIdx of the decoded image is derived as follows (in floating-point): SSIM = (pqm_ssim[cIdx] – 127) / 128
[0116] The value of the syntax element pqm_msssim[ cIdx ] specifies the MS-SSIM value. The corresponding MS-SSIM for the color component cIdx of the decoded image is derived as follows (in floating-point): MS SSIM = ( pqm_msssim[ cIdx ] – 127 ) / 128
[0117] The value of the syntax element pqm_vqm[ cIdx ] specifies the VQM value. The VQM corresponding to the color component cIdx of the decoded image is derived as follows (in floating-point): VQM = pqm_vqm[ cIdx ] / 50
[0118] Image quality metrics (SEI) messages can carry other quality-related metrics, such as perceptual evaluation of video quality (PEVQ), mean opinion score (MOS), and / or other image quality metrics.
[0119] In some examples, when an image is associated with a stereo frame encapsulation layout SEI message, the image quality metric SEI message can specify a quality metric for each component image. When an image is associated with a region-based frame encapsulation SEI message, the image quality metric SEI message can specify a quality metric for each region. Additional syntax elements can be added to the SEI message to indicate whether an image quality metric exists for each component image or each region.
[0120] In other examples, the Image Quality Indicator (SEI) message can carry quality metrics for one or more sub-images or regions of interest (ROIs) of the image associated with the SEI message. Syntax elements indicating the number of sub-images or ROIs, the location of sub-images or ROIs, and / or the size of sub-images or ROIs can also be specified in the SEI message.
[0121] Additional quality metrics, such as weighted PSNR (wPSNR) and weighted to spherical uniform PSNR (WS-PSNR), can be included in the SEI message to indicate the quality of high dynamic range (HDR) and 360 video content.
[0122] Table 6 provides another example of the SEI message format for image metrics: Table 6 Image Quality Measurement SEI Message Syntax Picture_quality_metrics( payloadSize ) { [Descriptor] [pqm_cnt_minus1] u(8) for( i = 0; i <= pqm_cnt_minus1; i++ ) { [pqm_type][i] u(8) [pqm_value][i] u(16) } }
[0123] The Image Quality Indicator (SEI) message above provides a quality metric for the currently decoded image.
[0124] The value of the syntax element pqm_cnt_minus1 plus 1 specifies the number of luminance component quality measures indicated by the SEI message.
[0125] The value of the syntax element pqm_type[i] indicates the i-th quality metric type associated with the decoded picture or video sequence as specified in Table 7. Table 7 Explanation of pqm_type [pqm_type] [measure] 0 PSNR 1 wPSNR 2 WS-PSNR 3 PSNR sequence 4 wPSNR sequence 5 WS-PSNR sequence
[0126] PSNR sequence, wPSNR sequence, and WS-PSNR sequence quality metric types indicate the PSNR, wPSNR, and WS-PSNR of multiple images in a sequence, respectively.
[0127] The value of the syntax element pqm_value[i] specifies the value of the i-th quality metric. When the value of the syntax element pqm_type is 0, the stored 16-bit unsigned integer pqm_value is interpreted as a PSNR value (in dB), as shown below (in floating point), except that for a pqm_value equal to 0, the PSNR is equal to infinity. Where M is an integer (e.g., 100).
[0128] When the value of the syntax element pqm_type is 1, the stored 16-bit unsigned integer pqm_value is interpreted as a wPSNR value (in dB), as shown below (in floating point), except that for a pqm_value equal to 0, wPSNR is equal to infinity. Where M is an integer (e.g., 100).
[0129] When the value of the syntax element pqm_type is 2, the stored 16-bit unsigned integer pqm_value is interpreted as a WS-PSNR value (in dB), as shown below (in floating point), except that for a pqm_value equal to 0, WS-PSNR is equal to infinity. Where M is an integer (e.g., 100).
[0130] When the value of the syntax element pqm_type is 3, the quality metric indicates the average brightness PSNR of the associated image within the CLVS. The 16-bit unsigned integer pqm_value is interpreted as the result of the sequence-level PSNR quality metric (in dB), and is deduced as follows (in floating-point terms), except that for a pqm_value of 0, the PSNR equals infinity. Where M is an integer (e.g., 100).
[0131] When the value of the syntax element pqm_type is 4, the quality metric indicates the average luminance-weighted PSNR of the associated image's CLVS. The 16-bit unsigned integer pqm_value is interpreted as a sequence-level wPSNR value (in dB), as shown below (in floating-point), except that for a pqm_value of 0, wPSNR equals infinity. Where M is an integer (e.g., 100).
[0132] When the value of the syntax element pqm_type is 5, the quality metric indicates the average WS-PSNR of the CLVS to which the associated image belongs. The 16-bit unsigned integer pqm_value is interpreted as a sequence-level WS-PSNR value (in dB), as shown below (in floating-point), except that for a pqm_value of 0, WS-PSNR equals infinity. Where M is an integer (e.g., 100).
[0133] In another example, an additional quality metric type can be included in the SEI message to indicate the average quality metric applicable to multiple video frames. A first syntax element can be specified in the SEI message to indicate that the quality metric specified in the SEI message applies to the associated picture and persists for all subsequent pictures of the current layer in output order. A second syntax element can be specified in the SEI message to cancel the persistence of any previous quality metrics in output order.
[0134] In another example, when an average quality metric (such as sequence-level PSNR) exists for any image in the CLVS, the first image in the CLVS should have an associated image quality SEI message. The average image metric for all SEI messages applicable to the same CLVS should have the same content.
[0135] Figure 5 is a flowchart illustrating an example procedure for decoding supplemental enhancement information messages for quality metrics. Figure 5 shows both the encoding and decoding procedures of this disclosure. As shown in Figure 5, the encoding procedure can be performed by a source device 102 including a video encoder 200. The decoding procedure can be performed by a destination device 116 including a video decoder 300.
[0136] In one example of this disclosure, source device 102 may be configured to receive an image (500). Source device 102 may also be configured (e.g., using video encoder 200) to encode the image and send the encoded video bitstream to destination device 116. Source device 102 may also be configured to determine a quality metric for the image (502). Source device 102 may also be configured to encode a quality metric message including quality metric syntax elements, wherein the quality metric syntax elements indicate the value of a quality metric associated with the image (504).
[0137] In another example of this disclosure, the source device 102 may also be configured to encode a quality metric type syntax element in a quality metric message, wherein the quality metric type syntax element indicates a quality metric of a type indicated by the quality metric syntax element among multiple types of quality metrics. In one example, the multiple types of quality metrics include Peak Signal-to-Noise Ratio (PSNR). In another example, the multiple types of quality metrics include two additional items from the following: Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), Multi-Scale Structural Similarity Index (MS-SSIM), Video Quality Metric (VQM), Weighted PSNR (wPSNR), Weighted to Spherical Uniform PSNR (WS-PSNR), Sequence PSNR, Sequence wPSNR, or Sequence WS-PSNR. In the examples above, the quality metric syntax element indicates the value of the quality metric indicated by the quality metric type syntax element.
[0138] Destination device 116 can be configured to receive an image (510). Destination device 116 can also be configured to decode the image (e.g., using video decoder 300). Destination device 116 can also decode a quality metric message including a quality metric syntax element, wherein the quality metric syntax element indicates the value of a quality metric associated with the image (512). Destination device 116 can also be configured to apply post-processing techniques to the image based on the value of the quality metric to form a processed image (514), and display the processed image (516).
[0139] In another example of this disclosure, the destination device 116 may also be configured to decode a quality metric type syntax element in a quality metric message, wherein the quality metric type syntax element indicates a quality metric of a type indicated by the quality metric syntax element among multiple types of quality metrics. In one example, the multiple types of quality metrics include Peak Signal-to-Noise Ratio (PSNR). In another example, the multiple types of quality metrics include two additional items from the following: Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), Multi-Scale Structural Similarity Index (MS-SSIM), Video Quality Metric (VQM), Weighted PSNR (wPSNR), Weighted to Spherical Uniform PSNR (WS-PSNR), Sequence PSNR, Sequence wPSNR, or Sequence WS-PSNR. In the examples above, the quality metric syntax element indicates the value of the quality metric indicated by the quality metric type syntax element.
[0140] In one example of this disclosure, the quality metric message includes a quality metric SEI message. In another example, the quality metric message includes a quality metric Open Bit Stream Unit (OBU).
[0141] In other examples of this disclosure, source device 102 and / or destination device 116 may be configured to decode a second quality metric message including a second quality metric syntax element, wherein the second quality metric syntax element indicates a second value of a second quality metric associated with a sub-picture or region of interest of a picture.
[0142] For illustrative purposes, the techniques described above are set in the context of VVC (ITU-T H.266, under development) and HEVC (ITU-T H.265). However, the techniques of this disclosure can be implemented by video coding devices configured for other video decoding standards and formats, such as AV1, future versions of AV1, and successors to the AV1 video decoding format. For example, instead of SEI messages, these messages can be encapsulated data, such as Open Bit Stream Units (OBUs) that include at least some of the metadata described in this disclosure. As an example, some or all of the syntaxes described above included in a picture orientation SEI message can be included in a picture orientation OBU, such that the picture orientation OBU includes one or more of a cancellation flag, a persistence flag, a constituent picture match flag, or a transformation type syntax element. As another example, some or all of the syntaxes described above included in a picture quality metric SEI message can be included in a picture quality metric OBU, such that the picture quality metric OBU includes one or more syntax elements indicating the quality metric of the picture.
[0143] Figure 6 is a block diagram illustrating an example video encoder 200 capable of performing the techniques of this disclosure. Figure 6 is provided for illustrative purposes and should not be construed as a limitation on the techniques as extensively illustrated and described in this disclosure. For illustrative purposes, this disclosure describes a video encoder 200 based on VVC (ITU-T H.266, under development) and HEVC (ITU-T H.265) technologies. However, the techniques of this disclosure can be performed by video encoding devices configured for other video decoding standards and formats, such as AV1 and its successors.
[0144] In the example of Figure 6, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a conversion processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse conversion processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy encoding unit 220. Any or all of the video data memory 230, mode selection unit 202, residual generation unit 204, conversion processing unit 206, quantization unit 208, inverse quantization unit 210, inverse conversion processing unit 212, reconstruction unit 214, filter unit 216, DPB 218, and entropy encoding unit 220 can be implemented in one or more processors or in processing circuitry. For example, the units of the video encoder 200 can be implemented as one or more circuit or logic elements as part of hardware circuitry, or as part of a processor, ASIC, or FPGA. Furthermore, the video encoder 200 may include additional or alternative processors or processing circuitry to perform these and other functions.
[0145] Video data memory 230 can store video data to be encoded by components of video encoder 200. Video encoder 200 can receive the video data stored in video data memory 230 from, for example, video source 104 (FIG. 1). DPB 218 can act as a reference picture memory, storing reference video data for use when the video encoder 200 predicts subsequent video data. Video data memory 230 and DPB 218 can be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. Video data memory 230 and DPB 218 can be provided by the same memory device or separate memory devices. In various examples, video data memory 230 can be on-chip (as shown) with other components of video encoder 200, or off-chip relative to those components.
[0146] In this disclosure, references to video data memory 230 should not be construed as being limited to memory within the video encoder 200 (unless specifically described as such) or memory outside the video encoder 200 (unless specifically described as such). Rather, references to video data memory 230 should be understood as reference memory storing video data received by the video encoder 200 for encoding (e.g., video data for the current block to be encoded). Memory 106 of FIG1 may also provide temporary storage for outputs from various units of the video encoder 200.
[0147] The various units in Figure 6 are illustrated to aid in understanding the operations performed by the video encoder 200. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide specific functions and are pre-configured regarding the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functionality in the operations that can be performed. For example, programmable circuits can execute software or firmware that allows the programmable circuits to operate in a manner defined by instructions from the software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by fixed-function circuits is generally immutable. In some examples, one or more of these units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units may be integrated circuits.
[0148] The video encoder 200 may include an arithmetic logic unit (ALU), an essential function unit (EFU), digital circuitry, analog circuitry, and / or a programmable core formed from programmable circuitry. In an example where the operation of the video encoder 200 is performed using software executed by programmable circuitry, memory 106 (FIG. 1) may store instructions (e.g., object codes) of the software received and executed by the video encoder 200, or another memory (not shown) within the video encoder 200 may store such instructions.
[0149] The video data memory 230 is configured to store the received video data. The video encoder 200 can retrieve images of the video data from the video data memory 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data memory 230 can be the original video data to be encoded.
[0150] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra-frame prediction unit 226. The mode selection unit 202 may include additional functional units that perform video prediction based on other prediction modes. As an example, the mode selection unit 202 may include a palette unit, an intra-block copying unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.
[0151] Mode selection unit 202 typically coordinates multiple coding passes to test combinations of coding parameters and the rate-distortion values obtained for such combinations. Coding parameters may include segmenting the CTU into CUs, the prediction mode used for the CUs, the transformation type of the residual data used for the CUs, and the quantization parameters of the residual data used for the CUs. Mode selection unit 202 can ultimately select a combination of coding parameters that yields a better rate-distortion value than other tested combinations.
[0152] The video encoder 200 can segment images retrieved from the video data memory 230 into a series of CTUs, and encapsulate one or more CTUs within a slice. The mode selection unit 202 can segment the CTUs of the image according to a tree structure (such as the MTT structure, QTBT structure, superblock structure, or quadtree structure described above). As mentioned above, the video encoder 200 can segment CTUs according to a tree structure to form one or more CUs. Such CUs can also be referred to as "video blocks" or "blocks".
[0153] Typically, mode selection unit 202 also controls its components (e.g., motion estimation unit 222, motion compensation unit 224, and intra-frame prediction unit 226) to generate predicted blocks for the current block (e.g., the current CU, or the overlapping portion of PU and TU in HEVC). For inter-frame prediction of the current block, motion estimation unit 222 may perform motion search to identify one or more closely matching reference blocks in one or more reference images (e.g., one or more previously decoded images stored in DPB 218). Specifically, motion estimation unit 222 may calculate values representing the similarity between the potential reference block and the current block, for example, based on the sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared error (MSD), etc. Motion estimation unit 222 may typically perform these calculations using the sample-by-sample difference between the current block and the reference block being considered. The motion estimation unit 222 can identify the reference block with the lowest value obtained from these calculations, which indicates the reference block that most closely matches the current block.
[0154] Motion estimation unit 222 can generate one or more motion vectors (MVs) that define the position of a reference block in a reference image relative to the position of the current block in the current image. Motion estimation unit 222 can then provide the motion vectors to motion compensation unit 224. For example, for unidirectional inter-frame prediction, motion estimation unit 222 can provide a single motion vector, while for bidirectional inter-frame prediction, motion estimation unit 222 can provide two motion vectors. Motion compensation unit 224 can then use the motion vectors to generate prediction blocks. For example, motion compensation unit 224 can use the motion vectors to retrieve data from the reference blocks. As another example, if the motion vectors have fractional sample accuracy, motion compensation unit 224 can interpolate the values used for predicting blocks according to one or more interpolation filters. Furthermore, for bidirectional inter-frame prediction, motion compensation unit 224 can retrieve data for the two reference blocks identified by the corresponding motion vectors and combine the retrieved data, for example, by means of sample-wise averaging or weighted averaging.
[0155] When operating according to the AV1 video decoding format, the motion estimation unit 222 and the motion compensation unit 224 can be configured to encode the decoded blocks of the video data (e.g., both luma and chroma decoded blocks) using translational motion compensation, affine motion compensation, overlapping block motion compensation (OBMC), and / or composite intra-inter-frame prediction.
[0156] As another example, for intra-prediction or intra-prediction decoding, intra-prediction unit 226 can generate a prediction block based on samples adjacent to the current block. For example, in directional mode, intra-prediction unit 226 can typically mathematically combine the values of adjacent samples and fill these calculated values across the current block in a defined direction to generate a prediction block. As another example, in DC mode, intra-prediction unit 226 can calculate the average of samples adjacent to the current block and generate a prediction block to include the obtained average for each sample in the prediction block.
[0157] When operating according to the AV1 video decoding format, the intra-frame prediction unit 226 can be configured to encode decoding blocks of video data (e.g., both luma and chroma decoding blocks) using directional intra-frame prediction, non-directional intra-frame prediction, recursive filter intra-frame prediction, chroma-based luma (CFL) prediction, intra-block copying (IBC), and / or palette modes. The mode selection unit 202 may include additional functional units for performing video prediction based on other prediction modes.
[0158] Mode selection unit 202 provides the prediction block to residual generation unit 204. Residual generation unit 204 receives the original, unencoded version of the current block from video data memory 230 and the prediction block from mode selection unit 202. Residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines the residual block for the current block. In some examples, residual generation unit 204 may also determine the difference between sample values in the residual block to generate the residual block using residual differential pulse decode-modulation (RDPCM). In some examples, one or more subtractor circuits performing binary subtraction may be used to form residual generation unit 204.
[0159] In the example where mode selection unit 202 divides a CU into PUs, each PU can be associated with a luma prediction unit and a corresponding chroma prediction unit. Video encoder 200 and video decoder 300 can support PUs of various sizes. As noted above, the size of a CU can refer to the size of its luma decoding block, while the size of a PU can refer to the size of its luma prediction unit. Assuming a specific CU size is 2Nx2N, video encoder 200 can support PU sizes of 2Nx2N or NxN for intra-frame prediction, and symmetrical PU sizes of 2Nx2N, 2NxN, Nx2N, NxN, or similar sizes for inter-frame prediction. Video encoder 200 and video decoder 300 can also support asymmetric segmentation for PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter-frame prediction.
[0160] In the example where mode selection unit 202 does not further divide the CU into PUs, each CU can be associated with a luma decoding block and a corresponding chroma decoding block. As mentioned above, the size of the CU can refer to the size of the luma decoding block of the CU. Video encoder 200 and video decoder 300 can support CU sizes of 2Nx2N, 2NxN, or Nx2N.
[0161] For other video decoding techniques (such as intra-block copy mode decoding, affine mode decoding, and linear model (LM) mode decoding), as some examples, mode selection unit 202 generates a prediction block for the current block being encoded via a corresponding unit associated with the decoding technique. In some examples (such as palette mode decoding), mode selection unit 202 may not generate a prediction block, but instead generate syntax elements indicating how the block should be reconstructed based on the selected palette. In such a mode, mode selection unit 202 may provide these syntax elements to entropy coding unit 220 for encoding.
[0162] As described above, the residual generation unit 204 receives video data for the current block and the corresponding prediction block. Then, the residual generation unit 204 generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block.
[0163] Transformation processing unit 206 applies one or more transformations to the residual block to produce a block of transformation coefficients (referred to herein as a "transformation coefficient block"). Transformation processing unit 206 may apply various transformations to the residual block to form the transformation coefficient block. For example, transformation processing unit 206 may apply a Discrete Cosine Transform (DCT), direction transformation, Karhunen-Loeve Transform (KLT), or conceptually similar transformations to the residual block. In some examples, transformation processing unit 206 may perform multiple transformations on the residual block, such as primary transformations and secondary transformations (e.g., rotation transformations). In some examples, transformation processing unit 206 does not apply any transformations to the residual block.
[0164] When operating according to AV1, the transformation processing unit 206 may apply one or more transformations to the residual block to produce a block of transformation coefficients (referred to herein as a "transformation coefficient block"). The transformation processing unit 206 may apply various transformations to the residual block to form the transformation coefficient block. For example, the transformation processing unit 206 may apply a combination of horizontal / vertical transformations, which may include the Discrete Cosine Transform (DCT), the Asymmetric Discrete Sine Transform (ADST), the Inverted ADST (e.g., the Reverse ADST), and the Identity Transform (IDTX). When using the Identity Transform, the transformation is skipped in either the vertical or horizontal direction. In some examples, the transformation processing may be skipped entirely.
[0165] Quantization unit 208 can quantize the conversion coefficients in the conversion coefficient block to produce a quantized conversion coefficient block. Quantization unit 208 can quantize the conversion coefficients of the conversion coefficient block based on the quantization parameter (QP) value associated with the current block. Video encoder 200 (e.g., via mode selection unit 202) can adjust the degree of quantization applied to the conversion coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may introduce information loss, and therefore, the quantized conversion coefficients may have lower accuracy compared to the original conversion coefficients generated by conversion processing unit 206.
[0166] The inverse quantization unit 210 and the inverse transformation processing unit 212 can apply inverse quantization and inverse transformation to the quantized transformation coefficient block to reconstruct the residual block from the transformation coefficient block. The reconstruction unit 214 can generate a reconstructed block corresponding to the current block based on the reconstructed residual block and the prediction block generated by the mode selection unit 202 (although potentially with some degree of distortion). For example, the reconstruction unit 214 can add samples of the reconstructed residual block to corresponding samples of the prediction block generated by the mode selection unit 202 to generate the reconstructed block.
[0167] Filter unit 216 can perform one or more filtering operations on the reconstructed blocks. For example, filter unit 216 can perform deblocking operations to reduce blocking artifacts along the edges of the CU. In some examples, the operations of filter unit 216 can be skipped.
[0168] When operating according to AV1, filter unit 216 can perform one or more filter operations on the reconstructed block. For example, filter unit 216 can perform a deblocking operation to reduce block artifacts along the edges of the CU. In other examples, filter unit 216 can apply a restricted direction enhancement filter (CDEF), which can be applied after deblocking, and can include an inseparable nonlinear low-pass directional filter applied based on the estimated edge direction. Filter unit 216 can also include a loop recovery filter, which is applied after CDEF, and can include a separable symmetric normalized Wiener filter or a dual self-guided filter.
[0169] The video encoder 200 stores the reconstructed blocks in the DPB 218. For example, in an example where the filter unit 216 is not operated, the reconstruction unit 214 can store the reconstructed blocks in the DPB 218. In an example where the filter unit 216 is operated, the filter unit 216 can store the filtered reconstructed blocks in the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 can retrieve reference images formed from the reconstructed (and potentially filtered) blocks from the DPB 218 to perform inter-frame prediction of blocks in subsequent encoded images. Additionally, the intra-frame prediction unit 226 can use the reconstructed blocks of the current image in the DPB 218 to perform intra-frame prediction of other blocks in the current image.
[0170] Typically, entropy coding unit 220 can entropy code syntax elements received from other functional components of video encoder 200. For example, entropy coding unit 220 can entropy code quantized transform coefficient blocks from quantization unit 208. As another example, entropy coding unit 220 can entropy code predictive syntax elements (e.g., motion information for inter-frame prediction or intra-frame mode information for intra-frame prediction) from mode selection unit 202. Entropy coding unit 220 can perform one or more entropy coding operations on syntax elements, another example of video data, to produce entropy-coded data. For example, entropy coding unit 220 can perform context-adaptive variable-length decoding (CAVLC), CABAC, variable-to-variable (V2V) length decoding, syntax-based context-adaptive binary arithmetic decoding (SBAC), probabilistic interval partitioned entropy (PIPE) decoding, exponential-Golomb coding, or another type of entropy coding operation on the data. In some examples, the entropy coding unit 220 can operate in a bypass mode where syntax elements are not entropy encoded.
[0171] The video encoder 200 can output a bitstream that includes the entropy-encoded syntax elements needed to reconstruct slices or blocks of images. Specifically, the entropy coding unit 220 can output a bitstream.
[0172] According to AV1, entropy coding unit 220 can be configured as a symbol-to-symbol adaptive multi-symbol arithmetic decoder. The syntax elements in AV1 include an N-element alphabet, and the context (e.g., a probability model) includes a set of N probabilities. Entropy coding unit 220 can store the probabilities as an n-bit (e.g., 15-bit) cumulative distribution function (CDF). Entropy coding unit 220 can perform recursive scaling to update the context using an update factor based on the alphabet size.
[0173] The above description pertains to the block operations. This description should be understood as referring to operations applied to the luma decoding block and / or chroma decoding block. As mentioned above, in some examples, the luma decoding block and chroma decoding block are the luma and chroma components of the CU. In some examples, the luma decoding block and chroma decoding block are the luma and chroma components of the PU.
[0174] In some examples, the operations performed on the luma decoding block do not need to be repeated for the chroma decoding block. As an example, the operations used to identify the motion vectors (MVs) and reference images for the luma decoding block do not need to be repeated for identifying the MVs and reference images for the chroma block. Specifically, the MVs used for the luma decoding block can be scaled to determine the MVs for the chroma block, and the reference images can be the same. As another example, the intra-frame prediction procedure can be the same for both the luma and chroma decoding blocks.
[0175] Based on the SEI technology discussed above, video encoder 200 represents an example of a device configured to encode video data, the device including: a memory configured to store video data; and one or more processing units implemented in a circuit and configured to: receive an image; and encode image orientation information including a transformation type syntax element, wherein the transformation type syntax element indicates a transformation from among multiple transformations to be applied to the image. Video encoder 200 can also be configured to encode quality metric information including a quality metric syntax element, wherein the quality metric syntax element indicates a value of a quality metric associated with the image.
[0176] Figure 7 is a block diagram illustrating an example video decoder 300 capable of performing the techniques of this disclosure. Figure 7 is provided for illustrative purposes and is not intended to limit the techniques as extensively illustrated and described in this disclosure. For illustrative purposes, this disclosure describes a video decoder 300 based on VVC (ITU-T H.266, under development) and HEVC (ITU-T H.265) technologies. However, the techniques of this disclosure can be performed by video decoding devices configured for other video decoding standards.
[0177] In the example of Figure 7, the video decoder 300 includes a decoded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse conversion processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 134. Any or all of the CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse conversion processing unit 308, reconstruction unit 310, filter unit 312, and DPB 134 can be implemented in one or more processors or in processing circuitry. For example, units of the video decoder 300 can be implemented as one or more circuit or logic elements as part of hardware circuitry, or as part of a processor, ASIC, or FPGA. Furthermore, the video decoder 300 may include additional or alternative processors or processing circuitry to perform these and other functions.
[0178] The prediction processing unit 304 includes a motion compensation unit 316 and an intra-frame prediction unit 318. The prediction processing unit 304 may include additional units for performing predictions based on other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, an intra-block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, the video decoder 300 may include more, fewer, or different functional components.
[0179] When operating according to AV1, compensation unit 316 can be configured to decode the decoded blocks of video data (e.g., both luma and chroma decoded blocks) using translational motion compensation, affine motion compensation, OBMC, and / or composite inter-frame intra-frame prediction, as described above. Intra-frame prediction unit 318 can be configured to decode the decoded blocks of video data (e.g., both luma and chroma decoded blocks) using directional intra-frame prediction, non-directional intra-frame prediction, recursive filter intra-frame prediction, CFL, intra-block copy (IBC), and / or palette mode.
[0180] CPB memory 320 may store video data, such as encoded video bitstreams, to be decoded by components of video decoder 300. The video data stored in CPB memory 320 may be obtained, for example, from computer-readable media 110 (FIG. 1). CPB memory 320 may include a CPB storing encoded video data (e.g., syntax elements) from the encoded video bitstream. Furthermore, CPB memory 320 may store video data other than the syntax elements of the decoded picture, such as temporary data representing the output from various units of video decoder 300. DPB 314 typically stores decoded pictures, which video decoder 300 may output, and / or uses as reference video data when decoding subsequent data or pictures of the encoded video bitstream. CPB memory 320 and DPB 314 may be formed from any of a variety of memory devices, such as DRAM (including SDRAM), MRAM, RRAM, or other types of memory devices. CPB memory 320 and DPB 314 can be provided by the same memory device or separate memory devices. In various examples, CPB memory 320 can be on-chip with other components of video decoder 300, or off-chip relative to those components.
[0181] Alternatively, in some examples, the video decoder 300 may retrieve the decoded video data from memory 120 (FIG. 1). That is, memory 120 may utilize CPB memory 320 to store data as discussed above. Similarly, when some or all of the functions of the video decoder 300 are implemented using software to be executed by the processing circuitry of the video decoder 300, memory 120 may store instructions to be executed by the video decoder 300.
[0182] The various units shown in Figure 7 are illustrated to aid in understanding the operations performed by the video decoder 300. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Similar to Figure 6, a fixed-function circuit refers to a circuit that provides a specific function and is pre-configured regarding the operations that can be performed. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provides flexible functionality in the operations that can be performed. For example, a programmable circuit can execute software or firmware that allows the programmable circuit to operate in a manner defined by instructions from the software or firmware. A fixed-function circuit can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is generally immutable. In some examples, one or more of these units can be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units can be integrated circuits.
[0183] The video decoder 300 may include an ALU, EFU, digital circuitry, analog circuitry, and / or a programmable core formed from programmable circuitry. In an example where the operation of the video decoder 300 is performed by software executed on the programmable circuitry, on-chip or off-chip memory may store instructions (e.g., object code) of the software received and executed by the video decoder 300.
[0184] Entropy decoding unit 302 can receive encoded video data from CPB and perform entropy decoding on the video data to regenerate syntax elements. Prediction processing unit 304, inverse quantization unit 306, inverse conversion processing unit 308, reconstruction unit 310, and filter unit 312 can generate decoded video data based on syntax elements extracted from the bitstream.
[0185] Typically, the video decoder 300 reconstructs the image on a block-by-block basis. The video decoder 300 can perform the reconstruction operation on each block individually (where the block currently being reconstructed (i.e., decoded) can be referred to as the "current block").
[0186] Entropy decoding unit 302 can entropy decode the syntax elements of the quantized conversion coefficients that define the quantized conversion coefficient block, as well as conversion information such as quantization parameters (QP) and / or conversion mode indications. Inverse quantization unit 306 can use the QP associated with the quantized conversion coefficient block to determine the quantization level, and similarly, determine the inverse quantization level to be applied by inverse quantization unit 306. Inverse quantization unit 306 can, for example, perform a bitwise left shift operation to inverse quantize the quantized conversion coefficients. Inverse quantization unit 306 can thus form a conversion coefficient block including the conversion coefficients.
[0187] After the inverse quantization unit 306 forms the transformation coefficient block, the inverse transformation processing unit 308 can apply one or more inverse transformations to the transformation coefficient block to generate a residual block associated with the current block. For example, the inverse transformation processing unit 308 can apply an inverse DCT, an inverse integer transformation, an inverse Karhunen-Loeve transformation (KLT), an inverse rotation transformation, an inverse direction transformation, or another inverse transformation to the transformation coefficient block.
[0188] Furthermore, the prediction processing unit 304 generates prediction blocks based on the prediction information syntax elements entropy-decoded by the entropy decoding unit 302. For example, if the prediction information syntax elements indicate that the current block is inter-frame predicted, the motion compensation unit 316 can generate prediction blocks. In this case, the prediction information syntax elements may indicate the reference picture from which the reference block is to be retrieved in the DPB 314, and identify the motion vector of the reference block's position in the reference picture relative to the position of the current block in the current picture. The motion compensation unit 316 can generally perform the inter-frame prediction procedure in a manner substantially similar to that described with respect to the motion compensation unit 224 (FIG. 6).
[0189] As another example, if the prediction information syntax element indicates that the current block is intra-predicted, then the intra-prediction unit 318 can generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Again, the intra-prediction unit 318 can generally perform the intra-prediction procedure in a manner substantially similar to that described with respect to the intra-prediction unit 226 (FIG. 6). The intra-prediction unit 318 can retrieve data from the DPB 314 of samples adjacent to the current block.
[0190] Reconstruction unit 310 can use the predicted block and the residual block to reconstruct the current block. For example, reconstruction unit 310 can add samples from the residual block to the corresponding samples from the predicted block to reconstruct the current block.
[0191] Filter unit 312 can perform one or more filtering operations on the reconstructed blocks. For example, filter unit 312 can perform a deblocking operation to reduce blocking artifacts along the edges of the reconstructed blocks. The operation of filter unit 312 is not necessarily performed in all examples.
[0192] The video decoder 300 can store the reconstructed blocks in the DPB 314. For example, in an example where the filter unit 312 is not operated, the reconstruction unit 310 can store the reconstructed blocks in the DPB 314. In an example where the filter unit 312 is operated, the filter unit 312 can store the filtered reconstructed blocks in the DPB 314. As discussed above, the DPB 314 can provide reference information (such as samples of the current image for intra-frame prediction and previously decoded images for subsequent motion compensation) to the prediction processing unit 304. Furthermore, the video decoder 300 can output decoded images (e.g., decoded video) from the DPB 314 for subsequent presentation on a display device such as the display device 118 of FIG. 1.
[0193] Based on the SEI technology discussed above, video decoder 300 represents an example of a device configured to decode video data, the device including: a memory configured to store the video data; and one or more processing units implemented in a circuit and configured to: receive an image; and decode image orientation information including a transformation type syntax element, wherein the transformation type syntax element indicates a transformation from among multiple transformations to be applied to the image. Video decoder 300 can also be configured to decode quality metric information including a quality metric syntax element, wherein the quality metric syntax element indicates a value of a quality metric associated with the image.
[0194] Figure 8 is a flowchart illustrating an example method for encoding a current block according to the technology of this disclosure. The current block may include the current CU. Although the video encoder 200 (Figures 1 and 6) is described in relation to the present disclosure, it should be understood that other devices may be configured to perform a method similar to that of Figure 8.
[0195] In this example, the video encoder 200 initially predicts the current block (350). For example, the video encoder 200 may form a prediction block for the current block. Then, the video encoder 200 may compute a residual block for the current block (352). To compute the residual block, the video encoder 200 may compute the difference between the original unencoded block and the prediction block for the current block. Then, the video encoder 200 may transform the residual block and quantize the transformation coefficients of the residual block (354). Next, the video encoder 200 may scan the quantized transformation coefficients of the residual block (356). During or after scanning, the video encoder 200 may entropy encode the transformation coefficients (358). For example, the video encoder 200 may use CAVLC or CABAC to encode the transformation coefficients. Then, the video encoder 200 may output the entropy-encoded data of the block (360).
[0196] Figure 9 is a flowchart illustrating an example method for decoding a current block of video data according to the technology of this disclosure. The current block may include the current CU. Although the video decoder 300 (Figures 1 and 7) is described in relation to the present disclosure, it should be understood that other devices may be configured to perform a method similar to that of Figure 9.
[0197] The video decoder 300 can receive entropy-encoded data for the current block (such as entropy-encoded prediction information and entropy-encoded data for the transformation coefficients of the residual block corresponding to the current block) (370). The video decoder 300 can entropy decode the entropy-encoded data to determine the prediction information for the current block and regenerate the transformation coefficients of the residual block (372). The video decoder 300 can predict the current block, for example, using an intra-frame prediction mode or inter-frame prediction mode indicated by the prediction information for the current block (374), to calculate the prediction block for the current block. Then, the video decoder 300 can inverse scan the regenerated transformation coefficients (376) to create a block of quantized transformation coefficients. Then, the video decoder 300 can inverse quantize the transformation coefficients and apply the inverse transformation to the transformation coefficients to generate the residual block (378). Finally, the video decoder 300 can decode the current block by combining the prediction block and the residual block (380).
[0198] Other illustrative aspects of the technology and apparatus described below.
[0199] Aspect 1A - A method for processing video data, the method comprising: receiving an image; and decoding an image orientation supplementation and enhancement information (SEI) message, the image orientation SEI message including one or more of the following: a deactivation flag, a persistent flag, a composition image matching flag, or a transformation type syntax element, wherein the transformation type syntax element indicates one or more of rotation or mirroring to be applied to the image.
[0200] Aspect 2A - The method according to aspect 1A, wherein decoding includes decoding, and wherein the method further includes: processing the image according to the image orientation SEI message.
[0201] Aspect 3A - The method according to aspect 1A or aspect 2A, wherein the image orientation information includes image orientation supplementation and enhancement information (SEI) information.
[0202] Aspect 4A - The method according to aspect 1A or aspect 2A, wherein the image orientation information includes an image orientation open bit stream unit (OBU).
[0203] Aspect 5A - The method according to claim 1A, wherein decoding includes encoding.
[0204] Aspect 6A - A method for processing video data, the method comprising: receiving an image; and decoding an image quality metric supplemental enhancement information (SEI) message, the image quality metric SEI message including one or more syntax elements indicating a quality metric of the image.
[0205] Aspect 7A - The method according to aspect 6A, wherein decoding includes decoding, and wherein the method further includes: processing the image according to the image quality metric SEI message.
[0206] Aspect 8A - The method described in aspect 6A or aspect 7A, wherein the image orientation information includes image quality metric supplementary enhancement information (SEI) information.
[0207] Aspect 9A - The method according to any one of Aspects 6A-8A, wherein the image quality measurement message includes one or more syntax elements, the one or more syntax elements indicating the quality measurement of one or more sub-images or regions of interest associated with the image quality measurement message.
[0208] Aspect 10A - The method according to aspect 9A, wherein the one or more syntax elements indicate a quality metric for high dynamic range (HDR) or 360 video content.
[0209] Aspect 11A - The method according to any one of aspects 6A-10A, wherein decoding includes encoding.
[0210] Aspect 12A - Any combination of the methods described in accordance with aspects 1A-10A.
[0211] Aspect 13A - An apparatus for processing video data, the apparatus comprising one or more components for performing the method according to any one of aspects 1A-12A.
[0212] Aspect 14A - The device according to aspect 13A, wherein the one or more components include one or more processors implemented in a circuit.
[0213] Aspect 15A - The device according to any one of aspects 13A and 14A further includes: a memory for storing the video data.
[0214] The device according to any one of aspects 13A-15A, further comprising: a display configured to display decoded video data.
[0215] Aspect 17A - The device according to any one of aspects 13A-16A, wherein the device includes one or more of the following: a camera, a computer, a mobile device, a broadcast receiver device or a set-top box.
[0216] Aspect 18A - The device according to any one of aspects 13A-17A, wherein the device includes a video decoder.
[0217] Aspect 19A - The device according to any one of aspects 13A-18A, wherein the device includes a video encoder.
[0218] Aspect 20A - A computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to perform the method according to any one of aspects 1A-12A.
[0219] Aspect 1B - A method for processing video data, the method comprising: receiving an image; and decoding a quality metric message including a quality metric syntax element, wherein the quality metric syntax element indicates a value of a quality metric associated with the image.
[0220] Aspect 2B - The method according to aspect 1B further includes: decoding a quality metric type syntax element in the quality metric message, wherein the quality metric type syntax element indicates a quality metric of a type indicated by the quality metric syntax element among a plurality of quality metrics.
[0221] Aspect 3B - The method according to aspect 2B, wherein the plurality of quality measures include peak signal-to-noise ratio (PSNR).
[0222] Aspect 4B - According to the method described in aspect 2B, wherein the plurality of quality metrics include two additional items from the following: peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), multi-scale structural similarity index (MS-SSIM), video quality metric (VQM), weighted PSNR (wPSNR), weighted to spherical uniform PSNR (WS-PSNR), sequence PSNR, sequence wPSNR, or sequence WS-PSNR.
[0223] Aspect 5B - According to the method described in aspect 2B, wherein the quality metric syntax element indicates the value of the quality metric indicated by the quality metric type syntax element.
[0224] Aspect 6B - The method according to aspect 1B further includes: decoding a second quality metric message including a second quality metric syntax element, wherein the second quality metric syntax element indicates a second value of a second quality metric associated with a sub-image of the image.
[0225] Aspect 7B - The method according to aspect 1B further includes: decoding a second quality metric message including a second quality metric syntax element, wherein the second quality metric syntax element indicates a second value of a second quality metric associated with a region of interest of the image.
[0226] Aspect 8B - The method according to aspect 1B, wherein decoding includes decoding, and wherein the method further includes: applying post-processing techniques to the image based on the value of the quality metric to form a processed image; and displaying the processed image.
[0227] Aspect 9B - The method according to aspect 1B, wherein the quality metric message includes quality metric supplementary enhancement information (SEI) message.
[0228] Aspect 10B - The method according to aspect 1B, wherein the quality metric message includes a quality metric open bit stream unit (OBU).
[0229] Aspect 11B - An apparatus configured to process video data, the apparatus comprising: a memory configured to store images; and one or more processors implemented in a circuit and communicating with the memory, the one or more processors being configured to: receive the images; and decode a quality metric message including quality metric syntax elements, wherein the quality metric syntax elements indicate values of quality metrics associated with the images.
[0230] Aspect 12B - The apparatus according to aspect 11B, wherein the one or more processors are further configured to: decode a quality metric type syntax element in the quality metric message, wherein the quality metric type syntax element indicates a quality metric of a type indicated by the quality metric syntax element among a plurality of quality metrics.
[0231] Aspect 13B - The apparatus according to aspect 12B, wherein the plurality of quality metrics include peak signal-to-noise ratio (PSNR).
[0232] Aspect 14B - The apparatus according to aspect 12B, wherein the plurality of quality metrics include two additional items from the following: peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), multi-scale structural similarity index (MS-SSIM), video quality metric (VQM), weighted PSNR (wPSNR), weighted to spherical uniform PSNR (WS-PSNR), sequence PSNR, sequence wPSNR, or sequence WS-PSNR.
[0233] Aspect 15B - The apparatus according to aspect 12B, wherein the quality metric syntax element indicates the value of the quality metric indicated by the quality metric type syntax element.
[0234] Aspect 16B - The apparatus according to aspect 11B, wherein the one or more processors are further configured to: decode a second quality metric message including a second quality metric syntax element, wherein the second quality metric syntax element indicates a second value of a second quality metric associated with a sub-image of the image.
[0235] Aspect 17B - The apparatus according to aspect 11B, wherein the one or more processors are further configured to: decode a second quality metric message including a second quality metric syntax element, wherein the second quality metric syntax element indicates a second value of a second quality metric associated with a region of interest of the image.
[0236] Aspect 18B - The apparatus according to aspect 11B, wherein the apparatus is configured to decode the quality metric message, and wherein the one or more processors are further configured to: apply post-processing techniques to the image based on the value of the quality metric to form a processed image; and display the processed image.
[0237] Aspect 19B - The apparatus according to aspect 11B, wherein the quality metric message includes quality metric supplemental enhancement information (SEI) message.
[0238] Aspect 20B - The apparatus according to aspect 11B, wherein the quality metric message includes a quality metric open bit stream unit (OBU).
[0239] Aspect 21B - An apparatus configured to process video data, the apparatus comprising: a component for receiving an image; and a component for decoding a quality metric message including a quality metric syntax element, wherein the quality metric syntax element indicates a value of a quality metric associated with the image.
[0240] Aspect 22B - The apparatus according to aspect 21B further includes: a component for decoding a quality metric type syntax element in the quality metric message, wherein the quality metric type syntax element indicates a quality metric of a type indicated by the quality metric syntax element from a plurality of quality metrics.
[0241] Aspect 23B - The apparatus according to aspect 22B, wherein the plurality of quality metrics include peak signal-to-noise ratio (PSNR).
[0242] Aspect 24B - The apparatus according to aspect 22B, wherein the plurality of quality metrics include two additional items from the following: peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), multi-scale structural similarity index (MS-SSIM), video quality metric (VQM), weighted PSNR (wPSNR), weighted to spherical uniform PSNR (WS-PSNR), sequence PSNR, sequence wPSNR, or sequence WS-PSNR.
[0243] Aspect 25B - The apparatus according to aspect 22B, wherein the quality metric syntax element indicates the value of the quality metric indicated by the quality metric type syntax element.
[0244] Aspect 26B - The apparatus according to aspect 21B further includes: a component for decoding a second quality metric message including a second quality metric syntax element, wherein the second quality metric syntax element indicates a second value of a second quality metric associated with a sub-picture of the picture.
[0245] Aspect 27B - The apparatus according to aspect 21B further includes: a component for decoding a second quality metric message including a second quality metric syntax element, wherein the second quality metric syntax element indicates a second value of a second quality metric associated with a region of interest of the image.
[0246] Aspect 28B - The apparatus according to aspect 21B, wherein the decoding component includes a decoding component, and wherein the apparatus further includes: a component for applying post-processing techniques to the image based on the value of the quality metric to form a processed image; and a component for displaying the processed image.
[0247] Aspect 29B - The apparatus according to aspect 21B, wherein the quality metric message includes quality metric supplemental enhancement information (SEI) message.
[0248] Aspect 30B - The apparatus according to aspect 21B, wherein the quality metric message includes a quality metric open bit stream unit (OBU).
[0249] Aspect 31B - A non-transitory computer-readable storage medium for storing instructions, which, when executed, cause one or more processors of a device configured to process video data to: receive a picture; and decode a quality metric message including a quality metric syntax element, wherein the quality metric syntax element indicates a value of a quality metric associated with the picture.
[0250] Aspect 32B - The non-transitory computer-readable storage medium according to aspect 31B, wherein the instructions further cause the one or more processors to perform the following operation: decoding a quality metric type syntax element in the quality metric message, wherein the quality metric type syntax element indicates a quality metric of a type indicated by the quality metric syntax element among a plurality of quality metrics.
[0251] Aspect 33B - Non-transitory computer-readable storage media according to aspect 32B, wherein the plurality of quality measures include peak signal-to-noise ratio (PSNR).
[0252] Aspect 34B - The non-transitory computer-readable storage medium according to aspect 32, wherein the plurality of quality metrics include two additional items of the following: peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), multi-scale structural similarity index (MS-SSIM), video quality metric (VQM), weighted PSNR (wPSNR), weighted to spherical uniform PSNR (WS-PSNR), sequence PSNR, sequence wPSNR, or sequence WS-PSNR.
[0253] Aspect 35B - A non-transitory computer-readable storage medium according to aspect 32B, wherein the quality metric syntax element indicates the value of the quality metric indicated by the quality metric type syntax element.
[0254] Aspect 36B - A non-transitory computer-readable storage medium according to aspect 31B, wherein the instructions further cause the one or more processors to: decode a second quality metric message including a second quality metric syntax element, wherein the second quality metric syntax element indicates a second value of a second quality metric associated with a sub-picture of the picture.
[0255] Aspect 37B - A non-transitory computer-readable storage medium according to aspect 31B, wherein the instructions further cause the one or more processors to: decode a second quality metric message including a second quality metric syntax element, wherein the second quality metric syntax element indicates a second value of a second quality metric associated with a region of interest of the image.
[0256] Aspect 38B - A non-transitory computer-readable storage medium according to aspect 31B, wherein the device is configured to decode the quality metric message, and wherein the instructions further cause the one or more processors to: apply post-processing techniques to the image based on the value of the quality metric to form a processed image; and display the processed image.
[0257] Aspect 39B - The non-transitory computer-readable storage medium according to aspect 31B, wherein the quality measurement message includes quality measurement supplementary enhancement information (SEI) message.
[0258] Aspect 40B - A non-transitory computer-readable storage medium according to aspect 31B, wherein the quality measurement message includes a quality measurement open bit stream unit (OBU).
[0259] Aspect 1C - A method for processing video data, the method comprising: receiving an image; and decoding a quality metric message including a quality metric syntax element, wherein the quality metric syntax element indicates a value of a quality metric associated with the image.
[0260] Aspect 2C - The method according to aspect 1C further includes: decoding a quality metric type syntax element in the quality metric message, wherein the quality metric type syntax element indicates a quality metric of a type indicated by the quality metric syntax element among a plurality of quality metrics.
[0261] Aspect 3C - The method according to aspect 2C, wherein the plurality of quality measures include peak signal-to-noise ratio (PSNR).
[0262] Aspect 4C - The method according to aspect 2C, wherein the plurality of quality metrics includes two additional items from the following: peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), multi-scale structural similarity index (MS-SSIM), video quality metric (VQM), weighted PSNR (wPSNR), weighted to spherical uniform PSNR (WS-PSNR), sequence PSNR, sequence wPSNR, or sequence WS-PSNR.
[0263] Aspect 5C - The method according to any one of Aspects 2C-4C, wherein the quality metric syntax element indicates the value of the quality metric indicated by the quality metric type syntax element.
[0264] Aspect 6C - The method according to any one of Aspects 1C-5C further includes: decoding a second quality metric message including a second quality metric syntax element, wherein the second quality metric syntax element indicates a second value of a second quality metric associated with a sub-picture of the picture.
[0265] Aspect 7C - The method according to any one of Aspects 1C-5C further includes: decoding a second quality metric message including a second quality metric syntax element, wherein the second quality metric syntax element indicates a second value of a second quality metric associated with a region of interest of the image.
[0266] Aspect 8C - The method according to any one of Aspects 1C-7C, wherein decoding includes decoding, and wherein the method further includes: applying post-processing techniques to the image based on the value of the quality metric to form a processed image; and displaying the processed image.
[0267] Aspect 9C - The method according to any one of Aspects 1C-8C, wherein the quality metric message includes quality metric supplementary enhancement information (SEI) message.
[0268] Aspect 10C - The method according to any one of Aspects 1C-8C, wherein the quality metric message includes a quality metric open bit stream unit (OBU).
[0269] Aspect 11C - An apparatus configured to process video data, the apparatus comprising: a memory configured to store images; and one or more processors implemented in a circuit and communicating with the memory, the one or more processors being configured to: receive the images; and decode a quality metric message including quality metric syntax elements, wherein the quality metric syntax elements indicate a value of a quality metric associated with the images.
[0270] Aspect 12C - The apparatus according to aspect 11C, wherein the one or more processors are further configured to: decode a quality metric type syntax element in the quality metric message, wherein the quality metric type syntax element indicates a quality metric of a type indicated by the quality metric syntax element among a plurality of quality metrics.
[0271] Aspect 13C - The apparatus according to aspect 12C, wherein the plurality of quality metrics include peak signal-to-noise ratio (PSNR).
[0272] Aspect 14C - The apparatus according to aspect 12C, wherein the plurality of quality metrics include two additional items from the following: peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), multi-scale structural similarity index (MS-SSIM), video quality metric (VQM), weighted PSNR (wPSNR), weighted to spherical uniform PSNR (WS-PSNR), sequence PSNR, sequence wPSNR, or sequence WS-PSNR.
[0273] Aspect 15C - The apparatus according to any one of aspects 12C-14C, wherein the quality metric syntax element indicates the value of the quality metric indicated by the quality metric type syntax element.
[0274] Aspect 16C - The apparatus according to any one of aspects 11C-15C, wherein the one or more processors are further configured to: decode a second quality metric message including a second quality metric syntax element, wherein the second quality metric syntax element indicates a second value of a second quality metric associated with a sub-picture of the picture.
[0275] Aspect 17C - The apparatus according to any one of aspects 11C-15C, wherein the one or more processors are further configured to: decode a second quality metric message including a second quality metric syntax element, wherein the second quality metric syntax element indicates a second value of a second quality metric associated with a region of interest of the image.
[0276] Aspect 18C - An apparatus according to any one of aspects 11C-17C, wherein the apparatus is configured to decode the quality metric message, and wherein the one or more processors are further configured to: apply post-processing techniques to the image based on the value of the quality metric to form a processed image; and display the processed image.
[0277] Aspect 19C - The apparatus according to any one of aspects 11C-18C, wherein the quality measurement message includes quality measurement supplementary enhancement information (SEI) message.
[0278] Aspect 20C - The apparatus according to any one of Aspects 11C-18C, wherein the quality metric message includes a quality metric open bit stream unit (OBU).
[0279] Aspect 21C - An apparatus configured to process video data, the apparatus comprising: a component for receiving an image; and a component for decoding a quality metric message including a quality metric syntax element, wherein the quality metric syntax element indicates a value of a quality metric associated with the image.
[0280] Aspect 22C - The apparatus according to aspect 21C further includes: a component for decoding a quality metric type syntax element in the quality metric message, wherein the quality metric type syntax element indicates a quality metric of a type indicated by the quality metric syntax element among a plurality of quality metrics.
[0281] Aspect 23C - The apparatus according to aspect 22C, wherein the plurality of quality metrics include peak signal-to-noise ratio (PSNR).
[0282] Aspect 24C - The apparatus according to aspect 22C, wherein the plurality of quality metrics include two additional items from the following: peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), multi-scale structural similarity index (MS-SSIM), video quality metric (VQM), weighted PSNR (wPSNR), weighted to spherical uniform PSNR (WS-PSNR), sequence PSNR, sequence wPSNR, or sequence WS-PSNR.
[0283] Aspect 25C - The apparatus according to any one of aspects 22C-24C, wherein the quality metric syntax element indicates the value of the quality metric indicated by the quality metric type syntax element.
[0284] Aspect 26C - The apparatus according to any one of aspects 21C-25C further includes: a component for decoding a second quality metric message including a second quality metric syntax element, wherein the second quality metric syntax element indicates a second value of a second quality metric associated with a sub-picture of the picture.
[0285] Aspect 27C - The apparatus according to any one of aspects 21C-25C further includes: a component for decoding a second quality metric message including a second quality metric syntax element, wherein the second quality metric syntax element indicates a second value of a second quality metric associated with a region of interest of the image.
[0286] Aspect 28C - The apparatus according to any one of Aspects 21C-27C, wherein the decoding component includes a decoding component, and wherein the apparatus further includes: a component for applying post-processing techniques to the image according to the value of the quality metric to form a processed image; and a component for displaying the processed image.
[0287] Aspect 29C - The apparatus according to any one of aspects 21C-28C, wherein the quality measurement information includes quality measurement supplementary enhancement information (SEI) information.
[0288] Aspect 30C - The apparatus according to any one of Aspects 21C-28C, wherein the quality metric message includes a quality metric open bit stream unit (OBU).
[0289] Aspect 31C - A non-transitory computer-readable storage medium storing instructions, which, when executed, cause one or more processors of a device configured to process video data to: receive a picture; and decode a quality metric message including a quality metric syntax element, wherein the quality metric syntax element indicates a value of a quality metric associated with the picture.
[0290] Aspect 32C - A non-transitory computer-readable storage medium according to aspect 31C, wherein the instructions further cause the one or more processors to perform the following operation: decoding a quality metric type syntax element in the quality metric message, wherein the quality metric type syntax element indicates a quality metric of a type indicated by the quality metric syntax element among a plurality of quality metrics.
[0291] Aspect 33C - Non-transitory computer-readable storage media according to aspect 32C, wherein the plurality of quality measures include peak signal-to-noise ratio (PSNR).
[0292] Aspect 34C - The non-transitory computer-readable storage medium according to aspect 32, wherein the plurality of quality metrics include two additional items of the following: peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), multi-scale structural similarity index (MS-SSIM), video quality metric (VQM), weighted PSNR (wPSNR), weighted to spherical uniform PSNR (WS-PSNR), sequence PSNR, sequence wPSNR, or sequence WS-PSNR.
[0293] Aspect 35C - A non-transitory computer-readable storage medium according to any one of aspects 32C-34C, wherein the quality metric syntax element indicates the value of the quality metric indicated by the quality metric type syntax element.
[0294] Aspect 36C - A non-transitory computer-readable storage medium according to any one of aspects 31C-35C, wherein the instructions further cause the one or more processors to perform the following operation: decoding a second quality metric message including a second quality metric syntax element, wherein the second quality metric syntax element indicates a second value of a second quality metric associated with a sub-picture of the picture.
[0295] Aspect 37C - A non-transitory computer-readable storage medium according to any one of aspects 31C-35C, wherein the instructions further cause the one or more processors to: decode a second quality metric message including a second quality metric syntax element, wherein the second quality metric syntax element indicates a second value of a second quality metric associated with a region of interest of the image.
[0296] Aspect 38C - A non-transitory computer-readable storage medium according to any one of aspects 31C-37C, wherein the device is configured to decode the quality metric message, and wherein the instructions further cause the one or more processors to: apply post-processing techniques to the image based on the value of the quality metric to form a processed image; and display the processed image.
[0297] Aspect 39C - A non-transitory computer-readable storage medium according to any one of aspects 31C-38C, wherein the quality measurement message includes quality measurement supplementary enhancement information (SEI) message.
[0298] Aspect 40C - A non-transitory computer-readable storage medium according to any one of aspects 31C-38C, wherein the quality measurement message includes a quality measurement open bit stream unit (OBU).
[0299] It should be recognized that, based on the examples, certain actions or events of any of the techniques described herein may be performed in a different order, may be added, combined, or may be omitted entirely (e.g., not all described actions or events are necessary for the implementation of the technique). Furthermore, in some examples, actions or events may be performed in parallel rather than sequentially, for example, through multithreaded processing, interrupt handling, or multiple processors.
[0300] In one or more examples, the described functionality can be implemented using hardware, software, firmware, or any combination thereof. If implemented in software, the functionality can be stored as one or more instructions or codes on a computer-readable medium or transmitted via a computer-readable medium and executed by a hardware-based processing unit. Computer-readable media can include computer-readable storage media (which corresponds to tangible media such as data storage media) or communication media (which includes, for example, any media that facilitates the transfer of a computer program from one place to another according to a communication protocol). In this way, computer-readable media can generally correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. Data storage media can be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures used to implement the techniques described in this disclosure. Computer program products can include computer-readable media.
[0301] For example, rather than limiting, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other media capable of storing desired program code in the form of instructions or data structures, and accessible by a computer. Furthermore, any connection is appropriately referred to as computer-readable media. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) is included in the definition of media. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer instead to non-transient tangible storage media. As used herein, magnetic disks and optical disks include CDs, laser discs, optical discs, DVDs, floppy disks, and Blu-ray discs, wherein magnetic disks typically copy data magnetically, while optical discs use lasers to copy data optically. Combinations of the above should also be included within the scope of computer-readable media.
[0302] Instructions can be executed by one or more processors (such as one or more DSPs, general-purpose microprocessors, ASICs, FPGAs, or other equivalent integrated or discrete logic circuits). Therefore, as used herein, the terms "processor" and "processing circuitry" can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Furthermore, the techniques can be fully implemented in one or more circuit or logic elements.
[0303] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless handheld devices, integrated circuits (ICs), or a set of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but implementation via different hardware units is not necessarily required. Specifically, as described above, various units can be combined in a codec hardware unit or provided by a collection of interoperable hardware units (including one or more processors as described above) combined with appropriate software and / or firmware.
[0304] Examples have been described. These and other examples are within the scope of the appended patent applications.
[0305] 100: System 102: Source equipment 104: Video Source 106: Memory 108: Output Interface 110: Computer-readable media 112: Storage device 114: Archive Server 116: Destination Equipment 118: Display device 120: Memory 122: Input Interface 150: Images 152: Converted image 160: Original Image 162: Images 164: Images 166: Images 168: Images 170: Image 172: Image 174: Images 200: Video Encoder 202: Mode Selection Unit 204: Residual Generation Unit 206: Conversion Processing Unit 208: Quantization unit 210: Inverse quantization unit 212: Inverse conversion processing unit 214: Reconstruction Unit 216: Filter Unit 218: Decoded Picture Buffer (DPB) 220: Entropy Coding Unit 222: Motion Estimation Unit 224: Motion Compensation Unit 226: Intra-prediction unit 230: Video data memory 300: Video Decoder 302: Entropy Decoding Unit 304: Predictive Processing Unit 306: Inverse quantization unit 308: Inverse conversion processing unit 310: Reconstruction Unit 312: Filter Unit 314: Decoded Picture Buffer (DPB) 316: Motion Compensation Unit 318: Intra-frame prediction unit 320: CPB Memory 350: Steps 352: Steps 354: Steps 356: Steps 358: Steps 360: Steps 370: Steps 372: Steps 374: Steps 376: Steps 378: Steps 380: Steps 400: Steps 402: Steps 404: Steps 410: Steps 412: Steps 414: Steps 416: Steps 500: Steps 502: Steps 504: Steps 510: Steps 512: Steps 514: Steps 516: Steps
Claims
1. A method for processing video data, the method comprising: Receives an image that includes multiple sub-images; The method also decodes Quality Measurement Supplemental Enhancement Information (SEI) messages associated with multiple sub-images of the image, the SEI messages including: a first syntax element indicating the number of sub-images of the multiple sub-images; a quality measurement type syntax element associated with at least one sub-image of the multiple sub-images, set to indicate a first value among multiple values of multiple types of multiple quality measurements, wherein the first value is set to 0; and a quality measurement syntax element specifying a second value; determining that the quality measurement type syntax element is set to 0; based on the determination that the quality measurement type syntax element is set to 0, determining that the second value specified by the quality measurement syntax element is interpreted as indicating a Peak Signal-to-Noise Ratio (PSNR) value associated with the at least one sub-image; and deriving the derived PSNR associated with the at least one sub-image based on the second value specified by the quality measurement syntax element. According to the method of claim 1, deriving the derived PSNR associated with the at least one sub-image includes determining a quotient based on an integer value and the second value specified by the quality measurement syntax element.
2.
3. The method according to claim 2, wherein, The second value specified by the quality metric syntax element includes a 16-bit unsigned integer and the integer value is 100.
4. The method according to request item 3, wherein, The multiple types of quality metrics include the PSNR and at least one of the following: structural similarity index (SSIM), weighted PSNR (wPSNR), or weighted to spherical uniform PSNR (WS-PSNR).
5. The method according to claim 3 further includes: Receive a second image that includes multiple second sub-images; The second quality metric SEI message associated with multiple second sub-images of the second image is decoded. The second quality metric SEI message includes: a second syntax element indicating the number of second sub-images of the multiple second sub-images; a second quality metric type syntax element associated with at least one second sub-image of the multiple second sub-images; a third value set to indicate a third value among multiple values of multiple types of the multiple quality metrics, wherein the third value is set to 1; and a second quality metric syntax element specifying a fourth value. The second quality metric type syntax element is set to 1; based on the decision that the second quality metric type syntax element is set to 1, the fourth value specified by the second quality metric syntax element is interpreted as an indication of the SSIM value associated with the at least one second sub-image; and based on the fourth value specified by the second quality metric syntax element, the deduced SSIM associated with the at least one second sub-image is derived. According to the method described in claim 5, deriving the derived SSIM associated with the at least one second sub-image includes determining the second quotient based on a second integer value and the fourth value specified by the second quality metric syntax element.
6. 。 7. The method according to claim 3, further comprising: Post-processing techniques are applied to the at least one sub-image based on the derived PSNR to form a processed image; And display the processed image.
8. An apparatus configured to process video data, the apparatus comprising: Memory configured to store images; And one or more processors implemented in the circuit and communicating with the memory, the one or more processors being configured to: receive the image comprising a plurality of sub-images; The process also includes decoding Quality Measurement Supplemental Enhancement Information (SEI) messages associated with multiple sub-images of the image, the SEI messages including: a first syntax element indicating the number of sub-images of the multiple sub-images; a quality measurement type syntax element associated with at least one sub-image of the multiple sub-images, set to indicate a first value among multiple values of multiple types of multiple quality measurements, wherein the first value is set to 0; and a quality measurement syntax element specifying a second value; determining that the quality measurement type syntax element is set to 0; and based on the determination that the quality measurement type syntax element is set to 0, determining that the second value specified by the quality measurement syntax element is interpreted as indicating a Peak Signal-to-Noise Ratio (PSNR) value associated with the at least one sub-image. And based on the second value specified by the quality metric syntax element, derive the derived PSNR associated with the at least one sub-image.
9. The apparatus according to claim 8, wherein, The one or more processors are further configured to: determine the quotient based on an integer value and the second value specified by the quality metric syntax element.
10. The apparatus according to claim 9, wherein, The second value specified by the quality metric syntax element includes a 16-bit unsigned integer and the integer value is 100.
11. The apparatus according to claim 10, wherein, The multiple types of quality metrics include the PSNR and at least one of the following: structural similarity index (SSIM), weighted PSNR (wPSNR), or weighted to spherical uniform PSNR (WS-PSNR).
12. The apparatus according to claim 10, wherein, The one or more processors are further configured to: receive a second image comprising a plurality of second sub-images; The second quality metric SEI message associated with multiple second sub-images of the second image is decoded. The second quality metric SEI message includes: a second syntax element indicating the number of second sub-images of the multiple second sub-images; a second quality metric type syntax element associated with at least one second sub-image of the multiple second sub-images; a third value set to indicate a third value among multiple values of multiple types of the multiple quality metrics, wherein the third value is set to 1; and a second quality metric syntax element specifying a fourth value. The second quality metric type syntax element is set to 1; based on the decision that the second quality metric type syntax element is set to 1, the fourth value specified by the second quality metric syntax element is interpreted as an indication of the SSIM value associated with the at least one second sub-image; and based on the fourth value specified by the second quality metric syntax element, the deduced SSIM associated with the at least one second sub-image is derived.
13. The apparatus according to claim 12, wherein, The one or more processors are further configured to: determine the second quotient based on the second integer value and the fourth value specified by the second quality metric syntax element.
14. The apparatus according to claim 10, wherein, The one or more processors are further configured to: apply post-processing techniques to the at least one sub-image based on the derived PSNR to form a processed image; and display the processed image.
15. An apparatus configured to process video data, the apparatus comprising: A component for receiving an image that includes multiple sub-images; The system also includes components for decoding Quality Metric Supplemental Enhancement Information (SEI) messages associated with multiple sub-images of the image, the SEI messages including: a first syntax element indicating the number of sub-images of the multiple sub-images; a quality metric type syntax element associated with at least one sub-image of the multiple sub-images, set to indicate a first value among multiple values of multiple types of multiple quality metrics, wherein the first value is set to 0; and a quality metric syntax element specifying a second value; components for determining that the quality metric type syntax element is set to 0; components for determining, based on the determination that the quality metric type syntax element is set to 0, to interpret the second value specified by the quality metric syntax element as indicating a Peak Signal-to-Noise Ratio (PSNR) value associated with the at least one sub-image; and components for deriving a derived PSNR associated with the at least one sub-image based on the second value specified by the quality metric syntax element. The apparatus according to claim 15, wherein the component for deriving the derived PSNR associated with the at least one sub-image includes a component for determining the second value quotient based on an integer value and the quality metric syntax element.
16. 。 17. The apparatus according to claim 16, wherein, The second value specified by the quality metric syntax element includes a 16-bit unsigned integer and the integer value is 100.
18. The apparatus according to claim 16, wherein, The multiple types of quality metrics include the PSNR and at least one of the following: structural similarity index (SSIM), weighted PSNR (wPSNR), or weighted to spherical uniform PSNR (WS-PSNR).
19. The apparatus according to claim 17, further comprising: A component for receiving a second image that includes multiple second sub-images; A means for decoding a second quality metric SEI message associated with a plurality of second sub-images of the second image, the second quality metric SEI message including: a second syntax element indicating the number of second sub-images of the plurality of second sub-images; a second quality metric type syntax element associated with at least one of the plurality of second sub-images; a third value set to indicate a plurality of values of a plurality of types of the plurality of quality metrics, wherein the third value is set to 1; and a second quality metric syntax element specifying a fourth value; a means for determining that the second quality metric type syntax element is set to 1; a means for determining, based on the determination that the second quality metric type syntax element is set to 1, to interpret the fourth value specified by the second quality metric syntax element as indicating an SSIM value associated with the at least one second sub-image; and a means for deriving a deduced SSIM associated with the at least one second sub-image based on the fourth value specified by the second quality metric syntax element. According to the apparatus of claim 19, wherein the component for deriving the derived SSIM associated with the at least one second sub-image includes a component for determining a second quotient based on a second integer value and the fourth value specified by the second quality metric syntax element.
20. 。 21. The apparatus according to claim 17, further comprising: A component for applying post-processing techniques to the at least one sub-image based on the derived PSNR to form a processed image; And components for displaying the processed image.
22. A non-transitory computer-readable storage medium storing instructions, which, when executed, cause one or more processors of a device configured to process video data to: receive an image comprising a plurality of sub-images; And decoding of Quality Metric Supplemental Enhancement Information (SEI) messages associated with multiple sub-images of the image, the quality metric supplemental enhancement information messages including: A first syntax element indicating the number of sub-images of the plurality of sub-images; a quality metric type syntax element associated with at least one sub-image of the plurality of sub-images, set to indicate a first value among a plurality of values of a plurality of types of a plurality of quality metrics, wherein the first value is set to 0; and a quality metric syntax element specifying a second value; determining that the quality metric type syntax element is set to 0; based on the determination that the quality metric type syntax element is set to 0, determining that the second value specified by the quality metric syntax element is interpreted as indicating a peak signal-to-noise ratio (PSNR) value associated with the at least one sub-image; And based on the second value specified by the quality metric syntax element, derive the derived PSNR associated with the at least one sub-image.
23. The non-transitory computer-readable storage medium according to claim 22, wherein, The instructions also cause the one or more processors to perform the following operation: determine the quotient based on the integer value and the second value specified by the quality metric syntax element.
24. The non-transitory computer-readable storage medium according to claim 23, wherein, The second value specified by the quality metric syntax element includes a 16-bit unsigned integer and the integer value is 100.
25. The non-transitory computer-readable storage medium according to claim 24, wherein, The multiple types of quality metrics include the PSNR and at least one of the following: structural similarity index (SSIM), weighted PSNR (wPSNR), or weighted to spherical uniform PSNR (WS-PSNR).
26. The non-transitory computer-readable storage medium according to claim 22, wherein, The instructions also cause the one or more processors to perform the following operation: receive a second image comprising a plurality of second sub-images; The second quality metric SEI message associated with multiple second sub-images of the second image is decoded. The second quality metric SEI message includes: a second syntax element indicating the number of second sub-images of the multiple second sub-images; a second quality metric type syntax element associated with at least one second sub-image of the multiple second sub-images; a third value set to indicate a third value among multiple values of multiple types of the multiple quality metrics, wherein the third value is set to 1; and a second quality metric syntax element specifying a fourth value. The second quality metric type syntax element is set to 1; based on the decision that the second quality metric type syntax element is set to 1, the fourth value specified by the second quality metric syntax element is interpreted as an indication of the SSIM value associated with the at least one second sub-image; and based on the fourth value specified by the second quality metric syntax element, the deduced SSIM associated with the at least one second sub-image is derived.
27. The non-transitory computer-readable storage medium according to claim 26, wherein, The instructions also cause the one or more processors to perform the following operation: determine a second quotient based on a second integer value and the fourth value specified by the second quality metric syntax element.
28. The non-transitory computer-readable storage medium according to claim 24, wherein, The device is configured to decode the quality metric message, and the instructions further cause the one or more processors to: apply post-processing techniques to the at least one sub-image based on the derived PSNR to form a processed image; and display the processed image.