Low-complexity adaptive quantization for video compression.

Adaptive quantization in video encoding reduces computational complexity by determining quantization offsets based on side information, enhancing compression efficiency and enabling parallel processing of transform coefficients.

JP7681028B2Active Publication Date: 2025-05-21QUALCOMM INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022546463
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-02-03
Filing Date
2021-02-04
Publication Date
2025-05-21
Estimated Expiration
2041-02-04

AI Technical Summary

Technical Problem

Modern video encoders face high computational complexity due to complex entropy coding contexts and multiple passes for quantizing transform coefficients, making it expensive and inefficient.

Method used

Adaptive quantization techniques are employed by determining quantization offset parameters based on side information associated with video blocks, allowing for parallel quantization of transform coefficients without bit cost estimates, reducing computational complexity.

Benefits of technology

This approach improves compression efficiency and reduces processing cycles while maintaining effective quantization, enabling parallel computation of transform coefficients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007681028000075
    Figure 0007681028000075
  • Figure 0007681028000076
    Figure 0007681028000076
  • Figure 0007681028000077
    Figure 0007681028000077
Patent Text Reader

Abstract

The video encoder may determine, based on side information associated with the block of video data, a set of quantization offset parameters for the group of scaled transform coefficients for the block of video data. The video encoder may further quantize the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data based at least in part on the set of quantization offset parameters. The video encoder may further generate an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data.
Need to check novelty before this filing date? Find Prior Art

Description

Claiming priority

[0001]

[0001] This application claims priority to U.S. Application No. 17 / 166,639, filed February 3, 2021, and U.S. Provisional Application No. 62 / 970,588, filed February 5, 2020, the entire contents of each of which are incorporated herein by reference. U.S. Application No. 17 / 166,639, filed February 3, 2021, claims the benefit of U.S. Provisional Application No. 62 / 970,588, filed February 5, 2020. [Technical field]

[0002]

[0002] This disclosure relates to video encoding and decoding. [Background technology]

[0003] Digital video capabilities may be incorporated into a wide range of devices, including digital televisions, digital direct broadcast systems, wireless broadcast systems, personal digital assistants (PDAs), laptop or desktop computers, tablet computers, e-book readers, digital cameras, digital recording devices, digital media players, video gaming devices, video game consoles, cellular or satellite radio telephones, so-called "smartphones," video teleconferencing devices, video streaming devices, and the like. Digital video devices implement video coding techniques, such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4, Part 10, Advanced Video Coding (AVC), ITU-T H.265 / High Efficiency Video Coding (HEVC), and extensions to such standards. By implementing such video coding techniques, video devices may more efficiently transmit, receive, encode, decode, and / or store digital video information.

[0004]

[0004] Video coding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in video sequences. In block-based video coding, a video slice (e.g., a video picture or a portion of a video picture) may be partitioned into video blocks, which may also be referred to as coding tree units (CTUs), coding units (CUs) and / or coding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction with respect to reference samples in neighboring blocks in the same picture. Video blocks in an inter-coded (P or B) slice of a picture may use spatial prediction with respect to reference samples in adjacent blocks in the same picture or temporal prediction with respect to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. Summary of the Invention

[0005]

[0005] Generally, this disclosure describes techniques for adaptive quantization of transform coefficients for encoding video by determining quantization offsets for quantizing the transform coefficients. In particular, instead of quantizing the transform coefficients based on bit cost estimates for optimally quantizing the transform coefficients determined for entropy coding of the quantized transform coefficients, a video encoder may determine a set of quantization offset parameters for a group of transform coefficients in a block of video data based on side information associated with the block of video data, and may quantize the group of transform coefficients based on the set of quantization offset parameters to generate near-optimal quantized transform coefficients.

[0006]

[0006] In accordance with the techniques of this disclosure, a video encoder may split transform coefficients of a block of video data into groups of transform coefficients. For each of the groups of transform coefficients, the video encoder may determine a set of quantization offset parameters associated with the group of transform coefficients based on side information about the block of video data, such as a slice type of the block of video data and / or an indication of whether the block of video data comprises a luminance or chrominance component. The video encoder may thus quantize each group of transform coefficients based on the quantization offset parameters associated with the group of transform coefficients, with performance improvement based on selection of the best offset.

[0007]

[0007] The technical problem solved by the techniques of this disclosure relates to the fact that entropy coding performed by modern video encoders can be very complex due to the large number of arithmetic coding contexts and complex context selection rules. Furthermore, modern video coding may encode transform coefficients in multiple passes. For example, entropy coding in modern video encoders may be performed in up to five passes. This makes it potentially complex and computationally expensive to calculate and use bit cost estimates for optimally quantizing the transform coefficients for each decision.

[0008]

[0008] In contrast, by refraining from utilizing bit cost estimates for quantizing transform coefficients determined during entropy coding of the quantized transform coefficients to quantize the transform coefficients, the techniques of this disclosure improve compression due to the lower computational complexity of quantizing the transform coefficients, thereby allowing the video encoder to utilize fewer processing cycles to quantize the transform coefficients. Furthermore, because the techniques of this disclosure may determine a single set of quantization offsets used to quantize each transform coefficient in the group of transform coefficients, the techniques of this disclosure allow the video encoder to quantize the transform coefficients in the group of transform coefficients in parallel (e.g., the video encoder may quantize the transform coefficients in a first group of transform coefficients simultaneously or overlapping in time with quantizing the transform coefficients in a second group of transform coefficients).

[0009]

[0009] One or more computer systems may be configured to perform particular operations or actions by having installed on the system software, firmware, hardware, or a combination thereof that, during operation, causes the system to perform the actions. One or more computer programs may be configured to perform particular operations or actions by including instructions that, when executed by a data processing device, cause the device to perform the actions.

[0010]

[0010] One general aspect includes a method of decoding video data. The method includes determining a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data based on side information associated with the block of video data. The method further includes quantizing the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data based at least in part on the set of quantization offset parameters. The method further includes generating an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data.

[0011]

[0011] One general aspect includes a device for encoding video data, the device including a memory, the device further including a processing circuit in communication with the memory, the processing circuit configured to determine a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data based on side information associated with the block of video data, quantize the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data based at least in part on the set of quantization offset parameters, and generate an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data.

[0012]

[0012] One general aspect includes an apparatus for decoding video data. The apparatus includes means for determining a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data based on side information associated with the block of video data. The apparatus further includes means for quantizing the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data based at least in part on the set of quantization offset parameters. The apparatus further includes means for generating an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data.

[0013]

[0013] One general aspect includes a computer-readable storage medium having stored thereon instructions that, when executed by one or more processors, cause the one or more processors to determine a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data based on side information associated with the block of video data, quantize the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data based at least in part on the set of quantization offset parameters, and generate an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data.

[0014] The details of one or more examples are set forth in the accompanying drawings and the description below. Other features, objects, and advantages will become apparent from the description, drawings, and claims. [Brief description of the drawings]

[0015] [Figure 1]

[0015] FIG. 1 is a block diagram illustrating an example video encoding and decoding system in which the techniques of this disclosure may be implemented. [Figure 2A]

[0016] 1 is a conceptual diagram illustrating an exemplary quad-tree binary tree (QTBT) structure. [Figure 2B] A conceptual diagram showing a corresponding coding tree unit (CTU). [Figure 3A]

[0017] 1 illustrates a video coding system for adaptive and / or rate-distortion optimal quantization. [Figure 3B]

[0018] 1 illustrates a video coding system for low-complexity adaptive quantization based on block classification and one or more sets of quantization offset parameters, in accordance with techniques of this disclosure. [Figure 4]

[0019] 4 illustrates a parallel implementation of adaptive quantization using a set of quantization offset parameters in accordance with techniques of this disclosure. [Diagram 5]

[0020] 13 illustrates coefficients for sign bit concealment to approximate rate-distortion cost. [Figure 6]

[0021] 4A-4C show techniques for determining a quantization offset parameter in accordance with techniques of this disclosure. [Figure 7]

[0022] 4 is a diagram showing an example of code for identifying the location of a sub-block within a block of video data. [Figure 8]

[0023] 1 is a block diagram illustrating an example video encoder that may implement the techniques of this disclosure. [Figure 9]

[0024] 1 is a block diagram illustrating an example video decoder that may implement the techniques of this disclosure. [Figure 10]

[0025] 1 is a flowchart illustrating an example method for encoding a current block. [Figure 11]

[0026] 4 is a flowchart illustrating an example method for decoding a current block of video data. [Figure 12]

[0027] 1 is a flowchart illustrating a method for encoding video data in accordance with techniques of this disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0016]

[0028] 1 is a block diagram illustrating an example video encoding and decoding system 100 that may implement the techniques of this disclosure. The techniques of this disclosure are generally directed to coding (encoding and / or decoding) video data. Generally, video data includes any data for processing video. Thus, video data may include raw uncoded video, coded video, decoded (e.g., reconstructed) video, and video metadata, such as signaling data.

[0017]

[0029] 1, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116, in this example. In particular, source device 102 provides the video data to destination device 116 via a computer-readable medium 110. Source device 102 and destination device 116 may comprise any of a wide range of devices, including desktop computers, notebook (i.e., laptop) computers, mobile devices, tablet computers, set-top boxes, telephone handsets such as smartphones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, broadcast receiver devices, and the like. In some cases, source device 102 and destination device 116 may be equipped for wireless communication and thus may be referred to as wireless communication devices.

[0018]

[0030] In the example of FIG. 1, source device 102 includes a video source 104, memory 106, video encoder 200, and output interface 108. Destination device 116 includes an input interface 122, a video decoder 300, memory 120, and a display device 118. According to this disclosure, the video encoder 200 of source device 102 and the video decoder 300 of destination device 116 may be configured to apply techniques for quantizing transform coefficients. Thus, source device 102 represents an example of a video encoding device, and destination device 116 represents an example of a video decoding device. In other examples, the source device and destination device may include other components or arrangements. For example, source device 102 may receive video data from an external video source, such as an external camera. Similarly, destination device 116 may interface with an external display device rather than including an integrated display device.

[0019]

[0031] The system 100 shown in FIG. 1 is merely an example. In general, any digital video encoding and / or decoding device may implement techniques for quantizing transform coefficients. The source device 102 and the destination device 116 are merely examples of coding devices, such that the source device 102 generates coded video data for transmission to the destination device 116. This disclosure refers to a "coding" device as a device that performs coding (encoding and / or decoding) of data. Thus, the video encoder 200 and the video decoder 300 represent examples of coding devices, particularly video encoders and video decoders, respectively. In some examples, the source device 102 and the destination device 116 may operate substantially symmetrically, such that each of the source device 102 and the destination device 116 includes video encoding and decoding components. Thus, the system 100 may support unidirectional or bidirectional video transmission between the source device 102 and the destination device 116, for example, video streaming, video playback, video broadcasting, or video telephony.

[0020]

[0032] Generally, the video source 104 represents a source of video data (i.e., raw, unencoded video data) and provides a continuous series of pictures (also called “frames”) of the video data to the video encoder 200, which encodes the data for the pictures. The video source 104 of the source device 102 may include a video capture device, such as a video camera, a video archive containing previously captured raw video, and / or a video feed interface for receiving video from a video content provider. As a further alternative, the video source 104 may generate computer graphics-based data as the source video, or a combination of live, archived, and computer-generated video. In each case, the video encoder 200 encodes the captured, pre-captured, or computer-generated video data. The video encoder 200 may reorder the pictures from a receiving order (sometimes called a “display order”) into a coding order for coding. The video encoder 200 may generate a bitstream including the encoded video data. The source device 102 may then output the encoded video data via the output interface 108 onto a computer-readable medium 110 for receipt and / or retrieval by, for example, an input interface 122 of the destination device 116 .

[0021]

[0033] The memory 106 of the source device 102 and the memory 120 of the destination device 116 represent general-purpose memories. In some examples, the memories 106, 120 may store raw video data, e.g., raw video from the video source 104 and raw decoded video data from the video decoder 300. Additionally or alternatively, the memories 106, 120 may store software instructions executable by, e.g., the video encoder 200 and the video decoder 300, respectively. Although the memories 106 and 120 are shown in this example separately from the video encoder 200 and the video decoder 300, it should be understood that the video encoder 200 and the video decoder 300 may also include internal memories for functionally similar or equivalent purposes. Additionally, the memories 106, 120 may store encoded video data, e.g., the output from the video encoder 200 and the input to the video decoder 300. In some examples, portions of the memory 106, 120 may be allocated as one or more video buffers, for example, to store raw decoded and / or encoded video data.

[0022]

[0034] The computer-readable medium 110 may represent any type of medium or device capable of transporting encoded video data from the source device 102 to the destination device 116. In one example, the computer-readable medium 110 represents a communication medium for enabling the source device 102 to transmit the encoded video data directly to the destination device 116 in real time, for example, via a radio frequency network or a computer-based network. The output interface 108 may modulate a transmission signal including the encoded video data, and the input interface 122 may demodulate a received transmission signal according to a communication standard such as a wireless communication protocol. The communication medium may comprise any wireless or wired communication medium, such as a radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network, such as a local area network, a wide area network, or a global network such as the Internet. The communication medium may include routers, switches, base stations, or any other equipment that may be useful to facilitate communication from the source device 102 to the destination device 116.

[0023]

[0035] In some examples, source device 102 may output the encoded data from output interface 108 to storage device 112. Similarly, destination device 116 may access the encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a variety of distributed or locally accessed data storage media, such as a hard drive, Blu-ray disc, DVD, CD-ROM, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.

[0024]

[0036] In some examples, the source device 102 may output the encoded video data to a file server 114 or another intermediate storage device that may store the encoded video data generated by the source device 102. The destination device 116 may access the stored video data from the file server 114 via streaming or download. The file server 114 may be any type of server device capable of storing the encoded video data and transmitting the encoded video data to the destination device 116. The file server 114 may represent a web server (e.g., for a website), a file transfer protocol (FTP) server, a content delivery network device, or a network attached storage (NAS) device. The destination device 116 may access the encoded video data from the file server 114 through any standard data connection, including an Internet connection. This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both, that is suitable for accessing the encoded video data stored on the file server 114. The file server 114 and the input interface 122 may be configured to operate according to a streaming transmission protocol, a download transmission protocol, or a combination thereof.

[0025]

[0037] The output interface 108 and the input interface 122 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples in which the output interface 108 and the input interface 122 comprise wireless components, the output interface 108 and the input interface 122 may be configured to transfer data, such as encoded video data, according to a cellular communication standard, such as 4G, 4G-LTE (Long Term Evolution), LTE Advanced, 5G, etc. In some examples in which the output interface 108 comprises a wireless transmitter, the output interface 108 and the input interface 122 may be configured to transfer data, such as encoded video data, according to other wireless standards, such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee), the Bluetooth standard, etc. In some examples, the source device 102 and / or the destination device 116 may include respective system-on-chip (SoC) devices. For example, the source device 102 may include a SoC device for performing functions attributed to the video encoder 200 and / or the output interface 108, and the destination device 116 may include a SoC device for performing functions attributed to the video decoder 300 and / or the input interface 122.

[0026]

[0038] The techniques of this disclosure may be applied to video coding supporting any of a variety of multimedia applications, such as over-the-air television broadcast, cable television transmission, satellite television transmission, Internet streaming video transmission such as Dynamic Adaptive Streaming over HTTP (DASH), digital video encoded on a data storage medium, decoding of digital video stored on a data storage medium, or other applications.

[0027]

[0039] The input interface 122 of the destination device 116 receives the encoded video bitstream from the computer-readable medium 110 (e.g., a communication medium, a storage device 112, a file server 114, etc.). The encoded video bitstream may include signaling information defined by the video encoder 200 that is also used by the video decoder 300, such as syntax elements having values ​​that describe characteristics and / or processing of video blocks or other coded units (e.g., slices, tiles, bricks, pictures, picture groups, sequences, etc.). The display device 118 displays decoded pictures of the decoded video data to a user. The display device 118 may represent any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light emitting diode (OLED) display, or another type of display device.

[0028]

[0040] 1, in some examples, the video encoder 200 and the video decoder 300 may each be integrated with an audio encoder and / or decoder and may include appropriate MUX-DEMUX units, or other hardware and / or software, to handle multiplexed streams that include both audio and video in a common data stream. Where applicable, the MUX-DEMUX units may conform to the ITU H.223 multiplexer protocol or other protocols, such as the User Datagram Protocol (UDP).

[0029]

[0041] The video encoder 200 and the video decoder 300 may each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the techniques are implemented partially in software, the device may store the software's instructions on a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to implement the techniques of the present disclosure. Each of the video encoder 200 and the video decoder 300 may be included within one or more encoders or decoders, any of which may be integrated as part of a combined encoder / decoder (CODEC) in the respective device. The device including the video encoder 200 and / or the video decoder 300 may comprise an integrated circuit, a microprocessor, and / or a wireless communication device such as a mobile phone.

[0030]

[0042] Video encoder 200 and video decoder 300 may operate according to a video coding standard, such as ITU-T H.265, also known as High Efficiency Video Coding (HEVC), or extensions thereof, such as multiview and / or scalable video coding extensions. Alternatively, video encoder 200 and video decoder 300 may operate according to other proprietary or industry standards, such as ITU-T H.266, also known as Generic Video Coding (VVC). A recent draft of the VVC standard is set forth in Bross et al., "Versatile Video Coding (Draft 10)," Joint Video Experts Team (JVET) of ITU-T SG16 WP3 and ISO / IEC JTC1 / SC29 / WG11, 18th Meeting: Teleconference, June 22-July 1, 2020, JVET-S2001-vA (hereinafter "VVC Draft 10"), available under https: / / jvet-experts.org / . However, the techniques of this disclosure are not limited to any particular coding standard.

[0031]

[0043] Generally, the video encoder 200 and the video decoder 300 may perform block-based coding of pictures. The term “block” generally refers to a structure that includes data to be processed (e.g., encoded, decoded, or otherwise used in an encoding and / or decoding process). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. Generally, the video encoder 200 and the video decoder 300 may code video data represented in a YUV (e.g., Y, Cb, Cr) format. That is, rather than coding red, green, and blue (RGB) data for a sample of a picture, the video encoder 200 and the video decoder 300 may code a luminance component and a chrominance component, where the chrominance component may include both red and blue hues of chrominance components. In some examples, the video encoder 200 converts received RGB format data to a YUV representation prior to encoding, and the video decoder 300 converts the YUV representation to the RGB format. Alternatively, pre-processing and post-processing units (not shown) may perform these conversions. However, the techniques of this disclosure are not limited to YUV representations, but may be applied to RGB representations with three color components.

[0032]

[0044] This disclosure may generally refer to coding (e.g., encoding and decoding) a picture to include processes of encoding or decoding data for a picture. Similarly, this disclosure may refer to coding a block of a picture to include processes of encoding or decoding data for the block, e.g., predictive and / or residual coding. A coded video bitstream generally includes a series of values ​​for syntax elements that represent coding decisions (e.g., coding modes) and partitioning of a picture into blocks. Thus, references to coding a picture or a block should be understood generally as coding values ​​for the syntax elements that form the picture or block.

[0033]

[0045] HEVC defines various blocks, including coding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video coder (such as video encoder 200) partitions coding tree units (CTUs) into CUs according to a quadtree structure. That is, the video coder partitions CTUs and CUs into four equal non-overlapping squares, and each node of the quadtree has either zero or four child nodes. A node without children may be called a "leaf node", and a CU of such a leaf node may include one or more PUs and / or one or more TUs. The video coder may further partition PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents a partition of TUs. In HEVC, a PU represents inter-predicted data, while a TU represents residual data. An intra-predicted CU includes intra-prediction information, such as an intra-mode indication.

[0034]

[0046] As another example, the video encoder 200 and the video decoder 300 may be configured to operate according to VVC. According to VVC, a video coder (such as the video encoder 200) partitions a picture into multiple coding tree units (CTUs). The video encoder 200 may partition the CTUs according to a tree structure, such as a quad-tree binary tree (QTBT) structure or a multi-type tree (MTT) structure. The QTBT structure eliminates the concept of multiple partition types, such as the distinction between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels: a first level partitioned according to a quad-tree partition, and a second level partitioned according to a binary tree partition. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to coding units (CUs).

[0035]

[0047] In the MTT partitioning structure, blocks may be partitioned using quad tree (QT) partitioning, binary tree (BT) partitioning, and one or more types of triple tree (TT) (also called ternary tree (TT)) partitioning. A triple tree or ternary tree partitioning is a partitioning in which a block is split into three sub-blocks. In some examples, a triple tree or ternary tree partitioning divides a block into three sub-blocks without splitting the original block through the center. The partitioning types in MTT (e.g., QT, BT, and TT) may be symmetric or asymmetric.

[0036]

[0048] In some examples, video encoder 200 and video decoder 300 may use a single QTBT or MTT structure to represent each of the luminance and chrominance components, and in other examples, video encoder 200 and video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luminance component and another QTBT / MTT structure for both chrominance components (or two QTBT / MTT structures for each chrominance component).

[0037]

[0049] Video encoder 200 and video decoder 300 may be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, or other partition structures according to HEVC. For purposes of explanation, the description of the techniques of this disclosure is presented with respect to QTBT partitioning. However, it should be understood that the techniques of this disclosure may also be applied to video coders configured to use quadtree partitioning, or other types of partitioning as well.

[0038]

[0050] In some examples, the CTU includes a coding tree block (CTB) of luma samples, two corresponding CTBs of chroma samples for a picture with three sample arrays, or a CTB of samples for a monochrome picture, or a picture coded using three separate color planes and syntax structures used to code the samples. The CTB can be an N×N block of samples for some value of N such that the division of the components into the CTB is partitioned. The components are arrays or single samples from one of the three arrays (luma and two chromas) that configure the picture in 4:2:0, 4:2:2, or 4:4:4 color format, or an array or single sample of an array that configures the picture in monochrome format. In some examples, the coding block is an M×N block of samples for some value of M and N such that the division of the CTB into the coding block is partitioned.

[0039]

[0051] Blocks (e.g., CTUs or CUs) may be grouped in various ways in a picture. As an example, a brick may refer to a rectangular region of a CTU row in a particular tile in a picture. A tile may be a rectangular region of a CTU in a particular tile column and in a particular tile row in a picture. A tile column refers to a rectangular region of a CTU having a height equal to the height of the picture and a width specified by a syntax element (e.g., in a picture parameter set). A tile row refers to a rectangular region of a CTU having a height specified by a syntax element (e.g., in a picture parameter set) and a width equal to the width of the picture.

[0040]

[0052] In some examples, a tile may be partitioned into multiple bricks, each of which may include one or more CTU rows within the tile. A tile that is not partitioned into multiple bricks may also be referred to as a brick. However, a brick that is a true subset of a tile may not be referred to as a tile.

[0041]

[0053] The bricks in a picture may also be arranged into slices. A slice may be an integer number of bricks of a picture that may be contained entirely in a single Network Abstraction Layer (NAL) unit. In some examples, a slice includes either several complete tiles or only a continuous sequence of complete bricks of one tile.

[0042]

[0054] In this disclosure, "N x N" and "N by N" may be used interchangeably to refer to the sample dimensions of a block (such as a CU or other video block) in terms of vertical and horizontal dimensions, e.g., 16 x 16 samples or 16 by 16 samples. Generally, a 16 x 16 CU has 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an N x N CU generally has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. Samples in a CU may be arranged in rows and columns. Moreover, a CU does not necessarily have the same number of samples in the horizontal direction as in the vertical direction. For example, a CU may comprise N x M samples, where M is not necessarily equal to N.

[0043]

[0055] The video encoder 200 encodes video data for a CU that represents prediction and / or residual information, as well as other information. The prediction information indicates how the CU should be predicted to form a predictive block for the CU. The residual information generally represents sample-by-sample differences between samples of the CU prior to encoding and the predictive block.

[0044]

[0056] To predict a CU, the video encoder 200 may generally form a predictive block for the CU through inter-prediction or intra-prediction. Inter-prediction generally refers to predicting a CU from data of a previously coded picture, while intra-prediction generally refers to predicting a CU from previously coded data of the same picture. To perform inter-prediction, the video encoder 200 may generate a predictive block using one or more motion vectors. The video encoder 200 may generally perform a motion search to identify a reference block that closely matches the CU, for example, with respect to the difference between the CU and the reference block. The video encoder 200 may calculate a difference metric using a sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculation to determine whether the reference block closely matches the current CU. In some examples, the video encoder 200 may predict the current CU using unidirectional or bidirectional prediction.

[0045]

[0057] Some examples of VVC also provide an affine motion compensation mode, which may be considered an inter-prediction mode. In an affine motion compensation mode, video encoder 200 may determine two or more motion vectors that represent non-translational motion, such as zooming in or out, rotation, perspective motion, or other irregular motion types.

[0046]

[0058] To perform intra prediction, the video encoder 200 may select an intra prediction mode to generate a predictive block. Some examples of VVC provide 67 intra prediction modes, including various directional modes, as well as planar and DC modes. In general, the video encoder 200 selects an intra prediction mode that describes neighboring samples relative to a current block (e.g., a block of a CU) from which samples of the current block should be predicted. Such samples may generally be above, above and to the left, or to the left of the current block in the same picture as the current block, assuming that the video encoder 200 codes the CTUs and CUs in raster scan order (left to right, top to bottom).

[0047]

[0059] Video encoder 200 encodes data representing a prediction mode for the current block. For example, in an inter prediction mode, video encoder 200 may encode data representing which of various available inter prediction modes is used, as well as motion information of the corresponding mode. For example, in unidirectional or bidirectional inter prediction, video encoder 200 may encode motion vectors using advanced motion vector prediction (AMVP) or merge mode. Video encoder 200 may use similar modes to encode motion vectors for affine motion compensation modes.

[0048]

[0060] Following prediction, such as intra- or inter-prediction, of a block, the video encoder 200 may calculate residual data for the block. The residual data, such as a residual block, represents sample-by-sample differences between the block and a predictive block for the block formed using a corresponding prediction mode. The video encoder 200 may apply one or more transforms to the residual block to produce a block, such as a transform block (TB) or a transform coefficient block, of data transformed into a transform domain rather than the sample domain. For example, the video encoder 200 may apply a discrete cosine transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. Additionally, the video encoder 200 may apply a secondary transform, such as a mode-dependent non-separable secondary transform (MDNSST), a signal-dependent transform, a Karhunen-Loeve transform (KLT), etc., following the first transform. The video encoder 200 produces the transform coefficients following application of the one or more transforms.

[0049]

[0061] As mentioned above, following any transformation to produce transform coefficients, the video encoder 200 may perform quantization of the transform coefficients. Quantization generally refers to a process in which transform coefficients are quantized to possibly reduce the amount of data used to represent the transform coefficients, providing further compression. By performing a quantization process, the video encoder 200 may reduce the bit depth associated with some or all of the transform coefficients. For example, the video encoder 200 may truncate an n-bit value to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder 200 may perform a bitwise right shift of the value to be quantized.

[0050]

[0062] According to aspects of this disclosure, video encoder 200 may split transform coefficients for a block of pixels or samples, such as a transform block, into blocks of transform coefficients. For each group of transform coefficients, video encoder 200 may determine a set of quantization offset parameters associated with the group of transform coefficients based on side information about the block of video pixels or samples. Video encoder 200 may quantize each group of transform coefficients based on the set of quantization offset parameters associated with the respective group of transform coefficients.

[0051]

[0063] Following quantization, the video encoder 200 may scan the transform coefficients to produce a one-dimensional vector from a two-dimensional matrix including the quantized transform coefficients. The scan may be designed to place higher energy (and therefore lower frequency) transform coefficients at the front of the vector and lower energy (and therefore higher frequency) transform coefficients at the rear of the vector. In some examples, the video encoder 200 may utilize a predefined scan order to scan the quantized transform coefficients to produce a serialized vector, and then entropy code the quantized transform coefficients of the vector. In other examples, the video encoder 200 may perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 200 may entropy code the one-dimensional vector, for example, according to context-adaptive binary arithmetic coding (CABAC). The video encoder 200 may also entropy code values ​​for syntax elements describing metadata associated with the encoded video data for use by the video decoder 300 in decoding the video data.

[0052]

[0064] To perform CABAC, the video encoder 200 may assign a context in a context model to a symbol to be transmitted. The context may relate, for example, to whether neighboring values ​​of the symbol are zero-valued. The probability determination may be based on the context assigned to the symbol.

[0053]

[0065] The video encoder 200 may further generate syntax data, such as block-based syntax data, picture-based syntax data, and sequence-based syntax data, for the video decoder 300, e.g., among other syntax data, such as a picture header, a block header, a slice header, or a sequence parameter set (SPS), a picture parameter set (PPS), or a video parameter set (VPS). The video decoder 300 may similarly decode such syntax data to determine how to decode the corresponding video data. The side information regarding the block of video data may be syntax data associated with the block of video data, e.g., block-based syntax data.

[0054]

[0066] In this manner, video encoder 200 may generate a bitstream including syntax elements that describe encoded video data, e.g., partitions of a picture into blocks (e.g., CUs) and predictive and / or residual information for the blocks. Finally, video decoder 300 may receive the bitstream and decode the encoded video data.

[0055]

[0067] In general, the video decoder 300 performs an inverse process to that performed by the video encoder 200 to decode the encoded video data of the bitstream. For example, the video decoder 300 may decode values ​​of syntax elements of the bitstream using CABAC in a manner that is inverse to, but substantially similar to, the CABAC encoding process of the video encoder 200. The syntax elements may define partition information for partitioning a picture into CTUs and partitions of each CTU according to a corresponding partition structure, such as a QTBT structure, to define CUs of the CTU. The syntax elements may further define prediction and residual information for blocks (e.g., CUs) of video data.

[0056]

[0068] The residual information may be represented, for example, by quantized transform coefficients. The video decoder 300 may dequantize and inverse transform the quantized transform coefficients of the block to reconstruct a residual block of the block. The video decoder 300 uses the prediction mode (intra- or inter-prediction) signaled in the bitstream and associated prediction information (e.g., motion information for inter-prediction) to form a predictive block for the block. The video decoder 300 may then combine (sample by sample) the predictive block and the residual block to reconstruct the original block. The video decoder 300 may perform additional processing, such as performing a deblocking process to reduce visual artifacts along block boundaries.

[0057]

[0069] According to the techniques of this disclosure, video encoder 200 may determine a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data based on side information associated with the block of video data, quantize the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data based at least in part on the set of quantization offset parameters, and generate an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data.

[0058]

[0070] This disclosure may generally refer to "signaling" certain information, such as syntax elements. The term "signaling" may generally refer to communication of values ​​for syntax elements and / or other data used to decode encoded video data. That is, video encoder 200 may signal values ​​for syntax elements in a bitstream. Generally, signaling refers to generating values ​​in the bitstream. As described above, source device 102 may transport the bitstream to destination device 116 in substantially real-time or may transport the bitstream to destination device 116 in non-real-time, such as may be done when storing syntax elements to storage device 112 for later retrieval by destination device 116.

[0059]

[0071] 2A and 2B are conceptual diagrams illustrating an exemplary quad-tree bi-tree (QTBT) structure 130 and a corresponding coding tree unit (CTU) 132. Solid lines represent quad-tree splitting, and dotted lines indicate bi-tree splitting. At each split (i.e., non-leaf) node of the bi-tree, one flag is signaled to indicate which splitting type (i.e., horizontal or vertical) is used, where in this example, 0 indicates horizontal splitting and 1 indicates vertical splitting. In quad-tree splitting, there is no need to indicate the splitting type, since the quad-tree node splits a block horizontally and vertically into four sub-blocks with equal size. Thus, video encoder 200 may encode, and video decoder 300 may decode, syntax elements (e.g., solid lines) (e.g., splitting information) for the region tree level (i.e., first level) of QTBT structure 130 and syntax elements (e.g., dashed lines) (e.g., splitting information) for the prediction tree level (i.e., second level) of QTBT structure 130. Video encoder 200 may encode, and video decoder 300 may decode, video data, such as prediction and transform data, for CUs represented by terminal leaf nodes of QTBT structure 130.

[0060]

[0072] 2B may be associated with parameters that define the sizes of blocks corresponding to nodes of the QTBT structure 130 at the first and second levels. These parameters may include a CTU size (representing the size of the CTU 132 in a sample), a minimum quadtree size (MinQTSize representing the minimum allowed quadtree leaf node size), a maximum bi-tree size (MaxBTSize representing the maximum allowed bi-tree root node size), a maximum bi-tree depth (MaxBTDepth representing the maximum allowed bi-tree depth), and a minimum bi-tree size (MinBTSize representing the minimum allowed bi-tree leaf node size).

[0061]

[0073] The root node of the QTBT structure corresponding to the CTU may have four child nodes at the first level of the QTBT structure, each of which may be partitioned according to a quadtree partition. That is, the nodes at the first level are either leaf nodes (without child nodes) or have four child nodes. The example QTBT structure 130 represents such a node including a parent node and a child node with solid lines for branching. If the nodes at the first level are not larger than the maximum allowed binary tree root node size (MaxBTSize), the nodes may be further partitioned by their respective binary trees. The binary tree splitting of one node may be repeated until the nodes resulting from the split reach the minimum allowed binary tree leaf node size (MinBTSize) or the maximum allowed binary tree depth (MaxBTDepth). The example QTBT structure 130 represents such a node with dashed lines for branching. The binary tree leaf nodes are called coding units (CUs) used for prediction (e.g., intra-picture or inter-picture prediction) and transformation. As explained above, a CU may also be referred to as a "video block" or a "block."

[0062]

[0074] In one example of a QTBT partitioning structure, the CTU size is set as 128x128 (luma samples and two corresponding 64x64 chroma samples), MinQTSize is set as 16x16, MaxBTSize is set as 64x64, MinBTSize is set as 4 (for both width and height), and MaxBTDepth is set as 4. Quadtree partitioning is first applied to the CTU to generate quadtree leaf nodes. The quadtree leaf nodes may have sizes from 16x16 (i.e., MinQTSize) to 128x128 (i.e., CTU size). If the quadtree leaf node is 128x128, it is not further split by the bi-tree because the size exceeds MaxBTSize (i.e., 64x64 in this example). Otherwise, the quadtree leaf node is further partitioned by the bi-tree. Therefore, the quadtree leaf node is also the root node for the binary tree and has the binary tree depth as 0. When the binary tree depth reaches MaxBTDepth (4 in this example), no further splitting is allowed. When a binary tree node has a width equal to MinBTSize (4 in this example), it implies that no further vertical splitting is allowed. Similarly, a binary tree node with a height equal to MinBTSize implies that no further horizontal splitting is allowed for that binary tree node. As mentioned above, the leaf nodes of the binary tree are called CUs and are further processed according to the prediction and transformation without further partitioning.

[0063]

[0075] As described throughout this disclosure, aspects of this disclosure describe determining quantization offset vectors for the transform coefficients based at least in part on the scaled transform coefficients and side information for the current block, and quantizing the transform coefficients to generate quantized transform coefficients for the current block based at least in part on the quantization offset vectors. One or more of these example techniques, as well as example techniques described below, may be implemented by the video encoder 200. In particular, the techniques described below may be implemented by, for example, the transform processing unit 206 and the quantization unit 208 of the video encoder 200 as shown in FIG. 8. In some examples, these techniques may be implemented by the video decoder 300.

[0064]

[0076] To minimize decoder cost, video compression standards such as H.264 / AVC, H.265 / HEVC, VP9, ​​and AV1 use a quantization style that is defined as having the simplest and cheapest implementation, where inverse quantization (inverse quantization unit 306 of video decoder 300 as shown in FIG. 9) is the norm alone. This allows encoder designers to choose among different quantization methods for video encoder 200, considering a compromise between optimizing compression performance and minimizing encoder cost. Details of some exemplary quantization schemes can be found in I.E.Richardson, "The H.264 Advanced Video Compression Standard," 2nd ed., John Wiley and Sons Ltd., 2010; M. Wien, "High Efficienty Video Coding: Coding Tools and Specification," Springer-Verlag, Berlin, 2015; D. Mukherjee, J. Bankoski, R.S.Bultje, A. Grange, J. Han, J. Koleszar, P. Wikins, and Y. Xu, "The latest open-source video codec VP9 - an overview and preliminary results," in Proceedings of the 30th Picture Coding Symp., San Jose, CA, USA. Jose, C.A., December 2013; and Y. Chen, D. Murherjee, J. Han, A. Grange, Y. Xu, Z. Liu, S. Parker, C. Chen, H. Su, U. Joshi, C. Chiang, Y. Wang, P. Wilkins, J. Bankoski, L. Trudeau, N. Egge, J. Valin, T. Davies, S. Midtskogen, A. Norking, P. de. Rivaz, "An Overview of core coding tools in the AV1 video codec" in Proceedings of the 33rd Picture Codin Symp., San Francisco, CA, June 2018.

[0065]

[0077] There are known techniques for improving quantization that can result in significant coding gains compared to simple scalar quantization. However, one or more of these known techniques for improving quantization may have high computational complexity. While such high computational complexity may be acceptable for a video-on-demand encoder, it may be computationally too expensive for a real-time video encoder. Furthermore, these techniques may be based on strictly sequential computations, which may strongly limit the throughput of the video encoder.

[0066]

[0078] This disclosure presents example techniques for potentially solving one or more of these problems, however, the example techniques described in this disclosure should not be considered as solving problems or limited to solving only problems of high computational complexity and strictly sequential computation.

[0067]

[0079] Aspects of this disclosure describe heuristic quantization that can provide significant coding gains with the low computational complexity required for real-time encoders. These techniques are based on a machine learning framework that "shortcuts" optimization decisions. Rather than basing quantization decisions on per-coefficient bitrate values ​​that are nearly exact but can be computationally very expensive to compute, the example techniques described in this disclosure can use training data to learn estimates of bitrates shared together by several coefficients, which can efficiently translate into improved quantization choices.

[0068]

[0080] Experimental simulations using a version of the fully implemented technique show results comparable to the technique used in the reference HEVC Test Model (HM) codec, but with lower computational complexity and easier and more efficient parallel implementation. Using the VVC common test conditions, the new technique described in this disclosure provides gains of 4.88%, 3.24%, 4.50%, and 4.44% in all intra, random access, low-latency B, and low-latency P, which is better than the rate-distorted optimized quantization (RDOQ) of the HEVC Test Model (HM) except for random access.

[0069]

[0081] Quantization is a component of all lossy video compression techniques, since it is the stage of the video coding process that defines the relationship between bit usage and acceptable "loss" of the signal being compressed. That is, quantization defines what parts of the original media signal should be discarded to achieve the desired compressed data size. Details of some exemplary quantization schemes can be found in I.E.Richardson, "The H.264 Advanced Video Compression Standard", 2nd ed., John Wiley and Sons Ltd., 2010; M. Wien, "High Efficienty Video Coding: Coding Tools and Specification", Springer-Verlag, Berlin, 2015; D. Mukherjee, J. Bankoski, R.S.Bultje, A. Grange, J. Han, J. Koleszar, P. Wikins, and Y. Xu, "The latest open-source video codec VP9 - an overview and preliminary results", in Proceedings of the 30th Picture Coding Symp., San Jose, CA, USA. Jose, CA, December 2013; and Y. Chen, D. Murherjee, J. Han, A. Grange, Y. Xu, Z. Liu, S. Parker, C. Chen, H. Su, U. Joshi, C. Chiang, Y. Wang, P. Wilkins, J. Bankoski, L. Trudeau, N. Egge, J. Valin, T. Davies, S. Midtskogen, A. Norking, P. de Rivaz, "An Overview of core coding tools in the AV1 video codec," in Proceedings of the 33rd Picture Coding Symp., San Francisco, CA, June 2018; J. E. Woods, "Multidimensional Signal, Image, and Video Processing and Coding," 2nd ed., Academic Press, 2011; W. A. ​​Pearlman and A.Said, "Digital Signal Compression: Principles and Pracice", Cambridge University Press, 2011; and D. Taubman, W. Marcellin, "JPEG2000: Image Compression Fundamentals, Standards and Practice", Springer, 2002.

[0070]

[0082] When used with an orthogonal transform, the video encoder 200 and the video decoder 300 may compute nearly-uniform-reconstruction quantization (NURQ) by simply using some form of rounding in the video encoder 200 and potentially using multiplications (scalar scaling) and possibly one addition in the video decoder 300. Some aspects of NURQ are described in GJ Sullivan and S. Sun, "On dead-one plus uniform threshold scalar quantization," in Proceedings SPIE5960, Visual Communications and Image processing, Beijing, China, 2005.

[0071]

[0083] 3A illustrates a video coding system with adaptive and / or rate-distortion optimal quantization. Specifically, FIG. 3A illustrates the main components of a video encoder, such as video encoder 200, in a hybrid video coding system, such as video encoding and decoding system 100, and provides examples of some of the terms and notations used throughout this disclosure. Although the video encoding process uses pixels organized in two-dimensional blocks, this disclosure describes, in some examples, data used in the video encoding process as N-dimensional vectors to simplify notation of the data described herein.

[0072]

[0084] The video encoder 200 may include a buffer 134 that receives and stores blocks of pixels 150, such as pixels of video data, for encoding. A prediction unit 136 may generate a predictive block b from the blocks of pixels 150. A residual generation unit 137 may generate residual data r based on the predictive block b and the blocks of pixels 150. A transform unit 138 may generate transform coefficients c based on the residual data. A scaling unit 140 may scale the transform coefficients c based on a quantizer step size s to determine scaled transform coefficients x. A quantization unit 142 may quantize the scaled transform coefficients x based on a bit cost estimate and side information from an entropy coding engine 144 to generate quantized transform coefficients q. The entropy coding engine 144 may generate an encoded video bitstream including compressed data bits 146 based on the quantized transform coefficients q.

[0073]

[0085] Using the notation shown in FIG. 3A, assuming a quantizer with a scaling factor (quantization step size) s, the vector representing the scaled transform coefficients

[0074]

number

[0075] teeth,

[0076]

number

[0077] where c represents the transform coefficients prior to scaling and s is the quantization step size.

[0078]

[0086] NURQ quantization is defined by two parameters p and u, where u is the quantization offset,

[0079]

number

[0080] is the floor function (·),

[0081]

number

[0082] and the inverse of NURQ quantization is

[0083]

number

[0084] where:

[0085]

number

[0086]

[0087] Because this simple but effective technique has a very low implementation cost, it has been adopted by all first media compression standards and is still a common choice for real-time encoders. The H.264 / AVC and H.265 / HEVC standards have a normative rule that equation (3) with p=0 must be used by the video decoder 300. Unless otherwise stated, this disclosure assumes that the video decoder 300 may use the parameter p=0 to dequantize the quantization value.

[0087]

[0088] Since quantization defines an aspect of lossy compression that is a trade-off between the resulting degradation in playback quality and the number of bits used (distortion and bitrate), a truly optimal form of quantization can take into account all factors that affect rate and distortion (RD) simultaneously for all vectors, not just for each scalar being quantized. This makes the problem of optimal quantization potentially very complex, and it may be necessary to identify and address only the most important and manageable factors in order to simplify quantization for practical application in real-life video encoders.

[0088]

[0089] To improve quantization, the use of vector quantization as described in W. A. ​​Pearlman and A. Said, "Digital Signal Compression: Principles and Practice", Cambridge University Press, 2011 and D. Taubman, W. Marcellin, "JPEG2000: Image Compression Fundamentals, Standards and Practice", Springer, 2002; A. Said, OG Guleryuz, and S. Yea, "Improving hybrid coding via control of quantization errors in the spatial and frequency domains", Proceedings of IEEE Int. Conf. Image Process, Paris, France, September 2014; and OG Guleryuz, A. Said, and S. Yea, "Non-causal encoding of predictively coded There are several practical techniques that have been shown to result in significant coding gains (compression improvements), such as effectively modeling quantization interactions with predictive coding as described in "Multimedia Applications of 3D Coding with Multiplexed Coding Samples," Paris, France, September 2014, and determining optimal integer-valued coefficient units using a rate-distortion cost function.

[0089]

[0090] However, in practice, compared to the simplest forms of quantization, these approaches may result in a relatively large increase in computational complexity and implementation costs. Indeed, some mathematical quantization optimization problems may be computationally challenging. Thus, solutions that are not necessarily optimal but are good enough for compression may be easier to determine and more practical to implement by a hybrid video coding system, such as the video encoding and decoding system 100.

[0090]

[0091] Aspects of the present disclosure describe a technique, which is in some sense a form of vector quantization, that uses information from several coefficients for better performance and then applies coefficient-wise rules to enable parallel computation and achieve computational complexity that may be low enough for real-time video encoders.

[0091]

[0092] Adaptive quantization is described below. One feature of video coding is that data statistics can change quite significantly depending on the type of prediction residual being coded, the quality setting, the video features, etc. Therefore, an adaptive form of quantization can improve compression by using different parameters according to the expected statistics type.

[0092]

[0093] Encoder settings and states are readily available, so the easiest adaptation may be based solely on these. For example, reference implementations of H.264 / AVC and H.265 / HEVC use equation (2) with offset parameter u=1 / 3 for slices (I slices) in the case of intra-frame only prediction, and offset parameter u=1 / 6 in other cases.

[0093]

[0094] Adaptation can be extended using statistical analysis of the data being coded, as described in GJ Sullivan and S. Sun, "On dead-one plus uniform threshold scalar quantization," in Proceedings SPIE5960, Visual Communications and Image processing, Beijing, China, 2005. However, those empirical techniques can be very limiting because they do not consider how the quantization decisions affect the bitrate.

[0094]

[0095] Rate distortion based on quantization is discussed below. Techniques have been proposed to modify quantization to account for both distortion and bit count that are compatible with some very early compression standards, such as JPEG and MPEG-2. Some of these techniques are described in K. Ramchandran and M. Vetterli, "Rate-distortion optimal fast thresholding with complete JPEG / MPEG decoder compatibility," IEEE Trans.on Image Processing, Vol. 3, No. 5, September 1994, and K. Ramchandran, A. Ortega and M. Vetterli, "Bit allocation for dependent quantization with applications to multiresolution and MPEG video coders," IEEE Trans.on Image Processing, Vol. 3, No. 5, September 1994.

[0095]

[0096] The optimal quantization problem is defined by the minimization of a distortion function averaged over all video blocks, constrained by an upper bound on the bit rate. Since blocks are quantized and coded independently, this problem can be solved using Lagrangian multiplication λ, as described in K. Ramchandran and M. Vetterli, "Rate-distortion optimal fast thresholding with complete JPEG / MPEG decoder compatibility," IEEE Trans. on Image Processing, Vol. 3, No. 5, September 1994. To that end, this disclosure describes the problem of directly optimizing the quantization in video blocks in that form.

[0096]

[0097] The vector x from equation (1), and the vector with the quantized transform coefficient values

[0097]

number

[0098] , we have a function D that measures the distortion resulting from quantization and the number of bits required to entropy code a vector q, respectively. s (x,q) and B(q) can be defined. When the transformation is orthogonal and the distortion corresponds to the squared error, the quantization is optimal in the rate-distortion sense if it is solved by the following optimization problem:

[0099]

number

[0100]

[0098] Due to complexity constraints, in order to exactly solve the optimization problem expressed by equation (5), any general method for optimization problems with integer variables may not be available in a practical video encoder, and it may be necessary to adopt a heuristic method to solve the optimization problem.

[0101] One useful tool for dealing with such problems is to test how much the objective function changes when a single element of the solution vector q is changed from one integer value to another. To express this formally, we use the element-wise permutation operator

[0102]

number

[0103] teeth,

[0104]

number

[0105] is defined as the difference operator

[0106]

number

[0107] is defined as follows:

[0108]

number

[0109]

number

[0110] and

[0111]

number

[0112]

[0100] In general, it is as follows.

[0113]

number

[0114] With these definitions, the heuristic optimization method does not use the exact value of B(q), but instead uses an approximate value whenever a negative value is encountered.

[0115]

number

[0116] but,

[0117]

number

[0118] and used to modify q accordingly. An algorithm of this type, implemented in the reference software for the H.265 / HEVC standard, is rate-distortion optimal quantization (RDOQ), as described in M. Karczewicz, P. Chen, Y. Ye, and R. Joshi, "RD based quantization in H.264," in Proceedings SPIE7443, Applications of Digital Image Processing XXXII, September 2009. RDOQ uses approximations and non-exact (heuristic) optimization techniques, so despite its name, the quantization cannot be truly optimized.

[0119]

[0102] Several properties can be used to reduce the number of times equation (10) is calculated. For example, since the statistical distribution of the transform coefficients decreases monotonically with magnitude, the expectation is

[0120]

number

[0121] where the notation E is used to indicate that this is expected for all blocks. a.e. {} is used, but there can be some rare exceptions.

[0122]

[0103] HEVC sign bit concealment is described below. The H.265 / HEVC video coding standard includes a technique called sign bit concealment as described in M. Wien, "High Efficiency Video Coding: Coding Tools and Specification", §8.2.4, Springer-Verlag, Berlin, 2015. Specifically, the H.265 / HEVC video coding standard specifies that for a quantized transform coefficient group that meets the condition for a minimum number of non-zero elements, the parity of the quantized magnitude sum must be equal to the sign bit of the first non-zero coefficient. In this way, the sign bit does not have to be coded, thereby reducing the total number of coding bits.

[0123]

[0104] In some instances, when a quantized coefficient does not satisfy the parity condition (on average, 50% of the cases), it may be possible to find a coefficient whose quantized value could have been incremented or decremented without a significant change in distortion. Thus, the scalar quantization for the HEVC HM software implements this technique by searching for the allowed change that corresponds to the smallest increase in squared error distortion.

[0124] When RDOQ is enabled, the search is based on a full RD cost estimate, and the index k of the coefficient whose quantized value must be modified is

[0125]

number

[0126] where A is the set of indices for which the quantization value may be changed. There may be additional implementation steps.

[0127]

[0106] Several techniques for potentially solving some of the problems mentioned above are now described: As explained above, quantization can be made very effectively adaptive and partially optimized if the quantizer uses some method for accurately measuring how the number of coded bits differs with changes in the quantizer decisions.

[0128]

[0107] One potential problem may be that modern encoders use very complex entropy coding, with many arithmetic coding contexts and complex context selection rules. Furthermore, transform coefficients are coded in more than one pass. For example, entropy coding in H.265 / HEVC may be done in up to five passes. This makes it potentially complex and computationally expensive to calculate and use their bit costs for each decision.

[0129]

[0108] Aspects of this disclosure describe techniques that potentially solve these problems by eliminating the need to derive, for each non-zero transform coefficient, the index of the arithmetic coding context used to entropy code the transform coefficient, and the need to access the state of those arithmetic coding contexts. Aspects of this disclosure describe techniques in which estimation rules are computed once and used in the quantization of several transform coefficients, thereby reducing the average per-pixel complexity and enabling parallel computation of the quantized transform coefficients.

[0130] In some aspects, video encoder 200 may determine a set of quantization offset parameters for quantizing a group of transform coefficients. The quantization offsets in the set of quantization offset parameters may be variable rather than fixed, may depend on a quantization interval, may be a function of values ​​of transform coefficients within the group of transform coefficients (e.g., magnitudes of the transform coefficients), side information associated with a block of transform coefficients that includes the group of transform coefficients, values ​​from other blocks of transform coefficients, etc.

[0131]

[0110] Additional examples of rate-distortion analysis are described below. The rate-distortion formulas analyzed for quantization optimization are generally of the form of equations (5) or (9), which are also used in the computational quantization calculations. One potential problem with an intuitive interpretation of those formulas is that they contain a Lagrangian multiplier coefficient λ, whose value can vary over a very wide range, depending on the playback quality choice.

[0132]

[0111] Although the value of λ is independent of other encoder decisions, it can be directly related to the selection of the quantizer step size s, as explained in T. Wiegand and B. Girod, "Lagrange multiplier selection in hybrid video coder control," in Proceedings of the IEEE Int. Conf. Image Processes, Thessaloniki, Greece, 2001, vol. 3, pp. 542-545. This is because, in order to have a quantization that matches λ, we have chosen

[0133]

number

[0134] where α varies over a relatively small range. For example, the HEVC HM software defines α as follows:

[0135]

number

[0136]

[0112] Replacing λ in equation (9) with λ defined in equation (12) changes the normalized form of the RD cost:

[0137]

number

[0138]

[0113] Note that equation (14) is derived directly from the objective function in equation (5), i.e., there are no approximations or special assumptions, only the normalization and explicit use of the squared error distortion, to allow for a more intuitive interpretation.

[0139]

[0114] Special case m=n+1

[0140]

number

[0141] shows in a straightforward and more intuitive way that the choice of quantizer value n+1 can be better than n if and only if equation (15) is non-negative, which is equal to:

[0142]

number

[0143]

[0115] From equation (16),

[0144]

number

[0145] If , the optimal quantization generally corresponds to the rounding operation (equation (2) with parameters p=0, u=½). In the following, this disclosure describes how equation (16) can be used to obtain a more general form that is also similar to the quantization defined by equation (2).

[0146]

[0116] Improved quantization using a trained quantization offset vector (QOV) is described below. One potential difficulty in using equation (14) in its exact form is the bit count loss from changing a quantized value from n to m.

[0147]

number

[0148] In the calculation of the change in bit count, this potential problem can potentially be mitigated if an approximation is used instead of calculating the change in bit count, such as shown in equation (10) and the RDOQ technique described in M. Karczewicz, P. Chen, Y. Ye, and R. Joshi, "RD based quantization in H.264," Applications of Digital Image Processing XXXII, September 2009, Proceedings of SPIE7443. However, another way to solve this problem is that for each transform coefficient being quantized, no information can be directly obtained from the arithmetic coding context (e.g., a probability estimation element).

[0149] According to aspects of this disclosure, instead of determining the difference in bit count as part of quantizing the transform coefficients, a video encoder such as video encoder 200 may quantize the transform coefficients based on quantization offsets that represent an estimate of the difference in bit count. The video encoder may determine a set of quantization offset parameters for a group of transform coefficients in the block of video data based on side information associated with the block of video data, and may quantize the group of transform coefficients based on the set of quantization offset parameters to generate quantized transform coefficients.

[0150]

[0118] FIG. 3B illustrates a technique for low-complexity adaptive quantization based on block classification and one or more sets of quantization offset parameters according to an aspect of the present disclosure. As shown in FIG. 3B, the number of bits to encode the quantization value q is

[0151]

number

[0152] Instead of using information from the entropy coding engine 144, such as an index of the arithmetic coding context used to entropy code each non-zero transform coefficient, to estimate a change in , and having information indicative of a bit cost estimate for quantizing the transform coefficients, or access to the state of those contexts, scaling and classification unit 152 of video encoder 200 may use the actual distribution of the scaled transform coefficients x and side information Γ associated with the block of video data (e.g., block size, type of prediction, etc.) to determine a quantization offset parameter v that quantization unit 142 of video encoder 200 may use to generate quantized transform coefficients.

[0153]

[0119] Video encoder 200 may use a bitwise operator to measure the change in the number of bits when a single element of q changes from integer value n to integer value m.

[0154]

number

[0155] It may be possible to use a quantization offset parameter to generate the quantized transform coefficients rather than using a modification of

[0156]

number

[0157] is defined as in equation (6) and as in equations (14) to (16): The video encoder 200 uses a difference in the number of bits to entropy code the quantized transform coefficients q, rather than relying directly on the exact number of bits B(q).

[0158]

number

[0159] It may be possible to determine optimal quantization values ​​for the transform coefficients based on the quantization coefficients. Thus, the video encoder 200 may identify valid patterns for groups of transform coefficients that may be independent of the exact number of bits to entropy code the quantized transform coefficients. Changing the bit depth

[0160]

number

[0161] is usually small. In fact, when m or n is equal to zero or near-zero, the maximum magnitude may be only a few bits, and the magnitude of those changes decreases very rapidly with larger magnitudes of m and n. For example,

[0162]

number

[0163] That is, The constant α may be relatively large, so the ratio

[0164]

number

[0165] may be relatively small. This may mean that the optimization decision in equation (16) (when parameter p=0) corresponds to an offset parameter u in equation (2) that is close to ½. For example, in the example in Table I below, offset parameter u=0.1 may be optimal only if there is an increase of approximately 9 bits when the magnitude of the quantization value of a single coefficient is incremented by 1, which is not what is expected.

[0166] Table I - Quantization offset u and increment in number of coded bits to optimize quantization using equation (2) (parameters p=0, non-negative values ​​of n, α=11.14)

[0167]

number

[0168] :

[0169] [Table 1]

[0170]

[0121] Example techniques described in this disclosure may take advantage of the properties described above and assume an L-dimensional subgroup of elements of vector x (e.g., a 4-dimensional subgroup or a 16-dimensional subgroup of a 4x4 pixel subgroup of a picture), and may use the approximation function

[0171]

number

[0172] teeth,

[0173]

number

[0174] where x represents a subgroup of scaled transform coefficients, n is a proposed quantization value for a scaled transform coefficient in the group of scaled transform coefficients x, g is an index of the group of scaled transform coefficients to which the scaled transform coefficient x belongs, and Γ represents a data structure having side information for a block of video data. In some examples, the side information Γ may include one or more of a slice type (I, P, or B) of the block of video data, residual data from intraframe or interframe prediction of the block of video data, a block size (e.g., 4×4, 8×8, 16×16, or 32×32) of the block of video data, and a luminance or chrominance component.

[0175]

[0122] To simplify notation, e.g., to make memory access more efficient or to make the approximation more accurate, the elements of the vector x representing the scaled transform coefficients may be previously rearranged so that subgroup elements in the vector x have consecutive indices.

[0176] In this notation, assume that B(1)=B(-q),

[0177]

number

[0178] If equality exists in equation (17), then the quantized value is

[0179]

number

[0180] which corresponds to equation (16), where

[0181]

number

[0182] is the conversion factor x i where u(n,g,x,Γ) represents a set of quantization offset parameters for a group g of transform coefficients of a scaled transform coefficient vector x of a block of video data, such as a transform block.

[0183]

[0124] In a practical application example,

[0184]

number

[0185] Therefore, equation (18) can often be used for mathematical consistency. If this condition is not met (possibly when n=0), a slightly more complex quantization rule can be used based on equation (14) instead of equation (15). Equation (19) can be the same type of low computational complexity quantization as equation (2), with p=0 and a quantization offset u that is modified according to the magnitude of the coefficient being quantized.

[0186] Another simplification for practical implementation is for larger |n|

[0187]

number

[0188] Therefore, a P-dimensional vector function v(g, x, Γ), called the quantization offset vector (QOV), can be expressed as follows:

[0189]

number

[0190] If it can be defined as

[0191]

number

[0192] The quantization offset vector thus represents the set of P quantization offset parameters that result from the application of a cutoff at the larger quantization value |n|, with the above assumptions regarding the change in the number of bits close to zero.

[0193]

number

[0194] If , then this cutoff is reflected in the application of a min-function to the QOV index in equation (21).

[0195] Based on these definitions, the following is an example of a technique for adaptive quantization, where for each transform coefficient vector: 1. Determine the side information Γ of the block of video data to which the transform coefficient vector x belongs, such as a transform block; 2. Split the transform coefficient vector x into K=N / L subgroups, where N is the number of pixels in the block and L is the number of pixels in each subgroup (e.g., L=16 for 4×4 subblocks); For each subgroup g, g=0,1,(...,N) / L-1, a. Determine a quantization offset vector v(g, x, Γ); b. Each transform coefficient x using index i=gL, gL+1,...,(g+1)L-1 i , i.e., for each transform coefficient in subgroup g, For example, according to equation (21), x i quantization offset vector v(g,x,Γ) based at least in part on the magnitude of i Calculate the quantized value of

[0196] According to aspects of this disclosure, the video encoder 200 may calculate a number of bits for a block of video data, such as a transform coefficient block or a transform block of video data.

[0197]

number

[0198] For a block of video data, video encoder 200 may generate a set of quantization offset parameters for quantizing scaled transform coefficients for the block of video data that correspond to the change in quantization offset parameters. For a block of video data, video encoder 200 may determine side information associated with the block of video data. The side information may include any combination of one or more of a slice type (e.g., I, P, B) of the block of video data, residual data resulting from intra or inter prediction of the block of video data, a block size (e.g., 4×4, 8×8, 16×16, or 32×32) of the block of video data, and / or luminance or chrominance components of the block of video data.

[0199]

[0128] The video encoder 200 may divide the scaled transform coefficients of a block of video data into multiple groups of scaled transform coefficients, particularly associated with sub-blocks of the block of video data. If the block of video data includes N pixels, the video encoder 200 may divide the block of video data into sub-blocks of L groups of pixels to result in N / L groups of scaled transform coefficients. For example, the video encoder 200 may divide the block of video data into 4x4 sub-blocks (groups of 16 pixels), 8x8 sub-blocks (groups of 64 pixels), 16x16 sub-blocks (groups of 256 pixels), etc. Each such sub-block of the block of video data may be indicated by an index g from 0 to N / L-1, where each sub-block includes a group of transform coefficients for the sub-block.

[0200]

[0129] For each subblock, video encoder 200 may determine a set of quantization offset parameters for the subblock such that video encoder 200 may quantize a group of scaled transform coefficients in the same subblock using the same set of quantization offset parameters determined for that subblock. As described above, the quantization offset parameters may be associated with a change in bit count based on a change in a single element of the quantized transform coefficients for a block of video data. In some examples, each quantization offset parameter in the set of quantization offset parameters may range from 0 to 0.5.

[0201]

[0130] The set of quantization offset parameters for the group of scaled transform coefficients in a subblock may be adaptive rather than fixed, such that the video encoder 200 may adaptively select different quantization offset parameters for quantizing different scaled transform coefficients in the group of scaled transform coefficients, rather than using the same fixed quantization offset parameter for quantizing the scaled transform coefficients in the group of scaled transform coefficients. One exemplary set of adaptive quantization offset parameters for the group of scaled transform coefficients may be [0.2, 0.3, 0.35, 0.4, 0.45, 0.5]. As can be seen, the quantization offsets in the set of adaptive quantization offset parameters are not fixed to a single value, and each element of the set of quantization offset parameters is not necessarily incremented by the same value. For example, a first element having a value of 0.2 is incremented by 0.1 to result in a second element having a value of 0.3, and the second element is incremented by 0.05 to result in a third element having a value of 0.35.

[0202]

[0131] Video encoder 200 may quantize the scaled transform coefficient, for example, by adding the value of the scaled transform coefficient to a quantization offset and truncating or flooring the resulting sum to an integer value. For example, given a scaled transform coefficient x and a quantization offset u, video encoder 200 may determine a quantization value q of the scaled transform coefficient as follows:

[0203]

number

[0204]

[0132] Video encoder 200 may determine a set of quantization offset parameters for a sub-block based on side information of the block of video data. In some examples, video encoder 200 may also determine a set of quantization offset parameters for a sub-block based on side information associated with the sub-block, such as a maximum magnitude range of the transform coefficients of the sub-block, a block size of the sub-block, a relative position of the sub-block within the block, etc.

[0205]

[0133] One example of a set of quantization offset parameters is a quantization offset vector (QOV). A QOV may include a list of quantization offsets, where an M-dimensional QOV may include M quantization offsets. Although aspects of the present disclosure are described with respect to QOVs, the techniques described herein are generally applicable to any form of set of quantization offset parameters, such as an array, list, stack, queue, table, graph, etc.

[0206] In some examples, the quantization offsets in an M-dimensional QOV are indexed from 0 to M−1. Thus, the video encoder 200 may implement the formula v n A quantization offset can be selected from QOV using -=V[min(M-1,n)], where v n is the nth element of the quantization offset if n is less than M-1. Otherwise, v n is the M-1th element of QOV. Thus, given a scaled transform coefficient x, video encoder 200 may take the absolute value of transform coefficient x and round down to the nearest integer, so that the equation becomes

[0207]

number

[0208] In the formula,

[0209]

number

[0210] is a quantization offset parameter selected from the QOV for transform coefficient x. Given a transform coefficient x and a quantization offset u, the video encoder 200 may calculate the quantized transform coefficient q as

[0211]

number

[0212] Therefore, the quantization offset selected from QOV for the transform coefficient x may be determined as

[0213]

number

[0214] Then, the video encoder 200 may denote the quantized transform coefficient x by

[0215]

number

[0216] It can be determined that:

[0217]

[0135] Figure 4 illustrates a parallel implementation of adaptive quantization using a set of quantization offset parameters according to the techniques of this disclosure. Specifically, Figure 4 illustrates an example implementation of the above technique for quantizing scaled variable coefficients, where quantization operations on the scaled variable coefficients of a block are performed in parallel upon determining a set of quantization offset parameters for the scaled variable coefficients of a block of video data. The components illustrated in Figure 4 may be, for example, part of the scaling and classification unit 152 and the quantization unit 142 of the video encoder 200 illustrated in Figure 3B.

[0218]

[0136] As shown in FIG. 4, the video encoder 200 may include a quantization offset parameter determination unit 154, a grouping unit 156, and offset quantization units 158A-158P. The grouping unit 156 may be a processing circuit configured to receive scaled transform coefficients x of a block (e.g., a transform block) of video data and divide the scaled transform coefficients into groups of scaled transform coefficients, such as by dividing the block of video data into sub-blocks. The grouping unit 156 may divide the N scaled transform coefficients into L groups of scaled transform coefficients to result in K=N / L groups of scaled transform coefficients indexed from 0 to N / L-1. For example, the grouping unit 156 may divide the scaled transform coefficients X 0 From X L -1 set, scaled conversion coefficients X L From X 2L-1 Such a set of scaled conversion coefficients X N-L From X N-1 The items can be grouped into sets of

[0219]

[0137] Quantization offset parameter determination unit 154 may be a processing circuit configured to receive scaled transform coefficients x of a block of video data and side information Γ for the block of video data, and to determine a set of quantization offset parameters for each of the groups of scaled transform coefficients determined by grouping unit 156. Quantization offset parameter determination unit 154 may determine a set of quantization offset parameters for each group of scaled transform coefficients based on the side information Γ for the block of video data and / or the values ​​of the scaled transform coefficients in the group of scaled transform coefficients.

[0220]

[0138] In the example of FIG. 4, the quantization offset parameter determination unit 154 may determine a QOV as a set of quantization offset parameters for a group of scaled transform coefficients. The QOV is presented in FIG. 4 in the form of v(g, x, Γ), where g is an index of the scaled transform coefficients 0 to N / L-1, x is a set of scaled transform coefficients, and Γ is side information for the block of video data. Thus, the quantization offset parameter determination unit 154 may determine a QOVv(0, x, Γ) for a group of scaled transform coefficients associated with an index g of 0, a QOVv(1, x, Γ) for a group of scaled transform coefficients associated with an index g of 1, up to a QOVv(N / L-1, x, Γ) for a group of scaled transform coefficients associated with an index g of N / L-1.

[0221]

[0139] The offset quantization units 158A-1-158A-M may be processing circuits configured to quantize the scaled transform coefficients based on a quantization offset parameter. For example, the offset quantization units 158A-1-158A-M may be configured to quantize the scaled transform coefficients based on a quantization offset parameter. 0 From q L-1 To generate the scaled transform coefficients x, use QOVv(0,x,Γ). 0 x L-1 The offset quantization units 158B-1 to 158B-M may quantize a group of quantization values ​​q L From q 2L-1 To generate the scaled transform coefficients x, use QOVv(1,x,Γ). L x 2L-1 The offset quantization units 158P-1 to 158P-M may quantize a group of quantization values ​​q N-L From q N-1 To generate the scaled conversion coefficients x, use QOVv(N / L,x,Γ). N-L x N-1 The group of may be quantized.

[0222]

[0140] Some practical techniques that may be used to determine a set of quantization offset parameters, such as QOV, are further described below. As can be observed from the above derivation, determining the optimal QOV involves:

[0223]

number

[0224] This may be mathematically equivalent to estimating

[0225] In some examples, aspects of the techniques described herein may be applicable to HEVC sign bit concealment. As presented in the description of HEVC sign bit concealment above, optimized sign bit concealment may be based on rate-distortion cost. As with quantization, a quantization offset parameter such as QOV may be used to improve the performance of the sign bit concealment.

[0226]

[0142] Using the notation of the alternative rate-distortion analysis description above,

[0227]

number

[0228] Then, the cost of changing the value of that quantized transform coefficient to satisfy the sign bit concealment parity constraint is

[0229]

number

[0230] which can be shown to be equal to:

[0231]

number

[0232]

[0143] The rule for identifying the index of the coefficient to be modified, which is equivalent to equation (11), is as follows:

[0233]

number

[0234] Using equations (17), (18), and (20), the change in RD cost can be approximated when the quantization values ​​are changed for sign bit concealment using the following function:

[0235]

number

[0236]

[0145] Figure 5 shows the coefficients in this formula for sign bit concealment to approximate the rate-distortion cost. Specifically, Figure 5 shows the functions used in the calculation of the sign bit concealment RD cost (assuming positive x). The RD cost is zero at the value used as the threshold in the quantization formula (21), which is consistent with the example where the exact coefficient values ​​of the two quantization levels have the same RD cost.

[0237] For example, graph 162A shows that the sign of the difference between a scaled transform coefficient x and its quantization value q, sign(xq), varies with quantization offset v as follows:

[0238]

number

[0239] from

[0240]

number

[0241] is positive,

[0242]

number

[0243] from

[0244]

number

[0245] Graph 162B shows that

[0246]

number

[0247] but,

[0248]

number

[0249] From 1-v to

[0250]

number

[0251] Graph 162C shows that

[0252]

number

[0253] but

[0254]

number

[0255] 1-v in

[0256]

number

[0257] is 0 at

[0258]

number

[0259] Show that v in .

[0260] Using these approximations, equation (24) becomes:

[0261]

number

[0262] such that an index k of a transform coefficient whose quantization value should be modified for sign bit concealment may be determined based on the side information Γ of the block of video data that includes the transform coefficient. Thus, video encoder 200 may determine quantization values ​​for transform coefficients in a block of video data that video encoder 200 may modify for sign bit concealment based on the side information Γ of the block of video data that includes the transform coefficient associated with the quantization value.

[0263]

[0148] Figure 6 illustrates a technique for determining quantization offset parameters according to the techniques of this disclosure. Given a criterion for mapping a triplet (g, x, Γ), where x is a vector of scaled transform coefficients, g is a group index of the scaled transform coefficients, and Γ is side information about a block of video data containing the group of scaled transform coefficients, to a set of quantization offset parameters such as QOV, the values ​​of the elements (i.e., quantization offsets) of the set of quantization offset parameters may be optimized using statistical or machine learning techniques such that the quantization offset parameters correspond to quantization offset values ​​that maximize the ratio of the average RD cost function obtained with those quantization offset values ​​to the actual (exact) cost function.

[0264]

[0149] There may be a variety of different techniques for determining a set of quantization offset parameters for a group of scaled transform coefficients (e.g., scaled variable coefficients in a sub-block) by mapping the parameters g, x, and Γ associated with the group of scaled transform coefficients to a QOV, and the techniques of this disclosure may encompass any suitable technique for mapping the parameters g, x, and Γ associated with the group of scaled transform coefficients to a QOV.

[0265]

[0150] Figure 6 illustrates some example techniques for determining a set of quantization offset parameters for a group of transform coefficients in accordance with the techniques of this disclosure. Although Figure 6 illustrates the set of quantization offset parameters as QOV, the techniques illustrated herein are applicable to any other suitable form of the set of quantization offset parameters.

[0266]

[0151] As shown in FIG. 6, in example 172A, the video encoder 200 may implement a QOV calculation unit 174 that may directly calculate a QOV for a set of scaled transform coefficients from a parameter group index g, a scaled transform coefficient x, and side information Γ associated with a block of video data containing the group of scaled transform coefficients.

[0267]

[0152] In example 172B, video encoder 200 may use a parameter approach to determine a QOV, where an element of the QOV is defined based on a vector p having a smaller magnitude. Video encoder 200 may implement a QOV parameter calculation unit 176 that may determine a parameter vector p based at least in part on parameters g, x, and Γ associated with a group of transform coefficients, where parameter vector p may be a vector having a smaller magnitude (i.e., fewer elements) than the QOV to be determined. Video encoder 200 may implement a QOV calculation unit 178 that may determine a QOV for a group of transform coefficients having associated parameters g, x, and Γ based on the parameter vector p.

[0268] For example, the parameter vector p may be a two-dimensional parameter vector having a smaller magnitude than the P-dimensional QOV, and the video encoder 200 may determine a P-dimensional QOV for parameters g, x, and Γ based on the parameter vector p using the following equation, where v n is the nth element of the QOV:

[0269]

number

[0270]

[0154] In some examples, the video encoder 200 may utilize a set of pre-computed QOVs to determine a QOV for quantizing a transform coefficient. The set of pre-computed QOVs may be in the form of an array of QOVs, and the video encoder 200 may index into the array of QOVs to select a QOV for quantizing a group of transform coefficients. In example 172C, the video encoder 200 may implement a QOV index calculation unit 180 that may map a group of transform coefficients with associated parameters g, x, and Γ to an index n, which may be an integer. The video encoder 200 may implement a QOV retrieval unit 182 that may index into the pre-computed array of QOVs using the index n to determine one QOV from the array of QOVs to use for quantizing x.

[0271]

[0155] In some examples, video encoder 200 may use a general method that can be used for both classification and regression, such as a neural network, to determine the quantization offset parameters in examples 172A-172C. For example, such a neural network may be trained with training data that includes a set of parameters g, x, and Γ, and an optimal quantization offset parameter value, such as QOV, along with an encoding performance objective function, such as a rate-distortion value that results in a quantization value for the performance of RDOQ in HEVC HM. In this manner, the neural network is trained to associate side information Γ and values ​​of scaled transform coefficients of a block of video data to a set of quantization offset parameter values ​​that optimize the rate-distortion cost of quantizing the scaled transform coefficients with the associated set of quantization offset parameters.

[0272]

[0156] Similarly, in some examples, the video encoder 200 may use general regression methods such as linear regression, logistic regression, Poisson regression, etc. to determine the relationship between the set of parameters g, x, and Γ and the optimal quantization offset parameter value using the neural network described above, as in examples 172A and 172B.

[0273]

[0157] In some examples, the video encoder 200 may use a classification method, such as a classification tree, to classify the set of parameters g, x, and Γ to select a quantization offset parameter for them from a set of pre-computed quantization offset parameters, as in example 172C. For example, the neural network described above may act as a classifier trained to classify a group of scaled transform coefficients based on side information of a block of video data that includes the group of scaled transform coefficients. By classifying the group of scaled transform coefficients, the video encoder 200 may select a set of quantization offset parameters for quantizing the group of scaled transform coefficients from a plurality of sets of quantization offset parameters. An example of such a classification method is described below.

[0274]

[0158] As described in example 172C, video encoder 200 may select a set of quantization offset parameters (e.g., QOVs) from a set of pre-computed quantization offset parameters (e.g., pre-computed QOVs) for a group of scaled transform coefficients x in a sub-block having associated side information Γ. Video encoder 200 may select a QOV from the set of pre-computed QOVs based at least in part on the side information for the sub-block that includes the group of scaled transform coefficients as well as side information for the block of video data that includes the sub-block. For example, video encoder 200 may select a QOV from the set of pre-computed QOVs based at least in part on a location of the sub-block within the block of video data, a size of the block of video data, etc.

[0275]

[0159] Figure 7 shows examples of codes numbered 0 through 9 for identifying the location of a sub-block within a block of video data. Such a block of video data may be, for example, an HEVC or VVC transform coefficient block. As shown in Figure 7, a sub-block may be a 4x4 HEVC transform coefficient block, an 8x8 HEVC transform coefficient block, a 16x16 HEVC transform coefficient block, or a 32x32 HEVC transform coefficient block.

[0276]

[0160] A subblock may have a code that is associated with the location of the subblock within the block, and also with the size of the block. When a 4x4 subblock is within a 4x4 block, the subblock may have a code of 0. When a 4x4 subblock is within an 8x8 block, the subblock may have a code of 1 when the subblock is the top left 4x4 subblock, and a code of 2 when the subblock is not the top left 4x4 subblock. When a 4x4 subblock is within a 16x16 block, the subblock may have a code of 3 when the subblock is the top left 4x4 subblock, a code of 4 when the subblock is the top left 8x8 of the block but not the top left 4x4 subblock, and a code of 5 when the subblock is not the top left 8x8 of the block. When a 4x4 sub-block is within a 32x32 block, the sub-block may have a code of 6 when the sub-block is the top left 4x4 sub-block, a code of 7 when the sub-block is the top left 8x8 of the block and not the top left 4x4 sub-block, a code of 8 when the sub-block is the top left 16x16 of the block and not the top left 8x8 of the block, and a code of 9 when the sub-block is not within the top left 16x16 of the block.

[0277]

[0161] Aspects of the present disclosure have been implemented and tested to create files that comply with the HEVC standard using a modified version of the HM reference software (eg, encoder modifications only). One exemplary implementation is based on example 172C of FIG. 6, where one QOV index is selected for each group of 4×4 transform coefficients in a block.

[0278]

[0162] In this implementation, video encoder 200 may determine a QOV for a group of scaled transform coefficients from a set of pre-computed QOVs using the following function:

[0279]

number

[0280]

[0163] Video encoder 200 may determine an index of a QOV for a 4x4 sub-block from the set of pre-computed QOVs based on the following parameters: For all conversion coefficient values ​​in the 4×4 conversion coefficient group, 0 =max(c(x i )∈{-1,0,1,2,3}, P 1 = min(2,k-1) ∈ {0,1,2}, where k is the number of times max(c(x i ))=P 0 That is, P 2 ∈{0,1,...,9} is the code used to indicate both the transform coefficient block size and the location of the 4×4 groups within the block, according to the scheme shown in FIG. P 3 ∈{0,3} is 0 if the block is part of an intra slice (according to the HEVC standard) and 1 otherwise.

[0281]

[0164] As can be seen, video encoder 200 may determine a QOV for a group of scaled transform coefficients in a sub-block of a block of video data based on side information associated with the sub-block. The side information associated with the sub-block used to determine the QOV may include the location of the sub-block within the block of video data, such as, for example, whether it is in a top-left sub-block in the block of video data.

[0282]

[0165] Case P 0 = -1 may correspond to a group with all coefficients quantized to zero. Thus, in this case, the video encoder 200 may set P 0Based on these definitions, video encoder 200 may calculate a set of 240 QOV indexes, for example, using the following equations:

[0283]

number

[0284] .

[0285]

[0166] In the above example, video encoder 200 may determine a QOV for a group of scaled transform coefficients in a sub-block of a block of video data based on whether the block of video data is part of an intra slice, the location of the sub-block within the block of video data, the size of the block of video data, and the maximum scaled transform coefficient value within the group of scaled transform coefficients and the number of times the maximum scaled transform coefficient value is within a 2x2 sub-block within the sub-block.

[0286] In some examples, the video encoder 200 may use the following equation: 3 +P 2 ∈{0,1,2,...,19}, video encoder 200 may calculate a set of 20 QOV indexes that are independent of the values ​​of the transform coefficients (e.g., vector x) in the sub-block. In this example, video encoder 200 may calculate the QOV index based on the location of the sub-block within the transform coefficient block, the size of the transform coefficient block, and whether the transform coefficient block is part of an intra slice. Determining a QOV index that is independent of the values ​​of the transform coefficients in the sub-block may allow video encoder 200 to perform quantization of the group of scaled transform coefficients in the sub-block in a single pass, thereby reducing the number of processing cycles used to quantize the group of scaled transform coefficients.

[0287]

[0168] As can be seen in the techniques described above, video encoder 200 may determine a set of quantization offset parameters for a group of scaled transform coefficients in a sub-block of a block of video data based at least in part on a location of the sub-block within the block of video data as well as a size of the block of video data. In some examples, video encoder 200 may also utilize values ​​of the scaled transform coefficients in the sub-block to determine a set of quantization offset parameters for the sub-block, and in other examples, video encoder 200 may be able to determine a set of quantization offset parameters for the sub-block without utilizing values ​​of the scaled transform coefficients in the sub-block.

[0288]

[0169] Figure 8 is a block diagram illustrating an example video encoder 200 that may implement the techniques of this disclosure. Figure 8 is provided for illustrative purposes and should not be considered as limiting the techniques as broadly illustrated and described in this disclosure. For purposes of explanation, this disclosure describes video encoder 200 in accordance with VVC (ITU-T H.266 under development) and HEVC (ITU-T H.265) techniques. However, the techniques of this disclosure may be implemented by video encoding devices configured for other video coding standards.

[0289] 8, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any or all of the video data memory 230, the mode selection unit 202, the residual generation unit 204, the transform processing unit 206, the quantization unit 208, the inverse quantization unit 210, the inverse transform processing unit 212, the reconstruction unit 214, the filter unit 216, the DPB 218, and the entropy coding unit 220 may be implemented in one or more processors or processing circuits. For example, the units of video encoder 200 may be implemented as one or more circuits or logic elements, as part of a hardware circuit, or as part of a processor in an FPGA, an ASIC, etc. Moreover, video encoder 200 may include additional or alternative processors or processing circuits for performing these and other functions.

[0290]

[0171] The video data memory 230 may store video data to be encoded by the components of the video encoder 200. The video encoder 200 may receive video data stored in the video data memory 230, for example, from the video source 104 (FIG. 1). The DPB 218 may act as a reference picture memory that stores reference video data for use in predicting subsequent video data by the video encoder 200. The video data memory 230 and the DPB 218 may be formed by any of a variety of memory devices, such as synchronous dynamic random access memory (DRAM), including DRAM (SDRAM), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The video data memory 230 and the DPB 218 may be provided by the same memory device or separate memory devices. In various examples, the video data memory 230 may be on-chip with the other components of the video encoder 200, as shown, or off-chip relative to those components.

[0291] In this disclosure, references to video data memory 230 should not be construed as limited to memory internal to video encoder 200, unless specifically so described, or to memory external to video encoder 200, unless specifically so described. Instead, references to video data memory 230 should be understood as a reference memory that stores video data that video encoder 200 receives for encoding (e.g., video data of a current block to be encoded). Memory 106 of FIG. 1 may also provide temporary storage of outputs from various units of video encoder 200.

[0292]

[0173] The various units in FIG. 8 are shown to aid in understanding the operations performed by the video encoder 200. The units may be implemented as fixed-function circuits, programmable circuits, or a combination thereof. A fixed-function circuit refers to a circuit that provides a specific function and is pre-configured with respect to the operations that may be performed. A programmable circuit refers to a circuit that may be programmed to perform various tasks and to provide flexible functionality in the operations that may be performed. For example, a programmable circuit may execute software or firmware that causes the programmable circuit to operate in a manner defined by the software or firmware instructions. A fixed-function circuit may execute software instructions (e.g., to receive parameters or output parameters), but the types of operations that the fixed-function circuit performs are generally unchanging. In some examples, one or more of the units may be separate circuit blocks (fixed function or programmable), and in some examples, one or more of the units may be integrated circuits.

[0293]

[0174] Video encoder 200 may include an arithmetic logic unit (ALU), a basic functional unit (EFU), a programmable core formed of digital circuits, analog circuits, and / or programmable circuits. In examples in which the operations of video encoder 200 are implemented using software executed by programmable circuits, memory 106 (FIG. 1) may store instructions (e.g., object code) of the software that video encoder 200 receives and executes, or another memory (not shown) within video encoder 200 may store such instructions.

[0294]

[0175] The video data memory 230 is configured to store the received video data. The video encoder 200 may retrieve pictures of the video data from the video data memory 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data memory 230 may be raw video data to be encoded.

[0295]

[0176] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra prediction unit 226. The mode selection unit 202 may include additional functional units for performing video prediction according to other prediction modes. As examples, the mode selection unit 202 may include a palette unit, an intra block copy unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, a combined inter-intra prediction (CIIP) unit, etc.

[0296]

[0177] The mode selection unit 202 generally coordinates multiple encoding passes to test combinations of encoding parameters and resulting rate-distortion values ​​for such combinations. The encoding parameters may include partitioning of the CTU into CUs, prediction modes for the CUs, transform types for residual data of the CUs, quantization parameters for residual data of the CUs, etc. The mode selection unit 202 may finally select a combination of encoding parameters that has a rate-distortion value that is better than other tested combinations.

[0297]

[0178] The video encoder 200 may partition a picture retrieved from the video data memory 230 into a series of CTUs, encapsulating one or more CTUs in a slice. The mode selection unit 202 may partition the CTUs of the picture according to a tree structure, such as the QTBT structure or quadtree structure of HEVC described above. As described above, the video encoder 200 may form one or more CUs from partitioning the CTUs according to the tree structure. Such a CU may also be generally referred to as a "video block" or "block."

[0298]

[0179] In general, the mode selection unit 202 also controls its components (e.g., the motion estimation unit 222, the motion compensation unit 224, and the intra prediction unit 226) to generate a prediction block for a current block (e.g., the current CU, or in HEVC, the overlapping portion of the PU and TU). For inter prediction of the current block, the motion estimation unit 222 may perform motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously coded pictures stored in the DPB 218). In particular, the motion estimation unit 222 may calculate a value representing how similar a potential reference block is to the current block according to, for example, a sum of absolute differences (SAD), a sum of squared differences (SSD), a mean absolute difference (MAD), a mean squared difference (MSD), etc. The motion estimation unit 222 may generally perform these calculations using a sample-by-sample difference between the current block and the reference block under consideration. Motion estimation unit 222 may identify the reference block having the lowest value resulting from these calculations, which indicates the reference block that most closely matches the current block.

[0299]

[0180] The motion estimation unit 222 may form one or more motion vectors (MVs) that define the position of a reference block in a reference picture relative to the position of a current block in a current picture. The motion estimation unit 222 may then provide the motion vectors to the motion compensation unit 224. For example, in unidirectional inter prediction, the motion estimation unit 222 may provide a single motion vector, while in bidirectional inter prediction, the motion estimation unit 222 may provide two motion vectors. The motion compensation unit 224 may then generate a predictive block using the motion vectors. For example, the motion compensation unit 224 may use the motion vectors to retrieve data of the reference block. As another example, if the motion vectors have partial sample precision, the motion compensation unit 224 may interpolate values ​​for the predictive block according to one or more interpolation filters. Furthermore, in bidirectional inter prediction, the motion compensation unit 224 may retrieve data for the two reference blocks identified by the respective motion vectors and combine the retrieved data, for example, through sample-wise averaging or weighted averaging.

[0300]

[0181] As another example, for intra prediction, or intra predictive coding, the intra prediction unit 226 may generate a predictive block from samples neighboring a current block. For example, in a directional mode, the intra prediction unit 226 may generally mathematically combine values ​​of neighboring samples and populate these calculated values ​​in a prescribed direction across the current block to produce a predictive block. As another example, in a DC mode, the intra prediction unit 226 may calculate an average of the neighboring samples for the current block and generate a predictive block to include this resulting average for each sample of the predictive block.

[0301]

[0182] The mode selection unit 202 provides the prediction block to the residual generation unit 204. The residual generation unit 204 receives a raw, uncoded version of the current block from the video data memory 230 and receives the prediction block from the mode selection unit 202. The residual generation unit 204 calculates sample-by-sample differences between the current block and the prediction block. The resulting sample-by-sample differences define a residual block for the current block. In some examples, the residual generation unit 204 may also determine differences between sample values ​​in the residual block to generate the residual block using residual differential pulse code modulation (RDPCM). In some examples, the residual generation unit 204 may be formed using one or more subtractor circuits that perform binary subtraction.

[0302]

[0183] In an example where the mode selection unit 202 partitions a CU into PUs, each PU may be associated with a luma prediction unit and a corresponding chroma prediction unit. The video encoder 200 and the video decoder 300 may support PUs having various sizes. As indicated above, the size of a CU may refer to the size of the luma coding block of the CU, and the size of a PU may refer to the size of the luma prediction unit of the PU. Assuming that the size of a particular CU is 2Nx2N, the video encoder 200 may support a PU size of 2Nx2N or NxN for intra prediction, and a symmetric PU size of 2Nx2N, 2NxN, Nx2N, NxN, or the like for inter prediction. The video encoder 200 and the video decoder 300 may also support asymmetric partitioning for PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter prediction.

[0303]

[0184] In examples where the mode selection unit 202 does not further partition the CUs into PUs, each CU may be associated with a luma coding block and a corresponding chroma coding block. As described above, the size of a CU may refer to the size of the luma coding block of the CU. The video encoder 200 and the video decoder 300 may support CU sizes of 2Nx2N, 2NxN, or Nx2N.

[0304]

[0185] In other video coding techniques, such as intra block copy mode coding, affine mode coding, and linear model (LM) mode coding, as some examples, the mode selection unit 202 generates a predictive block for the current block being coded via a respective unit associated with the coding technique. In some examples, such as palette mode coding, the mode selection unit 202 may not generate a predictive block, but instead may generate syntax elements that indicate how the block should be reconstructed based on a selected palette. In such modes, the mode selection unit 202 may provide these syntax elements to be coded to the entropy coding unit 220.

[0305]

[0186] As described above, the residual generation unit 204 receives the video data of a current block and the corresponding predictive block. The residual generation unit 204 then generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates sample-by-sample differences between the predictive block and the current block.

[0306]

[0187] Transform processing unit 206 applies one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 may apply various transforms to the residual block to form the transform coefficient block. For example, transform processing unit 206 may apply a discrete cosine transform (DCT), a directional transform, a Karhunen-Loeve transform (KLT), or a conceptually similar transform to the residual block. In some examples, transform processing unit 206 may perform multiple transforms on the residual block, e.g., a linear transform and a secondary transform, such as a rotation transform. In some examples, transform processing unit 206 does not apply a transform to the residual block.

[0307]

[0188] The quantization unit 208 may quantize the transform coefficients in the transform coefficient block to produce a quantized transform coefficient block. In some examples, the quantization unit 208 may quantize the transform coefficients of the transform coefficient block according to a quantization parameter (QP) value associated with the current block. The video encoder 200 (e.g., via the mode selection unit 202) may adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may result in loss of information, and thus the quantized transform coefficients may have less precision than the original transform coefficients produced by the transform processing unit 206.

[0308]

[0189] The quantization unit 208 may perform the techniques of this disclosure for quantizing transform coefficients, such as the techniques described with respect to Figures 3A-7. In particular, the quantization unit 208 may perform the functions described with respect to the scaling and classification unit 152 and the quantization unit 142 of Figure 3B, the quantization offset parameter unit 154, the grouping unit 156, and the offset quantization units 158A-158P of Figure 4, and the QOV calculation unit 174, the QOV parameter calculation unit 176, the QOV calculation unit 178, the QOV index calculation unit 180, and the QOV retrieval unit 182 of Figure 6.

[0309]

[0190] Quantization unit 208 may determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data based on side information associated with the block of video data. For example, quantization unit 208 may determine, for each sub-block of a transform coefficient block, such as each 4x4 sub-block of the transform coefficient block, a set of quantization offset parameters for a group of scaled transform coefficients in the sub-block based on side information associated with the transform coefficient block. The quantization offsets in the set of quantization offset parameters for a group of scaled transform coefficients may not be constant. Instead, the quantization offsets in the set of quantization offset parameters may vary depending on a quantization interval.

[0310]

[0191] Quantization unit 208 may implement some of the techniques described in this disclosure for determining a set of quantization offset parameters for a group of scaled transform coefficients in a block of video data based on side information associated with the block of video data. For example, the side information associated with the block of video data may include any combination of one or more of a slice type (e.g., I, P, B) of the block of video data, residual data resulting from intra or inter prediction of the block of video data, a block size (e.g., 4x4, 8x8, 16x16, or 32x32) of the block of video data, and / or a luma or chroma component of the block of video data. The side information associated with the block of video data may include side information associated with a sub-block that includes a group of scaled transform coefficients, such as one or more of a maximum magnitude range of the scaled transform coefficients of the sub-block, a block size of the sub-block, a relative position of the sub-block within the transform coefficient block, etc.

[0311]

[0192] In some examples, the quantization unit 208 may determine, for a group of scaled transform coefficients, a set of quantization offset parameters that may optimize a rate-distortion cost of quantizing the group of scaled transform coefficients based on side information for the block of video data. For example, the quantization unit 208 may use machine learning techniques, such as a neural network that may be trained on the side information, the scaled transform coefficient values, an optimal rate-distortion cost of quantization, etc., to determine a set of quantization offset parameters for the group of scaled transform coefficients. The quantization unit 208 may use such a neural network to implement regression and / or classification methods to determine a set of quantization offset parameters for the group of scaled transform coefficients based on side information associated with the block of video data. Thus, the quantization unit 208 may be able to determine a set of quantization offset parameters for a group of scaled transform coefficients without using bit cost estimates determined by the entropy encoding unit 220 and without deriving, for each particular non-zero transform coefficient, an index of the arithmetic coding context used to entropy code that particular non-zero transform coefficient.

[0312]

[0193] The quantization unit 208 may quantize each group of scaled transform coefficients for a block of video data to generate quantized transform coefficients for each subblock of the block of video data based at least in part on the set of quantization offset parameters. As described above, the quantization unit 208 may determine a set of quantization offset parameters for each group of scaled transform coefficients in the block of video data, and thus the quantization unit 208 may quantize each group of scaled transform coefficients using the set of quantization offset parameters associated with the corresponding group of scaled transform coefficients. Thus, in some examples, the quantization unit 208 may be capable of quantizing scaled transform coefficients within the same group of scaled transform coefficients in parallel. Furthermore, in some examples, the quantization unit 208 may be capable of quantizing multiple groups of scaled transform coefficients for a block of video data in parallel.

[0313]

[0194] The inverse quantization unit 210 and the inverse transform processing unit 212 may apply inverse quantization and inverse transform to the quantized transform coefficient block, respectively, to reconstruct a residual block from the transform coefficient block. The reconstruction unit 214 may produce a reconstructed block that corresponds to the current block (potentially with some distortion) based on the reconstructed residual block and the predictive block generated by the mode selection unit 202. For example, the reconstruction unit 214 may add samples of the reconstructed residual block to corresponding samples from the predictive block generated by the mode selection unit 202 to produce a reconstructed block.

[0314]

[0195] Filter unit 216 may perform one or more filter operations on the reconstructed blocks. For example, filter unit 216 may perform a deblocking operation to reduce blockiness artifacts along the edges of a CU. The operations of filter unit 216 may be skipped in some examples.

[0315]

[0196] The video encoder 200 stores the reconstructed blocks in the DPB 218. For example, in examples where the operation of the filter unit 216 is not required, the reconstruction unit 214 may store the reconstructed blocks in the DPB 218. In examples where the operation of the filter unit 216 is required, the filter unit 216 may store the filtered reconstructed blocks in the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 may retrieve reference pictures formed from the reconstructed (and potentially filtered) blocks from the DPB 218 to inter predict blocks of a later-encoded picture. Furthermore, the intra prediction unit 226 may use the reconstructed blocks in the DPB 218 of the current picture to intra predict other blocks in the current picture.

[0316]

[0197] In general, entropy encoding unit 220 may entropy encode syntax elements received from other functional components of video encoder 200. For example, entropy encoding unit 220 may entropy encode quantized transform coefficient blocks from quantization unit 208 to generate an encoded video bitstream based at least in part on the quantized transform coefficients for the blocks of video data.

[0317]

[0198] As another example, entropy encoding unit 220 may entropy encode predictive syntax elements (e.g., motion information for inter prediction or intra mode information for intra prediction) from mode selection unit 202. Entropy encoding unit 220 may perform one or more entropy encoding operations on syntax elements, which are another example of video data, to generate entropy encoded data. For example, entropy encoding unit 220 may perform a context-adaptive variable length coding (CAVLC) operation, a CABAC operation, a variable-to-variable (V2V) length coding operation, a syntax-based context-adaptive binary arithmetic coding (SBAC) operation, a probability interval partitioned entropy (PIPE) coding operation, an exponential-Golomb coding operation, or another type of entropy encoding operation on the data. In some examples, entropy encoding unit 220 may operate in a bypass mode in which syntax elements are not entropy encoded.

[0318]

[0199] Video encoder 200 may output a bitstream that includes entropy coding syntax elements needed to reconstruct blocks of a slice or picture. In particular, entropy coding unit 220 may output the bitstream.

[0319]

[0200] The operations described above are described with respect to blocks. Such descriptions should be understood as being operations for luma coding blocks and / or chroma coding blocks. As described above, in some examples, the luma coding blocks and chroma coding blocks are luma and chroma components of a CU. In some examples, the luma coding blocks and chroma coding blocks are luma and chroma components of a PU.

[0320]

[0201] In some examples, operations performed with respect to luma coding blocks do not need to be repeated for chroma coding blocks. As an example, operations for identifying motion vectors (MVs) and reference pictures for luma coding blocks do not need to be repeated to identify MVs and reference pictures for chroma blocks. Rather, MVs for luma coding blocks may be scaled to determine MVs for chroma blocks, and the reference pictures may be the same. As another example, the intra prediction process may be the same for luma coding blocks and chroma coding blocks.

[0321]

[0202] Video encoder 200 represents an example of a device configured to encode video data, including a memory configured to store video data and one or more processing units implemented in a circuit and configured to determine, based on side information associated with the block of video data, a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data, quantize the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data based at least in part on the set of quantization offset parameters, and generate an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data.

[0322]

[0203] Figure 9 is a block diagram illustrating an example video decoder 300 that may implement the techniques of this disclosure. Figure 9 is provided for purposes of illustration and not to limit the techniques as broadly illustrated and described in this disclosure. For purposes of illustration, this disclosure describes the video decoder 300 in accordance with VVC (ITU-T H.266 under development) and HEVC (ITU-T H.265) techniques. However, the techniques of this disclosure may be implemented by video coding devices configured for other video coding standards.

[0323]

[0204] In the example of FIG. 9, the video decoder 300 includes a coded picture buffer (CPB) 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 314. Any or all of the CPB memory 320, the entropy decoding unit 302, the prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, the filter unit 312, and the DPB 314 may be implemented in one or more processors or processing circuits. For example, the units of the video decoder 300 may be implemented as one or more circuits or logic elements, as part of a hardware circuit, or as part of a processor of an FPGA, an ASIC. Moreover, the video decoder 300 may include additional or alternative processors or processing circuits for performing these and other functions.

[0324]

[0205] The prediction processing unit 304 includes a motion compensation unit 316 and an intra prediction unit 318. The prediction processing unit 304 may include additional units for performing prediction according to other prediction modes. By way of example, the prediction processing unit 304 may include a palette unit, an intra block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, a composite inter-intra prediction (CIIP) unit, etc. In other examples, the video decoder 300 may include more, fewer, or differently functional components.

[0325]

[0206] The CPB memory 320 may store video data, such as an encoded video bitstream, to be decoded by components of the video decoder 300. The video data stored in the CPB memory 320 may be obtained, for example, from the computer-readable medium 110 (FIG. 1). The CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. The CPB memory 320 may also store video data other than syntax elements of a coded picture, such as temporary data representing output from various units of the video decoder 300. The DPB 314 generally stores decoded pictures that the video decoder 300 may output and / or use as reference video data when decoding subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and the DPB 314 may be formed by any of a variety of memory devices, such as DRAM, including SDRAM, MRAM, RRAM, or other types of memory devices. The CPB memory 320 and the DPB 314 may be provided by the same memory device or separate memory devices. In various examples, the CPB memory 320 may be on-chip with other components of the video decoder 300 or off-chip relative to those components.

[0326] Additionally or alternatively, in some examples, video decoder 300 may retrieve coded video data from memory 120 (FIG. 1). That is, memory 120 may store data as described above using CPB memory 320. Similarly, memory 120 may store instructions to be executed by video decoder 300 when some or all of the functionality of video decoder 300 is implemented in software to be executed by processing circuitry of video decoder 300.

[0327]

[0208] The various units shown in FIG. 9 are shown to aid in understanding the operations performed by the video decoder 300. The units may be implemented as fixed-function circuits, programmable circuits, or a combination thereof. As with FIG. 8, a fixed-function circuit refers to a circuit that provides a specific function and is preconfigured into the operations that may be performed. A programmable circuit refers to a circuit that may be programmed to perform various tasks and provide flexible functionality in the operations that may be performed. For example, a programmable circuit may execute software or firmware that causes the programmable circuit to operate in a manner defined by the software or firmware instructions. Although a fixed-function circuit may execute software instructions (e.g., to receive parameters or output parameters), the type of operations that the fixed-function circuit performs is generally unchanging. In some examples, one or more of the units may be separate circuit blocks (fixed function or programmable), and in some examples, one or more of the units may be an integrated circuit.

[0328]

[0209] The video decoder 300 may include a programmable core formed from ALUs, EFUs, digital circuits, analog circuits, and / or programmable circuits. In examples in which the operations of the video decoder 300 are performed by software executing on programmable circuits, on-chip or off-chip memory may store instructions (e.g., object code) of the software that the video decoder 300 receives and executes.

[0329]

[0210] The entropy decoding unit 302 may receive the encoded video data from the CPB and entropy decode the video data to recover the syntax elements. The prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, and the filter unit 312 may generate decoded video data based on the syntax elements extracted from the bitstream.

[0330]

[0211] In general, the video decoder 300 reconstructs a picture on a block-by-block basis. The video decoder 300 may perform a reconstruction operation on each block individually (wherein the currently reconstructed, i.e., currently decoded, block may be referred to as the "current block").

[0331]

[0212] The entropy decoding unit 302 may entropy decode syntax elements defining the quantized transform coefficients of the quantized transform coefficient block, as well as transform information such as a quantization parameter (QP) and / or a transform mode indication. The inverse quantization unit 306 may use the QP associated with the quantized transform coefficient block to determine the degree of quantization, and similarly, to determine the degree of inverse quantization that the inverse quantization unit 306 should apply. The inverse quantization unit 306 may perform, for example, a bitwise left shift operation to inverse quantize the quantized transform coefficients. The inverse quantization unit 306 may thereby form a transform coefficient block including the transform coefficients.

[0332]

[0213] After the inverse quantization unit 306 forms a transform coefficient block, the inverse transform processing unit 308 may apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 may apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotational transform, an inverse transform, or another inverse transform to the transform coefficient block.

[0333]

[0214] Furthermore, prediction processing unit 304 generates a prediction block according to the prediction information syntax element entropy decoded by entropy decoding unit 302. For example, if the prediction information syntax element indicates that the current block is inter predicted, motion compensation unit 316 may generate a prediction block. In this case, the prediction information syntax element may indicate a reference picture in DPB 314 from which to retrieve the reference block, as well as a motion vector that identifies the location of the reference block in the reference picture relative to the location of the current block in the current picture. Motion compensation unit 316 may generally perform the inter prediction process in a manner substantially similar to that described with respect to motion compensation unit 224 (FIG. 8).

[0334]

[0215] As another example, if the prediction information syntax element indicates that the current block is intra predicted, the intra prediction unit 318 may generate a prediction block according to the intra prediction mode indicated by the prediction information syntax element. Again, the intra prediction unit 318 may generally perform an intra prediction process in a manner substantially similar to that described with respect to the intra prediction unit 226 (FIG. 8). The intra prediction unit 318 may retrieve data of neighboring samples for the current block from the DPB 314.

[0335]

[0216] The reconstruction unit 310 may reconstruct the current block using the predictive block and the residual block. For example, the reconstruction unit 310 may add samples of the residual block to corresponding samples of the predictive block to reconstruct the current block.

[0336]

[0217] Filter unit 312 may perform one or more filter operations on the reconstructed block. For example, filter unit 312 may perform a deblocking operation to reduce blockiness artifacts along edges of the reconstructed block. The operations of filter unit 312 may not be performed in all instances.

[0337]

[0218] The video decoder 300 may store the reconstructed block in the DPB 314. For example, in an example where the operation of the filter unit 312 is not performed, the reconstruction unit 310 may store the reconstructed block in the DPB 314. In an example where the operation of the filter unit 312 is performed, the filter unit 312 may store the filtered reconstructed block in the DPB 314. As described above, the DPB 314 may provide reference information to the prediction processing unit 304, such as samples of the current picture for intra prediction and previously decoded pictures for subsequent motion compensation. Moreover, the video decoder 300 may output the decoded picture (e.g., decoded video) from the DPB 314 for subsequent presentation on a display device, such as the display device 118 of FIG. 1.

[0338]

[0219] Thus, the video decoder 300 represents an example of a video decoding device that includes a memory configured to store video data and one or more processing units implemented in a circuit and configured to implement the techniques of the present disclosure.

[0339]

[0220] Figure 10 is a flowchart illustrating an example method for encoding a current block. The current block may comprise a current CU. Although described with respect to video encoder 200 (Figures 1 and 8), it should be understood that other devices may be configured to implement a method similar to that of Figure 10.

[0340]

[0221] In this example, video encoder 200 first predicts the current block (350). For example, video encoder 200 may form a predictive block for the current block. Video encoder 200 may then calculate a residual block for the current block (352). To calculate the residual block, video encoder 200 may calculate the difference between the original uncoded block and the predictive block for the current block. Video encoder 200 may then transform the residual block and quantize transform coefficients of the residual block (354). Video encoder 200 may then scan the quantized transform coefficients of the residual block (356). Specifically, video encoder 200 may implement techniques of this disclosure including determining a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data based on side information associated with the block of video data, quantizing the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data based at least in part on the set of quantization offset parameters, and generating an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data.

[0341]

[0222] During or following the scan, video encoder 200 may entropy encode the transform coefficients (358). For example, video encoder 200 may encode the transform coefficients using CAVLC or CABAC. Video encoder 200 may then output entropy encoded data for the block (360).

[0342] 11 is a flow chart illustrating an example method for decoding a current block of video data. The current block may comprise a current CU. Although described with respect to video decoder 300 (FIGS. 1 and 9), it should be understood that other devices may be configured to implement a method similar to that of FIG.

[0343]

[0224] The video decoder 300 may receive entropy coded data for the current block, such as the entropy coded prediction information and the entropy coded data for the transform coefficients of the residual block that correspond to the current block (370). The video decoder 300 may entropy decode the entropy coded data to determine the prediction information for the current block and to reconstruct the transform coefficients of the residual block (372). The video decoder 300 may predict the current block, e.g., using an intra-prediction or inter-prediction mode indicated by the prediction information for the current block, to compute a predictive block for the current block (374). The video decoder 300 may then inverse scan the reconstructed transform coefficients to create a block of quantized transform coefficients (376). The video decoder 300 may then inverse quantize the transform coefficients and apply an inverse transform to the transform coefficients to produce a residual block (378). The video decoder 300 finally decodes the current block by combining the predictive block and the residual block (380).

[0344]

[0225] Figure 12 is a flow chart illustrating an example method for encoding video data in accordance with the techniques of this disclosure. Although described with respect to video encoder 200 (Figures 1 and 8), it should be understood that other devices may be configured to perform a method similar to that of Figure 12.

[0345] 12, video encoder 200 (e.g., quantization unit 208) may determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data based on side information associated with the block of video data (402). In some examples, to determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data, video encoder 200 may classify the group of scaled transform coefficients based at least in part on the side information associated with the block of video data, and select a set of quantization offset parameters for the group of scaled transform coefficients from a plurality of sets of quantization offset parameters based on the classification of the group of scaled transform coefficients.

[0346]

[0227] In some examples, to classify the group of scaled transform coefficients, video encoder 200 may determine an index for the group of scaled transform coefficients based at least in part on side information associated with the block of video data, and may select a set of quantization offset parameters for the group of scaled transform coefficients from the multiple sets of quantization offset parameters based on the classification of the group of scaled transform coefficients, and video encoder 200 may index into the multiple sets of quantization offset parameters using the determined index to select the set of quantization offset parameters for the group of scaled transform coefficients.

[0347]

[0228] In some examples, the group of scaled transform coefficients comprises scaled transform coefficients for sub-blocks of the block of video data, and the side information associated with the block of video data comprises a location of the sub-block within the block of video data and a block size of the block of video data. In some examples, classifying the group of scaled transform coefficients based at least in part on the side information associated with the block of video data is not based on values ​​of the scaled transform coefficients for the sub-blocks of the block of video data.

[0348]

[0229] In some examples, to determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data, the video encoder 200 may determine a set of parameters having a smaller magnitude than the set of quantization offset parameters based on side information associated with the block of video data, and may determine the set of quantization offset parameters for the group of scaled transform coefficients based on the set of parameters.

[0349]

[0230] In some examples, to determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data, video encoder 200 may use a neural network to determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data. In some examples, to determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data, video encoder 200 may use at least one of classification or regression techniques to determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data.

[0350]

[0231] The video encoder 200 (e.g., quantization unit 208) may quantize the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data based at least in part on the set of quantization offset parameters (404).

[0351]

[0232] Video encoder 200 (eg, entropy encoding unit 220) may generate an encoded video bitstream based at least in part on the quantized transform coefficients for the blocks of video data (406).

[0352]

[0233] In some examples, video encoder 200 may divide the multiple scaled transform coefficients into multiple groups of scaled transform coefficients associated with sub-blocks of the block of video data, where the multiple groups of scaled transform coefficients include groups of scaled transform coefficients for the sub-blocks of the block of video data. In some examples, to determine a set of quantization offset parameters for the groups of scaled transform coefficients for the block of video data, video encoder 200 may determine a corresponding set of quantization offset parameters for each of the multiple groups of scaled transform coefficients, where quantizing the groups of scaled transform coefficients for the sub-blocks of the block of video data includes quantizing each of the multiple groups of scaled transform coefficients based on the corresponding set of quantization offset parameters.

[0353]

[0234] In some examples, to quantize scaled transform coefficients for sub-blocks of a block of video data, the video encoder 200 may determine, for each scaled transform coefficient, a corresponding quantization offset parameter from a set of quantization offset parameters, and may quantize each scaled transform coefficient based at least in part on the corresponding quantization offset parameter.

[0354]

[0235] In some examples, the multiple sets of quantization offset parameters comprise a quantization offset vector.

[0355]

[0236] Illustrative examples of the first aspect of the present disclosure include the following:

[0356]

[0237] Aspect 1: A method for coding video data, the method comprising any combination of techniques described in this disclosure.

[0357]

[0238] Aspect 2: A method for coding video data, comprising: determining a plurality of quantization offset vectors for a plurality of transform coefficients for a current block of video data based on scaled transform coefficients for the current block and side information associated with the current block; and quantizing the transform coefficients for the current block to generate quantized transform coefficients for the current block based at least in part on the plurality of quantization offset vectors.

[0358]

[0239] Aspect 3: The method of aspect 2, further comprising splitting the transform coefficients into a plurality of subgroups of transform coefficients, and determining a plurality of quantization offset vectors for the plurality of transform coefficients for the current block comprises determining a quantization offset vector for each subgroup of the plurality of subgroups of transform coefficients.

[0359]

[0240] Aspect 4: A method according to any combination of aspects 2 and 3, wherein determining a plurality of quantization offset vectors further comprises determining a plurality of index values ​​and using the plurality of index values ​​to index into a table of quantization offset vectors to determine the plurality of quantization offset vectors.

[0360]

[0241] Aspect 5: The method of aspect 4, wherein determining the plurality of index values ​​is based at least in part on values ​​of a plurality of transform coefficients in the transform coefficient group.

[0361]

[0242] Aspect 6: The method of any combination of aspects 4 and 5, wherein determining the plurality of index values ​​is based at least in part on a size of the transform coefficient block and a location of the transform coefficient group within the transform coefficient block.

[0362]

[0243] Aspect 7: The method of any combination of aspects 4 to 6, wherein determining the plurality of index values ​​is based at least in part on whether the current block is part of an intra slice.

[0363]

[0244] Aspect 8: A method according to any combination of aspects 2 to 7, wherein determining a plurality of quantization offset vectors further comprises determining one or more of the plurality of quantization offset vectors without scaled transform coefficients for the current block, and determining other quantization offset vectors based on the scaled transform coefficients for the current block.

[0364]

[0245] Aspect 9: A method according to any combination of aspects 2 to 8, wherein determining a plurality of quantization offset vectors for a plurality of transform coefficients for a current block of video data comprises defining the quantization offset vectors using a parameter vector of a smaller magnitude.

[0365]

[0246] Aspect 10: A method according to any combination of aspects 2 to 9, wherein the side information comprises one or more of a slice type, a block size, a type of prediction, or an indication of whether the current block to be quantized comprises a luminance component or a chrominance component.

[0366]

[0247] Aspect 11: A method according to any combination of aspects 2 to 10, wherein quantizing the transform coefficients for a current block of video data comprises quantizing a first transform coefficient of the transform coefficients for the current block in parallel with a second transform coefficient of the transform coefficients for the current block.

[0367]

[0248] Aspect 12: A method according to any combination of aspects 2 to 11, wherein determining a plurality of quantization offset vectors for a plurality of transform coefficients for a current block comprises determining an estimated change in the number of bits for entropy coding the quantized transform coefficients for the current block of video data based on a change in a single element of the quantized transform coefficients for the current block.

[0368]

[0249] Aspect 13: The method described in aspect 12, wherein determining an estimated change in the number of bits for entropy coding the quantized transform coefficients for the current block based on a change in a single element of the quantized transform coefficients for the current block comprises using the same estimation rule calculated once to determine multiple quantization offset vectors for multiple transform coefficients for the current block.

[0369]

[0250] Aspect 14: A method according to any combination of aspects 2 to 13, wherein determining a plurality of quantization offset vectors for a plurality of transform coefficients for the current block comprises determining, for each particular non-zero transform coefficient, a plurality of quantization offset vectors for a plurality of transform coefficients for the current block without deriving an index of an arithmetic coding context used to entropy code that particular non-zero transform coefficient.

[0370]

[0251] Aspect 15: A method according to any combination of aspects 2 to 14, wherein determining a plurality of quantization offset vectors for a plurality of transform coefficients for the current block comprises optimizing values ​​of the plurality of quantization offset vectors using at least one of statistical techniques or machine learning techniques.

[0371]

[0252] Aspect 16: The method of aspect 15, wherein at least one of the statistical techniques or machine learning techniques comprises at least one of the classification techniques or regression techniques.

[0372]

[0253] Aspect 17: The method of aspect 16, wherein the regression technique comprises a general regression technique.

[0373]

[0254] Aspect 18: The method of any combination of aspects 16 and 17, wherein the classification technique comprises a classification tree.

[0374]

[0255] Aspect 19: The method of any combination of aspects 2 to 18, wherein the coding comprises decoding.

[0375]

[0256] Aspect 20: The method of any combination of aspects 2 to 18, wherein the coding comprises encoding.

[0376]

[0257] Aspect 21: A device for coding video data, comprising one or more means for performing the method according to any one of aspects 1 to 20.

[0377]

[0258] Aspect 22: The device of aspect 21, wherein the one or more means comprise one or more processors implemented in circuitry.

[0378]

[0259] Aspect 23: A device described in any combination of aspects 21 and 22, further comprising a memory for storing video data.

[0379]

[0260] Aspect 24: The device of any combination of aspects 21 to 23, further comprising a display configured to display the decoded video data.

[0380]

[0261] Aspect 25: A device described in any combination of aspects 21 to 24, wherein the device comprises one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.

[0381]

[0262] Aspect 26: The device of any combination of aspects 21 to 25, wherein the device comprises a video decoder.

[0382]

[0263] Aspect 27: The device of any combination of aspects 21 to 26, wherein the device comprises a video encoder.

[0383]

[0264] Aspect 28: A computer-readable storage medium storing instructions that, when executed, cause one or more processors to perform a method according to any combination of aspects 1 to 20.

[0384]

[0265] Illustrative examples of the second aspect of the present disclosure include the following:

[0385]

[0266] Aspect 1: A method for encoding video data, comprising: determining a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data based on side information associated with the block of video data; quantizing the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data based at least in part on the set of quantization offset parameters; and generating an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data.

[0386]

[0267] Aspect 2: The method of aspect 1, wherein determining a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data comprises selecting a set of quantization offset parameters for the group of scaled transform coefficients from a plurality of sets of quantization offset parameters based on side information associated with the block of video data.

[0387]

[0268] Aspect 3: The method of aspect 2, wherein selecting a set of quantization offset parameters for a group of scaled transform coefficients from a plurality of sets of quantization offset parameters comprises determining an index associated with the group of scaled transform coefficients based at least in part on side information associated with the block of video data, and indexing into the plurality of sets of quantization offset parameters using the index associated with the group of scaled transform coefficients to select a set of quantization offset parameters for the group of scaled transform coefficients.

[0388]

[0269] Aspect 4: The method described in aspect 2 or 3, wherein the group of scaled transform coefficients comprises scaled transform coefficients for a sub-block of a block of video data, and the side information associated with the block of video data comprises a location of the sub-block within the block of video data and a block size of the block of video data.

[0389]

[0270] Aspect 5: The method described in aspect 4, wherein selecting a set of quantization offset parameters for a group of scaled transform coefficients from a plurality of sets of quantization offset parameters comprises selecting a set of quantization offset parameters for a group of scaled transform coefficients from a plurality of sets of quantization offset parameters without using scaled transform coefficient values ​​of the group of scaled transform coefficients for a sub-block of a block of video data.

[0390]

[0271] Aspect 6: Determining a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data includes: The method of aspect 1, comprising: determining a set of parameters for parameterization of a set of quantization offset parameters based on side information associated with a block of video data, the set of parameters having a smaller size than the set of quantization offset parameters; and determining a set of quantization offset parameters for a group of scaled transform coefficients based on the set of parameters having the smaller size.

[0391]

[0272] Aspect 7: A method as described in any one of aspects 1 to 5, wherein determining a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data further comprises determining a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data based on side information associated with the block of video data using a neural network.

[0392]

[0273] Aspect 8: The method described in aspect 7, wherein determining a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data further comprises determining a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data using at least one of a classification technique or a regression technique.

[0393]

[0274] Aspect 9: The method of any one of aspects 1 to 8, further comprising dividing a plurality of scaled transform coefficients of a block of video data into a plurality of groups of scaled transform coefficients associated with sub-blocks of the block of video data, wherein the plurality of groups of scaled transform coefficients include groups of scaled transform coefficients for the block of video data, wherein determining a set of quantization offset parameters for the groups of scaled transform coefficients for the block of video data comprises determining a corresponding set of quantization offset parameters for each of the plurality of groups of scaled transform coefficients, and wherein quantizing the groups of scaled transform coefficients for the block of video data includes quantizing each of the plurality of groups of scaled transform coefficients based on the corresponding set of quantization offset parameters.

[0394]

[0275] Aspect 10: A method according to any one of aspects 1 to 9, wherein quantizing the group of scaled transform coefficients for a block of video data further includes, for each scaled transform coefficient of the group of scaled transform coefficients, determining a corresponding quantization offset parameter from a set of quantization offset parameters, and quantizing each scaled transform coefficient of the group of scaled transform coefficients based at least in part on the corresponding quantization offset parameter.

[0395]

[0276] Aspect 11: A method described in any one of aspects 1 to 10, wherein the side information includes one or more of a slice type of the block of video data, a block size of the block of video data, or an indication of whether the block of video data has a luminance component or a chrominance component.

[0396]

[0277] Aspect 12: A method as described in any one of aspects 1 to 11, wherein determining a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data comprises determining a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data without using one or more bit cost estimates determined via entropy coding.

[0397]

[0278] Aspect 13: The method of any one of aspects 1 to 12, wherein the set of quantization offset parameters comprises a quantization offset vector.

[0398]

[0279] Aspect 14: A device for encoding video data, comprising: a memory; and a processing circuit in communication with the memory, wherein the processing circuit is configured to determine a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data based on side information associated with the block of video data, quantize the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data based at least in part on the set of quantization offset parameters, and generate an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data.

[0399]

[0280] Aspect 15: The device described in aspect 14, wherein to determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data, the processing circuit is further configured to select a set of quantization offset parameters for the group of scaled transform coefficients from a plurality of sets of quantization offset parameters based on side information associated with the block of video data.

[0400]

[0281] Aspect 16: The device described in aspect 15, wherein to select a set of quantization offset parameters for the group of scaled transform coefficients from a plurality of sets of quantization offset parameters, the processing circuitry is further configured to determine an index associated with the group of scaled transform coefficients based at least in part on side information associated with the block of video data, and to index into the plurality of sets of quantization offset parameters using the index associated with the group of scaled transform coefficients to select a set of quantization offset parameters for the group of scaled transform coefficients.

[0401]

[0282] Aspect 17: A device as described in aspect 15 or 16, wherein the group of scaled transform coefficients comprises scaled transform coefficients for a sub-block of a block of video data, and the side information associated with the block of video data comprises a location of the sub-block within the block of video data and a block size of the block of video data.

[0402]

[0283] Aspect 18: The device described in aspect 17, wherein to select a set of quantization offset parameters for the group of scaled transform coefficients from the multiple sets of quantization offset parameters, the processing circuitry is further configured to select a set of quantization offset parameters for the group of scaled transform coefficients from the multiple sets of quantization offset parameters without using scaled transform coefficient values ​​of the group of scaled transform coefficients for a sub-block of a block of video data.

[0403]

[0284] Aspect 19: The device of aspect 14, wherein to determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data, the processing circuit is further configured to determine a set of parameters for parameterization of the set of quantization offset parameters based on side information associated with the block of video data, and to determine the set of quantization offset parameters for the group of scaled transform coefficients based on a set of parameters having a smaller size, the set of parameters having a smaller size than the set of quantization offset parameters.

[0404]

[0285] Aspect 20: A device described in any one of aspects 14 to 18, wherein to determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data, the processing circuit is further configured to use a neural network to determine a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data based on side information associated with the block of video data.

[0405]

[0286] Aspect 21: The device described in aspect 20, wherein the processing circuit is further configured to determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data using at least one of a classification technique or a regression technique to determine a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data.

[0406]

[0287] Aspect 22: The device of any one of aspects 14 to 21, wherein the processing circuit is further configured to divide a plurality of scaled transform coefficients of a block of video data into a plurality of groups of scaled transform coefficients associated with sub-blocks of the block of video data, wherein the plurality of groups of scaled transform coefficients include a group of scaled transform coefficients for the block of video data, wherein to determine a set of quantization offset parameters for the group of scaled transform coefficients for the block of video data, the processing circuit is further configured to determine a corresponding set of quantization offset parameters for each of the plurality of groups of scaled transform coefficients, and wherein quantizing the group of scaled transform coefficients for the block of video data includes quantizing each of the plurality of groups of scaled transform coefficients based on the corresponding set of quantization offset parameters.

[0407]

[0288] Aspect 23: A device described in any one of aspects 14 to 22, wherein, to quantize a group of scaled transform coefficients for a block of video data, the processing circuitry is further configured to determine, for each scaled transform coefficient of the group of scaled transform coefficients, a corresponding quantization offset parameter from a set of quantization offset parameters, and quantize each scaled transform coefficient of the group of scaled transform coefficients based at least in part on the corresponding quantization offset parameter.

[0408]

[0289] Aspect 24: A device described in any one of aspects 14 to 23, wherein the side information includes one or more of a slice type of the block of video data, a block size of the block of video data, or an indication of whether the block of video data has a luminance component or a chrominance component.

[0409]

[0290] Aspect 25: A device described in any one of aspects 14 to 24, wherein, to determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data, the processing circuit is further configured to determine a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data without using one or more bit cost estimates determined via entropy coding.

[0410]

[0291] Aspect 26: A device described in any one of aspects 14 to 25, wherein the multiple sets of quantization offset parameters comprise quantization offset vectors.

[0411]

[0292] Aspect 27: A device described in any one of aspects 14 to 26, wherein the device comprises one or more of a camera, a computer, or a mobile device.

[0412]

[0293] Aspect 28: An apparatus for encoding video data, comprising: means for determining a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data based on side information associated with the block of video data; means for quantizing the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data based at least in part on the set of quantization offset parameters; and means for generating an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data.

[0413]

[0294] Aspect 29: The apparatus described in aspect 28, wherein the means for determining a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data further comprises means for selecting a set of quantization offset parameters for the group of scaled transform coefficients from a plurality of sets of quantization offset parameters based on side information associated with the block of video data.

[0414]

[0295] Aspect 30: The apparatus described in aspect 29, wherein the means for selecting a set of quantization offset parameters for the group of scaled transform coefficients from a plurality of sets of quantization offset parameters further comprises: means for determining an index associated with the group of scaled transform coefficients based at least in part on side information associated with the block of video data; and means for indexing into the plurality of sets of quantization offset parameters to select a set of quantization offset parameters for the group of scaled transform coefficients using the index associated with the group of scaled transform coefficients.

[0415]

[0296] Aspect 31: The apparatus of aspect 29 or 30, wherein the group of scaled transform coefficients comprises scaled transform coefficients for a sub-block of a block of video data, and the side information associated with the block of video data comprises a location of the sub-block within the block of video data and a block size of the block of video data.

[0416]

[0297] Aspect 32: The apparatus described in aspect 31, wherein the means for selecting a set of quantization offset parameters for the group of scaled transform coefficients from a plurality of sets of quantization offset parameters further comprises means for selecting a set of quantization offset parameters for the group of scaled transform coefficients from the plurality of sets of quantization offset parameters without using scaled transform coefficient values ​​of the group of scaled transform coefficients for a sub-block of a block of video data.

[0417]

[0298] Aspect 33: The apparatus described in aspect 28, wherein the means for determining a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data further comprises: means for determining a set of parameters for parameterization of the set of quantization offset parameters based on side information associated with the block of video data; and means for determining the set of quantization offset parameters for the group of scaled transform coefficients based on a set of parameters having a smaller size, where the set of parameters has a smaller size than the set of quantization parameters.

[0418]

[0299] Aspect 34: An apparatus described in any one of aspects 28 to 32, wherein the means for determining a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data further comprises means for determining, using a neural network, a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data based on side information associated with the block of video data.

[0419]

[0300] Aspect 35: The apparatus described in aspect 34, wherein the means for determining a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data comprises means for determining a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data using at least one of a classification technique or a regression technique.

[0420]

[0301] Aspect 36: The apparatus of any one of aspects 28 to 35, further comprising means for dividing a plurality of scaled transform coefficients of a block of video data into a plurality of groups of scaled transform coefficients associated with sub-blocks of the block of video data, wherein the plurality of groups of scaled transform coefficients include a group of scaled transform coefficients for the block of video data, wherein the means for determining a set of quantization offset parameters for the group of scaled transform coefficients for the block of video data comprises means for determining a corresponding set of quantization offset parameters for each of the plurality of groups of scaled transform coefficients, and wherein the means for quantizing the group of scaled transform coefficients for the block of video data further comprises means for quantizing each of the plurality of groups of scaled transform coefficients based on the corresponding set of quantization offset parameters.

[0421]

[0302] Aspect 37: An apparatus described in any one of aspects 28 to 36, wherein the means for quantizing a group of scaled transform coefficients for a block of video data further comprises: means for determining, for each scaled transform coefficient of the group of scaled transform coefficients, a corresponding quantization offset parameter from a set of quantization offset parameters, and means for quantizing each scaled transform coefficient of the group of scaled transform coefficients based at least in part on the corresponding quantization offset parameter.

[0422]

[0303] Aspect 38: An apparatus described in any one of aspects 28 to 37, wherein the side information includes one or more of a slice type of the block of video data, a block size of the block of video data, or an indication of whether the block of video data has a luminance component or a chrominance component.

[0423]

[0304] Aspect 39: An apparatus described in any one of aspects 28 to 38, wherein means for determining a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data comprises means for determining a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data without using one or more bit cost estimates determined via entropy coding.

[0424]

[0305] Aspect 40: The apparatus of any one of aspects 28 to 39, wherein the set of quantization offset parameters comprises a quantization offset vector.

[0425]

[0306] Aspect 41: A computer-readable storage medium having instructions stored thereon that, when executed, cause one or more processors to determine a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data based on side information associated with the block of video data, quantize the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data based at least in part on the set of quantization offset parameters, and generate an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data.

[0426]

[0307] Aspect 42: A computer-readable storage medium as described in aspect 41, wherein instructions for causing one or more processors to determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data include instructions for causing one or more processors to select a set of quantization offset parameters for the group of scaled transform coefficients from a plurality of sets of quantization offset parameters based on side information associated with the block of video data.

[0427]

[0308] Aspect 43: A computer-readable storage medium as described in aspect 42, wherein the instructions for causing one or more processors to select a set of quantization offset parameters for a group of scaled transform coefficients from a plurality of sets of quantization offset parameters include instructions for causing the one or more processors to determine an index for the group of scaled transform coefficients based at least in part on side information associated with the block of video data, and to index into the plurality of sets of quantization offset parameters using the index to select a set of quantization offset parameters for the group of scaled transform coefficients.

[0428]

[0309] Aspect 44: A computer-readable storage medium as described in aspect 42 or 43, wherein the group of scaled transform coefficients comprises scaled transform coefficients for a sub-block of a block of video data, and the side information associated with the block of video data comprises a location of the sub-block within the block of video data and a block size of the block of video data.

[0429]

[0310] Aspect 45: A computer-readable storage medium as described in aspect 44, comprising instructions for causing one or more processors to select a set of quantization offset parameters for a group of scaled transform coefficients from a plurality of sets of quantization offset parameters, the instructions causing one or more processors to select a set of quantization offset parameters for a group of scaled transform coefficients from a plurality of sets of quantization offset parameters without using scaled transform coefficient values ​​of the group of scaled transform coefficients for a sub-block of a block of video data.

[0430]

[0311] Aspect 46: A computer-readable storage medium as described in aspect 41, wherein instructions for causing one or more processors to determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data include instructions for causing one or more processors to determine a set of parameters for parameterization of the set of quantization offset parameters based on side information associated with the block of video data, and determine the set of quantization offset parameters for the group of scaled transform coefficients based on a set of parameters having a smaller size, where the set of parameters has a smaller size than the set of quantization offset parameters.

[0431]

[0312] Aspect 47: A computer-readable storage medium as described in any one of aspects 41 to 45, wherein the instructions for causing one or more processors to determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data include instructions for causing one or more processors to determine, using a neural network, a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data based on side information associated with the block of video data.

[0432]

[0313] Aspect 48: A computer-readable storage medium as described in aspect 47, wherein the instructions for causing one or more processors to determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data include instructions for causing one or more processors to determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data using at least one of a classification technique or a regression technique.

[0433]

[0314] Aspect 49: A computer-readable storage medium of any one of aspects 41 to 48, wherein the instructions further cause the one or more processors to divide a plurality of scaled transform coefficients of a block of video data into a plurality of groups of scaled transform coefficients associated with sub-blocks of the block of video data, wherein the plurality of groups of scaled transform coefficients include a group of scaled transform coefficients for the block of video data, wherein the instructions causing the one or more processors to determine a set of quantization offset parameters for the groups of scaled transform coefficients for the block of video data comprise instructions that cause the one or more processors to determine a corresponding set of quantization offset parameters for each of the plurality of groups of scaled transform coefficients, and wherein the instructions causing the one or more processors to quantize the group of scaled transform coefficients for the block of video data comprise instructions that cause the one or more processors to quantize each of the plurality of groups of scaled transform coefficients based on the corresponding set of quantization offset parameters.

[0434]

[0315] Aspect 50: A computer-readable storage medium according to any one of aspects 41 to 49, wherein the instructions for causing one or more processors to quantize a group of scaled transform coefficients for a block of video data include instructions for causing one or more processors to determine, for each scaled transform coefficient of the group of scaled transform coefficients, a corresponding quantization offset parameter from a set of quantization offset parameters, and to quantize each scaled transform coefficient of the group of scaled transform coefficients based at least in part on the corresponding quantization offset parameter.

[0435]

[0316] Aspect 51: A computer-readable storage medium according to any one of aspects 41 to 50, wherein the side information includes one or more of a slice type of the block of video data, a block size of the block of video data, or an indication of whether the block of video data has a luminance component or a chrominance component.

[0436]

[0317] Aspect 52: A computer-readable storage medium as described in any one of aspects 41 to 51, wherein instructions for causing one or more processors to determine a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data include instructions for causing one or more processors to determine a set of quantization offset parameters for a group of scaled transform coefficients for the block of video data without using one or more bit cost estimates determined via entropy coding.

[0437]

[0318] Aspect 53: A computer-readable storage medium according to any one of aspects 41 or 52, wherein the set of quantization offset parameters comprises a quantization offset vector.

[0438]

[0319] In accordance with the above examples, it should be recognized that some acts or events of any of the techniques described herein may be performed in a different sequence, added, merged, or omitted entirely (e.g., not all described acts or events may be necessary to the practice of the techniques). Moreover, in some examples, acts or events may be performed simultaneously rather than sequentially, for example, through multi-threaded processing, interrupt processing, or multiple processors.

[0439]

[0320] In one or more examples, the described functions may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored on or transmitted via a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium, which corresponds to a tangible medium, such as a data storage medium, or may include a communication medium, which includes any medium that facilitates the transfer of a computer program from one place to another, for example, according to a communication protocol. In this manner, a computer-readable medium may generally correspond to (1) a tangible computer-readable storage medium that is non-transitory, or (2) a communication medium, such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code and / or data structures for implementation of the techniques described in this disclosure. A computer program product may include a computer-readable medium.

[0440]

[0321] By way of example and not limitation, such computer-readable storage media may comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, flash memory, or any other medium that may be used to store desired program code in the form of instructions or data structures and that may be accessed by a computer. Also, any connection is properly referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transitory media, but instead cover non-transitory tangible storage media. As used herein, disk and disc include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically and discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer readable media.

[0441]

[0322] The instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Thus, the terms "processor" and "processing circuitry" as used herein may refer to any of the above structures, or any other structures suitable for implementing the techniques described herein. Furthermore, in some aspects, the functions described herein may be provided in dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into a composite codec. Also, the techniques may be fully implemented in one or more circuits or logic elements.

[0442]

[0323] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including wireless handsets, integrated circuits (ICs) or sets of ICs (e.g., chipsets). This disclosure has described various components, modules, or units to highlight functional aspects of devices configured to perform the disclosed techniques, but they do not necessarily need to be realized by different hardware units. Rather, as described above, the various units may be combined in a codec hardware unit or provided by a collection of interoperable hardware units, including one or more processors described above, along with suitable software and / or firmware.

[0443]

[0324] Various examples have been described. These and other examples are within the scope of the following claims. The invention as described in the claims of the original application is set forth below. [C1] A method for encoding video data, comprising the steps of: determining a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data based on side information associated with the block of video data; quantizing the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data based at least in part on the set of quantization offset parameters; generating an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data; A method for providing the above. [C2] Determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data comprises: selecting a set of quantization offset parameters for the group of scaled transform coefficients from a plurality of sets of quantization offset parameters based on the side information associated with the block of video data; The method of claim C1, comprising: [C3] Selecting the set of quantization offset parameters for the group of scaled transform coefficients from the plurality of sets of quantization offset parameters comprises: determining an index associated with the group of scaled transform coefficients based at least in part on the side information associated with the block of video data; indexing into the plurality of sets of quantization offset parameters to select the set of quantization offset parameters for the group of scaled transform coefficients using the index associated with the group of scaled transform coefficients; The method of C2, comprising: [C4] The group of scaled transform coefficients comprises scaled transform coefficients for sub-blocks of the block of video data; the side information associated with the block of video data comprises a location of the sub-block within the block of video data and a block size of the block of video data. The method according to C2. [C5] Selecting the set of quantization offset parameters for the group of scaled transform coefficients from the plurality of sets of quantization offset parameters comprises: selecting a set of quantization offset parameters for the group of scaled transform coefficients from the multiple sets of quantization offset parameters without using scaled transform coefficient values ​​of the group of scaled transform coefficients for the sub-block of the block of video data; The method of claim C4, comprising: [C6] Determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data comprises: determining a set of parameters for parameterization of the set of quantization offset parameters based on the side information associated with the block of video data, the set of parameters having a smaller size than the set of quantization offset parameters; determining the set of quantization offset parameters for the group of scaled transform coefficients based on the set of parameters having the smaller size; The method of claim C1, comprising: [C7] Determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data comprises: determining, using a neural network, the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data based on the side information associated with the block of video data. [C8] Determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data comprises: determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data using at least one of a classification technique or a regression technique; The method of C7, further comprising: [C9] dividing a plurality of scaled transform coefficients of the block of video data into a plurality of groups of scaled transform coefficients associated with sub-blocks of the block of video data; and wherein the plurality of groups of scaled transform coefficients includes the group of scaled transform coefficients for the block of video data; Determining the set of quantization offset parameters for the groups of scaled transform coefficients for the block of video data includes determining, for each of the plurality of groups of scaled transform coefficients, a corresponding set of quantization offset parameters; quantizing the groups of scaled transform coefficients for the block of video data includes quantizing each of the groups of scaled transform coefficients based on the corresponding set of quantization offset parameters. The method according to C1. [C10] Quantizing the group of scaled transform coefficients for the block of video data comprises: determining, for each scaled transform coefficient of the group of scaled transform coefficients, a corresponding quantization offset parameter from the set of quantization offset parameters; quantizing each scaled transform coefficient of the group of scaled transform coefficients based at least in part on the corresponding quantization offset parameter. [C11] The method of C1, wherein the side information includes one or more of a slice type for the block of video data, a block size for the block of video data, or an indication of whether the block of video data comprises a luminance component or a chrominance component. [C12] Determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data comprises: determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data without using one or more bit cost estimates determined via entropy coding; The method of claim C1, comprising: [C13] The method of C1, wherein the set of quantization offset parameters comprises a quantization offset vector. [C14] A device for encoding video data, comprising: Memory, a processing circuit in communication with the memory unit; The processing circuitry comprises: determining a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data based on side information associated with the block of video data; quantizing the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data based at least in part on the set of quantization offset parameters; generating an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data. Devices that are configured to: [C15] To determine the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data, the processing circuitry selecting a set of quantization offset parameters for the group of scaled transform coefficients from a plurality of sets of quantization offset parameters based on the side information associated with the block of video data. The device of C14, further configured as follows: [C16] To select the set of quantization offset parameters for the group of scaled transform coefficients from the plurality of sets of quantization offset parameters, the processing circuitry comprises: determining an index associated with the group of scaled transform coefficients based at least in part on the side information associated with the block of video data; indexing into a plurality of sets of quantization offset parameters to select the set of quantization offset parameters for the group of scaled transform coefficients using the index associated with the group of scaled transform coefficients. The device of C15, further configured as follows: [C17] The group of scaled transform coefficients comprises scaled transform coefficients for sub-blocks of the block of video data, the side information associated with the block of video data comprises a location of the sub-block within the block of video data and a block size of the block of video data. The device described in C15. [C18] To select the set of quantization offset parameters for the group of scaled transform coefficients from the plurality of sets of quantization offset parameters, the processing circuitry comprises: selecting the set of quantization offset parameters for the group of scaled transform coefficients from the multiple sets of quantization offset parameters without using scaled transform coefficient values ​​of the group of scaled transform coefficients for the sub-block of the block of video data. The device of C17, further configured as follows: [C19] To determine the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data, the processing circuitry determining a set of parameters for parameterization of the set of quantization offset parameters based on the side information associated with the block of video data, the set of parameters having a smaller size than the set of quantization offset parameters; determining the set of quantization offset parameters for the group of scaled transform coefficients based on the set of parameters having the smaller size; The device of C14, further configured to: [C20] To determine the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data, the processing circuitry determining, using a neural network, the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data based on the side information associated with the block of video data. The device of C14, further configured as follows: [C21] To determine the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data, the processing circuitry determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data using at least one of a classification technique or a regression technique; The device of C20, further configured as follows: [C22] The processing circuitry comprises: and further configured to divide the plurality of scaled transform coefficients of the block of video data into a plurality of groups of scaled transform coefficients associated with sub-blocks of the block of video data, the plurality of groups of scaled transform coefficients including the group of scaled transform coefficients for the block of video data; To determine the set of quantization offset parameters for the groups of scaled transform coefficients for the block of video data, the processing circuitry is further configured to determine, for each of the plurality of groups of scaled transform coefficients, a corresponding set of quantization offset parameters; quantizing the groups of scaled transform coefficients for the block of video data includes quantizing each of the groups of scaled transform coefficients based on the corresponding set of quantization offset parameters. The device described in C14. [C23] To quantize the group of scaled transform coefficients for the block of video data, the processing circuitry determining, for each scaled transform coefficient of the group of scaled transform coefficients, a corresponding quantization offset parameter from the set of quantization offset parameters; quantizing each scaled transform coefficient of the group of scaled transform coefficients based at least in part on the corresponding quantization offset parameter; The device of C14, further configured as follows: [C24] The device of C14, wherein the side information includes one or more of: a slice type for the block of video data, a block size for the block of video data, or an indication of whether the block of video data comprises a luminance component or a chrominance component. [C25] To determine the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data, the processing circuitry determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data without using one or more bit cost estimates determined via entropy coding. The device of C14, further configured as follows: [C26] The device of C14, wherein the set of quantization offset parameters comprises a quantization offset vector. [C27] The device of C14, wherein the device comprises one or more of a camera, a computer, or a mobile device. [C28] Apparatus for encoding video data, comprising: means for determining a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data based on side information associated with the block of video data; means for quantizing the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data based at least in part on the set of quantization offset parameters; means for generating an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data; An apparatus comprising: [C29] The means for determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data comprises: means for selecting a set of quantization offset parameters for the group of scaled transform coefficients from a plurality of sets of quantization offset parameters based on the side information associated with the block of video data; The apparatus of C28, further comprising: [C30] The means for selecting the set of quantization offset parameters for the group of scaled transform coefficients from the plurality of sets of quantization offset parameters comprises: means for determining an index associated with the group of scaled transform coefficients based at least in part on the side information associated with the block of video data; means for indexing into a plurality of sets of quantization offset parameters to select the set of quantization offset parameters for the group of scaled transform coefficients using the index associated with the group of scaled transform coefficients; The apparatus of C29, comprising: [C31] The group of scaled transform coefficients comprises scaled transform coefficients for sub-blocks of the block of video data, the side information associated with the block of video data comprises a location of the sub-block within the block of video data and a block size of the block of video data. The apparatus described in C29. [C32] The means for selecting the set of quantization offset parameters for the group of scaled transform coefficients from the plurality of sets of quantization offset parameters comprises: means for selecting the set of quantization offset parameters for the group of scaled transform coefficients from the plurality of sets of quantization offset parameters without using scaled transform coefficient values ​​of the group of scaled transform coefficients for the sub-block of the block of video data; The apparatus of C31, further comprising: [C33] The means for determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data comprises: means for determining a set of parameters for parameterization of the set of quantization offset parameters based on the side information associated with the block of video data, the set of parameters having a smaller size than the set of quantization offset parameters; means for determining the set of quantization offset parameters for the group of scaled transform coefficients based on the set of parameters having the smaller size; The apparatus of C28, further comprising: [C34] The means for determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data comprises: means for determining, using a neural network, the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data based on the side information associated with the block of video data; The apparatus of C28, further comprising: [C35] The means for determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data comprises: means for determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data using at least one of a classification technique or a regression technique; The apparatus described in C34, comprising: [C36] means for dividing a plurality of scaled transform coefficients of said block of video data into a plurality of groups of scaled transform coefficients associated with sub-blocks of said block of video data; and wherein the plurality of groups of scaled transform coefficients includes the group of scaled transform coefficients for the block of video data; the means for determining the set of quantization offset parameters for the groups of scaled transform coefficients for the block of video data comprises means for determining, for each of the plurality of groups of scaled transform coefficients, a corresponding set of quantization offset parameters; the means for quantizing the groups of scaled transform coefficients for the block of video data further comprises means for quantizing each of the plurality of groups of scaled transform coefficients based on the corresponding set of quantization offset parameters. The apparatus described in C28. [C37] The means for quantizing the group of scaled transform coefficients for the block of video data comprises: means for determining, for each scaled transform coefficient of the group of scaled transform coefficients, a corresponding quantization offset parameter from the set of quantization offset parameters; means for quantizing each scaled transform coefficient of the group of scaled transform coefficients based at least in part on the corresponding quantization offset parameter; The apparatus of C28, further comprising: [C38] The apparatus of C28, wherein the side information includes one or more of a slice type for the block of video data, a block size for the block of video data, or an indication of whether the block of video data comprises a luminance component or a chrominance component. [C39] The means for determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data comprises: means for determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data without using one or more bit cost estimates determined via entropy coding; The apparatus of C28, comprising: [C40] The apparatus of C28, wherein the set of quantization offset parameters comprises a quantization offset vector. [C41] A computer-readable storage medium having stored thereon instructions that, when executed, cause one or more processors to: determining a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data based on side information associated with the block of video data; quantizing the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data based at least in part on the set of quantization offset parameters; and generating an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data. A computer-readable storage medium for causing a computer to perform the above steps. [C42] The instructions for causing the one or more processors to determine the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data may further include causing the one or more processors to: selecting a set of quantization offset parameters for the group of scaled transform coefficients from a plurality of sets of quantization offset parameters based on the side information associated with the block of video data; 20. The computer-readable storage medium of claim 19, comprising instructions for causing a computer to: [C43] The instructions for causing the one or more processors to select the set of quantization offset parameters for the group of scaled transform coefficients from the multiple sets of quantization offset parameters may further comprise causing the one or more processors to: determining an index for the group of scaled transform coefficients based at least in part on the side information associated with the block of video data; indexing into a plurality of sets of quantization offset parameters to select the set of quantization offset parameters for the group of scaled transform coefficients using the index; 20. A computer-readable storage medium as described in C42, comprising instructions for causing a computer to: [C44] The group of scaled transform coefficients comprises scaled transform coefficients for sub-blocks of the block of video data, the side information associated with the block of video data comprises a location of the sub-block within the block of video data and a block size of the block of video data. A computer-readable storage medium as described in C42. [C45] The instructions for causing the one or more processors to select the set of quantization offset parameters for the group of scaled transform coefficients from the multiple sets of quantization offset parameters may further include causing the one or more processors to: selecting a set of quantization offset parameters for the group of scaled transform coefficients from the multiple sets of quantization offset parameters without using scaled transform coefficient values ​​of the group of scaled transform coefficients for the sub-block of the block of video data; 20. A computer readable storage medium as set forth in claim 19, comprising instructions for causing a computer to: [C46] The instructions for causing the one or more processors to determine the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data may further include causing the one or more processors to: determining a set of parameters for parameterization of the set of quantization offset parameters based on the side information associated with the block of video data, the set of parameters having a smaller size than the set of quantization offset parameters; determining the set of quantization offset parameters for a group of scaled transform coefficients based on the set of parameters having the smaller size; 20. The computer-readable storage medium of claim 19, comprising instructions for causing a computer to: [C47] The instructions for causing the one or more processors to determine the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data may further include causing the one or more processors to: 5. The computer-readable storage medium of claim 41, comprising instructions for: determining, using a neural network, the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data based on the side information associated with the block of video data. [C48] The instructions for causing the one or more processors to determine the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data may further include causing the one or more processors to: determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data using at least one of a classification technique or a regression technique; 20. A computer readable storage medium as set forth in claim 19, comprising instructions for causing a computer to: [C49] The instructions cause the one or more processors to: dividing a plurality of scaled transform coefficients of the block of video data into a plurality of groups of scaled transform coefficients associated with sub-blocks of the block of video data; wherein the plurality of groups of scaled transform coefficients includes the group of scaled transform coefficients for the block of video data; The instructions to cause the one or more processors to determine the set of quantization offset parameters for the groups of scaled transform coefficients for the block of video data comprise instructions to cause the one or more processors to determine, for each of the multiple groups of scaled transform coefficients, a corresponding set of quantization offset parameters; The computer-readable storage medium of C41, wherein the instructions to cause the one or more processors to quantize the groups of scaled transform coefficients for the block of video data comprise instructions to cause the one or more processors to quantize each of the multiple groups of scaled transform coefficients based on the corresponding set of quantization offset parameters. [C50] The instructions to cause the one or more processors to quantize the group of scaled transform coefficients for the block of video data may further include instructions to the one or more processors to: determining, for each scaled transform coefficient of the group of scaled transform coefficients, a corresponding quantization offset parameter from the set of quantization offset parameters; quantizing each scaled transform coefficient of the group of scaled transform coefficients based at least in part on the corresponding quantization offset parameter. [C51] The computer-readable storage medium of C41, wherein the side information includes one or more of a slice type of the block of video data, a block size of the block of video data, or an indication of whether the block of video data comprises a luminance component or a chrominance component. [C52] The instructions for causing the one or more processors to determine the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data may further include causing the one or more processors to: determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data without using one or more bit cost estimates determined via entropy coding; 20. The computer-readable storage medium of claim 19, comprising instructions for causing a computer to: [C53] The computer-readable storage medium of C41, wherein the set of quantization offset parameters comprises a quantization offset vector.

Claims

1. 1. A method for encoding video data, comprising the steps of: determining a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data based on side information associated with the block of video data, where the block of video data comprises a plurality of scaled transform coefficients, the plurality of scaled transform coefficients being transform coefficients of the block scaled by a scaling factor, the group of scaled transform coefficients being a subset of the plurality of scaled transform coefficients of the block, and the side information includes one or more of a slice type of the block of video data, a block size of the block of video data, or an indication of whether the block of video data comprises a luminance component or a chrominance component. quantizing the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data by summing each quantization offset parameter of the set of quantization offset parameters with a respective scaled transform coefficient of the group of scaled transform coefficients and truncating each resulting sum to an integer value; generating an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data; A method for providing the above.

2. Determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data comprises: selecting a set of quantization offset parameters for the group of scaled transform coefficients from a plurality of sets of quantization offset parameters based on the side information associated with the block of video data; The method of claim 1 , comprising:

3. Selecting the set of quantization offset parameters for the group of scaled transform coefficients from the plurality of sets of quantization offset parameters comprises: determining an index associated with the group of scaled transform coefficients based at least in part on the side information associated with the block of video data; indexing into the plurality of sets of quantization offset parameters to select the set of quantization offset parameters for the group of scaled transform coefficients using the index associated with the group of scaled transform coefficients; The method of claim 2 comprising:

4. the group of scaled transform coefficients comprises scaled transform coefficients for sub-blocks of the block of video data; the side information associated with the block of video data comprises a location of the sub-block within the block of video data and a block size of the block of video data. The method of claim 2.

5. Selecting the set of quantization offset parameters for the group of scaled transform coefficients from the plurality of sets of quantization offset parameters comprises: selecting a set of quantization offset parameters for the group of scaled transform coefficients from the multiple sets of quantization offset parameters without using scaled transform coefficient values ​​of the group of scaled transform coefficients for the sub-block of the block of video data; The method of claim 4 comprising:

6. Determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data comprises: determining a set of parameters for a parameterization of the set of quantization offset parameters based on the side information associated with the block of video data, the set of parameters having a smaller size than the set of quantization offset parameters, the set of parameters for the parameterization of the set of quantization offset parameters having a value p 0 and p 1 is a two-dimensional parameter vector having the following structure: determining the set of quantization offset parameters for the group of scaled transform coefficients based on the set of parameters having a size smaller than the set of quantization offset parameters, where the determined set of quantization offset parameters includes elements v n is a P-dimensional vector with n The value of is expressed by the formula [0010] is given by, The method of claim 1 , comprising:

7. Determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data comprises: determining, using a neural network, the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data based on the side information associated with the block of video data; determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data using at least one of a classification technique or a regression technique; The method of claim 1 further comprising:

8. the group of scaled transform coefficients is one of a plurality of groups of scaled transform coefficients each associated with a respective sub-block of the block of video data; a corresponding set of quantization offset parameters is determined for each of the plurality of groups of scaled transform coefficients; quantizing the groups of scaled transform coefficients for the block of video data includes quantizing each of the plurality of groups of scaled transform coefficients based on the corresponding set of quantization offset parameters determined for the group of scaled transform coefficients. The method of claim 1.

9. Determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data comprises: determining the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data without using one or more bit cost estimates determined via entropy coding; The method of claim 1 , comprising:

10. 1. A device for encoding video data, comprising: Memory, a processing circuit in communication with the memory; The processing circuitry comprises: determining a set of quantization offset parameters for a group of scaled transform coefficients for a block of video data based on side information associated with the block of video data, where the block of video data comprises a plurality of scaled transform coefficients, the plurality of scaled transform coefficients being transform coefficients of the block scaled by a scaling factor, the group of scaled transform coefficients being a subset of the plurality of scaled transform coefficients of the block, and the side information includes one or more of a slice type of the block of video data, a block size of the block of video data, or an indication of whether the block of video data comprises a luminance component or a chrominance component. quantizing the group of scaled transform coefficients for the block of video data to generate quantized transform coefficients for the block of video data by summing each quantization offset parameter of the set of quantization offset parameters with a respective scaled transform coefficient of the group of scaled transform coefficients and truncating each resulting sum to an integer value; generating an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data. Devices that are configured to:

11. To determine the set of quantization offset parameters for the group of scaled transform coefficients for the block of video data, the processing circuitry comprises: selecting a set of quantization offset parameters for the group of scaled transform coefficients from a plurality of sets of quantization offset parameters based on the side information associated with the block of video data. The device of claim 10 further configured to:

12. To select the set of quantization offset parameters for the group of scaled transform coefficients from the plurality of sets of quantization offset parameters, the processing circuitry comprises: determining an index associated with the group of scaled transform coefficients based at least in part on the side information associated with the block of video data; indexing into a plurality of sets of quantization offset parameters to select the set of quantization offset parameters for the group of scaled transform coefficients using the index associated with the group of scaled transform coefficients. The device of claim 11 further configured to:

13. the group of scaled transform coefficients comprises scaled transform coefficients for sub-blocks of the block of video data; the side information associated with the block of video data comprises a location of the sub-block within the block of video data and a block size of the block of video data. The device of claim 11.

14. The device of claim 13 , wherein the device comprises one or more of a camera, a computer, or a mobile device.

15. 10. A computer readable storage medium having stored thereon instructions which, when executed by one or more processors, cause the one or more processors to perform the method of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Accelerated techniques for rate-distortion optimized quantization

    JP2013504934A

  • Chroma quantization in video coding

    JP2015053680A

  • Systems and methods for quantization parameter-based video processing

    JP2019512938A

  • Image coding device, image decoding device, image coding method, and image decoding method

    WO2011064926A1

  • System and method for video processing based on quantization parameter

    WO2017155786A1