Low complexity adaptive quantization for video compression

By determining the set of quantization offset parameters based on side information in the video encoder, the quantization process of the transform coefficients is simplified, solving the problem of computational complexity in modern video encoders and achieving video compression with low computational complexity.

CN115053524BActive Publication Date: 2026-04-28QUALCOMM INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QUALCOMM INC
Filing Date
2021-02-04
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

The entropy decoding process of modern video encoders is complex, especially when encoding transform coefficients in multiple paths. The computation is complex and expensive, resulting in video encoders requiring a large number of processing cycles when quantizing transform coefficients.

Method used

By avoiding the use of bit cost estimation, the video encoder determines the set of quantization offset parameters based on the side information of the video data block, quantizes the transform coefficient group, simplifies the quantization process, and allows for parallel processing of transform coefficients.

Benefits of technology

This reduces the computational complexity of the video encoder, shortens the processing cycle, and improves the efficiency of the quantization transform coefficients, thus achieving video compression with lower computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115053524B_ABST
    Figure CN115053524B_ABST
Patent Text Reader

Abstract

A video encoder can determine a set of quantization offset parameters for a set of scaled transform coefficients for a block of video data based on side information associated with the block of video data. The video encoder can also quantize the set of scaled transform coefficients for the block of video data based at least in part on the set of quantization offset parameters to generate quantized transform coefficients for the block of video data. The video encoder can generate an encoded video bitstream based at least in part on the quantized transform coefficients for the block of video data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to U.S. Application No. 17 / 166,639, filed February 3, 2021, and U.S. Provisional Application No. 62 / 970,588, filed February 5, 2020, the entire contents of each of which are incorporated herein by reference. U.S. Application No. 17 / 166,639, filed February 3, 2021, claims the benefit of U.S. Provisional Application No. 62 / 970,588, filed February 5, 2020. Technical Field

[0002] This disclosure relates to video encoding and video decoding. Background Technology

[0003] Digital video capabilities can be incorporated into a wide variety of devices, including digital televisions, digital live broadcast systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, tablets, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radio phones (so-called "smartphones"), video conferencing equipment, video streaming devices, and more. Digital video devices implement video decoding technologies (such as those described in standards defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 (Part 10, Advanced Video Decoding (AVC)), ITU-T H.265 / High Efficiency Video Decoding (HEVC), and extensions to such standards). By implementing such video decoding technologies, video devices can more efficiently send, receive, encode, decode, and / or store digital video information.

[0004] Video decoding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in video sequences. For block-based video decoding, video slices (e.g., video pictures or portions of video pictures) can be segmented into video blocks, which may also be referred to as decoding tree units (CTUs), decoding units (CUs), and / or decoding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction relative to reference samples in adjacent blocks within the same picture. Video blocks in an inter-coded (P or B) slice of a picture can use spatial prediction relative to reference samples in adjacent blocks within the same picture or temporal prediction relative to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. Summary of the Invention

[0005] In summary, this disclosure describes a technique for adaptively quantizing transform coefficients used to encode video data by determining a quantization offset for quantizing the transform coefficients. Specifically, instead of quantizing the transform coefficients based on bit cost estimation to achieve optimal quantization of transform coefficients determined by entropy decoding of the quantized transform coefficients, the video encoder can determine a set of quantization offset parameters for a group of transform coefficients in a video data block based on side information associated with the video data block, and can quantize the group of transform coefficients based on the set of quantization offset parameters to produce near-optimal quantized transform coefficients.

[0006] According to the technology of this disclosure, a video encoder can decompose the transform coefficients of a video data block into transform coefficient groups. For each transform coefficient group, the video encoder can determine a set of quantization offset parameters associated with the transform coefficient group based on side information of the video data block (such as the slice type of the video data block) and / or indications about whether the video data block includes luma or chroma components. Therefore, the video encoder can quantize each transform coefficient group based on the quantization offset parameters associated with the transform coefficient group, where performance is improved based on the selection of the optimal offset.

[0007] The technical problem addressed by the technology disclosed herein involves the fact that entropy decoding performed by modern video encoders can be highly complex, with numerous arithmetic decoding contexts and intricate context selection rules. Furthermore, modern video encoders can encode transform coefficients in multiple paths. For example, entropy decoding in a modern video encoder can be performed in up to five paths. This makes calculating and using bit cost estimation to optimally quantize the transform coefficients can be complex and computationally expensive for each decision.

[0008] Conversely, by avoiding the use of bit cost estimation to quantize the transform coefficients determined during entropy decoding of the quantized transform coefficients for further quantization, the techniques of this disclosure improve compression with lower computational complexity for transform coefficient quantization, thereby enabling the video encoder to quantize the transform coefficients using fewer processing cycles. Furthermore, because the techniques of this disclosure can determine a single set of quantization offsets for quantizing each transform coefficient within a transform coefficient group, the techniques of this disclosure enable the video encoder to quantize transform coefficients within a transform coefficient group in parallel (e.g., the video encoder can quantize transform coefficients in a second transform coefficient group simultaneously with or temporally overlap with the quantization of transform coefficients in a first transform coefficient group).

[0009] A system of one or more computers can be configured to perform specific operations or actions because software, firmware, hardware, or combinations thereof are installed on the system to operate and cause the system to perform actions. One or more computer programs can be configured to perform specific operations or actions because they include instructions that, when executed by a data processing device, cause the device to perform actions.

[0010] One general aspect includes a method for encoding video data. The method includes: determining a set of quantization offset parameters for a scaled transform coefficient group for the video data block based on side information associated with the video data block. The method further includes: quantizing the scaled transform coefficient group for the video data block at least partially based on the set of quantization offset parameters to generate quantized transform coefficients for the video data block. The method further includes: generating an encoded video bitstream at least partially based on the quantized transform coefficients for the video data block.

[0011] One general aspect includes an apparatus for encoding video data. The apparatus includes a memory. The apparatus also includes processing circuitry in communication with the memory, the processing circuitry being configured to: determine a set of quantization offset parameters for a scaled transform coefficient set for the video data block based on side information associated with the video data block; quantize the scaled transform coefficient set for the video data block at least in part based on the quantization offset parameter set to generate quantized transform coefficients for the video data block; and generate an encoded video bitstream at least in part based on the quantized transform coefficients for the video data block.

[0012] One general aspect includes an apparatus for decoding video data. The apparatus includes: unit for determining a set of quantization offset parameters for a scaled transform coefficient group of the video data block based on side information associated with the video data block. The apparatus further includes: unit for quantizing the scaled transform coefficient group of the video data block at least partially based on the set of quantization offset parameters to generate quantized transform coefficients of the video data block. The apparatus further includes: unit for generating an encoded video bitstream at least partially based on the quantized transform coefficients of the video data block.

[0013] One general aspect includes a computer-readable storage medium having instructions stored thereon, which, when executed by one or more processors, cause the one or more processors to: determine a set of quantization offset parameters for a scaled transform coefficient set for the video data block based on side information associated with the video data block; quantize the scaled transform coefficient set for the video data block at least in part based on the set of quantization offset parameters to generate quantized transform coefficients for the video data block; and generate an encoded video bitstream at least in part based on the quantized transform coefficients for the video data block.

[0014] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description, drawings, and claims. Attached Figure Description

[0015] Figure 1 This is a block diagram illustrating an example video encoding and decoding system that can perform the techniques described in this disclosure.

[0016] Figure 2A and Figure 2B This is a conceptual diagram showing an example quadtree binary tree (QTBT) structure and its corresponding decoding tree unit (CTU).

[0017] Figure 3A A video decoding system for adaptive and / or rate-distortion optimized quantization is shown.

[0018] Figure 3B A video decoding system for low-complexity adaptive quantization based on block classification and one or more sets of quantization offset parameters, according to the techniques of this disclosure, is shown.

[0019] Figure 4 A parallel implementation of adaptive quantization using a set of quantization offset parameters, based on the techniques described herein, is illustrated.

[0020] Figure 5 The factors for symbol bit hiding used to approximate rate-distortion cost are shown.

[0021] Figure 6 The present disclosure illustrates a technique for determining quantization offset parameters.

[0022] Figure 7 An example of a code used to identify the location of a sub-block within a video data block is shown.

[0023] Figure 8 This is a block diagram illustrating an example video encoder that can perform the techniques described in this disclosure.

[0024] Figure 9 This is a block diagram illustrating an example video decoder that can perform the techniques described in this disclosure.

[0025] Figure 10 This is a flowchart illustrating an example method for encoding the current block.

[0026] Figure 11 This is a flowchart illustrating an example method for decoding the current block of video data.

[0027] Figure 12 This is a flowchart illustrating a method for encoding video data according to the technology described in this disclosure. Detailed Implementation

[0028] Figure 1 This is a block diagram illustrating an example video encoding and decoding system 100 capable of performing the techniques of this disclosure. In summary, the techniques of this disclosure relate to decoding (encoding and / or decoding) video data. Typically, video data includes any data used for processing video. Therefore, video data can include raw, unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata (e.g., signaling data).

[0029] like Figure 1 As shown, in this example, system 100 includes source device 102, which provides encoded video data to be decoded and displayed by destination device 116. Specifically, source device 102 provides the video data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 can include any of a wide variety of devices, including desktop computers, laptop computers, mobile devices, tablet computers, set-top boxes, mobile phones (such as smartphones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, broadcast receivers, etc. In some cases, source device 102 and destination device 116 may be equipped for wireless communication and may therefore be referred to as wireless communication devices.

[0030] exist Figure 1In the example, source device 102 includes a video source 104, a memory 106, a video encoder 200, and an output interface 108. Destination device 116 includes an input interface 122, a video decoder 300, a memory 120, and a display device 118. According to this disclosure, the video encoder 200 of source device 102 and the video decoder 300 of destination device 116 can be configured to apply techniques for quantizing variation coefficients. Therefore, source device 102 represents an example of a video encoding device, while destination device 116 represents an example of a video decoding device. In other examples, the source and destination devices may include other components or arrangements. For example, source device 102 may receive video data from an external video source such as an external camera. Similarly, destination device 116 may interface with an external display device, rather than including an integrated display device.

[0031] exist Figure 1 The system 100 shown is merely an example. Typically, any digital video encoding and / or decoding device can perform techniques for quantizing variation coefficients. Source device 102 and destination device 116 are merely examples of such decoding devices, where source device 102 generates decoded video data for transmission to destination device 116. In this disclosure, "decoding device" refers to a device that performs the decoding (e.g., encoding and / or decoding) of data. Thus, video encoder 200 and video decoder 300 represent examples of decoding devices (specifically, video encoder and video decoder). In some examples, source device 102 and destination device 116 may operate in a substantially symmetrical manner, such that each of source device 102 and destination device 116 includes video encoding and decoding components. Therefore, system 100 can support one-way or two-way video transmission between source device 102 and destination device 116, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0032] Typically, video source 104 represents the source of video data (i.e., raw, unencoded video data) and provides a sequential series of pictures (also referred to as "frames") of the video data to video encoder 200, which encodes the data used for the pictures. Video source 104 of source device 102 may include video capture devices such as cameras, video archive units containing previously captured raw video, and / or video feed interfaces for receiving video from video content providers. Alternatively, video source 104 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In each case, video encoder 200 may encode the captured, pre-captured, or computer-generated video data. Video encoder 200 may rearrange the pictures from their received order (sometimes referred to as "display order") to a decoding order for decoding. Video encoder 200 may generate a bitstream comprising encoded video data. Then, the source device 102 can output the encoded video data to the computer-readable medium 110 via the output interface 108 so that it can be received and / or retrieved by, for example, the input interface 122 of the destination device 116.

[0033] The memory 106 of source device 102 and the memory 120 of destination device 116 represent general-purpose memory. In some examples, memories 106 and 120 may store raw video data, such as raw video from video source 104 and raw decoded video data from video decoder 300. Alternatively, memories 106 and 120 may store software instructions executable by, for example, video encoder 200 and video decoder 300 respectively. Although memories 106 and 120 are shown separately from video encoder 200 and video decoder 300 in this example, it should be understood that video encoder 200 and video decoder 300 may also include internal memory for functionally similar or equivalent purposes. Furthermore, memories 106 and 120 may store, for example, encoded video data output from video encoder 200 and input to video decoder 300. In some examples, portions of memories 106 and 120 may be allocated as one or more video buffers, for example, to store raw, decoded, and / or encoded video data.

[0034] Computer-readable medium 110 can represent any type of medium or device capable of transmitting encoded video data from source device 102 to destination device 116. In one example, computer-readable medium 110 represents a communication medium that enables source device 102 to directly transmit encoded video data to destination device 116 in real time, for example, via a radio frequency network or a computer-based network. Output interface 108 can modulate the transmitted signal including encoded video data according to a communication standard such as a wireless communication protocol, and input interface 122 can demodulate the received transmitted information according to a communication standard such as a wireless communication protocol. The communication medium can include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include a router, switch, base station, or any other device that may be useful for facilitating communication from source device 102 to destination device 116.

[0035] In some examples, source device 102 can output encoded data from output interface 108 to storage device 112. Similarly, destination device 116 can access encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a variety of distributed or locally accessible data storage media, such as hard disk drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.

[0036] In some examples, source device 102 may output encoded video data to file server 114 or to another intermediate storage device that can store the encoded video data generated by source device 102. Destination device 116 may access the stored video data from file server 114 via streaming or downloading. File server 114 may be any type of server device capable of storing encoded video data and sending such encoded video data to destination device 116. File server 114 may represent a web server (e.g., for a website), a file transfer protocol (FTP) server, a content delivery network device, or a network attached storage (NAS) device. Destination device 116 may access the encoded video data from file server 114 via any standard data connection (including an Internet connection). This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both, suitable for accessing the encoded video data stored on file server 114. File server 114 and input interface 122 may be configured to operate according to a streaming protocol, a download protocol, or a combination thereof.

[0037] Output interface 108 and input interface 122 can represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 can be configured to transmit data (such as encoded video data) according to cellular communication standards (such as 4G, 4G-LTE (Long Term Evolution), improved LTE, 5G, etc.). In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 can be configured to operate according to other wireless standards (such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee)). TM Bluetooth TM The source device 102 and / or destination device 116 may include corresponding system-on-chip (SoC) devices. For example, source device 102 may include an SoC device for performing the functions assigned to video encoder 200 and / or output interface 108, and destination device 116 may include an SoC device for performing the functions assigned to video decoder 300 and / or input interface 122.

[0038] The technology disclosed herein can be applied to video decoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission (such as HTTP-based Dynamic Adaptive Streaming (DASH)), digital video encoded onto data storage media, decoding digital video stored on data storage media, or other applications.

[0039] The input interface 122 of the destination device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., a communication medium, storage device 112, file server 114, etc.). The encoded video bitstream may include signaling information defined by the video encoder 200, such as syntax elements (also used by the video decoder 300), which have values ​​describing the characteristics and / or processing of video blocks or other decoding units (e.g., slices, tiles, bricks, pictures, picture groups, sequences, etc.). The display device 118 displays a decoded image of the decoded video data to the user. The display device 118 may represent any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.

[0040] Despite Figure 1 Not shown, but in some examples, the video encoder 200 and video decoder 300 may each be integrated with the audio encoder and / or audio decoder, and may include appropriate MUX-DEMUX units or other hardware and / or software to process the multiplexed streams of both audio and video included in a common data stream. Where applicable, the MUX-DEMUX unit may conform to the ITU H.223 multiplexer protocol or other protocols (such as User Datagram Protocol (UDP)).

[0041] The video encoder 200 and video decoder 300 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is implemented in part in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the technology of this disclosure. Each of the video encoder 200 and video decoder 300 may be included in one or more encoders or decoders, and either encoder or decoder may be integrated as part of a combined encoder / decoder (CODEC) in the respective device. Devices including the video encoder 200 and / or video decoder 300 may include integrated circuits, microprocessors, and / or wireless communication devices (such as cellular phones).

[0042] The video encoder 200 and video decoder 300 can operate according to video decoding standards such as ITU-T H.265 (also known as the High Efficiency Video Coding (HEVC) standard) or extensions thereof such as MultiView or Scalable Video Coding Extensions. Alternatively, the video encoder 200 and video decoder 300 can operate according to other proprietary or industry standards such as ITU-T H.266 (also known as Versatile Video Coding (VVC)). A recent draft of the VVC standard is described in: Bross et al., “Versatile Video Coding (Draft 10)”, ITU-T SG 16WP 3 and ISO / IEC JTC 1 / SC 29 / WG 11 Joint Video Experts Group (JVET), 18th meeting by teleconference, 22 June–1 July 2020, JVET-S2001-v17 (hereinafter referred to as “VVC Draft 10”), available for review. https: / / jvet- experts.org / However, the technology disclosed herein is not limited to any particular decoding standard.

[0043] Typically, video encoder 200 and video decoder 300 can perform block-based decoding of images. The term "block" generally refers to a structure that includes data to be processed (e.g., encoded, decoded, or otherwise used during encoding and / or decoding). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. Typically, video encoder 200 and video decoder 300 can decode video data represented in YUV (e.g., Y, Cb, Cr) format. That is, instead of decoding the red, green, and blue (RGB) data used for images, video encoder 200 and video decoder 300 can decode both luminance and chrominance components, where chrominance components may include both red hue and blue hue chrominance components. In some examples, video encoder 200 converts the received RGB-formatted data to a YUV representation before encoding, and video decoder 300 converts the YUV representation to RGB format. Alternatively, preprocessing and post-processing units (not shown) can perform these conversions. However, the technology disclosed herein is not limited to YUV representation, but can be applied to RGB representation with three color components.

[0044] In summary, this disclosure may relate to the decoding (e.g., encoding and decoding) of images to include the process of encoding or decoding the data of an image. Similarly, this disclosure may relate to the decoding of blocks of an image to include the process of encoding or decoding the data used for the blocks (e.g., prediction and / or residual decoding). Encoded video bitstreams typically include a series of values ​​for representing decoding decisions (e.g., decoding modes) and syntax elements that segment the image into blocks. Therefore, references to decoding images or blocks should generally be understood as decoding the values ​​of the syntax elements used to form images or blocks.

[0045] HEVC defines various blocks, including decoding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video decoder (such as a video encoder 200) partitions the decoding tree unit (CTU) into CUs according to a quadtree structure. That is, the video decoding device partitions the CTU and CU into four equal, non-overlapping squares, and each node of the quadtree has zero or four child nodes. Nodes without child nodes can be called "leaf nodes," and the CU of such leaf nodes can include one or more PUs and / or one or more TUs. The video decoding device can further partition the PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents a partition of a TU. In HEVC, PUs represent inter-frame prediction data, while TUs represent residual data. CUs with intra-frame prediction include intra-frame prediction information, such as intra-frame mode indication.

[0046] As another example, video encoder 200 and video decoder 300 can be configured to operate according to VVC. According to VVC, the video decoder (such as video encoder 200) segments the image into multiple decoding tree units (CTUs). Video encoder 200 can segment CTUs according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure eliminates the concept of multiple segmentation types, such as the distinction between CU, PU, ​​and TU in HEVC. The QTBT structure includes two levels: a first level based on quadtree segmentation and a second level based on binary tree segmentation. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to decoding units (CUs).

[0047] In the MTT partitioning structure, blocks can be partitioned using quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) partitioning (also known as triplet tree (TT)) partitioning. A ternary tree or triplet tree partitioning is a partition in which a block is split into three sub-blocks. In some examples, a ternary tree or triplet tree partitioning divides a block into three sub-blocks without splitting the original block by a center. The partitioning types in MTT (e.g., QT, BT, and TT) can be symmetric or asymmetric.

[0048] In some examples, the video encoder 200 and the video decoder 300 may use a single QTBT or MTT structure to represent each of the luma and chroma components, while in other examples, the video encoder 200 and the video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luma component and another QTBT / MTT structure for the two chroma components (or two QTBT / MTT structures for the respective chroma components).

[0049] The video encoder 200 and video decoder 300 can be configured to use quadtree segmentation, QTBT segmentation, MTT segmentation, or other segmentation structures based on HEVC. For illustrative purposes, a description of the techniques of this disclosure is given with respect to QTBT segmentation. However, it should be understood that the techniques of this disclosure can also be applied to video decoding apparatuses configured to use quadtree segmentation or other types of segmentation.

[0050] In some examples, a CTU includes a decoded tree block (CTB) of luminance samples, two corresponding CTBs of chrominance samples of an image with three sample arrays, or a CTB of samples of a monochrome image or an image decoded using three separate color planes and a syntax structure for decoding the samples. A CTB can be an N×N block of samples (for some value of N) such that dividing a component into a CTB is a partition. A component is an array or a single sample of one of the three arrays (one luminance and two chrominance) that make up an image in a 4:2:0, 4:2:2, or 4:4:4 color format, or an array or a single sample of an array or array that makes up an image in monochrome format. In some examples, a decoded block is an M×N block of samples (for some values ​​of M and N) such that dividing a CTB into a decoded block is a partition.

[0051] Blocks (e.g., CTUs or CUs) can be grouped in various ways within an image. As an example, a brick can refer to a rectangular area of ​​a CTU row within a specific tile in an image. A tile can be a rectangular area of ​​a CTU within a specific tile column and a specific tile row in an image. A tile column refers to a rectangular area of ​​a CTU with a height equal to the height of the image and a width specified by syntax elements (e.g., as in an image parameter set). A tile row refers to a rectangular area of ​​a CTU with a height specified by syntax elements (e.g., as in an image parameter set) and a width equal to the width of the image.

[0052] In some examples, a tile can be divided into multiple bricks, each brick potentially including one or more CTU rows within the tile. A tile that is not divided into multiple bricks can also be referred to as a brick. However, bricks that are a true subset of a tile may not be referred to as tiles.

[0053] The bricks in an image can also be arranged as slices. A slice can be an integer number of bricks in the image, which can be uniquely contained within a single Network Abstraction Layer (NAL) unit. In some examples, a slice consists of multiple complete tiles or a continuous sequence of complete bricks that contain only one tile.

[0054] This disclosure uses "NxN" and "N by N" interchangeably to refer to the sample size of a block (such as a CU or other video block) in the vertical and horizontal dimensions, for example, 16x16 samples or 16 by 16 samples. Typically, a 16x16 CU will have 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an NxNCU typically has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. Samples in a CU can be arranged in rows and columns. Furthermore, a CU does not necessarily need to have the same number of samples in the horizontal direction as it does in the vertical direction. For example, a CU can include NxM samples, where M is not necessarily equal to N.

[0055] The video encoder 200 encodes video data for use in predicting and / or residual information, as well as other information, for the CU. The prediction information indicates how the CU will be predicted to form a prediction block for the CU. The residual information typically represents the sample-by-sample difference between a sample of the CU before encoding and the prediction block.

[0056] To predict the Cubic Frame (CU), the video encoder 200 typically forms prediction blocks for the CU using either inter-frame prediction or intra-frame prediction. Inter-frame prediction generally refers to predicting the CU based on data from previously decoded images, while intra-frame prediction generally refers to predicting the CU based on data from previously decoded images of the same frame. To perform inter-frame prediction, the video encoder 200 can generate prediction blocks using one or more motion vectors. The video encoder 200 can typically perform a motion search to identify, for example, a reference block that closely matches the CU in terms of the difference between the CU and a reference block. The video encoder 200 can calculate a difference metric using the sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to determine whether the reference block closely matches the current CU. In some examples, the video encoder 200 can use unidirectional or bidirectional prediction to predict the current CU.

[0057] Some examples of VVC also provide an affine motion compensation mode, which can be considered an inter-frame prediction mode. In affine motion compensation mode, the video encoder 200 can determine two or more motion vectors representing non-translational motion (such as zooming in or out, rotation, perspective motion, or other irregular types of motion).

[0058] To perform intra-frame prediction, the video encoder 200 can select an intra-frame prediction mode to generate prediction blocks. Some examples of VVC provide sixty-seven intra-frame prediction modes, including various directional modes, as well as planar and DC modes. Typically, the video encoder 200 selects an intra-frame prediction mode that describes the samples of the current block (e.g., a block of a CU) to be predicted based on, which are the neighboring samples of the current block. Assuming that the video encoder 200 encodes the CTU and CU in raster scan order (from left to right, from top to bottom), such samples can typically be located above, to the upper left, or to the left of the current block within the same image.

[0059] The video encoder 200 encodes data representing the prediction mode used for the current block. For example, for inter-frame prediction modes, the video encoder 200 may encode data indicating which of the various available inter-frame prediction modes is used, as well as motion information for the corresponding mode. For unidirectional or bidirectional inter-frame prediction, for example, the video encoder 200 may use Advanced Motion Vector Prediction (AMVP) or merging modes to encode motion vectors. The video encoder 200 may use similar modes to encode motion vectors used for affine motion compensation modes.

[0060] Following a prediction, such as intra-frame or inter-frame prediction of a block, the video encoder 200 can compute residual data for that block. The residual data (such as a residual block) represents the sample-by-sample difference between the block and the prediction block used to form the block, which is formed using the corresponding prediction mode. The video encoder 200 can apply one or more transforms to the residual block to produce blocks of transformed data (such as transform blocks (TBs) or transform coefficient blocks) in the transform domain rather than the sample domain. For example, the video encoder 200 can apply discrete cosine transform (DCT), integer transform, wavelet transform, or conceptually similar transforms to the residual video data. Additionally, the video encoder 200 can apply a second transform after the first transform, such as mode-dependent inseparable quadratic transform (MDNSST), signal-dependent transform, Karhunen-Loeve transform (KLT), etc. The video encoder 200 produces transform coefficients after applying one or more transforms.

[0061] As described above, after any transformation to produce transform coefficients, the video encoder 200 can perform quantization of the transform coefficients. Quantization generally refers to the process of quantizing the transform coefficients to reduce the amount of data used to represent them, thereby providing further compression. By performing the quantization process, the video encoder 200 can reduce the bit depth associated with some or all of the transform coefficients. For example, the video encoder 200 can round an n-bit value down to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder 200 can perform a bitwise right shift of the values ​​to be quantized.

[0062] According to various aspects of this disclosure, the video encoder 200 can decompose the transform coefficients of a block of pixels or samples (such as a transform block) into transform coefficient groups. For each transform coefficient group, the video encoder 200 can determine a set of quantization offset parameters associated with the transform coefficient group based on edge information of the block used for the video pixels or samples. The video encoder 200 can quantize each transform coefficient group based on the set of quantization offset parameters associated with the corresponding transform coefficient group.

[0063] After quantization, the video encoder 200 can scan the transform coefficients to generate a one-dimensional vector from a two-dimensional matrix including the quantized transform coefficients. The scan can be designed to place higher-energy (and therefore lower-frequency) transform coefficients before the vector and lower-energy (and therefore higher-frequency) transform coefficients after the vector. In some examples, the video encoder 200 can utilize a predefined scan order to scan the quantized transform coefficients to produce a serialized vector, and then entropy-encode the quantized transform coefficients of the vector. In other examples, the video encoder 200 can perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 200 can entropy-encode the one-dimensional vector, for example, according to context-adaptive binary arithmetic decoding (CABAC). The video encoder 200 can also entropy-encode the values ​​of syntax elements used to describe metadata associated with the encoded video data for use by the video decoder 300 when decoding the video data.

[0064] To perform CABAC, the video encoder 200 can assign context within a context model to the symbols to be transmitted. Context may involve, for example, whether the neighboring values ​​of a symbol are zero. Probability determination can be based on the context assigned to the symbols.

[0065] The video encoder 200 can also generate syntax data (such as block-based syntax data, image-based syntax data, and sequence-based syntax data), or other syntax data (such as sequence parameter sets (SPS), image parameter sets (PPS), or video parameter sets (VPS)) for the video decoder 300, for example, in image headers, block headers, or slice headers. Similarly, the video decoder 300 can decode such syntax data to determine how to decode the corresponding video data. Side information for video data blocks can be syntax data associated with the video data blocks, such as block-based syntax data.

[0066] In this way, the video encoder 200 can generate a bitstream that includes encoded video data, such as syntax elements describing the segmentation of an image into blocks (e.g., CUs) and prediction and / or residual information for those blocks. Finally, the video decoder 300 can receive the bitstream and decode the encoded video data.

[0067] Typically, the video decoder 300 performs the reverse process of the video encoder 200 to decode the encoded video data of the bitstream. For example, the video decoder 300 can use CABAC to decode the values ​​of syntax elements used for the bitstream in a manner substantially similar to, but reversed, the CABAC encoding process of the video encoder 200. Syntax elements can define segmentation information for segmenting images into CTUs, and for segmenting each CTU according to a corresponding segmentation structure (such as a QTBT structure) to define the CUs of the CTU. Syntax elements can also define prediction and residual information for blocks (e.g., CUs) of the video data.

[0068] The residual information can be represented, for example, by quantized transform coefficients. The video decoder 300 can inversely quantize and inversely transform the quantized transform coefficients of the block to reconstruct the residual block used for that block. The video decoder 300 uses a prediction mode (intra-frame prediction or inter-frame prediction) that can be signaled in the bitstream and associated prediction information (e.g., motion information for inter-frame prediction) to form a prediction block for that block. The video decoder 300 can then combine the prediction block and the residual block (on a sample-by-sample basis) to reconstruct the original block. The video decoder 300 can perform additional processing, such as performing a deblocking process to reduce visual artifacts along the block boundaries.

[0069] According to the technology of this disclosure, a video encoder 200 can determine a set of quantization offset parameters for a scaled transform coefficient group of a video data block based on side information associated with the video data block; quantize the scaled transform coefficient group of the video data block at least in part based on the set of quantization offset parameters to generate quantized transform coefficients of the video data block; and generate an encoded video bitstream at least in part based on the quantized transform coefficients of the video data block.

[0070] In summary, this disclosure may involve "signaling" certain information (such as syntax elements). The term "signaling" can generally refer to the transmission of values ​​for syntax elements and / or other data for decoding encoded video data. That is, video encoder 200 can signal values ​​for syntax elements in the bitstream. Generally, signaling refers to generating values ​​in the bitstream. As described above, source device 102 can transmit the bitstream to destination device 116 substantially in real time or not in real time (such as when syntax elements are stored in storage device 112 for later retrieval by destination device 116).

[0071] Figure 2A and Figure 2B This is a conceptual diagram illustrating an example Quadtree Binary Tree (QTBT) structure 130 and its corresponding Decoding Tree Unit (CTU) 132. Solid lines represent quadtree splits, while dashed lines indicate binary tree splits. In each split (i.e., non-leaf) node of the binary tree, a flag is signaled to indicate which split type (i.e., horizontal or vertical) is used, where, in this example, 0 indicates a horizontal split and 1 indicates a vertical split. For quadtree splits, since the quadtree node splits the block horizontally and vertically into four sub-blocks of equal size, there is no need to indicate the split type. Therefore, the video encoder 200 can encode the following, and the video decoder 300 can decode the following: syntax elements (such as split information) for the region tree level (i.e., the first level) (i.e., solid lines) of the QTBT structure 130, and syntax elements (such as split information) for the prediction tree level (i.e., the second level) (i.e., dashed lines) of the QTBT structure 130. The video encoder 200 can encode video data (such as prediction and transform data) for a CU represented by the terminal leaf nodes of the QTBT structure 130, while the video decoder 300 can decode the video data.

[0072] generally, Figure 2BThe CTU 132 can be associated with parameters that define the size of the blocks corresponding to the nodes at the first and second levels of the QTBT structure 130. These parameters may include the CTU size (representing the size of the CTU 132 in the sample), the minimum quadtree size (MinQTSize, which represents the minimum allowed quadtree leaf node size), the maximum binary tree size (MaxBTSize, which represents the maximum allowed binary tree root node size), the maximum binary tree depth (MaxBTDepth, which represents the maximum allowed binary tree depth), and the minimum binary tree size (MinBTSize, which represents the minimum allowed binary tree leaf node size).

[0073] The root node corresponding to a CTU in a QTBT structure can have four child nodes at the first level of the QTBT structure, each child node being partitioned according to a quadtree. That is, the nodes at the first level are leaf nodes (without child nodes) or have four child nodes. An example of QTBT structure 130 represents such a node as including a parent node and child nodes with solid-line branches. If the nodes at the first level are not larger than the maximum allowed binary tree root node size (MaxBTSize), these nodes can be further partitioned by the corresponding binary tree. The binary tree split of a node can be iterated until the nodes resulting from the split reach the minimum allowed binary tree leaf node size (MinBTSize) or the maximum allowed binary tree depth (MaxBTDepth). An example of QTBT structure 130 represents such a node as having dashed-line branches. The binary tree leaf nodes are called coding units (CUs), which are used for prediction (e.g., intra-picture or inter-picture prediction) and transformation. As discussed above, CUs can also be referred to as “video chunks” or “blocks”.

[0074] In one example of a QTBT segmentation structure, the CTU size is set to 128x128 (luminance samples and two corresponding 64x64 chrominance samples), MinQTSize is set to 16x16, MaxBTSize is set to 64x64, MinBTSize (for both width and height) is set to 4, and MaxBTDepth is set to 4. First, a quadtree segmentation is applied to the CTU to generate quadtree leaf nodes. Quadtree leaf nodes can have sizes ranging from 16x16 (i.e., MinQTSize) to 128x128 (i.e., the CTU size). If a quadtree leaf node is 128x128, it will not be further split by the binary tree because this size exceeds MaxBTSize (i.e., 64x64 in this example). Otherwise, the quadtree leaf node will be further split by the binary tree. Therefore, the quadtree leaf node is also used as the root node of the binary tree and has a binary tree depth of 0. When the depth of the binary tree reaches MaxBTDepth (4 in this example), further splitting is not allowed. Similarly, when a binary tree node has a width equal to MinBTSize (4 in this example), further vertical splitting is not allowed. Likewise, a binary tree node with a height equal to MinBTSize means that further horizontal splitting is not allowed for that node. As mentioned above, the leaf nodes of the binary tree are called CUs and are further processed according to prediction and transformation without further splitting.

[0075] As described throughout this disclosure, various aspects of this disclosure describe determining multiple quantization offset vectors of multiple transform coefficients based at least in part on scaled transform coefficients and side information of the current block, and quantizing the transform coefficients at least in part on the multiple quantization offset vectors to generate quantized transform coefficients of the current block. One or more of these example techniques, as well as the example techniques described below, can be performed by the video encoder 200. Specifically, as Figure 8 As shown, the techniques described below can be performed by, for example, the transform processing unit 206 and quantization unit 208 of the video encoder 200. In some examples, this technique can be performed by the video decoder 300.

[0076] To minimize decoder costs, video compression standards such as H.264 / AVC, H.265 / HEVC, VP9, ​​and AV1 use quantization schemes, where only dequantization (e.g., ...) is performed. Figure 9The inverse quantization unit 306 of the video decoder 300 shown is specification and defined to have the simplest and cheapest implementation. This allows encoder designers to consider the trade-off between optimizing compression performance and minimizing encoder cost when choosing between different quantization methods for the video encoder 200. Details of some example quantization schemes can be found in: I.E. Richardson, “The H.264 Advanced Video Compression Standard”, 2nd edition, John Wiley and Sons Ltd., 2010; M. Wien, “High Efficiency Video Coding: Coding Tools and Specification”, Springer Verlag, Berlin, 2015; D. Mukherjee, J. Bankoski, RS Bultje, A. Grange, J. Han, J. Koleszar, P. Wilkins and Y. Xu, “The latest open-source video codec VP9 – an overview and preliminary "Results", Proceedings of the 30th Symposium on Image Coding, San Jose, California, December 2013; and Y. Chen, D. Murherjee, J. Han, A. Grange, Y. Xu, Z. Liu, S. Parker, C. Chen, H. Su, U. Joshi, C. Chiang, Y. Wang, P. Wilkins, J. Bankoski, L. Trudeau, N. Egge, J. Valin, T. Davies, S. Midtskogen, A. Norkin, P. de Rivaz, "An Overview of core coding tools in the AV1 video codec", Proceedings of the 33rd Symposium on Image Coding, San Francisco, California, June 2018.

[0077] There are known techniques for improving quantization that can yield considerable decoding gains compared to direct scalar quantization. However, one or more of these known techniques may have high computational complexity. Such high computational complexity may be acceptable for video-on-demand encoders, but may be too computationally expensive for real-time video encoders. Furthermore, these techniques may be based on strictly sequential computations, which could significantly limit the throughput of the video encoder.

[0078] This disclosure provides example techniques that may solve one or more of these problems. However, the example techniques described in this disclosure should not be construed as being limited to solving this problem or only solving problems with high computational complexity and strict sequential computation.

[0079] This disclosure describes various aspects of heuristic quantization techniques that can provide significant decoding gains with the low computational complexity required by real-time encoders. These techniques are based on a machine learning framework that “simplifies” optimization decisions. Instead of basing quantization decisions on nearly accurate but computationally very expensive per-coefficient bit rate values, the example techniques described in this disclosure can use training data to learn an estimate of the bit rate shared by several coefficients, which can be efficiently translated into improved quantization choices.

[0080] Experimental simulations using a fully implemented version of the technique demonstrate that the results are comparable to those used in the reference HEVC test model (HM) codec, but with significantly lower computational complexity and a direct and efficient parallel implementation. Using the general VVC test conditions, the novel technique described in this paper provides gains of 4.88%, 3.24%, 4.50%, and 4.44% across all intra-frame, random access, low-latency B, and low-latency P modes, respectively, which outperforms the Rate-Distortion Optimized Quantization (RDOQ) technique in the HEVC test model (HM) (except for random access).

[0081] Quantization is an integral part of all lossy video compression techniques because it is a stage in the video decoding process that defines bit usage and the acceptable "loss" of the compressed signal. In other words, quantization defines which parts of the original media signal should be discarded to obtain the desired compressed data size.Some aspects of quantization are described in the following: I.E. Richardson, “The H.264 Advanced Video Compression Standard,” 2nd edition, John Wiley and Sons Ltd., 2010; M. Wien, “High Effciency Video Coding: Coding Tools and Specification,” Springer-Verlag, Berlin, 2015; D. Mukherjee, J. Bankoski, RS Bultje, A. Grange, J. Han, J. Koleszar, P. Wilkins, and Y. Xu, “The latest open-source video codec VP9 – an overview and preliminary…” "Results", Proceedings of the 30th Image Coding Workshop, San Jose, California, December 2013; Y. Chen, D. Murherjee, J. Han, A. Grange, Y. Xu, Z. Liu, S. Parker, C. Chen, H. Su, U. Joshi, C. Chiang, Y. Wang, P. Wilkins, J. Bankoski, L. Trudeau, N. Egge, J. Valin, T. Davies, S. Midtskogen, A. Norkin, P. de Rivaz, "An Overview of Core Coding Tools in the AV1 Video Codec", Proceedings of the 33rd Image Coding Workshop, San Francisco, California, June 2018; J. W. Woods, "Multidimensional Signal, Image, and Video Processing and Coding", 2nd ed., Academic Press, 2011; W. A. ​​Earlman and A. Said, "Digital Signal Compression: Principles and "Practice", Cambridge University Press, 2011; D. Taubman and W. Marcellin, "JPEG2000: Image Compression Fundamentals, Standards and Practice", Springer, 2002.

[0082] When used with orthogonal transforms, video encoder 200 and video decoder 300 can compute approximate uniform reconstruction quantization (NURQ) by simply using some form of rounding at video encoder 200 and possibly only using multiplication (scalar scaling) and possibly one addition at video decoder 300. Some aspects of NURQ are described in the following entry: G.S. Jullivan and S. Sun, “On dead-zone plus uniform threshold scalarquantization”, SPIE 5960, Visual Communication and Image Processing, Beijing, China, 2005.

[0083] Figure 3A A video decoding system with adaptive and / or rate-distortion optimized quantization is shown. Specifically, Figure 3A The main components of a video encoder (such as video encoder 200) in a hybrid video decoding system (such as video encoding and decoding system 100) are shown, and examples of some terms and notations used throughout this disclosure are provided. Even though the video encoding process uses pixels organized as two-dimensional blocks, in some examples, this disclosure may describe the data used in the video encoding process as an N-dimensional vector to simplify the notation of the data described herein.

[0084] The video encoder 200 may include a buffer 134 that receives and stores a block 150 of pixels (such as pixels of video data) for encoding. A prediction unit 136 may generate a prediction block b from the pixel block 150. A residual generation unit 137 may generate residual data r based on the prediction block b and the pixel block 150. A transform unit 138 may generate transform coefficients c based on the residual data r. A scaling unit 140 may scale the transform coefficients c based on a quantizer step size s to determine scaled transform coefficients x. A quantization unit 142 may quantize the scaled transform coefficients x based on a bit cost estimate and side information from an entropy decoding engine 144 to generate quantized transform coefficients q. The entropy decoding engine 144 may generate an encoded video bitstream including compressed data bits 146 based on the quantized transform coefficients q.

[0085] use Figure 3A The notation shown assumes the quantizer has a scaling factor (quantization step size) s, and represents a vector of scaled transform coefficients. It can be defined as:

[0086]

[0087] Where c represents the transformation coefficients before scaling, and s is the quantizer step size.

[0088] NURQ quantization is defined using two parameters, p and u, where u is the quantization offset, and It is the function floor(.), and is calculated using the following formula:

[0089]

[0090] Furthermore, the inverse of NURQ quantization is defined as:

[0091]

[0092] in

[0093]

[0094] Because this simple yet effective method has a very low implementation cost, it was adopted by all first-generation media compression standards and is still commonly chosen for real-time encoders. The H.264 / AVC and H.265 / HEVC standards have a standard rule that the video decoder 300 must use equation (3) where p = 0. Unless otherwise stated, this disclosure assumes that the video decoder 300 can use the parameter p = 0 to dequantize the quantized values.

[0095] Because quantization defines one aspect of lossy compression—the trade-off between the resulting degradation in reproduction quality and the number of bits used (distortion and bit rate)—a truly optimal form of quantization can take into account all factors affecting rate and distortion (RD), not only for each scalar being quantized, but simultaneously for all vectors. This potentially makes the optimal quantization problem exceptionally complex and may require identifying and addressing only the most important and manageable factors to simplify quantization for practical applications in real-world video encoders.

[0096] Several practical methods have been proposed to improve quantization and have been shown to produce significant decoding gains (compression improvements), such as: using vector quantization, as described in: W. A. ​​Earlman and A. Said, “Digital Signal Compression: Principles and Practice,” Cambridge University Press, 2011; and D. Taubman and W. Marcellin, “JPEG2000: Image Compression Fundamentals, Standards and Practice,” Springer, 2002; efficient modeling of quantization interactions with predictive decoding, as described in: A. Said, O. Guleryuz and S. Yea, “Improving hybrid coding via control of quantization errors in the spatial and frequency domains,” IEEE International Conference on Image Processing, Paris, France, September 2014; and O. Guleryuz, A. Said and S. Yea, “Non-causal encoding of predictively coded "Samples", IEEE International Conference on Image Processing, Paris, France, September 2014; and the determination of the optimal integer value of each coefficient for the usage rate distortion cost function.

[0097] However, in practice, these methods can lead to a relatively large increase in computational complexity and implementation cost compared to the simplest form of quantization. In fact, some mathematical quantization optimization problems can be computationally difficult to solve. Therefore, solutions that are not necessarily optimal but good enough for compression can be more easily determined and can be more practically implemented by hybrid video decoding systems such as video encoding and decoding system 100.

[0098] This disclosure describes a technique that uses information from several coefficients to achieve better performance (in a sense, a form of vector quantization), and then applies a per-coefficient rule to allow parallel computation and achieve computational complexity that is likely low enough for a real-time video encoder.

[0099] The following describes adaptive quantization. A characteristic of video decoding is that data statistics can change significantly based on the type of prediction residuals being decoded, quality settings, video features, etc. Therefore, adaptive quantization can improve compression by using different parameters according to the desired statistical type.

[0100] The simplest adaptive form can be based solely on encoder settings and states, since these are readily available. For example, the reference implementations of H.264 / AVC and H.265 / HEVC use an offset parameter u = 1 / 3 in Equation (2) for slices with intra-frame prediction only (I-slices), and an offset parameter u = 1 / 6 otherwise.

[0101] Adaptive approaches can be extended using statistical analysis of the data being decoded, as described in: G.S. Jullivan and S. Sun, “On dead-zone plus uniform threshold scalarquantization,” SPIE 5960 Proceedings, Visual Communications & Image Processing, Beijing, China, 2005. However, these empirical techniques can be very limited because they do not take into account how quantization decisions affect the bit rate.

[0102] Rate-distortion-based quantization is described below. Techniques have been proposed to modify quantization to account for both distortion and bit count, and these techniques are even compatible with some very early compression standards such as JPEG and MPEG-2. Some of these techniques are described in the following: K. Ramchandran and M. Vetterli, “Rate-distortion optimal fast thresholding with complete JPEG / MPEG decoder compatibility,” IEEE Trans. on Image Processing, Vol. 3, No. 5, September 1994; and K. Ramchandran, A. Ortega and M. Vetterli, “Bit allocation for dependent quantization with applications to multiresolution and MPEG video coders,” IEEE Trans. on Image Processing, Vol. 3, No. 5, September 1994.

[0103] The optimal quantization problem is defined by minimizing the distortion function, constrained by an upper bound on the bit rate, and averaged over all video blocks. Since blocks are quantized and decoded independently, the problem can be solved using a Lagrange multiplier λ, as described in: K. Ramchandran and M. Vetterli, “Rate-distortion optimal fastthresholding with complete JPEG / MPEG decoder compatibility,” IEEE Trans. on Image Processing, Vol. 3, No. 5, September 1994. Therefore, this disclosure describes the problem of directly optimizing quantization in video blocks in this form.

[0104] Given a vector x from equation (1) and a vector with quantized transform coefficient values. Function D can be defined s B(x,q) and B(q) respectively measure the distortion produced by quantization and the number of bits required for entropy encoding of vector q. When the transforms are orthogonal and the distortion corresponds to squared errors, quantization is optimal in a rate-distortion sense if it is solved by the following optimization problem:

[0105]

[0106] Due to complexity constraints, there may not be a general method available in practical video encoders for solving optimization problems with integer variables precisely for the optimization problem described by equation (5), and heuristic methods may be required to solve the optimization problem.

[0107] A useful tool for handling such problems is to test how the objective function changes when a single element of the solution vector q is changed from one integer value to another. To formally represent this, the element-wise substitution operator is defined. Make:

[0108]

[0109] and difference operators Make

[0110]

[0111]

[0112] and

[0113]

[0114] generally,

[0115]

[0116] Based on these definitions, heuristic optimization methods can be designed to use approximate values ​​instead of the exact value of B(q). To calculate

[0117]

[0118] Furthermore, q is changed accordingly whenever a negative value is encountered. This type of algorithm (such as that described in: M. Karczewicz, P. Chen, Y. Ye, and R. Joshi, “RD based quantization in H.264”, SPIE 7443 Proceedings, Applications of Digital Image Processing XXXII, September 2009) is implemented in the reference software of the H.265 / HEVC standard and is called Rate-Distortion Optimized Quantization (RDOQ). Note that despite the name, the quantization may not be truly optimized because RDOQ uses approximate and inaccurate (heuristic) optimization techniques.

[0119] Some properties can be used to reduce the number of calculations in formula (10). For example, since the statistical distribution of the transformation coefficients monotonically decreases with amplitude, the expected value will make...

[0120]

[0121] Among them, the notation E a.e. The curly braces {·} are used to indicate that this is expected in all blocks, but there may be some rare exceptions.

[0122] The following describes HEVC symbol bit hiding. The H.265 / HEVC video decoding standard includes a technique called symbol bit hiding, as described in, for example, §8.2.4 of M. Wien, “High Effciency Video Coding: Coding Tools and Specification,” Springer-Verlag, Berlin, 2015. Specifically, the H.265 / HEVC video decoding standard specifies that for a quantized transform coefficient set satisfying the condition of a minimum number of non-zero elements, the parity of the sum of the quantized amplitudes must be equal to the symbol bit of the first non-zero coefficient. In this way, the symbol bit does not necessarily need to be encoded, thus reducing the total number of encoded bits.

[0123] In some examples, when the quantized coefficients do not meet the parity check condition (in 50% of cases on average), it may be possible to find a coefficient whose quantized value can be increased or decreased by 1 without a significant change in distortion. Therefore, scalar quantization in HEVC-HM software achieves this technique by searching for the allowable variation corresponding to the minimum increase in squared error distortion.

[0124] When RDOQ is enabled, the search is based on the full RD cost estimate, and the index k of the coefficients whose quantized values ​​will be modified can be calculated using the following formula:

[0125]

[0126] Here, A is a set of indices whose quantization values ​​can be changed. Additional implementation steps may exist.

[0127] Techniques for potentially solving some of the problems described above are now described. As mentioned above, quantization can be very effectively adaptively and partially optimized if it uses a method to accurately measure how the number of encoded bits changes with the quantizer's decisions.

[0128] One potential problem is that modern encoders use very fine-grained entropy decoding with numerous arithmetic decoding contexts and complex context selection rules. Furthermore, transform coefficients are encoded in more than one path. For example, entropy decoding in H.265 / HEVC can be done in up to five paths. This makes calculating and using those bit cost estimates complex and computationally expensive for each decision.

[0129] Aspects of this disclosure describe techniques that potentially solve these problems by eliminating the need to derive an index of the arithmetic decoding context for entropy decoding of the transform coefficients for each non-zero transform coefficient, and the need to access the state of those arithmetic decoding contexts. Aspects of this disclosure describe techniques in which the estimation rule is computed once and used in the quantization of multiple transform coefficients, thereby reducing the average complexity per pixel and enabling parallel computation of quantized transform coefficients.

[0130] In some aspects, the video encoder 200 may determine a set of quantization offset parameters for quantizing the transform coefficient group. The quantization offset in the set of quantization offset parameters may be variable rather than fixed, and may depend on the quantization interval, and may be a function of: the values ​​of the transform coefficients in the transform coefficient group (e.g., the magnitude of the transform coefficients), side information associated with the transform coefficient block that includes the transform coefficient group, values ​​from other transform coefficient blocks, etc.

[0131] The following describes additional examples of rate-distortion analysis. The rate-distortion equations analyzed for quantization optimization are typically in the form of equations (5) or (9), which are also used in the computational implementation of quantization. A potential problem with intuitively interpreting those equations is that they contain the Lagrange multiplier factor λ, which has a value that varies over a wide range depending on the choice of reproducibility quality.

[0132] Although the value of λ is independent of other encoder decisions, it may be directly related to the choice of the quantizer step size s, as described in: T. Wiegand and B. Girod, “Lagrange multiplier selection in hybrid video coder control,” IEEE International Conference on Image Processing, Thessaloniki, Greece, 2001, Vol. 3, pp. 542-545. This may mean that, in order to achieve rate distortion compatible with λ, we can define it as:

[0133]

[0134] Here, α varies within a relatively small range. For example, the HEVC HM software defines α as:

[0135]

[0136] By replacing λ in equation (9) with λ as defined in equation (12), we obtain the normalized form of the RD cost change:

[0137]

[0138] Note that equation (14) is derived directly from the objective function of equation (5), that is, without approximation or special assumptions, only normalization and explicit use of squared error distortion to achieve a more intuitive interpretation.

[0139] Special case m=n+1

[0140]

[0141] To show in a more direct and intuitive way, this is equivalent to only when equation (15) is nonnegative.

[0142]

[0143] Choosing the quantization value n+1 might be better than choosing n.

[0144] According to equation (16), if The optimal quantization typically corresponds to a rounding operation (equation (2), where parameters p = 0 and u = 1 / 2). Below, this disclosure describes how to use equation (16) to obtain a more general form, which is also similar to the quantization defined by equation (2).

[0145] The improved quantization with trained quantized offset vectors (QOV) is described below. One of the potential difficulties in using Equation (14) in its exact form is to compute the change in bit count as the quantized value changes from n to m. This potential problem can be potentially mitigated if an approximate estimate is used instead of the calculation of the bit count changes (as noted in Equation (10) and in the RDOQ technique described in the following: M. Karczewicz, P. Chen, Y. Ye and R. Joshi, “RD based quantization in H.264”, SPIE 7443 Proceedings, Applications of Digital Image Processing XXXII, September 2009). However, another approach to addressing this problem could be that, for each transform coefficient that is quantized, information is not obtained directly from the arithmetic decoding context (e.g., the probability estimation element).

[0146] According to various aspects of this disclosure, a video encoder (such as video encoder 200) can quantize transform coefficients based on an estimated quantization offset representing a bit count difference, rather than determining the bit count difference as part of the transform coefficient quantization process. The video encoder can determine a set of quantization offset parameters for a group of transform coefficients in a video data block based on side information associated with the video data block, and can quantize the group of transform coefficients based on the set of quantization offset parameters to generate quantized transform coefficients.

[0147] Figure 3B This disclosure illustrates a low-complexity adaptive quantization technique based on block classification and one or more sets of quantization offset parameters, according to various aspects of this disclosure. Figure 3B As shown, instead of using information from the entropy decoding engine 144 (such as the index of the arithmetic decoding context used for entropy decoding of each non-zero transform coefficient and the state of those contexts that must be accessed) or information indicating the bit cost estimate for quantizing the transform coefficients, in order to estimate the change in the number of bits used to encode the quantized value q... The scaling and classification unit 152 of the video encoder 200 can use the side information Γ (e.g., block size, prediction type, etc.) associated with the video data block and the actual distribution of the scaled transform coefficients x to determine the quantization unit 142 of the video encoder 200, which can be used to generate the quantization offset parameter v of the quantized transform coefficients.

[0148] The video encoder 200 is able to generate quantized transform coefficients using a quantization offset parameter, rather than using a change in the number of bits. The change in the number of bits is measured as a single element of q changes from an integer value n to an integer value m. To formally represent this, the element-wise substitution operator is defined as in equation (6) and in equations (14) through (16).

[0149] • The video encoder 200 can be based on the difference in bit count Instead of directly determining the optimal quantization value of the transform coefficients based on the precise number of bits B(q) for entropy coding of the quantized transform coefficients q, the video encoder 200 can identify patterns valid for the transform coefficient set independently of the precise number of bits used for entropy coding of the quantized transform coefficients.

[0150] • Changes in the number of bits Typically very small. In fact, when m or n is equal to or close to zero, the largest size may only be a few bits, and the magnitude of those variations decreases rapidly as m and n become larger. For example, if |m|>16, |n|>16, then and

[0151] Since the constant α may be relatively large, the ratio It may be relatively small. This could mean that the optimization decision in equation (16) corresponds to the offset parameter u in equation (2) (when parameter p = 0), which is close to 1 / 2. For example, in the example in Table I below, the offset parameter u = 0.1 is likely optimal only when the quantization value of a single coefficient increases by about 9 bits (which is not expected).

[0152] Table I – Using Equation (2) (parameter p = 0, non-negativity of n, α = 11.14) to optimize the quantization offset u and the increment of the number of coded bits

[0153]

[0154] The example techniques described in this disclosure can utilize the aforementioned properties and assume that for an L-dimensional subgroup of the elements of a vector x (e.g., a 4-dimensional or 16-dimensional subgroup of a 4x4 pixel block of an image), an approximate function can be found. Make

[0155]

[0156] Where x represents a subgroup of scaled transform coefficients, n is the proposed quantization value for the scaled transform coefficients in the group of scaled transform coefficients x, g is the index of the group of scaled transform coefficients to which the scaled transform coefficient x belongs, and Γ represents a data structure with side information for video data blocks. In some examples, the side information Γ may include one or more of the following: the slice type of the video data block (I, P, or B); residual data from intra-frame or inter-frame predictions of the video data block; the block size of the video data block (e.g., 4×4, 8×8, 16×16, or 32×32); and the luma or chroma component.

[0157] To simplify the notation, the elements of the vector x representing the scaled transformation coefficients may have been rearranged previously, for example, to make memory access more efficient or to make the approximation more accurate, and thus subgroups of elements in vector x may have consecutive indices.

[0158] Using this notation, and assuming B(q) = B(-q), if defined...

[0159]

[0160] If the equality sign is taken in equation (17), then the quantized value will correspond to equation (16), using

[0161]

[0162] in The transformation coefficient x i The quantized value of , and u(n,g,x,Γ) represents the set of quantization offset parameters of the scaled transform coefficient vector x of a video data block (such as a transform block).

[0163] Equation (18) can be primarily used for mathematical consistency because, in practical applications, If this condition is not met (possibly for the case of n=0), a slightly more complex quantization rule can be used based on equation (14) instead of equation (15). Equation (19) can be a low-computational-complexity quantization of the same type as in equation (2), where p=0 and has a quantization offset u that varies depending on the size of the quantized coefficients.

[0164] Another simplification in actual implementation can be based on The expectation (for larger values ​​of |n|). Therefore, if a P-dimensional vector function v(g,x,Γ) (which is called the quantization offset vector (QOV)) can be defined such that

[0165] v n (g,x,Γ)=u(n,g,x,Γ),n=0,1,2,…,P-1, (20)

[0166] The following formula can be used to approximate equation (19).

[0167]

[0168] Therefore, the quantization offset vector can represent a set of P quantization offset parameters, which are generated by applying a cutoff at larger quantization values ​​|n| due to the assumption that the change in the number of bits is close to zero. For example, for The cutoff is reflected in the application of the minimum function of the index of QOV in equation (21).

[0169] Based on these definitions, here is an example for adaptive quantization techniques: for each transform coefficient vector,

[0170] 1. Determine the edge information Γ of the video data block (such as a transform block) to which the transform coefficient vector x belongs;

[0171] 2. Split the transform coefficient vector x into K = N / L subgroups, where N is the number of pixels in the block, and L is the number of pixels in each subgroup (e.g., for a 4x4 sub-block, L = 16); and

[0172] 3. For each subgroup g, where g = 0, 1, ..., n / L-1:

[0173] a. Determine the quantization offset vector v(g,x,Γ); and

[0174] b. For each transformation coefficient x i Where the index i = gL, gL+1, ..., (g+1)L-1, that is, each transformation coefficient in the subgroup g:

[0175] i. For example, according to equation (21), at least in part based on x i The magnitude of x is calculated using the corresponding elements of the quantization offset vector v(g,x,Γ). i The quantified value.

[0176] According to various aspects of this disclosure, the video encoder 200 can generate a set of quantization offset parameters for quantizing scaled transform coefficients of video data blocks (such as transform coefficient blocks or transform blocks of video data), the scaled quantization offset parameters corresponding to a change in the number of bits. The video encoder 200 can determine side information associated with a video data block. The side information may include any combination of one or more of the following: the slice type of the video data block (e.g., I, P, or B), residual data generated by intra-frame prediction or inter-frame prediction of the video data block, the block size of the video data block (e.g., 4×4, 8×8, 16×16, or 32×32), and / or the luminance or chrominance components of the video data block.

[0177] The video encoder 200 can divide the scaled transform coefficients of a video data block into multiple sets of scaled transform coefficients, specifically associated with sub-blocks of the video data block. If the video data block comprises N pixels, the video encoder 200 can divide the video data block into sub-blocks of groups of L pixels to produce N / L sets of scaled transform coefficients. For example, the video encoder 200 can divide the video data block into 4x4 sub-blocks (groups of 16 pixels), 8x8 sub-blocks (groups of 64 pixels), 16x16 sub-blocks (groups of 256 pixels), and so on. Therefore, each sub-block of the video data block can be represented by an index g from 0 to N / L–1, where each sub-block includes a set of transform coefficients for the sub-block.

[0178] The video encoder 200 can determine a set of quantization offset parameters for each sub-block, allowing the video encoder 200 to quantize a scaled group of transform coefficients within the same sub-block using the same set of quantization offset parameters determined for that sub-block. As described above, the quantization offset parameters can be associated with a change in the bit count based on a change in a single element of the quantized transform coefficients for the data block. In some examples, each quantization offset parameter in the set of quantization offset parameters can be in the range of 0 to 0.5.

[0179] The set of quantization offset parameters used for the scaled transform coefficient group in a sub-block can be adaptive rather than fixed, allowing the video encoder 200 to adaptively select different quantization offset parameters to quantize different scaled transform coefficients in the scaled transform coefficient group, instead of using the same fixed quantization offset parameters to quantize the scaled transform coefficients in the scaled transform coefficient group. An example set of adaptive quantization offset parameters for the scaled transform coefficient group could be [0.2, 0.3, 0.35, 0.4, 0.45, 0.5]. It can be seen that the quantization offsets in the adaptive quantization offset parameter set are not fixed to a single value, and each element in the quantization offset parameter set does not necessarily increase by the same value. For example, increasing the first element with a value of 0.2 by 0.1 produces a second element with a value of 0.3, and increasing the second element by 0.05 produces a third element with a value of 0.35.

[0180] The video encoder 200 can quantize the scaled transform coefficients, for example, by adding the value of the scaled transform coefficient to a quantization offset and rounding the sum down or to its lower bound as an integer value. For example, given a scaled transform coefficient x and a quantization offset u, the video encoder 200 can determine the quantized value q of the scaled transform coefficient as follows:

[0181] The video encoder 200 can determine the set of quantization offset parameters for a sub-block based on the side information of the video data block. In some examples, the video encoder 200 can also determine the set of quantization offset parameters for a sub-block based on the side information associated with the sub-block, such as the maximum range of transform coefficients of the sub-block, the block size of the sub-block, and the relative position of the sub-block within the block.

[0182] An example of a set of quantization offset parameters is a quantization offset vector (QOV). A QOV can contain a list of quantization offsets, where an M-dimensional QOV can contain M quantization offsets. Although aspects of this disclosure are described in accordance with QOV, the techniques described herein are generally applicable to any form of set of quantization offset parameters, such as arrays, lists, stacks, queues, tables, graphs, etc.

[0183] In some examples, the quantization offsets in the M-dimensional QOV are indexed from 0 to M–1. Therefore, the video encoder 200 can use equation v n =V[min(M-1,n)] to select the quantization offset from QOV, where if n is less than M-1, then v n It is the quantization offset of the nth element of QOV. Otherwise, v n It is the M–1th element of QOV. Therefore, given the scaled transform coefficient x, the video encoder 200 can take the absolute value of the transform coefficient x and round it down to the nearest integer, such that the equation can become in This refers to the quantization offset selected from QOV for the transform coefficient x. Given the transform coefficient x and the quantization offset u, the video encoder 200 can determine the quantized transform coefficient q as... Therefore, given As a quantization offset selected from QOV for the transform coefficient x, the video encoder 200 can determine the quantized transform coefficient x as...

[0184] Figure 4 A parallel implementation of adaptive quantization using a set of quantization offset parameters, based on the techniques described herein, is illustrated. Specifically, Figure 4An example implementation of the above technique for quantizing scaled transform coefficients is shown, wherein the scaled transform coefficients of the block are quantized in parallel while determining the set of quantization offset parameters for the scaled transform coefficients of the video data block. Figure 4 The components shown can be, for example Figure 3B A portion of the scaling and classification unit 152 and the quantization unit 142 of the video encoder 200 shown.

[0185] like Figure 4 As shown, the video encoder 200 may include a quantization offset parameter determination unit 154, a grouping unit 156, and offset quantization units 158A-158P. The grouping unit 156 may be a processing circuit configured to receive scaled transform coefficients x of a video data block (e.g., a transform block) and divide the scaled transform coefficients into scaled transform coefficient groups, such as by dividing the video data block into sub-blocks. The grouping unit 156 may divide N scaled transform coefficients into groups of L scaled transform coefficients to produce K = N / L scaled transform coefficient groups indexed from 0 to N / L–1. For example, the grouping unit 156 may group the following: scaled transform coefficients x0 to x... L-1 The set of scaled transformation coefficients x L To x 2L-1 The set, and so on, up to the scaled transform coefficients x. N-L To x N-1 A set of.

[0186] The quantization offset parameter determination unit 154 may be a processing circuit configured to receive scaled transform coefficients x and side information Γ of a video data block, and determine a set of quantization offset parameters for each group in the scaled transform coefficient group determined by the grouping unit 156. The quantization offset parameter determination unit 154 may determine the set of quantization offset parameters for each group of scaled transform coefficients based on the side information Γ of the video data block and / or the values ​​of the scaled transform coefficients in the scaled transform coefficient group.

[0187] exist Figure 4 In the example, the quantization offset parameter determination unit 154 can determine the QOV as a set of quantization offset parameters for the scaled transform coefficient group. Figure 4The QOV is given in the form v(g,x,Γ), where g is an index of a scaled transform coefficient set from 0 to N / L–1, x is the set of scaled transform coefficients, and Γ is the side information of the video data block. Therefore, the quantization offset parameter determination unit 154 can determine the QOV v(0,x,Γ) for the scaled transform coefficient set associated with index g=0, and the QOV v(1,x,Γ) for the scaled transform coefficient set associated with index g=1, up to the QOV v(N / L–1,x,Γ) for the scaled transform coefficient set associated with index g=N / L–1.

[0188] Offset quantization units 158A-158P can be processing circuitry configured to quantize scaled transform coefficients based on quantization offset parameters. For example, offset quantization units 158A-1-158A-M can use QOV v(0,x,Γ) to quantize scaled transform coefficients x0 to x... L-1 The group is quantized to generate quantized values ​​q0 to q L-1 The offset quantization unit 158B-1-158B-M can use QOV v(1,x,Γ) to quantize the scaled transform coefficients x. L To x 2L-1 The group is quantized to generate a quantized value q. L to q 2L-1 The offset quantization unit 158P-1-158P-M can use QOV v(N / L–1,x,Γ) to quantize the scaled transform coefficients x. N-L To x N-1 The group is quantized to generate a quantized value q. N-L to q N-1 .

[0189] The following describes several practical techniques that can be used to determine the set of quantization offset parameters (such as QOV). As can be seen from the derivation above, determining the optimal QOV is mathematically equivalent to estimating...

[0190] In some examples, aspects of the techniques described herein can be applied to HEVC symbol bit hiding. As given in the description of HEVC symbol bit hiding above, optimized symbol bit hiding can be based on rate-distortion cost. Similar to quantization, quantization offset parameters (such as QOV) can also be used to improve the performance of symbol bit hiding.

[0191] Using the notation described in the above description of the substitution rate distortion analysis, and assuming The cost of changing the value of the quantized transform coefficient to satisfy the symbol bit hidden parity constraint is given by the following formula:

[0192]

[0193] It can be proven to be equal to

[0194]

[0195] The rule for identifying the index of the coefficient to be changed (equivalent to equation (11)) becomes

[0196]

[0197] Using equations (17), (18), and (20), the following function can be used to approximate the change in RD cost when the quantization value is changed for symbol bit hiding:

[0198]

[0199] Figure 5 The factor used for symbol bit hiding in the equation to account for the approximation rate-distortion cost is shown. Specifically, Figure 5 The function used in calculating the symbol bit-hiding RD cost (assuming a positive x) is shown. The RD cost is zero at the value used as a threshold in the quantization equation (21), consistent with the example where the exact coefficient values ​​for two quantization levels have the same RD cost.

[0200] For example, Figure 162A shows that, for a quantization offset v, sign(xq) (which is the sign of the difference between the scaled transform coefficient x and its quantized value q) from... arrive It is positive, and from arrive It is negative. Figure 162B shows that, The range of 1-xv is from 1-v at the location -v at the location. Figure 162C shows, exist At position 1-v, in The value is 0, and... The value at position v.

[0201] Using these approximations, equation (24) can be replaced with:

[0202]

[0203] The index k of the transform coefficients whose quantized values ​​need to be modified for symbol bit hiding purposes can be determined based on the side information Γ of the video data block including the transform coefficients. Therefore, the video encoder 200 can determine the quantized values ​​of the transform coefficients in the video data block that the video encoder 200 can modify for symbol bit hiding purposes based on the side information Γ of the video data block including the transform coefficients associated with the quantized values.

[0204] Figure 6 A technique for determining quantization offset parameters according to this disclosure is illustrated. Given a criterion for mapping a triple (g,x,Γ) (where x is a vector of scaled transform coefficients, g is a group index of the scaled transform coefficient group, and Γ is side information for a video data block including the scaled transform coefficient group) to a set of quantization offset parameters (such as QOV), statistical or machine learning techniques can be used to optimize the values ​​of the elements of the quantization offset parameter set (i.e., quantization offsets) such that the quantization offset parameters correspond to those quantization offset values ​​that maximize the ratio of the average RD cost function obtained using the quantization offset values ​​to the actual (precise) cost function.

[0205] There can be a wide variety of different techniques for determining the set of quantization offset parameters for a set of transform coefficients by mapping parameters g, x, and Γ associated with a scaled set of transform coefficients (e.g., scaled transform coefficients in a sub-block) to a QOV, and the techniques of this disclosure can include any suitable techniques for mapping parameters g, x, and Γ associated with a scaled set of transform coefficients to a QOV.

[0206] Figure 6 Some example techniques based on the present disclosure are shown for determining the set of quantization offset parameters used for a transform coefficient group. Although Figure 6 The set of quantization offset parameters is represented as QOV, but the techniques shown in this paper are applicable to any other suitable form of the set of quantization offset parameters.

[0207] like Figure 6 As shown, in Example 172A, the video encoder 200 can implement a QOV calculation unit 174, which can directly calculate the QOV of the scaled transform coefficient set from the parameter group index g, the scaled transform coefficient x, and the side information Γ containing the scaled transform coefficient set associated with the video data block.

[0208] In Example 172B, the video encoder 200 can determine the QOV using a parameterization method, where the elements of the QOV are defined based on a vector p with a smaller dimension. The video encoder 200 can implement a QOV parameter calculation unit 176, which can determine the parameter vector p based at least in part on parameters g, x, and Γ associated with the transform coefficient set, where the parameter vector p can be a vector with a smaller dimension (i.e., fewer elements) compared to the QOV to be determined. The video encoder 200 can implement a QOV calculation unit 178, which can determine the QOV for a transform coefficient set with associated parameters g, x, and Γ based on the parameter vector p.

[0209] For example, the parameter vector p can be a two-dimensional parameter vector with a smaller dimension than a p-dimensional QOV, and the video encoder 200 can determine the p-dimensional QOV for parameters g, x, and Γ based on the parameter vector p using the following equation, where v n It is the nth element of QOV:

[0210]

[0211] In some examples, the video encoder 200 can utilize a pre-computed set of QOVs to determine the QOV used for quantizing the transform coefficients. The pre-computed set of QOVs can be in the form of an array of QOVs, and the video encoder 200 can index into the array of QOVs to select the QOV used for quantizing the transform coefficient set. In Example 172C, the video encoder 200 can implement a QOV index calculation unit 180, which maps a transform coefficient set with associated parameters g, x, and Γ to an index n, where n can be an integer. The video encoder 200 can implement a QOV retrieval unit 182, which can use the index n to index into the pre-computed array of QOVs to determine the QOV used for quantizing x from the array of QOVs.

[0212] In some examples, the video encoder 200 may use general methods (such as neural networks) that can be used for both classification and regression to determine the quantization offset parameters in Examples 172A-172C. For example, such a neural network may be trained using training data comprising a set of parameters g, x, and Γ, and optimal quantization offset parameter values ​​(such as QOV), with an objective function for encoding performance (such as rate-distortion values), producing quantized values ​​of RDOQ performance relative to HEVC-HM. In this way, the neural network is trained to associate the side information Γ of the video data block and the values ​​of scaled transform coefficients with the set of quantization offset parameter values, which optimizes the rate-distortion cost of quantizing the scaled transform coefficients and the associated set of quantization offset parameters.

[0213] Similarly, in some examples, the video encoder 200 can use general regression methods (such as linear regression, logistic regression, Poisson regression, etc.) to determine the relationship between the set of parameters g, x, and Γ and the optimal quantization offset parameter values ​​using the neural network described above (e.g., in Examples 172A and 172B).

[0214] In some examples, the video encoder 200 can use classification methods (such as classification trees) to classify the set of parameters g, x, and Γ in order to select quantization offset parameters for those parameters from a pre-computed set of quantization offset parameters (e.g., in Example 172C). For example, the neural network described above can act as a classifier, trained to classify the scaled transform coefficient set based on side information of video data blocks containing the scaled transform coefficient set. By classifying the scaled transform coefficient set, the video encoder 200 can select from multiple sets of quantization offset parameters a set of quantization offset parameters for quantizing the scaled transform coefficient set. An example of such a classification method is described below.

[0215] As described in Example 172C, video encoder 200 can select a set of quantization offset parameters (e.g., a pre-computed QOV) for a set of scaled transform coefficients x in a sub-block having associated side information Γ. Video encoder 200 can select a QOV from the pre-computed QOV set based at least in part on the side information for the sub-block containing the scaled transform coefficient set and the side information for the video data block containing the sub-block. For example, video encoder 200 can select a QOV from the pre-computed QOV set based at least in part on the position of the sub-block within the video data block, the size of the video data block, etc.

[0216] Figure 7 An example of codes numbered 0 to 9 is shown for identifying the positions of sub-blocks within a video data block. Such a video data block could be, for example, a HEVC or VVC transform coefficient block. Figure 7 As shown, the sub-block can be a 4x4 sub-block of a 4x4 HEVC transform coefficient block, an 8x8 HEVC transform coefficient block, a 16x16 HEVC transform coefficient block, or a 32x32 HEVC transform coefficient block.

[0217] A sub-block can have a code associated with its position within a block and also with the size of the block. When a 4x4 sub-block is within a 4x4 block, it can have a code of 0. When a 4x4 sub-block is within an 8x8 block, it can have a code of 1 if it is a top-left 4x4 sub-block, and a code of 2 if it is not a top-left 4x4 sub-block. When a 4x4 sub-block is within a 16x16 block, it can have a code of 3 if it is a top-left 4x4 sub-block, a code of 4 if it is not a top-left 4x4 sub-block but is within a top-left 8x8 block, and a code of 5 if it is not within a top-left 8x8 block. When a 4x4 sub-block is within a 32x32 block, if the sub-block is a top-left 4x4 sub-block, it can have a code of 6; if the sub-block is not a top-left 4x4 sub-block but is within the top-left 8x8 block, it can have a code of 7; if the sub-block is not within the top-left 8x8 block but is within the top-left 16x16 block, it can have a code of 8; and if the sub-block is not within the top-left 16x16 block, it can have a code of 9.

[0218] Various aspects of this disclosure have been implemented and tested using modified versions of the HM reference software (e.g., with only the encoder changed) to create HEVC-compliant files. One example implementation is based on... Figure 6 Example 172C, where an index of QOV is selected for each group of 4×4 transform coefficients in the block.

[0219] In this implementation, the video encoder 200 can use the following function to determine the QOV for the scaled transform coefficient set from the pre-computed set of QOVs:

[0220]

[0221] The video encoder 200 can determine the index of the QOV of the 4x4 sub-block in the pre-computed QOV set based on the following parameters:

[0222] For all transformation coefficient values ​​in the 4×4 transformation coefficient set, P0 = max(c(x) i ))∈{-1,0,1,2,3};

[0223] • For the transformation coefficients in the 2×2 subgroups within the 4×4 group, P1 = min(2,k-1)∈{0,1,2}, where k is the maximum value of c(x). i The number of times P0 is equal to the number of times P0 is reached.

[0224] P2∈{0,1,…,9} is based on Figure 7 The scheme shown is used to indicate both the block size of the transform coefficients and the position of the 4×4 group within the block; and

[0225] • If the block is part of an intra-slice (according to the HEVC standard), then P3∈{0,1} is 0, otherwise it is 1.

[0226] As can be seen, the video encoder 200 can determine the QOV for a scaled set of transform coefficients within a sub-block of a video data block based on side information associated with the sub-block. The side information associated with the sub-block used to determine the QOV may include, for example, the position of the sub-block within the video data block, such as whether it is the top-left sub-block in the video data block.

[0227] The case P0 = –1 can correspond to the group where all coefficients are quantized to zero. Therefore, in this case, the video encoder 200 may not determine the QOV index for the sub-block with P0 = –1. Based on these definitions, the video encoder 200 can calculate the set of 240 QOV indices using, for example, the following equation:

[0228] n=10×(3×(2×P0+P3)+P1)+P2∈{0,1,2,…,239},

[0229] In the example above, the video encoder 200 can determine the QOV for the scaled transform coefficient group in the sub-block of the video data block based on whether the video data block is part of an intra-frame slice, the position of the sub-block within the video data block, the size of the video data block, and the values ​​of the scaled transform coefficient group (such as the maximum scaled transform coefficient value in the scaled transform coefficient group and the number of times the maximum scaled transform coefficient value is in the 2x2 sub-group within the sub-block).

[0230] In some examples, the video encoder 200 can compute a set of 20 QOV indices that are independent of the values ​​of the transform coefficients (e.g., vector x) within a sub-block, such as by using the following equation: n = 10 × P3 + P2 ∈ {0, 1, 2, ..., 19}. In this example, the video encoder 200 can compute the QOV indices based on the position of the sub-block within the transform coefficient block, the size of the transform coefficient block, and whether the transform coefficient block is part of an intra-slice. Determining QOV indices that are independent of the values ​​of the transform coefficients within a sub-block allows the video encoder 200 to perform quantization of a scaled group of transform coefficients in a sub-block in a single pass, thereby reducing the number of processing cycles required to quantize the scaled group of transform coefficients.

[0231] As can be seen in the above techniques, the video encoder 200 can determine the set of quantization offset parameters for the scaled transform coefficient group in the sub-block based at least in part on the position of the sub-block within the data block and the size of the data block. In some examples, the video encoder 200 can also use the values ​​of the scaled transform coefficients in the sub-block to determine the set of quantization offset parameters for the sub-block, while in other examples, the video encoder 200 can determine the set of quantization offset parameters for the sub-block without using the values ​​of the scaled transform coefficients in the sub-block.

[0232] Figure 8 This is a block diagram illustrating an example video encoder 200 that can perform the techniques described in this disclosure. Figure 8 This disclosure is provided for illustrative purposes and should not be construed as limiting the techniques illustrated and described in this disclosure in a general manner. For illustrative purposes, this disclosure describes a video encoder 200 based on VVC (ITU-T H.266, in deployment) and HEVC (ITU-TH.265) technologies. However, the technologies of this disclosure can be implemented by video encoding devices configured for other video decoding standards.

[0233] exist Figure 8 In the example, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any or all of the video data memory 230, mode selection unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, DPB 218, and entropy coding unit 220 can be implemented in one or more processors or in processing circuitry. For example, the units of the video encoder 200 can be implemented as one or more circuit or logic elements as part of hardware circuitry, or as part of a processor, ASIC, or FPGA. Furthermore, the video encoder 200 may include additional or alternative processors or processing circuitry to perform these and other functions.

[0234] The video data storage device 230 can store video data to be encoded by the components of the video encoder 200. The video encoder 200 can obtain data from, for example, a video source 104 (…). Figure 1The video data memory 230 receives video data stored in the video data memory 230. The DPB 218 can act as a reference picture memory, storing reference video data for use when the video encoder 200 predicts subsequent video data. The video data memory 230 and DPB 218 can be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The video data memory 230 and DPB 218 can be provided by the same memory device or separate memory devices. In various examples, the video data memory 230 can be on-chip (as shown) with other components of the video encoder 200, or off-chip relative to those components.

[0235] In this disclosure, references to video data memory 230 should not be construed as limited to memory within video encoder 200 (unless so specifically described) or to memory outside video encoder 200 (unless so specifically described). Rather, references to video data memory 230 should be understood as reference memory storing video data received by video encoder 200 for encoding (e.g., video data for the current block to be encoded). Figure 1 The memory 106 can also provide temporary storage for the outputs from the various units of the video encoder 200.

[0236] It shows Figure 8 The various units help to understand the operations performed by the video encoder 200. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide a specific function and are pre-configured regarding the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functionality in terms of the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is generally immutable. In some examples, one or more of these units can be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units can be integrated circuits.

[0237] The video encoder 200 may include an arithmetic logic unit (ALU), an essential function unit (EFU), digital circuitry, analog circuitry, and / or a programmable core, all formed by programmable circuitry. In an example where software executed by programmable circuitry is used to perform the operation of the video encoder 200, memory 106 ( Figure 1 The video encoder 200 may store instructions (e.g., object code) of the software received and executed by the video encoder 200, or another memory (not shown) within the video encoder 200 may store such instructions.

[0238] The video data storage unit 230 is configured to store the received video data. The video encoder 200 can retrieve images of the video data from the video data storage unit 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data storage unit 230 can be the raw video data to be encoded.

[0239] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra-frame prediction unit 226. The mode selection unit 202 may include additional functional units that perform video prediction based on other prediction modes. As an example, the mode selection unit 202 may include a palette unit, an intra-block copy unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, a combined inter-frame intra-frame prediction (CIIP) unit, etc.

[0240] Mode selection unit 202 typically coordinates multiple coding paths to test combinations of coding parameters and the rate-distortion values ​​obtained for such combinations. Coding parameters may include segmenting the CTU into CUs, the prediction mode for the CUs, the transformation type of the residual data for the CUs, and the quantization parameters for the residual data for the CUs. Mode selection unit 202 can ultimately select a combination of coding parameters that yields a better rate-distortion value than other tested combinations.

[0241] The video encoder 200 can segment images retrieved from the video data storage 230 into a series of CTUs and encapsulate one or more CTUs within a slice. The mode selection unit 202 can segment the CTUs of the image according to a tree structure (such as the QTBT structure or quadtree structure of HEVC described above). As mentioned above, the video encoder 200 can form one or more CUs by segmenting CTUs according to a tree structure. Such CUs can also generally be referred to as "video blocks" or "blocks".

[0242] Typically, mode selection unit 202 also controls its components (e.g., motion estimation unit 222, motion compensation unit 224, and intra-prediction unit 226) to generate prediction blocks for the current block (e.g., the current CU, or the overlapping portion of PU and TU in HEVC). To perform inter-frame prediction for the current block, motion estimation unit 222 may perform a motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously decoded pictures stored in DPB 218). Specifically, motion estimation unit 222 may calculate values ​​representing the similarity between a potential reference block and the current block, for example, based on sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), etc. Motion estimation unit 222 may typically perform these calculations using sample-by-sample differences between the current block and the considered reference blocks. Motion estimation unit 222 may identify the reference block with the lowest value obtained from these calculations, indicating the reference block that most closely matches the current block.

[0243] Motion estimation unit 222 can generate one or more motion vectors (MVs) that define the position of a reference block in a reference image relative to the position of the current block in the current image. Motion estimation unit 222 can then provide the motion vectors to motion compensation unit 224. For example, for unidirectional inter-frame prediction, motion estimation unit 222 can provide a single motion vector, while for bidirectional inter-frame prediction, motion estimation unit 222 can provide two motion vectors. Motion compensation unit 224 can then use the motion vectors to generate prediction blocks. For example, motion compensation unit 224 can use the motion vectors to retrieve data for the reference blocks. As another example, if the motion vectors have fractional-sample precision, motion compensation unit 224 can interpolate the values ​​used for the prediction blocks according to one or more interpolation filters. Furthermore, for bidirectional inter-frame prediction, motion compensation unit 224 can retrieve data for two reference blocks identified by corresponding motion vectors and combine the retrieved data, for example, by per-sample averaging or weighted averaging.

[0244] As another example, for intra-prediction or intra-prediction decoding, intra-prediction unit 226 can generate a prediction block based on samples adjacent to the current block. For example, in directional mode, intra-prediction unit 226 can typically mathematically combine the values ​​of adjacent samples and fill these calculated values ​​across the current block in a defined direction to generate a prediction block. As another example, in DC mode, intra-prediction unit 226 can calculate the average of the adjacent samples of the current block and generate a prediction block to include the obtained average for each sample of the prediction block.

[0245] Mode selection unit 202 provides a prediction block to residual generation unit 204. Residual generation unit 204 receives the original, uncoded version of the current block from video data memory 230 and the prediction block from mode selection unit 202. Residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines the residual block for the current block. In some examples, residual generation unit 204 may also determine the difference between sample values ​​in the residual block to generate the residual block using residual differential pulse code modulation (RDPCM). In some examples, one or more subtractor circuits performing binary subtraction may be used to form residual generation unit 204.

[0246] In the example where mode selection unit 202 divides a CU into PUs, each PU can be associated with a luma prediction unit and a corresponding chroma prediction unit. Video encoder 200 and video decoder 300 can support PUs of various sizes. As noted above, the size of a CU can refer to the size of the luma decoding block of the CU, and the size of a PU can refer to the size of the luma prediction unit of the PU. Assuming a particular CU size is 2Nx2N, video encoder 200 can support PU sizes of 2Nx2N or NxN for intra-frame prediction, and 2Nx2N, 2NxN, Nx2N, NxN, or similar symmetrical PU sizes for inter-frame prediction. Video encoder 200 and video decoder 300 can also support asymmetric segmentation for PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter-frame prediction.

[0247] In an example where the mode selection unit 202 does not further divide the CU into PUs, each CU can be associated with a luminance decoding block and a corresponding chrominance decoding block. As mentioned above, the size of the CU can refer to the size of the luminance decoding block of the CU. The video encoder 200 and the video decoder 300 can support CU sizes of 2Nx2N, 2NxN, or Nx2N.

[0248] For other video decoding techniques (such as block-based copy mode decoding, affine mode decoding, and linear model (LM) mode decoding, mode selection unit 202 generates a prediction block for the current block being encoded via a corresponding unit associated with the decoding technique. In some examples (such as palette mode decoding), mode selection unit 202 may not generate a prediction block, but instead generate syntax elements indicating how the block should be reconstructed based on the selected palette. In such a mode, mode selection unit 202 can provide these syntax elements to entropy coding unit 220 for encoding.

[0249] As described above, the residual generation unit 204 receives video data for the current block and the corresponding prediction block. Then, the residual generation unit 204 generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block.

[0250] Transform processing unit 206 applies one or more transformations to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 can apply various transformations to the residual block to form the transform coefficient block. For example, transform processing unit 206 can apply a discrete cosine transform (DCT), direction transform, Karhunen-Loeve transform (KLT), or conceptually similar transformations to the residual block. In some examples, transform processing unit 206 can perform multiple transformations on the residual block, such as primary and secondary transformations (e.g., rotation transformations). In some examples, transform processing unit 206 does not apply any transformations to the residual block.

[0251] Quantization unit 208 can quantize the transform coefficients in the transform coefficient block to produce a quantized transform coefficient block. In some examples, quantization unit 208 can quantize the transform coefficients of the transform coefficient block based on the quantization parameter (QP) value associated with the current block. Video encoder 200 (e.g., via mode selection unit 202) can adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may cause information loss, and therefore, the quantized transform coefficients may have lower accuracy compared to the original transform coefficients produced by transform processing unit 206.

[0252] Quantization unit 208 can perform the techniques of this disclosure to quantize the transform coefficients, such as regarding Figures 3A-7 The described technology. Specifically, the quantization unit 208 can perform operations related to... Figure 3B Scaling and classification unit 152 and quantization unit 142, Figure 4 The quantization offset parameter unit 154, grouping unit 156, and offset quantization units 158A-158P, and Figure 6 The functions described in the QOV calculation unit 174, QOV parameter calculation unit 176, QOV calculation unit 178, QOV index calculation unit 180, and QOV retrieval unit 182.

[0253] Quantization unit 208 can determine a set of quantization offset parameters for a scaled transform coefficient group for a video data block based on side information associated with the transform coefficient block. For example, quantization unit 208 can determine a set of quantization offset parameters for a scaled transform coefficient group within a sub-block (such as each 4x4 sub-block of the transform coefficient block) based on side information associated with the transform coefficient block. The quantization offset in the set of quantization offset parameters for the scaled transform coefficient group may not be constant. Instead, the quantization offset in the set of quantization offset parameters may vary according to the quantization interval.

[0254] Quantization unit 208 can perform any of the techniques disclosed in this disclosure to determine a set of quantization offset parameters for a scaled transform coefficient group within a video data block based on side information associated with the video data block. For example, the side information associated with a video data block may include any combination of one or more of the following: the slice type of the video data block (e.g., I, P, or B), residual data generated by intra-frame or inter-frame prediction of the video data block, the block size of the video data block (e.g., 4×4, 8×8, 16×16, or 32×32), and / or the luma or chroma components of the video data block. The side information associated with the video data block may also include side information associated with sub-blocks containing the scaled transform coefficient group, such as one or more of the following: the maximum size range of the scaled transform coefficients of the sub-block, the block size of the sub-block, the relative position of the sub-block within the transform coefficient block, etc.

[0255] In some examples, quantization unit 208 can determine a set of quantization offset parameters for a scaled transform coefficient set based on side information associated with the video data block. This set of quantization offset parameters can optimize the rate-distortion cost of quantizing the scaled transform coefficient set. For example, quantization unit 208 can use machine learning techniques (such as neural networks that can be trained on side information, scaled transform coefficient values, optimal rate-distortion cost of quantization, etc.) to determine the set of quantization offset parameters for the scaled transform coefficient set. Quantization unit 208 can use such neural networks to perform regression and / or classification methods to determine the set of quantization offset parameters for the scaled transform coefficient set based on side information associated with the video data block. Therefore, quantization unit 208 is able to determine the set of quantization offset parameters for the scaled transform coefficient set without using the bit cost estimate determined by entropy coding unit 220 and without deriving an index of the arithmetic decoding context for entropy coding of each particular non-zero transform coefficient.

[0256] Quantization unit 208 can quantize each scaled transform coefficient group of a video data block at least in part based on a set of quantization offset parameters to generate quantized transform coefficients for each sub-block of the video data block. As described above, because quantization unit 208 can determine the set of quantization offset parameters for each scaled transform coefficient group in the video data block, quantization unit 208 can use the set of quantization offset parameters associated with the corresponding scaled transform coefficient group to quantize each scaled transform coefficient group. Therefore, in some examples, quantization unit 208 is able to quantize scaled transform coefficients within the same scaled transform coefficient group in parallel. Furthermore, in some examples, quantization unit 208 is able to quantize multiple scaled transform coefficient groups for a video data block in parallel.

[0257] The inverse quantization unit 210 and the inverse transform processing unit 212 can apply inverse quantization and inverse transform, respectively, to the quantized transform coefficient block to reconstruct the residual block from the transform coefficient block. The reconstruction unit 214 can generate a reconstructed block corresponding to the current block based on the reconstructed residual block and the prediction block generated by the mode selection unit 202 (although potentially with some degree of distortion). For example, the reconstruction unit 214 can add samples of the reconstructed residual block to corresponding samples of the prediction block generated by the mode selection unit 202 to generate the reconstructed block.

[0258] Filter unit 216 can perform one or more filter operations on the reconstructed block. For example, filter unit 216 can perform a deblocking operation to reduce block artifacts along the edges of the CU. In some examples, the operation of filter unit 216 can be skipped.

[0259] The video encoder 200 stores the reconstructed blocks in the DPB 218. For example, in an example where the operation of the filter unit 216 is not required, the reconstruction unit 214 can store the reconstructed blocks in the DPB 218. In an example where the operation of the filter unit 216 is required, the filter unit 216 can store the filtered reconstructed blocks in the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 can retrieve a reference picture formed by the reconstructed (and potentially filtered) blocks from the DPB 218 to perform inter-frame prediction of blocks in subsequently encoded pictures. Additionally, the intra-frame prediction unit 226 can use the reconstructed blocks of the current picture in the DPB 218 to perform intra-frame prediction of other blocks in the current picture.

[0260] Typically, entropy coding unit 220 can entropy code syntax elements received from other functional components of video encoder 200. For example, entropy coding unit 220 can entropy code quantized transform coefficient blocks from quantization unit 208 to generate an encoded video bitstream based at least in part on the quantized transform coefficients for the video data blocks.

[0261] As another example, entropy coding unit 220 can entropy-encode predicted syntax elements (e.g., motion information for inter-frame prediction or intra-frame mode information for intra-frame prediction) from mode selection unit 202. Entropy coding unit 220 can perform one or more entropy coding operations on syntax elements, another example of video data, to generate entropy-encoded data. For example, entropy coding unit 220 can perform context-adaptive variable-length decoding (CAVLC), CABAC, variable-to-variable (V2V) length decoding, syntax-based context-adaptive binary arithmetic decoding (SBAC), probabilistic interval partitioned entropy (PIPE) decoding, exponential Golomb coding, or another type of entropy coding operation on the data. In some examples, entropy coding unit 220 can operate in a bypass mode where syntax elements are not entropy-encoded.

[0262] The video encoder 200 can output a bitstream that includes entropy-encoded syntax elements required for reconstructing slices or blocks of images. Specifically, the entropy coding unit 220 can output a bitstream.

[0263] The above operations are described in relation to the blocks. Such a description should be understood as referring to the operations used for the luma decoding block and / or the chroma decoding block. As mentioned above, in some examples, the luma decoding block and the chroma decoding block are the luma and chroma components of the CU. In some examples, the luma decoding block and the chroma decoding block are the luma and chroma components of the PU.

[0264] In some examples, it is not necessary to repeat the operations performed for the luma decoding block for the chroma decoding block. As an example, it is not necessary to repeat the operations used to identify the motion vector (MV) and reference image for the luma decoding block to identify the MV and reference image for the chroma block. Specifically, the MV for the luma decoding block can be scaled to determine the MV for the chroma block, and the reference image can be the same. As another example, the intra-frame prediction process can be the same for both the luma and chroma decoding blocks.

[0265] Video encoder 200 represents an example of a device configured to encode video data, the device including: a memory configured to store the video data; and one or more processing units implemented in circuitry and configured to: determine a set of quantization offset parameters for a scaled transform coefficient group of the video data block based on side information associated with the video data block; quantize the scaled transform coefficient group of the video data block at least in part based on the set of quantization offset parameters to generate quantized transform coefficients for the video data block; and generate an encoded video bitstream at least in part based on the quantized transform coefficients for the video data block.

[0266] Figure 9 This is a block diagram illustrating an example video decoder 300 capable of performing the techniques described herein. Figure 9 This disclosure is provided for illustrative purposes and does not limit the techniques illustrated and described in this disclosure in a general manner. For illustrative purposes, this disclosure describes a video decoder 300 based on VVC (ITU-T H.266, in deployment) and HEVC (ITU-T H.265) technologies. However, the technologies of this disclosure can be implemented by video decoding devices configured for other video decoding standards.

[0267] exist Figure 9 In the example, the video decoder 300 includes a decoded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 134. Any or all of the CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filter unit 312, and DPB 134 can be implemented in one or more processors or in processing circuitry. For example, units of the video decoder 300 can be implemented as one or more circuit or logic elements as part of hardware circuitry, or as part of a processor, ASIC, or FPGA. Furthermore, the video decoder 300 may include additional or alternative processors or processing circuitry to perform these and other functions.

[0268] The prediction processing unit 304 includes a motion compensation unit 316 and an intra-prediction unit 318. The prediction processing unit 304 may include an addition unit that performs predictions based on other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, an intra-block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, a combined inter-intra-prediction (CIIP) unit, etc. In other examples, the video decoder 300 may include more, fewer, or different functional components.

[0269] CPB memory 320 can store video data to be decoded by components of video decoder 300, such as encoded video bitstreams. For example, it can be stored from computer-readable medium 110 ( Figure 1 The video data stored in the CPB memory 320 is obtained. The CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Additionally, the CPB memory 320 may store video data other than the syntax elements of the decoded picture, such as temporary data representing the output from various units of the video decoder 300. The DPB 314 typically stores decoded pictures that the video decoder 300 may output and / or use as reference video data when decoding subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and DPB 314 may be formed of any of a variety of memory devices, such as DRAM, including SDRAM, MRAM, RRAM, or other types of memory devices. The CPB memory 320 and DPB 314 may be provided by the same memory device or separate memory devices. In various examples, the CPB memory 320 may be on-chip with other components of the video decoder 300, or off-chip relative to those components.

[0270] Alternatively or concurrently, in some examples, the video decoder 300 can be derived from the memory 120 ( Figure 1 The decoded video data is retrieved. In other words, memory 120 can utilize CPB memory 320 to store data, as discussed above. Similarly, when some or all of the functions of video decoder 300 are implemented using software to be executed by the processing circuitry of video decoder 300, memory 120 can store instructions to be executed by video decoder 300.

[0271] It shows Figure 9 The various units shown help to understand the operations performed by the video decoder 300. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Similar to... Figure 8Fixed-function circuits refer to circuits that provide a specific function and are pre-configured regarding the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functionality in terms of the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is typically immutable. In some examples, one or more of these units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units may be integrated circuits.

[0272] The video decoder 300 may include an ALU, EFU, digital circuitry, analog circuitry, and / or a programmable core formed by programmable circuitry. In an example where the operation of the video decoder 300 is performed by software executing on the programmable circuitry, on-chip or off-chip memory may store instructions (e.g., object code) of the software received and executed by the video decoder 300.

[0273] Entropy decoding unit 302 can receive encoded video data from the CPB and perform entropy decoding on the video data to reconstruct the syntax elements. Prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, and filter unit 312 can generate decoded video data based on the syntax elements extracted from the bitstream.

[0274] Typically, the video decoder 300 reconstructs the image on a block-by-block basis. The video decoder 300 can perform the reconstruction operation on each block individually (where the block currently being reconstructed (i.e., decoded) can be referred to as the "current block").

[0275] Entropy decoding unit 302 can entropy decode the syntax elements of the quantized transform coefficients that define the quantized transform coefficient block, as well as transform information such as quantization parameters (QPs) and / or transform mode indications. Inverse quantization unit 306 can use the QPs associated with the quantized transform coefficient block to determine the degree of quantization, and similarly, determine the degree of inverse quantization to be applied by inverse quantization unit 306. Inverse quantization unit 306 can, for example, perform a bitwise left shift operation to inverse quantize the quantized transform coefficients. Inverse quantization unit 306 can thus form a transform coefficient block including the transform coefficients.

[0276] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 can apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 can apply the inverse DCT, inverse integer transform, inverse Karhunen-Loeve transform (KLT), inverse rotation transform, inverse direction transform, or another inverse transform to the transform coefficient block.

[0277] Furthermore, the prediction processing unit 304 generates a prediction block based on the prediction information syntax elements entropy-decoded by the entropy decoding unit 302. For example, if the prediction information syntax elements indicate that the current block is inter-frame predicted, the motion compensation unit 316 can generate the prediction block. In this case, the prediction information syntax elements may indicate the reference picture from which the reference block is to be retrieved in the DPB 314, and a motion vector identifying the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion compensation unit 316 can typically be configured with respect to the motion compensation unit 224 ( Figure 8 The method described is basically similar to the way the inter-frame prediction process is performed.

[0278] As another example, if the prediction information syntax element indicates that the current block is intra-predicted, then intra-prediction unit 318 can generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Again, intra-prediction unit 318 can typically be configured with respect to intra-prediction unit 226 ( Figure 8 The intra-prediction process is performed in a manner substantially similar to that described above. The intra-prediction unit 318 can retrieve data from neighboring samples of the current block from the DPB 314.

[0279] Reconstruction unit 310 can reconstruct the current block using the prediction block and the residual block. For example, reconstruction unit 310 can reconstruct the current block by adding the samples of the residual block to the corresponding samples of the prediction block.

[0280] Filter unit 312 can perform one or more filter operations on the reconstructed block. For example, filter unit 312 can perform a deblocking operation to reduce block artifacts along the edges of the reconstructed block. The operation of filter unit 312 is not necessarily performed in all examples.

[0281] The video decoder 300 can store the reconstructed blocks in the DPB 314. For example, in an example where the operation of the filter unit 312 is not performed, the reconstruction unit 310 can store the reconstructed blocks in the DPB 314. In an example where the operation of the filter unit 312 is performed, the filter unit 312 can store the filtered reconstructed blocks in the DPB 314. As discussed above, the DPB 314 can provide reference information (such as samples of the current image for intra-frame prediction and samples of previously decoded images for subsequent motion compensation) to the prediction processing unit 304. Furthermore, the video decoder 300 can output decoded images (e.g., decoded video) from the DPB 314 for use in applications such as... Figure 1 The subsequent presentation on display devices such as display device 118.

[0282] In this way, video decoder 300 represents an example of a video decoding device, which includes: a memory configured to store video data; and one or more processing units implemented in circuitry and configured to perform the techniques of this disclosure.

[0283] Figure 10 This is a flowchart illustrating an example method for encoding the current block. The current block may include the current CU. Although regarding video encoder 200 ( Figure 1 and Figure 8 The description is provided, but it should be understood that other devices can be configured to perform the same actions. Figure 10 Similar to the method.

[0284] In this example, the video encoder 200 initially predicts the current block (350). For example, the video encoder 200 may form a prediction block for the current block. Then, the video encoder 200 may compute a residual block for the current block (352). To compute the residual block, the video encoder 200 may compute the difference between the original unencoded block and the prediction block for the current block. Then, the video encoder 200 may transform the residual block and quantize the transform coefficients of the residual block (354). Next, the video encoder 200 may scan the quantized transform coefficients of the residual block (356). Specifically, the video encoder 200 may perform the techniques of this disclosure, including: determining a set of quantization offset parameters for a scaled transform coefficient group for the video data block based on side information associated with the video data block; quantizing the scaled transform coefficient group of the video data block at least in part based on the set of quantization offset parameters to generate quantized transform coefficients of the video data block; and generating an encoded video bitstream at least in part based on the quantized transform coefficients of the video data block.

[0285] During or after scanning, the video encoder 200 can entropy encode the transform coefficients (358). For example, the video encoder 200 can use CAVLC or CABAC to encode the transform coefficients. The video encoder 200 can then output the entropy-encoded data of the block (360).

[0286] Figure 11 This is a flowchart illustrating an example method for decoding the current block of video data. The current block may include the current CU. Although regarding video decoder 300 ( Figure 1 and Figure 9 The description is provided, but it should be understood that other devices can be configured to perform the same actions. Figure 11 Similar to the method.

[0287] The video decoder 300 can receive entropy-encoded data for the current block (such as entropy-encoded prediction information and entropy-encoded data for the transform coefficients of the residual block corresponding to the current block) (370). The video decoder 300 can entropy decode the entropy-encoded data to determine the prediction information for the current block and reproduce the transform coefficients of the residual block (372). The video decoder 300 can predict the current block, for example, using an intra-frame or inter-frame prediction mode indicated by the prediction information for the current block (374), to compute a prediction block for the current block. The video decoder 300 can then perform an inverse scan on the reproduced transform coefficients (376) to create a block of quantized transform coefficients. The video decoder 300 can then inverse quantize the transform coefficients and apply the inverse transform to the transform coefficients to produce a residual block (378). Finally, the video decoder 300 can decode the current block by combining the prediction block and the residual block (380).

[0288] Figure 12 This is a flowchart illustrating a method for encoding video data according to the technology of this disclosure. Although regarding video encoder 200 ( Figure 1 and Figure 8 The description is provided, but it should be understood that other devices can be configured to perform the same actions. Figure 12 Similar to the method.

[0289] like Figure 12As shown, video encoder 200 (e.g., quantization unit 208) can determine a set of quantization offset parameters (402) for a scaled transform coefficient group for a video data block based on side information associated with the video data block. In some examples, to determine the set of quantization offset parameters for a scaled transform coefficient group for a video data block, video encoder 200 can classify the scaled transform coefficient group at least in part based on the side information associated with the video data block, and select the set of quantization offset parameters for the scaled transform coefficient group from multiple sets of quantization offset parameters based on the classification of the scaled transform coefficient group.

[0290] In some examples, in order to classify the scaled transform coefficient groups, the video encoder 200 may determine the index of the scaled transform coefficient groups based at least in part on the side information associated with the video data blocks. In order to select the set of quantization offset parameters for the scaled transform coefficient groups from multiple sets of quantization offset parameters based on the classification of the scaled transform coefficient groups, the video encoder 200 may use the determined index to index multiple sets of quantization offset parameters to select the set of quantization offset parameters for the scaled transform coefficient groups.

[0291] In some examples, the scaled transform coefficient set includes scaled transform coefficients for sub-blocks of the video data block, and the side information associated with the video data block includes the position of the sub-blocks within the video data block and the block size of the video data block. In some examples, the classification of the scaled transform coefficient set is based at least in part on the side information associated with the video data block, rather than on the values ​​of the scaled transform coefficients for the sub-blocks of the data block.

[0292] In some examples, in order to determine the set of quantization offset parameters for a scaled transform coefficient group for a video data block, the video encoder 200 may determine a parameter set with a smaller dimension than the quantization offset parameter set based on the side information associated with the video data block, and may determine the set of quantization offset parameters for the scaled transform coefficient group based on the parameter set.

[0293] In some examples, to determine the set of quantization offset parameters for a scaled transform coefficient group used for a video data block, the video encoder 200 may use a neural network to determine the set of quantization offset parameters for the scaled transform coefficient group used for the video data block. In some examples, to determine the set of quantization offset parameters for a scaled transform coefficient group used for a video data block, the video encoder 200 may use at least one of classification or regression techniques to determine the set of quantization offset parameters for the scaled transform coefficient group used for the video data block.

[0294] The video encoder 200 (e.g., quantization unit 208) can quantize the scaled transform coefficient set of the video data block at least in part based on the quantization offset parameter set to generate the quantized transform coefficients (404) of the video data block.

[0295] The video encoder 200 (e.g., entropy coding unit 220) can generate an encoded video bitstream (406) based at least in part on the quantized transform coefficients of the video data blocks.

[0296] In some examples, the video encoder 200 may divide a plurality of scaled transform coefficients into a plurality of scaled transform coefficient groups associated with sub-blocks of a video data block, wherein the plurality of scaled transform coefficient groups include scaled transform coefficient groups for sub-blocks of the video data block. In some examples, to determine the set of quantization offset parameters for the scaled transform coefficient groups of the video data block, the video encoder 200 may determine a corresponding set of quantization offset parameters for each of the plurality of scaled transform coefficient groups, wherein quantizing the scaled transform coefficient groups of the sub-blocks of the video data block includes quantizing each of the plurality of scaled transform coefficient groups based on the corresponding set of quantization offset parameters.

[0297] In some examples, in order to quantize the scaled transform coefficients of sub-blocks of a video data block, the video encoder 200 may determine the corresponding quantization offset parameter from the set of quantization offset parameters for each scaled transform coefficient, and may quantize each scaled transform coefficient at least in part based on the corresponding quantization offset parameter.

[0298] In some examples, multiple sets of quantization offset parameters include a quantization offset vector.

[0299] Illustrative examples of the first aspect of this disclosure include:

[0300] Aspect 1: A method for decoding video data, the method comprising any combination of the techniques described in this disclosure.

[0301] Aspect 2: A method for decoding video data, the method comprising: determining a plurality of quantization offset vectors for a plurality of transform coefficients for a current block of the video data based at least in part on scaled transform coefficients for a current block and side information associated with the current block; and quantizing the transform coefficients for the current block based at least in part on the plurality of quantization offset vectors to generate quantized transform coefficients for the current block.

[0302] Aspect 3: The method according to aspect 2 further includes: splitting the transform coefficients into multiple transform coefficient subgroups, wherein determining the multiple quantization offset vectors for the multiple transform coefficients of the current block includes: determining the quantization offset vector for each of the multiple transform coefficient subgroups.

[0303] Aspect 4: The method according to any combination of aspects 2 and 3, wherein determining the plurality of quantization vectors further includes: determining a plurality of index values; using the plurality of index values ​​to index into a quantization offset vector table to determine the plurality of quantization offset vectors.

[0304] Aspect 5: According to the method of aspect 4, the determination of the plurality of index values ​​is based at least in part on the values ​​of the plurality of transform coefficients in the transform coefficient group.

[0305] Aspect 6: The method according to any combination of aspects 4 and 5, wherein the determination of the plurality of index values ​​is based at least in part on the size of the transform coefficient block and the position of the transform coefficient group within the transform coefficient block.

[0306] Aspect 7: The method according to any combination of aspects 4-6, wherein determining the plurality of index values ​​is based at least in part on whether the current block is part of an intra-slice.

[0307] Aspect 8: The method according to any combination of aspects 2-7, wherein determining the plurality of quantization offset vectors further comprises: determining one or more quantization offset vectors among the plurality of quantization offset vectors that do not have the scaled transform coefficients for the current block; and determining other quantization offset vectors based on the scaled transform coefficients for the current block.

[0308] Aspect 9: The method according to any combination of aspects 2-8, wherein determining the plurality of quantization offset vectors for the plurality of transform coefficients for the current block of video data comprises: defining the quantization offset vectors using a smaller dimension parameter vector.

[0309] Aspect 10: The method according to any combination of aspects 2-9, wherein the edge information includes one or more of the following: slice type, block size, prediction type, or an indication of whether the current block to be quantized includes a luminance component or a chrominance component.

[0310] Aspect 11: The method according to any combination of aspects 2-10, wherein quantizing the transform coefficients for the current block of video data comprises: quantizing the first transform coefficients for the current block in parallel with the second transform coefficients for the current block of transform coefficients.

[0311] Aspect 12: The method according to any combination of aspects 2-11, wherein determining the plurality of quantization offset vectors for the plurality of transform coefficients of the current block comprises: determining an estimated change in the number of bits used for entropy decoding of the quantized transform coefficients of the current block of data based on a change in a single element of the quantized transform coefficients for the current block.

[0312] Aspect 13: According to the method of aspect 12, wherein determining the estimated change in the number of bits for entropy decoding of the quantized transform coefficients of the current block of data based on the change of the individual element of the quantized transform coefficients for the current block comprises: determining the plurality of quantization offset vectors of the plurality of quantization offset vectors of the plurality of transform coefficients for the current block using the same estimation rule computed once.

[0313] Aspect 14: The method according to any combination of aspects 2-13, wherein determining the plurality of quantization offset vectors for the plurality of transform coefficients of the current block comprises: determining the plurality of quantization offset vectors for the plurality of transform coefficients of the current block without deriving an index of an arithmetic decoding context for entropy decoding of the particular non-zero transform coefficient for each particular non-zero transform coefficient.

[0314] Aspect 15: The method according to any combination of aspects 2-14, wherein determining the plurality of quantization offset vectors for the plurality of transform coefficients of the current block comprises: optimizing the values ​​of the plurality of quantization offset vectors using at least one of statistical techniques or machine learning techniques.

[0315] Aspect 16: The method according to aspect 15, wherein the at least one of the statistical techniques or machine learning techniques includes at least one of the classification techniques or regression techniques.

[0316] Aspect 17: The method according to aspect 16, wherein the regression technique includes a general regression technique.

[0317] Aspect 18: The method according to any combination of aspects 16 and 17, wherein the classification technique includes a classification tree.

[0318] Aspect 19: The method according to any combination of aspects 2-18, wherein decoding includes decoding.

[0319] Aspect 20: The method according to any combination of aspects 2-18, wherein decoding includes encoding.

[0320] Aspect 21: An apparatus for decoding video data, the apparatus comprising one or more units for performing one or more of the methods according to any combination of aspects 1-20.

[0321] Aspect 22: The device according to aspect 21, wherein the one or more units include one or more processors implemented in a circuit.

[0322] Aspect 23: The device according to any combination of aspects 21 and 22 further includes: a memory for storing the video data.

[0323] Aspect 24: The device according to any combination of 21-23 further includes: a display configured to display decoded video data.

[0324] Aspect 25: The device according to any combination of aspects 21-24, wherein the device includes one or more of the following: camera, computer, mobile device, broadcast receiver device, or set-top box.

[0325] Aspect 26: The device according to any combination of aspects 21-25, wherein the device includes a video decoder.

[0326] Aspect 27: The device according to any combination of aspects 21-26, wherein the device includes a video encoder.

[0327] Aspect 28: A computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to perform the method according to any combination of aspects 1-20.

[0328] Illustrative examples of the second aspect of this disclosure include:

[0329] Aspect 1: A method for encoding video data, the method comprising: determining a set of quantization offset parameters for a scaled transform coefficient set for the video data block based on side information associated with the video data block; quantizing the scaled transform coefficient set for the video data block at least in part based on the set of quantization offset parameters to generate quantized transform coefficients for the video data block; and generating an encoded video bitstream at least in part based on the quantized transform coefficients for the video data block.

[0330] Aspect 2: According to the method of aspect 1, wherein determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block includes: selecting the set of quantization offset parameters for the scaled transform coefficient group from a plurality of quantization offset parameter sets based on the side information associated with the video data block.

[0331] Aspect 3: According to the method of aspect 2, wherein selecting the set of quantization offset parameters for the scaled transform coefficient group from the plurality of quantization offset parameter sets comprises: determining an index associated with the scaled transform coefficient group based at least in part on the side information associated with the video data block; and using the index associated with the scaled transform coefficient group to index to the plurality of quantization offset parameter sets to select the set of quantization offset parameters for the scaled transform coefficient group.

[0332] Aspect 4: The method according to aspect 2 or 3, wherein: the scaled transform coefficient set includes scaled transform coefficients for sub-blocks of the video data block; and the edge information associated with the video data block includes the position of the sub-block within the video data block and the block size of the video data block.

[0333] Aspect 5: According to the method of aspect 4, wherein selecting the set of quantization offset parameters for the scaled transform coefficient group from the plurality of quantization offset parameter sets comprises: selecting the set of quantization offset parameters for the scaled transform coefficient group from the plurality of quantization offset parameter sets without using the scaled transform coefficient values ​​of the scaled transform coefficient group for the sub-block of the video data block.

[0334] Aspect 6: According to the method of aspect 1, wherein determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block includes: determining a set of parameters for parameterizing the set of quantization offset parameters based on the edge information associated with the video data block, the set of parameters having a smaller size compared to the set of quantization offset parameters; and determining the set of quantization offset parameters for the scaled transform coefficient group based on the smaller size of the set of parameters.

[0335] Aspect 7: The method according to any one of Aspects 1 to 5, wherein determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block further comprises: using a neural network and based on the side information associated with the video data block to determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block.

[0336] Aspect 8: According to the method of aspect 7, determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block further includes: using at least one of classification techniques or regression techniques to determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block.

[0337] Aspect 9: The method according to any one of Aspects 1 to 8 further includes: dividing a plurality of scaled transform coefficients of the video data block into a plurality of scaled transform coefficient groups associated with sub-blocks of the video data block, wherein the plurality of scaled transform coefficient groups includes scaled transform coefficient groups for the video data block; wherein determining the quantization offset parameter set for the scaled transform coefficient groups for the video data block includes: determining a corresponding quantization offset parameter set for each of the plurality of scaled transform coefficient groups; and wherein quantizing the scaled transform coefficient groups for the video data block includes: quantizing each of the plurality of scaled transform coefficient groups based on the corresponding quantization offset parameter set.

[0338] Aspect 10: The method according to any one of Aspects 1 to 9, wherein quantizing the scaled transform coefficient group for the video data block further comprises: determining a corresponding quantization offset parameter from the quantization offset parameter set for each scaled transform coefficient of the scaled transform coefficient group; and quantizing each scaled transform coefficient of the scaled transform coefficient group based at least in part on the corresponding quantization offset parameter.

[0339] Aspect 11: The method according to any one of Aspects 1 to 10, wherein the side information includes one or more of the following: the slice type of the video data block, the block size of the video data block, or an indication of whether the video data block includes a luminance component or a chrominance component.

[0340] Aspect 12: The method according to any one of Aspects 1 to 11, wherein determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block comprises: determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block without using one or more bit cost estimates determined via entropy decoding.

[0341] Aspect 13: The method according to any one of Aspects 1 to 12, wherein the set of quantization offset parameters includes a quantization offset vector.

[0342] Aspect 14: An apparatus for encoding video data, the apparatus comprising: a memory; and processing circuitry in communication with the memory, the processing circuitry being configured to: determine a set of quantization offset parameters for a scaled transform coefficient set for the video data block based on side information associated with the video data block; quantize the scaled transform coefficient set for the video data block at least in part based on the set of quantization offset parameters to generate quantized transform coefficients for the video data block; and generate an encoded video bitstream at least in part based on the quantized transform coefficients for the video data block.

[0343] Aspect 15: The apparatus according to aspect 14, wherein, in order to determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block, the processing circuitry is further configured to: select the set of quantization offset parameters for the scaled transform coefficient group from a plurality of quantization offset parameter sets based on the side information associated with the video data block.

[0344] Aspect 16: The apparatus according to aspect 15, wherein, in order to select the set of quantization offset parameters for the scaled transform coefficient group from the plurality of quantization offset parameter sets, the processing circuitry is further configured to: determine an index associated with the scaled transform coefficient group based at least in part on the side information associated with the video data block; and use the index associated with the scaled transform coefficient group to index to the plurality of quantization offset parameter sets to select the set of quantization offset parameters for the scaled transform coefficient group.

[0345] Aspect 17: The device according to aspect 15 or 16, wherein: the scaled transform coefficient set includes scaled transform coefficients for sub-blocks of the video data block; and the edge information associated with the video data block includes the position of the sub-block within the video data block and the block size of the video data block.

[0346] Aspect 18: The apparatus according to aspect 17, wherein, in order to select the set of quantization offset parameters for the scaled transform coefficient group from the plurality of quantization offset parameter sets, the processing circuit is further configured to: select the set of quantization offset parameters for the scaled transform coefficient group from the plurality of quantization offset parameter sets without using the scaled transform coefficient values ​​of the scaled transform coefficient group for the sub-block of the video data block.

[0347] Aspect 19: The apparatus according to aspect 14, wherein, in order to determine the set of quantization offset parameters for the scaled transform coefficient group of the video data block, the processing circuitry is further configured to: determine a set of parameters for parameterizing the set of quantization offset parameters based on the side information associated with the video data block, the set of parameters having a smaller size than the set of quantization offset parameters; and determine the set of quantization offset parameters for the scaled transform coefficient group based on the smaller size of the set of parameters.

[0348] Aspect 20: The apparatus according to any one of aspects 14 to 18, wherein, in order to determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block, the processing circuitry is further configured to: use a neural network and determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block based on the side information associated with the video data block.

[0349] Aspect 21: The apparatus according to aspect 20, wherein, in order to determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block, the processing circuitry is further configured to: determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block using at least one of a classification technique or a regression technique.

[0350] Aspect 22: The apparatus according to any one of Aspects 14 to 21, wherein the processing circuitry is further configured to: divide a plurality of scaled transform coefficients of the video data block into a plurality of scaled transform coefficient groups associated with sub-blocks of the video data block, wherein the plurality of scaled transform coefficient groups includes the scaled transform coefficient groups for the video data block; wherein, in order to determine the set of quantization offset parameters for the scaled transform coefficient groups for the video data block, the processing circuitry is further configured to: determine a corresponding set of quantization offset parameters for each of the plurality of scaled transform coefficient groups; and wherein quantizing the scaled transform coefficient groups for the video data block includes: quantizing each of the plurality of scaled transform coefficient groups based on the corresponding set of quantization offset parameters.

[0351] Aspect 23: The apparatus according to any one of Aspects 14 to 22, wherein, in order to quantize the scaled transform coefficient group for the video data block, the processing circuitry is further configured to: determine a corresponding quantization offset parameter from the quantization offset parameter set for each scaled transform coefficient of the scaled transform coefficient group; and quantize each scaled transform coefficient of the scaled transform coefficient group based at least in part on the corresponding quantization offset parameter.

[0352] Aspect 24: The device according to any one of Aspects 14 to 23, wherein the side information includes one or more of the following: the slice type of the video data block, the block size of the video data block, or an indication of whether the video data block includes a luminance component or a chrominance component.

[0353] Aspect 25: The apparatus according to any one of aspects 14 to 24, wherein, in order to determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block, the processing circuitry is further configured to determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block without using one or more bit cost estimates determined via entropy decoding.

[0354] Aspect 26: The device according to any one of aspects 14 to 25, wherein the plurality of quantization offset parameter sets includes a quantization offset vector.

[0355] Aspect 27: The device according to any one of aspects 14 to 26, wherein the device includes one or more of the following: a camera, a computer, or a mobile device.

[0356] Aspect 28: An apparatus for encoding video data, the apparatus comprising: unit for determining a set of quantization offset parameters for a scaled transform coefficient group for the video data block based on side information associated with the video data block; unit for quantizing the scaled transform coefficient group for the video data block at least in part based on the set of quantization offset parameters to generate quantized transform coefficients for the video data block; and unit for generating an encoded video bitstream at least in part based on the quantized transform coefficients for the video data block.

[0357] Aspect 29: The apparatus according to aspect 28, wherein the unit for determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block further comprises: a unit for selecting the set of quantization offset parameters for the scaled transform coefficient group from a plurality of quantization offset parameter sets based on the side information associated with the video data block.

[0358] Aspect 30: The apparatus according to aspect 29, wherein the unit for selecting the set of quantization offset parameters for the scaled transform coefficient group from the plurality of quantization offset parameter sets further comprises: a unit for determining an index associated with the scaled transform coefficient group based at least in part on the side information associated with the video data block; and a unit for indexing to the plurality of quantization offset parameter sets using the index associated with the scaled transform coefficient group to select the set of quantization offset parameters for the scaled transform coefficient group.

[0359] Aspect 31: The apparatus according to aspect 29 or 30, wherein: the scaled transform coefficient set includes scaled transform coefficients for sub-blocks of the video data block; and the edge information associated with the video data block includes the position of the sub-block within the video data block and the block size of the video data block.

[0360] Aspect 32: The apparatus according to aspect 31, wherein the unit for selecting the set of quantization offset parameters for the scaled transform coefficient group from the plurality of quantization offset parameter sets further comprises: a unit for selecting the set of quantization offset parameters for the scaled transform coefficient group from the plurality of quantization offset parameter sets without using the scaled transform coefficient values ​​of the scaled transform coefficient group for the sub-block of the video data block.

[0361] Aspect 33: The apparatus according to aspect 28, wherein the unit for determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block further comprises: a unit for determining a set of parameters for parameterizing the set of quantization offset parameters based on the side information associated with the video data block, the set of parameters having a smaller size than the set of quantization offset parameters; and a unit for determining the set of quantization offset parameters for the scaled transform coefficient group based on the smaller size of the set of parameters.

[0362] Aspect 34: The apparatus according to any one of aspects 28 to 32, wherein the unit for determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block further comprises: a unit for determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block using a neural network and based on the side information associated with the video data block.

[0363] Aspect 35: The apparatus according to aspect 34, wherein the unit for determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block comprises: a unit for determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block using at least one of a classification technique or a regression technique.

[0364] Aspect 36: The apparatus according to any one of Aspects 28 to 35 further includes: a unit for dividing a plurality of scaled transform coefficients for the video data block into a plurality of scaled transform coefficient groups associated with sub-blocks of the video data block, wherein the plurality of scaled transform coefficient groups includes the scaled transform coefficient groups for the video data block; wherein the unit for determining the set of quantization offset parameters for the scaled transform coefficient groups for the video data block includes: a unit for determining a corresponding set of quantization offset parameters for each of the plurality of scaled transform coefficient groups; and wherein the unit for quantizing the scaled transform coefficient groups for the video data block further includes: a unit for quantizing each of the plurality of scaled transform coefficient groups based on the corresponding set of quantization offset parameters.

[0365] Aspect 37: The apparatus according to any one of Aspects 28 to 36, wherein the unit for quantizing the scaled transform coefficient group for the video data block further comprises: a unit for determining a corresponding quantization offset parameter from the quantization offset parameter set for each scaled transform coefficient of the scaled transform coefficient group; and a unit for quantizing each scaled transform coefficient of the scaled transform coefficient group based at least in part on the corresponding quantization offset parameter.

[0366] Aspect 38: The apparatus according to any one of aspects 28 to 37, wherein the side information includes one or more of the following: the slice type of the video data block, the block size of the video data block, or an indication of whether the video data block includes a luminance component or a chrominance component.

[0367] Aspect 39: The apparatus according to any one of Aspects 28 to 38, wherein the unit for determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block comprises: a unit for determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block without using one or more bit cost estimates determined via entropy decoding.

[0368] Aspect 40: The apparatus according to any one of aspects 28 to 39, wherein the set of quantization offset parameters includes a quantization offset vector.

[0369] Aspect 41: A computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to: determine a set of quantization offset parameters for a scaled transform coefficient set for the video data block based on side information associated with the video data block; quantize the scaled transform coefficient set for the video data block at least in part based on the set of quantization offset parameters to generate quantized transform coefficients for the video data block; and generate an encoded video bitstream at least in part based on the quantized transform coefficients for the video data block.

[0370] Aspect 42: The computer-readable storage medium according to aspect 41, wherein the instruction causing the one or more processors to determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block includes instructions causing the one or more processors to: select the set of quantization offset parameters for the scaled transform coefficient group from a plurality of sets of quantization offset parameters based on the side information associated with the video data block.

[0371] Aspect 43: A computer-readable storage medium according to aspect 42, wherein the step of causing the one or more processors to select from the plurality of quantization offset parameter sets for the scaled transform coefficient set includes instructions causing the one or more processors to: determine an index of the scaled transform coefficient set based at least in part on the side information associated with the video data block; and use the index to index to the plurality of quantization offset parameter sets to select the quantization offset parameter set for the scaled transform coefficient set.

[0372] Aspect 44: A computer-readable storage medium according to aspect 42 or 43, wherein: the scaled transform coefficient set includes scaled transform coefficients for sub-blocks of the video data block; and the edge information associated with the video data block includes the position of the sub-block within the video data block and the block size of the video data block.

[0373] Aspect 45: A computer-readable storage medium according to aspect 44, wherein the step of causing the one or more processors to select from the plurality of quantization offset parameter sets for the scaled transform coefficient group includes instructions causing the one or more processors to: select from the plurality of quantization offset parameter sets for the scaled transform coefficient group without using the scaled transform coefficient values ​​of the scaled transform coefficient group for the sub-block of the video data block.

[0374] Aspect 46: A computer-readable storage medium according to aspect 41, wherein the instructions for causing the one or more processors to determine the set of quantization offset parameters for the scaled transform coefficient group of the video data block include instructions for causing the one or more processors to: determine a set of parameters for parameterizing the set of quantization offset parameters based on the side information associated with the video data block, the set of parameters having a smaller size than the set of quantization offset parameters; and determine the set of quantization offset parameters for the scaled transform coefficient group based on the smaller size of the set of parameters.

[0375] Aspect 47: A computer-readable storage medium according to any one of aspects 41 to 45, wherein the instructions for causing the one or more processors to determine the set of quantization offset parameters for the scaled transform coefficient group of the video data block include instructions for causing the one or more processors to: determine the set of quantization offset parameters for the scaled transform coefficient group of the video data block using a neural network and based on the side information associated with the video data block.

[0376] Aspect 48: The computer-readable storage medium according to aspect 47, wherein the instructions for causing the one or more processors to determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block include instructions for causing the one or more processors to: determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block using at least one of a classification technique or a regression technique.

[0377] Aspect 49: A computer-readable storage medium according to any one of aspects 41 to 48, wherein the instructions further cause the one or more processors to: divide a plurality of scaled transform coefficients of the video data block into a plurality of scaled transform coefficient groups associated with sub-blocks of the video data block, wherein the plurality of scaled transform coefficient groups includes the scaled transform coefficient groups for the video data block; wherein the instructions causing the one or more processors to determine the set of quantization offset parameters for the scaled transform coefficient groups of the video data block include instructions causing the one or more processors to: determine a corresponding set of quantization offset parameters for each of the plurality of scaled transform coefficient groups; and wherein the instructions causing the one or more processors to quantize the scaled transform coefficient groups for the video data block include instructions causing the one or more processors to: quantize each of the plurality of scaled transform coefficient groups based on the corresponding set of quantization offset parameters.

[0378] Aspect 50: A computer-readable storage medium according to any one of aspects 41 to 49, wherein the instructions for causing the one or more processors to quantize the scaled transform coefficient set for the video data block include instructions for causing the one or more processors to: determine a corresponding quantization offset parameter from the set of quantization offset parameters for each scaled transform coefficient of the scaled transform coefficient set; and quantize each scaled transform coefficient of the scaled transform coefficient set at least in part based on the corresponding quantization offset parameter.

[0379] Aspect 51: A computer-readable storage medium according to any one of aspects 41 to 50, wherein the side information includes one or more of the following: the slice type of the video data block, the block size of the video data block, or an indication of whether the video data block includes a luminance component or a chrominance component.

[0380] Aspect 52: A computer-readable storage medium according to any one of aspects 41 to 51, wherein the instructions for causing the one or more processors to determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block include instructions for causing the one or more processors to: determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block without using one or more bit cost estimates determined via entropy decoding.

[0381] Aspect 53: A computer-readable storage medium according to any one of aspects 41 to 52, wherein the set of quantization offset parameters includes a quantization offset vector.

[0382] It should be recognized that, based on the examples, certain actions or events of any technique described herein may be performed in a different order, and may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for implementing the technique). Furthermore, in some examples, actions or events may be performed concurrently rather than sequentially, for example, through multithreaded processing, interrupt handling, or multiple processors.

[0383] In one or more examples, the described functionality can be implemented using hardware, software, firmware, or any combination thereof. If implemented in software, the functionality can be stored or transmitted as one or more instructions or code on or through a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium can include a computer-readable storage medium, which corresponds to a tangible medium such as a data storage medium or a communication medium, including, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol. In this way, a computer-readable medium can generally correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. A data storage medium can be any available medium that can be accessed by one or more computers or one or more processors to obtain instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products can include computer-readable media.

[0384] For example, rather than limiting, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, flash memory, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (e.g., infrared, radio, and microwave), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (e.g., infrared, radio, and microwave) is included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer instead to non-transient tangible storage media. As used herein, disks and optical discs include compact optical discs (CDs), laser optical discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically magnetically copy data, while optical discs utilize lasers to optically copy data. Combinations of the above items should also be included within the scope of computer-readable media.

[0385] Instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the terms "processor" and "processing circuitry" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Furthermore, the techniques can be implemented entirely within one or more circuit or logic elements.

[0386] The technologies disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or a set of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed technologies, but they do not necessarily need to be implemented through different hardware units. Specifically, as described above, the various units can be combined in a codec hardware unit, or provided by a collection of interoperable hardware units (including one or more processors as described above) combined with appropriate software and / or firmware.

[0387] Various examples have been described. These and other examples are within the scope of the appended claims.

Claims

1. A method for encoding video data, the method comprising: The set of quantization offset parameters for a scaled transform coefficient group for the video data block is determined at least in part based on side information associated with the video data block, wherein the side information includes one or more of the following: the slice type of the video data block, the block size of the video data block, or an indication of whether the video data block includes a luminance component or a chrominance component. The scaled transform coefficient set for the video data block is quantized at least in part by adding the scaled transform coefficient set for the video data block to the quantization offset parameter set to generate quantized transform coefficients for the video data block; and The encoded video bitstream is generated at least in part based on the quantized transform coefficients for the video data blocks.

2. The method according to claim 1, wherein, Determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block includes: The set of quantization offset parameters for the scaled transform coefficient group is selected from multiple sets of quantization offset parameters based on the side information associated with the video data block.

3. The method according to claim 2, wherein, Selecting the set of quantization offset parameters for the scaled transform coefficient group from the plurality of quantization offset parameter sets includes: The index associated with the scaled transform coefficient set is determined at least in part based on the side information associated with the video data block; and The index associated with the scaled transform coefficient set is used to index to the plurality of quantization offset parameter sets to select the quantization offset parameter set for the scaled transform coefficient set.

4. The method according to claim 2, wherein: The scaled transform coefficient set includes scaled transform coefficients for sub-blocks of the video data block; and The edge information associated with the video data block includes the position of the sub-block within the video data block and the block size of the video data block.

5. The method according to claim 4, wherein, The set of quantization offset parameters selected from the plurality of quantization offset parameter sets for the scaled transform coefficient group includes: Without using the scaled transform coefficient values ​​of the scaled transform coefficient set for the sub-block of the video data block, the quantization offset parameter set for the scaled transform coefficient set is selected from the plurality of quantization offset parameter sets.

6. The method according to claim 1, wherein, Determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block includes: A parameter set for parameterizing the quantization offset parameter set is determined based on the side information associated with the video data block; the parameter set having a smaller size compared to the quantization offset parameter set. The quantization offset parameter set for the scaled transform coefficient set is determined based on the smaller size of the parameter set.

7. The method according to claim 1, wherein, Determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block further includes: The set of quantization offset parameters for the scaled transform coefficient group for the video data block is determined using a neural network and based on the side information associated with the video data block.

8. The method according to claim 7, wherein, Determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block further includes: The set of quantization offset parameters for the scaled transform coefficient group for the video data block is determined using at least one of classification or regression techniques.

9. The method according to claim 1, further comprising: The multiple scaled transform coefficients of the video data block are divided into multiple scaled transform coefficient groups associated with sub-blocks of the video data block, wherein the multiple scaled transform coefficient groups include the scaled transform coefficient groups for the video data block. Determining the set of quantization offset parameters for the scaled transform coefficient groups of the video data block includes: determining a corresponding set of quantization offset parameters for each of the plurality of scaled transform coefficient groups; and Specifically, quantizing the scaled transform coefficient group for the video data block includes: quantizing each of the plurality of scaled transform coefficient groups based on the corresponding quantization offset parameter set.

10. The method according to claim 1, wherein, Quantizing the scaled transform coefficient set for the video data block further includes: For each scaled transform coefficient in the scaled transform coefficient set, a corresponding quantization offset parameter is determined from the quantization offset parameter set; and Each scaled transform coefficient of the scaled transform coefficient group is quantized at least in part based on the corresponding quantization offset parameter.

11. The method according to claim 1, wherein, Determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block includes: Without using one or more bit cost estimates determined via entropy decoding, the set of quantization offset parameters for the scaled transform coefficient group for the video data block is determined.

12. The method according to claim 1, wherein, The set of quantization offset parameters includes a quantization offset vector.

13. An apparatus for encoding video data, the apparatus comprising: Memory; A processing circuit that communicates with the memory, the processing circuit being configured to: The set of quantization offset parameters for a scaled transform coefficient group for the video data block is determined at least in part based on side information associated with the video data block, wherein the side information includes one or more of the following: the slice type of the video data block, the block size of the video data block, or an indication of whether the video data block includes a luminance component or a chrominance component. The scaled transform coefficient set for the video data block is quantized at least in part by adding the scaled transform coefficient set for the video data block to the quantization offset parameter set to generate quantized transform coefficients for the video data block; and The encoded video bitstream is generated at least in part based on the quantized transform coefficients for the video data blocks.

14. The device according to claim 13, wherein, To determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block, the processing circuitry is further configured to: The set of quantization offset parameters for the scaled transform coefficient group is selected from multiple sets of quantization offset parameters based on the side information associated with the video data block.

15. The device according to claim 14, wherein, In order to select the set of quantization offset parameters for the scaled transform coefficient group from the plurality of quantization offset parameter sets, the processing circuit is further configured to: The index associated with the scaled transform coefficient set is determined at least in part based on the side information associated with the video data block; and The index associated with the scaled transform coefficient set is used to index to the plurality of quantization offset parameter sets to select the quantization offset parameter set for the scaled transform coefficient set.

16. The device according to claim 14, wherein: The scaled transform coefficient set includes scaled transform coefficients for sub-blocks of the video data block; and The edge information associated with the video data block includes the position of the sub-block within the video data block and the block size of the video data block.

17. The device according to claim 16, wherein, In order to select the set of quantization offset parameters for the scaled transform coefficient group from the plurality of quantization offset parameter sets, the processing circuit is further configured to: Without using the scaled transform coefficient values ​​of the scaled transform coefficient set for the sub-block of the video data block, the quantization offset parameter set for the scaled transform coefficient set is selected from the plurality of quantization offset parameter sets.

18. The device according to claim 13, wherein, To determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block, the processing circuitry is further configured to: A set of parameters for parameterizing the quantization offset parameter set is determined based on the side information associated with the video data block, the set of parameters having a smaller size compared to the quantization offset parameter set; as well as The quantization offset parameter set for the scaled transform coefficient set is determined based on the smaller size of the parameter set.

19. The device according to claim 13, wherein, To determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block, the processing circuitry is further configured to: The set of quantization offset parameters for the scaled transform coefficient group for the video data block is determined using a neural network and based on the side information associated with the video data block.

20. The device according to claim 19, wherein, To determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block, the processing circuitry is further configured to: The set of quantization offset parameters for the scaled transform coefficient group for the video data block is determined using at least one of classification or regression techniques.

21. The device according to claim 13, wherein, The processing circuit is further configured to: The multiple scaled transform coefficients of the video data block are divided into multiple scaled transform coefficient groups associated with sub-blocks of the video data block, wherein the multiple scaled transform coefficient groups include the scaled transform coefficient groups for the video data block. In order to determine the set of quantization offset parameters for the scaled transform coefficient groups for the video data block, the processing circuit is further configured to: determine a corresponding set of quantization offset parameters for each of the plurality of scaled transform coefficient groups; and Specifically, quantizing the scaled transform coefficient group for the video data block includes: quantizing each of the plurality of scaled transform coefficient groups based on the corresponding quantization offset parameter set.

22. The device according to claim 13, wherein, In order to quantize the scaled transform coefficient set for the video data block, the processing circuit is further configured to: For each scaled transform coefficient in the scaled transform coefficient set, a corresponding quantization offset parameter is determined from the quantization offset parameter set; as well as Each scaled transform coefficient of the scaled transform coefficient group is quantized at least in part based on the corresponding quantization offset parameter.

23. The device according to claim 13, wherein, To determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block, the processing circuitry is further configured to: Without using one or more bit cost estimates determined via entropy decoding, the set of quantization offset parameters for the scaled transform coefficient group for the video data block is determined.

24. The device according to claim 13, wherein, The set of quantization offset parameters includes a quantization offset vector.

25. The device according to claim 13, wherein, The device includes one or more of the following: a camera, a computer, or a mobile device.

26. An apparatus for encoding video data, the apparatus comprising: A unit for determining a set of quantization offset parameters for a scaled transform coefficient group for a video data block based at least in part on side information associated with the video data block, wherein the side information includes one or more of the following: the slice type of the video data block, the block size of the video data block, or an indication of whether the video data block includes a luminance component or a chrominance component. A unit for quantizing the scaled transform coefficient set for the video data block, at least in part, by adding the scaled transform coefficient set for the video data block to the quantization offset parameter set, to generate quantized transform coefficients for the video data block; and A unit for generating an encoded video bitstream based at least in part on the quantized transform coefficients for the video data block.

27. The apparatus according to claim 26, wherein, The unit for determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block further includes: A unit for selecting, from a plurality of quantization offset parameter sets, the set of quantization offset parameters for the scaled transform coefficient group, based on the side information associated with the video data block.

28. The apparatus according to claim 27, wherein, The unit for selecting the set of quantization offset parameters for the scaled transform coefficient group from the plurality of quantization offset parameter sets further includes: A unit for determining an index associated with the scaled transform coefficient set based at least in part on the side information associated with the video data block; and A unit for indexing to the plurality of quantization offset parameter sets using the index associated with the scaled transform coefficient set to select the quantization offset parameter set for the scaled transform coefficient set.

29. The apparatus according to claim 27, wherein: The scaled transform coefficient set includes scaled transform coefficients for sub-blocks of the video data block; and The edge information associated with the video data block includes the position of the sub-block within the video data block and the block size of the video data block.

30. The apparatus according to claim 29, wherein, The unit for selecting the set of quantization offset parameters for the scaled transform coefficient group from the plurality of quantization offset parameter sets further includes: A unit for selecting, from the plurality of quantization offset parameter sets, the quantization offset parameter set for the scaled transform coefficient set without using the scaled transform coefficient set's scaled transform coefficient values ​​for the sub-blocks of the video data block.

31. The apparatus according to claim 26, wherein, The unit for determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block further includes: A unit for determining a set of parameters for parameterizing the quantization offset parameter set based on the side information associated with the video data block, the parameter set having a smaller size compared to the quantization offset parameter set; and A unit for determining the quantization offset parameter set for the scaled transform coefficient set based on the smaller size of the parameter set.

32. The apparatus according to claim 26, wherein, The unit for determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block further includes: A unit for determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block using a neural network and based on the side information associated with the video data block.

33. The apparatus according to claim 32, wherein, The unit for determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block includes: A unit for determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block using at least one of classification or regression techniques.

34. The apparatus of claim 26, further comprising: A unit for dividing a plurality of scaled transform coefficients of the video data block into a plurality of scaled transform coefficient groups associated with sub-blocks of the video data block, wherein the plurality of scaled transform coefficient groups include the scaled transform coefficient groups for the video data block. The unit for determining the set of quantization offset parameters for the scaled transform coefficient groups for the video data block includes: a unit for determining a corresponding set of quantization offset parameters for each of the plurality of scaled transform coefficient groups; and The unit for quantizing the scaled transform coefficient group for the video data block further includes a unit for quantizing each of the plurality of scaled transform coefficient groups based on the corresponding quantization offset parameter set.

35. The apparatus according to claim 26, wherein, The unit for quantizing the scaled transform coefficient set for the video data block further includes: A unit for determining a corresponding quantization offset parameter from the set of quantization offset parameters for each scaled transform coefficient in the scaled transform coefficient set; and A unit for quantizing each scaled transform coefficient of the scaled transform coefficient group based at least in part on the corresponding quantization offset parameter.

36. The apparatus according to claim 26, wherein, The unit for determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block includes: A unit for determining the set of quantization offset parameters for the scaled transform coefficient group for the video data block without using one or more bit cost estimates determined via entropy decoding.

37. The apparatus according to claim 26, wherein, The set of quantization offset parameters includes a quantization offset vector.

38. A computer-readable storage medium having instructions stored thereon, said instructions, when executed, causing one or more processors to perform the following operations: The set of quantization offset parameters for a scaled transform coefficient group for the video data block is determined at least in part based on side information associated with the video data block, wherein, The edge information includes one or more of the following: the slice type of the video data block, the block size of the video data block, or an indication of whether the video data block includes a luminance component or a chrominance component; The scaled transform coefficient set for the video data block is quantized at least in part by adding the scaled transform coefficient set for the video data block to the quantization offset parameter set to generate the quantized transform coefficients of the video data block. as well as The encoded video bitstream is generated at least in part based on the quantized transform coefficients for the video data blocks.

39. The computer-readable storage medium according to claim 38, wherein, The instructions that cause the one or more processors to determine the set of quantization offset parameters for the scaled transform coefficient group of the video data block include instructions that cause the one or more processors to perform the following operations: The set of quantization offset parameters for the scaled transform coefficient group is selected from multiple sets of quantization offset parameters based on the side information associated with the video data block.

40. The computer-readable storage medium according to claim 39, wherein, The instruction that causes the one or more processors to select the set of quantization offset parameters for the scaled transform coefficient group from the plurality of quantization offset parameter sets includes instructions that cause the one or more processors to perform the following operations: The index of the scaled transform coefficient group is determined at least in part based on the side information associated with the video data block; as well as The index is used to index to the plurality of quantization offset parameter sets to select the quantization offset parameter set for the scaled transform coefficient group.

41. The computer-readable storage medium according to claim 39, wherein: The scaled transform coefficient set includes scaled transform coefficients for sub-blocks of the video data block; and The edge information associated with the video data block includes the position of the sub-block within the video data block and the block size of the video data block.

42. The computer-readable storage medium according to claim 41, wherein, The instruction that causes the one or more processors to select the set of quantization offset parameters for the scaled transform coefficient group from the plurality of quantization offset parameter sets includes instructions that cause the one or more processors to perform the following operations: Without using the scaled transform coefficient values ​​of the scaled transform coefficient set for the sub-block of the video data block, the quantization offset parameter set for the scaled transform coefficient set is selected from the plurality of quantization offset parameter sets.

43. The computer-readable storage medium according to claim 38, wherein, The instructions that cause the one or more processors to determine the set of quantization offset parameters for the scaled transform coefficient group of the video data block include instructions that cause the one or more processors to perform the following operations: A set of parameters for parameterizing the quantization offset parameter set is determined based on the side information associated with the video data block, the set of parameters having a smaller size compared to the quantization offset parameter set; as well as The quantization offset parameter set for the scaled transform coefficient set is determined based on the smaller size of the parameter set.

44. The computer-readable storage medium according to claim 38, wherein, The instructions that cause the one or more processors to determine the set of quantization offset parameters for the scaled transform coefficient group of the video data block include instructions that cause the one or more processors to perform the following operations: The set of quantization offset parameters for the scaled transform coefficient group for the video data block is determined using a neural network and based on the side information associated with the video data block.

45. The computer-readable storage medium according to claim 44, wherein, The instructions that cause the one or more processors to determine the set of quantization offset parameters for the scaled transform coefficient group of the video data block include instructions that cause the one or more processors to perform the following operations: The set of quantization offset parameters for the scaled transform coefficient group for the video data block is determined using at least one of classification or regression techniques.

46. ​​The computer-readable storage medium of claim 38, wherein the instructions further cause the one or more processors to perform the following operations: The multiple scaled transform coefficients of the video data block are divided into multiple scaled transform coefficient groups associated with sub-blocks of the video data block, wherein, The plurality of scaled transform coefficient sets include the scaled transform coefficient sets for the video data blocks; Wherein, the instruction causing the one or more processors to determine the set of quantization offset parameters for the scaled transform coefficient groups for the video data block includes instructions causing the one or more processors to: determine a corresponding set of quantization offset parameters for each of the plurality of scaled transform coefficient groups; and The instruction that causes the one or more processors to quantize the scaled transform coefficient group for the video data block includes instructions that cause the one or more processors to quantize each of the plurality of scaled transform coefficient groups based on the corresponding quantization offset parameter set.

47. The computer-readable storage medium according to claim 38, wherein, The instructions that cause the one or more processors to quantize the scaled transform coefficient set for the video data block include instructions that cause the one or more processors to perform the following operations: For each scaled transform coefficient in the scaled transform coefficient set, a corresponding quantization offset parameter is determined from the quantization offset parameter set; as well as Each scaled transform coefficient of the scaled transform coefficient group is quantized at least in part based on the corresponding quantization offset parameter.

48. The computer-readable storage medium according to claim 38, wherein, The instructions that cause the one or more processors to determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block include instructions that cause the one or more processors to determine the set of quantization offset parameters for the scaled transform coefficient group for the video data block without using one or more bit cost estimates determined via entropy decoding.

49. The computer-readable storage medium according to claim 38, wherein, The set of quantization offset parameters includes a quantization offset vector.

Citation Information

Patent Citations

  • Chroma quantization in video coding

    US20150071344A1