Adaptive video filter

By applying adaptive video filters during video encoding and decoding, the problems of insufficient video quality and bandwidth consumption in existing technologies are solved, achieving more efficient video processing results.

CN121970318APending Publication Date: 2026-05-01QUALCOMM INC
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
QUALCOMM INC
Filing Date
2024-09-09
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing video encoding and decoding technologies have shortcomings in improving video quality, reducing bandwidth consumption and processing power usage, especially in the application of adaptive filters, which have not been fully optimized.

Method used

Adaptive video filters are used to apply the outputs of various filters, including fixed filters, adaptive loop filters, sample adaptive offset filters, bilateral filters, and cross-component SAO filters, to improve the processing effect of video data.

Benefits of technology

The application of adaptive video filters improves video quality, reduces bandwidth consumption and processing power usage, and enhances the efficiency of video encoding and decoding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121970318A_ABST
    Figure CN121970318A_ABST
Patent Text Reader

Abstract

A video coder may receive video data and apply a filter to a plurality of types of samples of the video data. The plurality of types of samples may include two or more of: an input of a fixed filter, an output of a fixed filter, an input of a signaled filter, an output of a signaled filter, an input of an adaptive loop filter, an output of an adaptive loop filter, an input of a sample adaptive offset (SAO) filter, the method comprises the following steps: receiving the input of an SAO filter, the output of an SAO filter, the input of a bilateral filter, the output of the bilateral filter, the input of a cross component SAO filter, the output of the cross component SAO filter, the input of a deblocking filter, the output of the deblocking filter, filtered reconstructed residual data, a dequantization coefficient, a filtered dequantization coefficient, a predictor or a filtered predictor.
Need to check novelty before this filing date? Find Prior Art

Description

Adaptive video filter

[0001] This application claims the benefit of U.S. Patent Application No. 18 / 826,793, filed September 6, 2024, and U.S. Provisional Patent Application No. 63 / 588,071, filed October 5, 2023, the entire contents of each of which are incorporated herein by reference. U.S. Patent Application No. 18 / 826,793, filed September 6, 2024, claims the benefit of U.S. Provisional Patent Application No. 63 / 588,071, filed October 5, 2023. Technical Field

[0002] This disclosure relates to video encoding and video decoding. Background Technology

[0003] Digital video capabilities can be incorporated into a wide variety of devices, including digital televisions, digital live broadcast systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, tablets, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite wireless phones (so-called "smartphones"), video conferencing equipment, video streaming devices, and more. Digital video devices implement video decoding technologies, such as those defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 (Part 10, Advanced Video Decoding (AVC)), ITU-T H.265 / High Efficiency Video Decoding (HEVC), ITU-T H.266 / Variety Video Decoding (VVC) and extensions to these standards, as well as proprietary video codecs / formats such as AOMedia Video1 (AV1) developed by the Open Media Alliance. By implementing such video decoding technologies, video devices can more efficiently send, receive, encode, decode, and / or store digital video information.

[0004] Video decoding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in video sequences. For block-based video decoding, video slices (e.g., video pictures or portions of video pictures) can be divided into video blocks, which may also be referred to as decoding tree units (CTUs), decoding units (CUs), and / or decoding nodes. Video blocks in a slice after intra-frame decoding (I) of a picture are encoded using spatial prediction relative to reference samples in adjacent blocks within the same picture. Video blocks in a slice after inter-frame decoding (P or B) of a picture can use spatial prediction relative to reference samples in adjacent blocks within the same picture or temporal prediction relative to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. Summary of the Invention

[0005] Generally, this disclosure describes techniques for adaptive video filtering. Specifically, it describes adaptive video filters that can be applied to the output of one or more other filters. The techniques described herein can improve video quality, reduce bandwidth consumption, and / or reduce the use of processing power, etc.

[0006] In one example, this disclosure describes a method comprising: receiving video data; and applying filters to samples of multiple classes of the video data. The multiple classes of samples may include two or more of the following: an input to a fixed filter, an output to a fixed filter, an input to a signaling filter, an output to a signaling filter, an input to an adaptive loop filter, an output to an adaptive loop filter, an input to a sample adaptive offset (SAO) filter, an output to a SAO filter, an input to a bilateral filter, an output to a bilateral filter, an input to a cross-component SAO filter, an output to a cross-component SAO filter, an input to a deblocking filter, an output to a deblocking filter, filtered reconstructed residual data, dequantization coefficients, filtered dequantization coefficients, a predictor, or a filtered predictor.

[0007] In another example, this disclosure describes an apparatus comprising: a memory; and processing circuitry in communication with the memory, the processing circuitry being configured to: receive video data; and apply filters to samples of multiple types of the video data.

[0008] In another example, this disclosure describes an apparatus comprising: a component for receiving video data; and a component for applying filters to samples of multiple types of video data.

[0009] In another example, this disclosure describes a non-transitory computer-readable storage medium that stores instructions, when executed, causing one or more processors configured to decode video data: receive the video data; and apply filters to samples of multiple types of the video data.

[0010] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description, drawings, and claims. Attached Figure Description

[0011] Figure 1 is a block diagram illustrating an example video encoding and decoding system that can perform the techniques of this disclosure.

[0012] Figure 2 illustrates a block diagram of an example filter according to an example of this disclosure.

[0013] Figure 3 illustrates a block diagram of an example filter according to an example of this disclosure.

[0014] Figure 4 illustrates a block diagram of an example filter according to an example of this disclosure.

[0015] Figure 5 is a block diagram illustrating an example video encoder that can perform the techniques of this disclosure.

[0016] Figure 6 is a block diagram illustrating an example video decoder that can perform the techniques of this disclosure.

[0017] Figure 7 is a flowchart illustrating an example method for encoding the current block according to the technology of this disclosure.

[0018] Figure 8 is a flowchart illustrating an example method for decoding the current block according to the technology of this disclosure.

[0019] Figure 9 is a flowchart illustrating an example method according to the technology of this disclosure.

[0020] Figure 10 is a flowchart illustrating another example method according to the technology of this disclosure. Detailed Implementation

[0021] The techniques described herein can improve adaptive video filtering. Specifically, this disclosure describes adaptive video filters that can be applied to the output of one or more other filters. These techniques can improve video quality, reduce bandwidth consumption, and / or reduce the use of processing power, etc.

[0022] Figure 1 is a block diagram illustrating an example video encoding and decoding system 100 capable of performing the techniques of this disclosure. The techniques of this disclosure generally relate to decoding (encoding and / or decoding) video data. In general, video data includes any data used for processing video. Thus, video data may include unencoded raw video, encoded video, decoded (e.g., reconstructed) video, and video metadata, such as signaling data.

[0023] As shown in Figure 1, in this example, system 100 includes a source device 102 that provides encoded video data to be decoded and displayed by a destination device 116. Specifically, source device 102 provides the video data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 can be or may include any of a wide range of devices, such as desktop computers, laptop computers, mobile devices, tablet computers, set-top boxes, mobile phones (such as smartphones), televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, broadcast receiver devices, etc. In some cases, source device 102 and destination device 116 may be configured for wireless communication and are therefore referred to as wireless communication devices.

[0024] In the example of Figure 1, source device 102 includes a video source 104, memory 106, video encoder 200, and output interface 108. Destination device 116 includes an input interface 122, video decoder 300, memory 120, and display device 118. According to this disclosure, the video encoder 200 of source device 102 and the video decoder 300 of destination device 116 can be configured to apply techniques for adaptive video filtering. Thus, source device 102 represents an example of a video encoding device, while destination device 116 represents an example of a video decoding device. In other examples, the source device and destination device may include other components or arrangements. For example, source device 102 may receive video data from an external video source such as an external camera. Similarly, destination device 116 may interface with an external display device instead of including an integrated display device.

[0025] The system 100 shown in Figure 1 is merely an example. Generally, any digital video encoding and / or decoding device can perform techniques for adaptive video filtering. The source device 102 and destination device 116 are merely examples of such decoding devices, where the source device 102 generates decoded video data for transmission to the destination device 116. This disclosure refers to a “decoding” device as a device that performs the decoding (e.g., encoding and / or decoding) of data. Thus, video encoder 200 and video decoder 300 represent examples of decoding devices, specifically, a video encoder and a video decoder, respectively. In some examples, the source device 102 and destination device 116 may operate in a substantially symmetrical manner, such that each of the source device 102 and destination device 116 includes video encoding and decoding components. Therefore, system 100 may support one-way or two-way video transmission between the source device 102 and destination device 116, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0026] Generally, video source 104 represents the source of video data (i.e., unencoded raw video data) and provides a sequential series of pictures (also referred to as "frames") of video data to video encoder 200, which encodes the data for the pictures. Video source 104 of source device 102 may include video capture devices such as cameras, video archives containing previously captured raw video, and / or video feed interfaces for receiving video from video content providers. Alternatively, video source 104 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In each case, video encoder 200 encodes the captured, pre-captured, or computer-generated video data. Video encoder 200 may rearrange the pictures from the received order (sometimes referred to as "display order") to a decoding order for decoding. Video encoder 200 may generate a bitstream comprising the encoded video data. Then, the source device 102 can output the encoded video data to the computer-readable medium 110 via the output interface 108 for reception and / or retrieval by, for example, the input interface 122 of the destination device 116.

[0027] The memory 106 of source device 102 and the memory 120 of destination device 116 represent general-purpose memory. In some examples, memories 106 and 120 may store raw video data, such as raw video from video source 104 and raw decoded video data from video decoder 300. Additionally or alternatively, memories 106 and 120 may store software instructions executable by, for example, video encoder 200 and video decoder 300. Although memories 106 and 120 are shown separately from video encoder 200 and video decoder 300 in this example, it should be understood that video encoder 200 and video decoder 300 may also include internal memory for functionally similar or equivalent purposes. Furthermore, memories 106 and 120 may store encoded video data, such as output from video encoder 200 and input to video decoder 300. In some examples, portions of memories 106 and 120 may be allocated as one or more video buffers, for example, to store raw decoded and / or encoded video data.

[0028] Computer-readable medium 110 may represent any type of medium or device capable of transmitting encoded video data from source device 102 to destination device 116. In one example, computer-readable medium 110 represents a communication medium that enables source device 102 to directly transmit encoded video data to destination device 116 in real time, for example, via a radio frequency network or a computer-based network. According to a communication standard such as a wireless communication protocol, output interface 108 may modulate the transmitted signal including the encoded video data, and input interface 122 may demodulate the received transmitted signal. The communication medium may include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium may form part of a packet-based network such as a local area network, a wide area network, or a global network (such as the Internet). The communication medium may include a router, switch, base station, or any other equipment that may be useful for facilitating communication from source device 102 to destination device 116.

[0029] In some examples, source device 102 can output encoded data from output interface 108 to storage device 112. Similarly, destination device 116 can access encoded data from storage device 112 via input interface 122. Storage device 112 may include any data storage medium of various distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.

[0030] In some examples, source device 102 can output encoded video data to file server 114 or another intermediate storage device that can store the encoded video data generated by source device 102. Destination device 116 can access the stored video data from file server 114 via streaming or download.

[0031] File server 114 can be any type of server device capable of storing encoded video data and sending the encoded video data to destination device 116. File server 114 may represent a web server (e.g., for a website), a server configured to provide file transfer protocol services (such as File Transfer Protocol (FTP) or FLUTE-based file delivery protocol), a Content Delivery Network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Service (MBMS) or Enhanced MBMS (eMBMS) server, and / or a Network Attached Storage (NAS) device. File server 114 may additionally or alternatively implement one or more HTTP streaming protocols, such as HTTP-based Dynamic Adaptive Streaming (DASH), HTTP Live Streaming (HLS), Real-Time Streaming Protocol (RTSP), HTTP Dynamic Streaming, etc.

[0032] Destination device 116 can access encoded video data from file server 114 via any standard data connection, including an internet connection. This may include a wireless channel (e.g., Wi-Fi connection), a wired connection (e.g., digital subscriber line (DSL), cable modem, etc.), or a combination of both, suitable for accessing encoded video data stored on file server 114. Input interface 122 can be configured to operate according to any or more of the various protocols discussed above for retrieving or receiving media data from file server 114, or other such protocols for retrieving media data.

[0033] Output interface 108 and input interface 122 can represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 can be configured to transmit data, such as encoded video data, according to cellular communication standards such as 4G, 4G-LTE (Long Term Evolution), Advanced LTE, 5G, etc. In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 can be configured to operate according to other wireless standards such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee), etc. ™ ),Bluetooth ™Standards are used to transmit data, such as encoded video data. In some examples, source device 102 and / or destination device 116 may include corresponding system-on-chip (SoC) devices. For example, source device 102 may include an SoC device for performing functions belonging to video encoder 200 and / or output interface 108, and destination device 116 may include an SoC device for performing functions belonging to video decoder 300 and / or input interface 122.

[0034] The technology disclosed herein can be applied to video decoding to support any multimedia application in a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission (such as HTTP-based Dynamic Adaptive Streaming (DASH)), digital video encoded onto data storage media, decoding of digital video stored on data storage media, or other applications.

[0035] The input interface 122 of the destination device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., a communication medium, storage device 112, file server 114, etc.). The encoded video bitstream may include signaling information defined by the video encoder 200 and also used by the video decoder 300, such as syntax elements having values ​​describing the characteristics and / or processing of video blocks or other decoded units (e.g., slices, pictures, picture groups, sequences, etc.). The display device 118 displays decoded images of the decoded video data to the user. The display device 118 may represent any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.

[0036] Although not shown in Figure 1, in some examples, both the video encoder 200 and the video decoder 300 may be integrated with the audio encoder and / or audio decoder (e.g., audio codec), and may include appropriate MUX-DEMUX units or other hardware and / or software to process multiplexed streams that include both audio and video in a common data stream. Example audio codecs may include AAC, AC-3, AC-4, ALAC, ALS, AMBE, AMR, AMR-WB (G.722.2), AMR-WB+, aptX (various versions), ATRAC, BroadVoice (BV16, BV32), CELT, Enhanced AC-3 (E-AC-3), EVS, FLAC, G.711, G.722, G.722.1, G.722.2 (AMR-WB), G.723.1, G.726, G.728, G.729, G.729.1, GSM-FR, HE-AAC, iLBC, iSAC, LA Lyra, Monkey's Audio, MP1, MP2 (MPEG-1, 2 Audio Layer II), MP3, Musepack, Nellymoser Asao, OptimFROG, Opus, Sac, Satin, SBC, SILK, Siren 7, Speex, SVOPC, True Audio (TTA), TwinVQ, USAC, Vorbis (Ogg), WavPack and Windows Media Aud.

[0037] Both the video encoder 200 and the video decoder 300 can be implemented as any of a variety of suitable encoder and / or decoder circuits comprising a processing system, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is implemented in part in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and execute the instructions in hardware using one or more processors to perform the technology of this disclosure. Each of the video encoder 200 and the video decoder 300 may be included in one or more encoders or decoders, either of which may be integrated as part of a combined encoder / decoder (CODEC) in the respective device. Devices including the video encoder 200 and / or the video decoder 300 may implement the video encoder 200 and / or the video decoder 300 in processing circuitry such as integrated circuits and / or microprocessors. Such devices may be wireless communication devices (such as cellular phones) or any other type of device described herein.

[0038] Video encoder 200 and video decoder 300 may operate according to video decoding standards such as ITU-T H.265 (also known as High Efficiency Video Decoding (HEVC)) or extensions thereof such as Multi-View and / or Scalable Video Decoding Extensions. Alternatively, video encoder 200 and video decoder 300 may operate according to other proprietary or industry standards such as ITU-T H.266 (also known as Multi-Functional Video Decoding (VVC)). In other examples, video encoder 200 and video decoder 300 may operate according to proprietary video codecs / formats such as AOMedia Video 1 (AV1), extensions to AV1, and / or subsequent versions of AV1 (e.g., AV2)). In other examples, video encoder 200 and video decoder 300 may operate according to other proprietary formats or industry standards. However, the techniques disclosed herein are not limited to any particular decoding standard or format. In general, video encoder 200 and video decoder 300 may be configured to perform the techniques of this disclosure in conjunction with any video decoding technique using adaptive video filtering.

[0039] Generally, the video encoder 200 and video decoder 300 perform block-based decoding of images. The term "block" typically refers to a structure that includes data to be processed (e.g., encoded, decoded, or otherwise used during encoding and / or decoding). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. Generally, the video encoder 200 and video decoder 300 decode video data represented in YUV (e.g., Y, Cb, Cr) format. That is, instead of decoding the red, green, and blue (RGB) data used for images, the video encoder 200 and video decoder 300 decode both the luminance and chrominance components, where the chrominance components may include both red and blue hue chrominance components. In some examples, the video encoder 200 converts the received RGB format data to a YUV representation before encoding, and the video decoder 300 converts that YUV representation to RGB format. Alternatively, preprocessing and postprocessing units (not shown) may perform these conversions.

[0040] This disclosure generally relates to the decoding (e.g., encoding and decoding) of images to include processes of encoding or decoding data of the image. Similarly, this disclosure may relate to the decoding of blocks of images to include processes of encoding or decoding data for the blocks (e.g., prediction and / or residual decoding). Encoded video bitstreams typically include a series of values ​​for syntax elements representing decoding decisions (e.g., decoding modes) and the partitioning of images into blocks. Therefore, references to the decoding of images or blocks should generally be understood as the decoded values ​​of the syntax elements that form the images or blocks.

[0041] HEVC defines various blocks, including decoding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video decoder (such as a video encoder 200) divides the decoding tree unit (CTU) into CUs based on a quadtree structure. That is, the video decoder divides the CTU and CU into four equal, non-overlapping squares, and each node of the quadtree has zero or four child nodes. Nodes without child nodes are called "leaf nodes," and the CU of such leaf nodes may include one or more PUs and / or one or more TUs. The video decoder may further divide the PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents the partitioning of the TU. In HEVC, the PU represents inter-frame prediction data, while the TU represents residual data. The CU after intra-frame prediction includes intra-frame prediction information, such as intra-frame mode indication.

[0042] As another example, video encoder 200 and video decoder 300 can be configured to operate according to VVC. According to VVC, the video decoder (such as video encoder 200) partitions the image into multiple CTUs. Video encoder 200 can partition the CTUs according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure removes the concept of multiple partition types, such as the separation between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels: a first level partitioned according to a quadtree and a second level partitioned according to a binary tree. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to CUs.

[0043] In the MTT partitioning structure, blocks can be divided using quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) partitioning (also known as triplet tree (TT)). A ternary tree or triplet tree partition is a partition in which a block is divided into three sub-blocks. In some examples, a ternary tree or triplet tree partition divides a block into three sub-blocks without dividing the original block through the center. Partition types in MTT (e.g., QT, BT, and TT) can be symmetric or asymmetric.

[0044] When operating according to the AV1 codec, the video encoder 200 and video decoder 300 can be configured to decode video data in blocks. In AV1, the largest decoded block that can be processed is called a superblock. In AV1, a superblock can be 128x128 luma samples or 64x64 luma samples. However, in subsequent video decoding formats (e.g., AV2), superblocks can be defined by different (e.g., larger) luma sample sizes. In some examples, the superblock is the top level of a block quadtree. The video encoder 200 can further divide the superblock into smaller decoded blocks. The video encoder 200 can use square or non-square partitions to divide the superblock and other decoded blocks into smaller blocks. Non-square blocks can include N / 2xN blocks, NxN / 2 blocks, N / 4xN blocks, and NxN / 4 blocks. The video encoder 200 and video decoder 300 can perform separate prediction and transform processing for each decoded block.

[0045] AV1 also defines video data tiles. A tile is a rectangular array of superblocks that can be decoded independently of other tiles. That is, the video encoder 200 and video decoder 300 can encode and decode the decoding blocks within a tile separately without using video data from other tiles. However, the video encoder 200 and video decoder 300 can perform filtering across tile boundaries. The tile size can be uniform or non-uniform. Tile-based decoding enables parallel processing and / or multithreading in the encoder and decoder implementations.

[0046] In some examples, the video encoder 200 and the video decoder 300 may use a single QTBT or MTT structure to represent each of the luma and chroma components, while in other examples, the video encoder 200 and the video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luma component and another QTBT / MTT structure for the two chroma components (or two QTBT / MTT structures for the respective chroma components).

[0047] The video encoder 200 and the video decoder 300 can be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, superblock partitioning or other partitioning structures.

[0048] In some examples, a CTU includes a decoded tree block (CTB) of luminance samples, two corresponding CTBs of chrominance samples of an image with three sample arrays, or a CTB of samples of an image decoded using three separate color planes and a syntax structure for decoding the samples. A CTB can be an NxN sample block of some value N, such that a partitioning method divides the components into CTBs. A component is an array or a single sample from one of the three arrays (luminance and two chrominance) constituting a 4:2:0, 4:2:2, or 4:4:4 color format image, or an array or a single sample constituting an array or array constituting a monochrome format image. In some examples, a decoded block is an MxN sample block of values ​​M and N, such that a partitioning method divides the CTB into decoded blocks.

[0049] Blocks (e.g., CTUs or CUs) can be grouped in various ways within an image. As an example, a brick can refer to a rectangular area of ​​a row of CTUs within a specific tile in an image. A tile can be a rectangular area of ​​CTUs within a specific tile column and a specific tile row in an image. A tile column refers to a rectangular area of ​​a CTU having a height equal to the height of the image and a width specified by syntax elements (e.g., such as in an image parameter set). A tile row refers to a rectangular area of ​​a CTU having a height specified by syntax elements (e.g., such as in an image parameter set) and a width equal to the width of the image.

[0050] In some examples, a tile can be divided into multiple bricks, each of which may include one or more CTU rows within the tile. A tile that is not divided into multiple bricks can also be called a brick. However, bricks that are a true subset of a tile cannot be called a tile. Bricks in an image can also be arranged in slices. A slice can be an integer number of bricks in an image that can be uniquely contained within a single Network Abstraction Layer (NAL) unit. In some examples, a slice includes multiple complete tiles or a continuous sequence of complete bricks that include only one tile.

[0051] This disclosure uses "NxN" and "N by N" interchangeably to refer to the sample size of a block (such as a CU or other video block) in the vertical and horizontal dimensions, for example, 16x16 samples or 16 by 16 samples. Generally, a 16x16 CU will have 16 samples in the vertical direction (y=16) and 16 samples in the horizontal direction (x=16). Similarly, an NxNCU typically has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. Samples in a CU can be arranged in rows and columns. Furthermore, a CU does not necessarily need to have the same number of samples in the horizontal direction as it does in the vertical direction. For example, a CU may include NxM samples, where M is not necessarily equal to N.

[0052] The video encoder 200 encodes video data representing prediction and / or residual information, as well as other information, for use in the control unit (CU). The prediction information indicates how the CU should be predicted to form a prediction block for the CU. The residual information typically represents the sample-by-sample difference between a sample of the CU before encoding and the prediction block.

[0053] To predict the Cubic Frame (CU), the video encoder 200 typically forms prediction blocks for the CU through inter-frame prediction or intra-frame prediction. Inter-frame prediction typically refers to predicting the CU from data in a previously decoded image, while intra-frame prediction typically refers to predicting the CU from data in a previously decoded image within the same frame. To perform inter-frame prediction, the video encoder 200 can use one or more motion vectors to generate prediction blocks. The video encoder 200 can typically perform a motion search to identify reference blocks that closely match the CU, for example, based on the differences between the CU and a reference block. The video encoder 200 can use sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to compute difference metrics to determine whether a reference block closely matches the current CU. In some examples, the video encoder 200 can use unidirectional or bidirectional prediction to predict the current CU.

[0054] Some examples of VVC also provide an affine motion compensation mode, which can be viewed as an inter-frame prediction mode. In affine motion compensation mode, the video encoder 200 can determine two or more motion vectors representing non-translational motion, such as zooming in or out, rotation, perspective motion, or other irregular motion types.

[0055] To perform intra-frame prediction, the video encoder 200 can select an intra-frame prediction mode to generate prediction blocks. Some examples of VVC provide sixty-seven intra-frame prediction modes, including various directional modes, as well as planar and DC modes. Generally, the video encoder 200 selects an intra-frame prediction mode that describes the neighboring samples of the current block (e.g., the block of the CU) and predicts samples of the current block from them. Assuming the video encoder 200 decodes the CTU and CU in raster scan order (from left to right, from top to bottom), such samples can typically be located above, to the upper left, or to the left of the current block in the same image as the current block.

[0056] The video encoder 200 encodes data representing the prediction mode of the current block. For example, for inter-frame prediction modes, the video encoder 200 may encode data indicating which of the various available inter-frame prediction modes is used, as well as the motion information for the corresponding mode. For example, for unidirectional or bidirectional inter-frame prediction, the video encoder 200 may encode motion vectors using Advanced Motion Vector Prediction (AMVP) or merging modes. The video encoder 200 may use similar modes to encode motion vectors used for affine motion compensation modes.

[0057] AV1 includes two common techniques for encoding and decoding blocks of video data. These two common techniques are intra-frame prediction (e.g., intra-frame prediction or spatial prediction) and inter-frame prediction (e.g., inter-frame prediction or temporal prediction). In the context of AV1, when using intra-frame prediction modes to predict blocks of video data for the current frame, the video encoder 200 and video decoder 300 do not use video data from other frames of the video data. For most intra-frame prediction modes, the video encoder 200 encodes blocks of the current frame based on the difference between sample values ​​in the current block and predicted values ​​generated from reference samples in the same frame. The video encoder 200 determines the predicted values ​​generated from the reference samples based on the intra-frame prediction mode.

[0058] After prediction (such as intra-frame or inter-frame prediction for a block), the video encoder 200 can compute residual data for the block. The residual data (such as a residual block) represents the sample-by-sample difference between the block and the prediction block used to form the block, which is formed using the corresponding prediction mode. The video encoder 200 can apply one or more transforms to the residual block to produce transformed data in the transform domain rather than the sample domain. For example, the video encoder 200 can apply a Discrete Cosine Transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. Additionally, the video encoder 200 can apply a secondary transform after the first transform, such as the Mode Correlated Inseparable Secondary Transform (MDNSST), the Signal Correlation Transform, the Karhunen-Loeve Transform (KLT), etc. The video encoder 200 produces transform coefficients after applying one or more transforms.

[0059] As noted above, after any transformation that produces the transform coefficients, the video encoder 200 may perform quantization on the transform coefficients. Quantization generally refers to a process in which the transform coefficients are quantized to reduce the amount of data used to represent them, thereby providing further compression. By performing the quantization process, the video encoder 200 may reduce the bit depth associated with some or all of the transform coefficients. For example, the video encoder 200 may round an n-bit value down to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder 200 may perform a bit-by-bit right shift on the value to be quantized.

[0060] After quantization, the video encoder 200 can scan the transform coefficients to generate a one-dimensional vector from a two-dimensional matrix containing the quantized transform coefficients. The scan can be designed to place higher-energy (and therefore lower-frequency) transform coefficients before the vector and lower-energy (and therefore higher-frequency) transform coefficients after the vector. In some examples, the video encoder 200 can utilize a predefined scan order to scan the quantized transform coefficients to produce a serialized vector, and then entropy-encode the quantized transform coefficients of that vector. In other examples, the video encoder 200 can perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 200 can entropy-encode the one-dimensional vector, for example, according to context-adaptive binary arithmetic decoding (CABAC). The video encoder 200 can also entropy-encode the values ​​of syntax elements describing metadata associated with the encoded video data, which is used by the video decoder 300 when decoding the video data.

[0061] To perform CABAC, the video encoder 200 can assign context within a context model to the symbols to be transmitted. Context may involve, for example, whether the neighboring values ​​of a symbol are zero. Probability determination can be based on the context assigned to the symbols.

[0062] The video encoder 200 may further generate syntax data for the video decoder 300, such as block-based syntax data, image-based syntax data, and sequence-based syntax data, for example, in image headers, block headers, and slice headers, or generate other syntax data such as sequence parameter sets (SPS), image parameter sets (PPS), or video parameter sets (VPS). The video decoder 300 may also decode such syntax data to determine how to decode the corresponding video data.

[0063] In this way, the video encoder 200 can generate a bitstream that includes encoded video data, such as syntax elements describing the partitioning of images into blocks (e.g., CUs) and prediction and / or residual information for the blocks. Finally, the video decoder 300 can receive the bitstream and decode the encoded video data.

[0064] Generally, the video decoder 300 performs the reverse process of the video encoder 200 to decode the encoded video data of the bitstream. For example, the video decoder 300 can use CABAC to decode the values ​​of syntax elements used for the bitstream in a manner substantially similar to but reversed by the CABAC encoding process of the video encoder 200. Syntax elements can define partitioning information for dividing a picture into CTUs and defining the CUs of each CTU according to a corresponding partitioning structure such as a QTBT structure. Syntax elements can further define prediction and residual information for video data blocks (e.g., CUs).

[0065] The residual information can be represented, for example, by quantized transform coefficients. The video decoder 300 can inversely quantize and inverse transform the quantized transform coefficients of the block to reconstruct the residual block for the block. The video decoder 300 uses a signaling prediction mode (intra-frame prediction or inter-frame prediction) and associated prediction information (e.g., motion information for inter-frame prediction) to form a prediction block for the block. The video decoder 300 can then combine the prediction block and the residual block (on a sample-by-sample basis) to reconstruct the original block. The video decoder 300 can perform additional processing, such as performing a deblocking process to reduce visual artifacts along the block boundaries.

[0066] This disclosure may generally relate to "signaling" certain information (such as syntax elements). The term "signaling" can generally refer to the communication of values ​​and / or other data of syntax elements used to decode encoded video data. That is, video encoder 200 may signal the values ​​of syntax elements in the bitstream. Generally speaking, signaling refers to generating values ​​in the bitstream. As noted above, source device 102 may transmit the bitstream to destination device 116 substantially in real time or not in real time (such as when syntax elements are stored in storage device 112 for later retrieval by destination device 116).

[0067] According to the technology disclosed herein, as will be explained in more detail below, the video encoder 200 and the video decoder 300 can be configured to receive video data and apply filters to samples of multiple types of the video data.

[0068] In the example video encoder, frames of the original video sequence are divided into rectangular regions or blocks, which can be encoded in intra-frame mode (I-mode) or inter-frame mode. This block can be decoded using some form of transform decoding (such as DCT decoding). However, purely transform-based decoding only reduces inter-pixel correlation within a specific block, without considering inter-block correlation. Purely transform-based decoding can still produce a relatively high transmission bit rate. Current digital image decoding standards also utilize certain techniques to reduce the correlation of pixel values ​​between blocks.

[0069] Generally, blocks encoded in inter-frame mode are predicted from multiple previously decoded frames. The prediction information for a block can be represented, for example, by two-dimensional (2D) motion vectors. For blocks encoded in intra-frame mode, spatial predictions from already encoded neighboring blocks within the same frame can be used to form the predicted block. The prediction error (e.g., the difference between the block being encoded and the predicted block) can be represented as a set of weighted basis functions of a discrete transform. This transform is typically performed on a block-by-block basis. The weights, such as transform coefficients, can then be quantized. Quantization introduces information loss. Therefore, the quantized transform coefficients may have lower precision than the original transform coefficients.

[0070] The quantized transform coefficients, along with the motion vector and some control information, form a decoded sequence representation, which can be referred to as syntax elements. Before being sent from the video encoder to the video decoder, the syntax elements can be entropy-decoded to further reduce the number of bits required for their representation (e.g., syntax elements).

[0071] In a video decoder, blocks in the current frame are obtained by first constructing a prediction of the block in the same manner as in a video encoder, and then adding the compressed prediction error to that prediction. The compressed prediction error can be found by weighting the transform basis function using the quantized transform coefficients. The difference between the reconstructed frame and the original frame is called the reconstruction error.

[0072] In video decoding, filtering is often applied to enhance the quality of the decoded video signal. A filter can be applied as a post-filter, where the filtered frame is not used to predict future frames. In other examples, the filter can be applied as an in-loop filter, where the filtered frame is used to predict future frames. For example, the filter can be designed to minimize the error between the original signal and the decoded filtered signal. Similar to transform coefficients, the coefficients of the filter can be quantized, for example, as follows. , :

[0073] .

[0074] The filter coefficients can then be decoded and sent to the video decoder. Usually equals . The larger the value, the more accurate the quantization, and the better the quantized filter coefficients. The better the performance offered, the better. On the other hand, Larger values ​​will result in more bits being needed to send. .

[0075]

[0076] In the video decoder, the reconstructed image is processed as follows. Apply decoded filter coefficients :in and These are the coordinates of pixels within the frame. They can also be used for the samples to be filtered. Difference between its neighboring samples Applying filter coefficients:

[0077]

[0078] In this case, the sum obtained can be compared with the reconstructed sample. Add to obtain the sample The difference can be modified, for example, by applying clipping. .

[0079] In VVC, block-based adaptive adaptive loop filters (ALFs) (see M. Karczewicz et al., “VVC in-loop filters,” IEEE Transactions on Video Technology, Vol. 31, No. 10, pp. 3907-3925, October 2021) are considered state-of-the-art in-loop filters. Using such ALF filters, sub-block or pixel-level filter adaptation can be applied. Blocks can be based on the directionality of the block. And the quantitative value of activity It is classified into one of 25 categories:

[0080] .

[0081] Each category can have its own assigned filter.

[0082] A Laplacian-based classifier can be used to deduce the class of a sample in a target block. A window covering the target block can be used to classify that specific target block. Activity and directionality are derived using the values ​​of the horizontal, vertical, and two diagonal gradients calculated using the 1-D Laplacian operator:

[0083]

[0084] The sum of the horizontal gradient, vertical gradient, and two diagonal gradients within the window can be expressed as follows: , , and Directionality This can be determined by comparing the following with a set of thresholds.

[0085]

[0086] Activity It can be calculated and The sum and activities The derivation is made by comparing with a set of thresholds.

[0087] Before filtering, certain geometric transformations (such as rotation, diagonal flip, and vertical flip) can be applied to pixels in the filter support region (e.g., pixels multiplied by the filtered coefficients), depending on the orientation of the gradient of the filtered pixels. These transformations increase the similarity between different regions within the image, such as their orientation. This reduces the number of filters that must be sent to the video decoder, and thus reduces the number of bits required to represent the filters, or additionally or alternatively, reduces reconstruction errors. Applying transformations to the filter support region is equivalent to applying transformations directly to the filter coefficients.

[0088] To reduce the number of bits used to represent filter coefficients, different classes can be merged. Information about which classes will be merged is provided to the video decoder by the video encoder by sending an index i_C for each of the 25 classes. Classes with the same index i_C can share the same filter.

[0089] The ALF coefficients of a reference image can be stored and reused as ALF coefficients for the current image. For the current image, the video encoder can use the ALF coefficients stored for the reference image and bypass signaling the ALF coefficients to the video decoder. In this case, the video encoder can simply signal the index of one of the reference images, and the stored ALF coefficients of the indicated reference image can be easily inherited for the current image.

[0090] In the paper "Compression efficiency methods beyond VVC" (JVET-U0100), presented at the 21st JVET Conference, January 2021, by Y.-J. Chang, C.-C. Chen, J. Chen, J. Dong, HE Egilmez, N. Hu, H. Huang, M. Karczewicz, J. Li, B. Ray, K. Reuze, V. Seregin, N. Shlyakhov, L. PhamVan, H. Wang, Y. Zhang, and Z. Zhang, at the 21st JVET Conference, January 2021 (hereinafter referred to as "Chang et al."), a method using three different classifiers (C0, C1, and C2) and three different filter sets (F0, F1, and F2) was proposed. Sets F0 and F1 contain fixed filters with coefficients trained for classifiers C0 and C1. The coefficients of the filters in F2 are signaled. For a given sample, the filters from set F0 are used to generate the desired output. i Which filter is used by classifier C? i The category assigned to the sample The decision is made. All three classifiers are based on the Laplacian operator and differ from the classifiers used in VVC by using windows with different numbers of samples and multiple thresholds to determine activity and orientation.

[0091] In Enhanced Compression Model version 9.0 (ECM-9.0), a third fixed filter (also known as a Gaussian filter) is applied to the samples before the deblocking filter. Additionally, a signaling filter is applied to the output of the Gaussian filter.

[0092] Cascaded filtering is proposed in co-pending U.S. Patent Application No. 18 / 605,416, filed March 14, 2024. One filter can be applied to the output of one filter. A classifier can be applied to the output sample values ​​of another filter.

[0093] To further improve decoding efficiency, U.S. Patent Application No. 18 / 605,416 describes the following techniques. First, it describes a difference-based classifier. Second, it describes a classifier based on multiple features. Such classifiers can be applied to both fixed filters and signaling filters. Third, it describes cascaded filtering. When cascaded filtering is applied, the output sample values ​​of one filter can be used as the input sample values ​​of another filter. Fourth, it describes filters applied to sample values ​​in multiple reconstruction stages. When sample values ​​from one stage are unavailable and / or unused, sample values ​​from another stage can be used. Fifth, when a filter is applied to sample values ​​from only one stage, the difference derivation of the input samples can be disabled. These described techniques can be applied individually or in any combination.

[0094] We will now describe a difference-based classifier. For reconstructing a target patch in an image, in order to deduce the class index... The video encoder 200 or video decoder 300 can calculate values. .

[0095] For example, value It can be a window The standard deviation, variance, median, or mean of the value; this window may include the target block. In another example, the value... It can be a sample value from the window.

[0096] Each sample value and The difference can be calculated within a window that may include the target block. This window can be the same A single window or different windows. For example, a video encoder 200 or a video decoder 300 can determine the difference.

[0097] In one example, the window size can vary for different fixed filters. For instance, the same sample window can be used for a difference-based classifier when applying a Laplacian operator classifier.

[0098] Another value The value can be calculated based on the derived difference. It can be the sum of absolute differences, the sum of squared differences, the square root of the sum of squared differences, or any other combination based on the derived differences. For example, a video encoder 200 or a video decoder 300 can determine the value. .

[0099] Category Index Available from value Derivation. Before deriving the category index, the value can be scaled using a scaling factor. Further quantification For example, video encoder 200 or video decoder 300 can determine the category index. And scaling factors can be applied (if used).

[0100] In one example, the scaling factor may be derived based on the activity value, window size, and / or the bit depth of the sample values ​​within the window. In another example, the activity value is derived as the sum of the values ​​of the horizontal and vertical gradients calculated using a 1-D Laplacian operator. For example, a video encoder 200 or a video decoder 300 may determine the scaling factor and / or the activity value.

[0101] Further quantification of values ​​is possible Prune the values ​​to fit within the allowed category index range. The pruned values ​​can then be used as category indexes. For example, a video encoder 200 or a video decoder 300 can process quantized values. Cut it.

[0102] For example, It can represent the reconstructed image. The coordinates of the sample in the window. Average value of the window It can be calculated as

[0103]

[0104] The average value can be determined by the video encoder 200 or the video decoder 300. .

[0105] value It can be calculated as Each sample in the window and The square root of the sum of the squared differences between them.

[0106]

[0107] The video encoder 200 or video decoder 300 can determine the value. .

[0108] Based on a Laplace operator-based classifier, the activity value of the current block can be... .

[0109] Category Index It can be deduced as

[0110]

[0111] Where the scaling factor array s[]={2, 2, 4, 4, 8, 8, 8, 8, 16, 16, 16, 16, 32, 32, 32, 32}, and the number of categories The video encoder 200 or video decoder 300 can determine the category index. .

[0112] The scaling factor can be further scaled based on the bit depth of the sample. In one example, This can be further modified by multiplying the following two: and category index

[0113]

[0114] In one example

[0115]

[0116] This can be derived by accessing sample values ​​once. As follows:

[0117]

[0118] For example, if Then the variance can be calculated as

[0119]

[0120] Therefore, when value Calculated as Each sample in the window and To find the square root of the sum of the squared differences between samples, one can calculate the sum and square of the sample values ​​within the window to derive the result. and further category indexes .

[0121] When calculated value At that time, the implementation scheme of this technology can Approximate value (called (This is incorporated into the calculation.) The goal could be to allow software / hardware designs for technologies to be implemented without partitioning. In other words, division ( / r) can be replaced by bit shifting.

[0122] As an example, for using Specific implementation, accurate value In the formula Can be Substitution. In this case, the formula becomes:

[0123]

[0124] Division can be implemented using bit shifting:

[0125]

[0126] Additional calculation The arithmetic operations involved can be broken down into multiple stages to reduce the number of bits required for intermediate values. As an example, the approximation shown in the previous example can be derived from... Revised to:

[0127]

[0128] We will now discuss classifiers based on multiple features. In this technique, a final class index can be derived from several (e.g., multiple) individually derived class indices. A mapping process can be introduced to derive the final class index from a combination of individual class indices. For example, a video encoder 200 or a video decoder 300 can determine the final class index.

[0129] In one example, the mapping process can be represented as follows. For example, C i (where i = 0…N–1) can represent the i-th individual classifier, and M i (where i = 0…N–1) can represent the total number of categories for the i-th individual classifier. For a given sample, It can be represented from the i-th individual classifier C i The category index is derived from (where i = 0…N–1). Category Index It can be deduced as

[0130]

[0131] For example, video encoder 200 or video decoder 300 can determine the category index. .

[0132] The derived category index C can be further mapped to the final category index. For example, two or more category index C values ​​can be mapped to the same final category index indicating the same filter to be used.

[0133] In another example of a classifier based on two features, and This can represent the class indices derived from the previously described difference-based classifier and the Laplacian operator-based classifier, respectively. The final class index can be derived as follows:

[0134]

[0135] Where M1 is the total number of classes in classifier C1.

[0136] Cascaded filtering is now described. When multiple filters are applied, another filter can be applied to the output of one filter. A classifier can be applied to the output sample values ​​of another filter. For example, a video encoder 200 or a video decoder 300 can apply another filter and / or a classifier to the output of one filter.

[0137] For example, in the literature of Chang et al., two fixed filters can be applied simultaneously to the input sample values ​​of an ALF filter. In one example, another fixed filter can be applied to the output sample values ​​of one fixed filter. After applying both filters, the output can be used as the input for ALF filtering.

[0138] In another example, a fixed filter can be applied to several types of input. In one example, the output sample values ​​of another fixed filter can be used as input, and sample values ​​prior to the ALF (e.g., samples prior to the deblocking filter, reconstructed residual samples, or predictors) can be used as another input. In some examples, a geometric transpose can be applied to one type of input. In some examples, a geometric transpose can not be applied to one type of input. In some examples, the type of geometric transpose can be the same for all types of input. In some examples, the type of geometric transpose can be determined by one type of input, and the determined transpose can be applied to other (or all) types of input.

[0139] The class index of each fixed filter can be derived from the reconstructed samples (e.g., the input to the ALF). In another example, the class index of another fixed filter can be derived from the output sample values ​​after a fixed filter has been applied.

[0140] The filters applied to the reconstructed residual sample values ​​are now described. Filters can be applied simultaneously to sample inputs obtained at different stages of the reconstruction process. In other words, the video encoder 200 or video decoder 300 can apply a filter to the samples at one stage of the reconstruction process and reapply the same filter to the samples at another stage of the reconstruction process. For example, the video encoder 200 or video decoder 300 can apply filters to the sample values ​​before and after some of the in-loop filters (such as deblocking filters, bilateral filters, SAO, cross-component SAO, etc.).

[0141] In another example, sample values ​​for which a filter is applied can be obtained after intra-frame prediction or inter-frame prediction. In yet another example, sample values ​​for which a filter is applied can be obtained after an inverse transform, which is a reconstruction residual.

[0142] If an input from a reconstruction stage is unavailable or unused, it can be replaced by another input for filtering purposes. In a sense, such input samples are used twice in the filtering process. The clipping value or clipping index applied to the same samples can be averaged (firstly) before the filtering process. In another example, when an input is unavailable or unused, the portion of the filter corresponding to that input is not applied, or alternatively, that portion of the filter can be applied to zero input.

[0143] We will now discuss derivation of the stop difference for a single input. In some cases, filters can be applied to the differences between samples from different stages; however, when only the residual input is used, multiple input stages are replaced by a single residual input for filtering, and this can produce zero difference when deriving the difference, because the residual is subtracted from itself. Such computations can waste processing resources.

[0144] In such cases, the video encoder 200 or video decoder 300 can use zero as the difference, for example, without performing any calculations. Alternatively, the difference may not be applied, and only the residual input may be used without deriving the difference. In another example, the residual input multiplied by a factor may be used (in one example, the factor may be equal to –1, in a previous sample (or in another example, the factor may be equal to 1).

[0145] Example

[0146] This disclosure describes techniques in which the video encoder 200 and video decoder 300 can be configured to apply filters to the inputs or outputs of one or more other filters. That is, the video encoder 200 and video decoder 300 can be configured to apply filters to the inputs and / or outputs of multiple filters. The filters disclosed herein can be post-filters or loop filters.

[0147] Additionally, in some examples, the video encoder 200 and video decoder 300 may be configured to apply a classifier m to the output sample values ​​of one or more other filters. The filters that can be used with the techniques of this disclosure can be any filters described above. The filters can be fixed filters or signaled filters. Furthermore, the classifier can be a classifier of a fixed filter or a classifier of a signaled filter.

[0148] Therefore, the video encoder 200 and the video decoder 300 can be configured to apply a classifier to the output sample values ​​of samples of multiple classes of video data. The classifier can be a classifier for a fixed filter or a classifier for a signaling filter.

[0149] For example, video encoder 200 and video decoder 300 can be configured to apply filters to samples of one or more classes generated from any combination of the following sources:

[0150] - Inputs and / or outputs of other fixed filters (e.g., fixed filters with Laplacian-based classifiers and / or Gaussian filters).

[0151] - Inputs and / or outputs of other signaling filters

[0152] - Input and / or output of an adaptive loop filter (ALF)

[0153] - Inputs and / or outputs of Sample Adaptive Shift (SAO) filters, bilateral filters, and / or cross-component sample adaptive shift filters (CC-SAO).

[0154] - Input and / or output of the deblocking filter

[0155] -Reconstruct residual data

[0156] - Filtered reconstructed residual data

[0157] -Dequantization coefficients

[0158] -Dequantization coefficients after filtering

[0159] -Predictors

[0160] -Filtered predictor

[0161] In a general example of this disclosure, the video encoder 200 and video decoder 300 can be configured to receive video data and apply filters to samples of multiple categories of the video data. That is, the video encoder 200 and video decoder 300 can apply filters to input samples of multiple categories, as shown above. Furthermore, the video encoder 200 and video decoder 300 can apply filters to output samples of multiple categories. In other words, the filters of this disclosure can accept more than one category of samples as input.

[0162] Multiple types of samples may include two or more of the following: input to a fixed filter, output to a fixed filter, input to a signaling filter, output to a signaling filter, input to an adaptive loop filter, output to an adaptive loop filter, input to a SAO filter, output to a SAO filter, input to a bilateral filter, output to a bilateral filter, input to a cross-component SAO filter, output to a cross-component SAO filter, input to a deblocking filter, output to a deblocking filter, filtered reconstructed residual data, dequantization coefficients, filtered dequantization coefficients, predictors, or filtered predictors.

[0163] As a more specific example, the video encoder 200 and video decoder 300 may receive a first output of a signal-notified filter, a second output of a fixed filter, and apply filters to both the first and second outputs. Figure 2 illustrates various examples of this filter structure. Filter 600A receives input from both the signal-notified filter 610A and the fixed filter 612A. Filter 600B receives input from both the signal-notified filter 610A and a different signal-notified filter 610B. Filter 600C receives input from both the fixed filter 612A and a different fixed filter 612B. In the example of Figure 2, the signal-notified filter may include any of an adaptive loop filter, a SAO filter, a bilateral filter, a cross-component SAO filter, or a deblocking filter. Any type of fixed filter can be used.

[0164] In another example of this disclosure, the video encoder 200 and video decoder 300 may receive first corresponding outputs of two or more signaling filters and second corresponding outputs of one or more fixed filters. The video encoder 200 and video decoder 300 may apply filters to the first and second corresponding outputs. Figure 3 illustrates various examples of this filter structure, but other combinations are also within the scope of this disclosure. Filter 700A receives input from signaling filter 710A, different signaling filters 710B, fixed filter 712A, and different fixed filters 712B. Filter 700B receives input from signaling filter 710A, different signaling filters 710B, and a single fixed filter 712A. Filter 700C receives input from signaling filter 710A, different signaling filters 710B, another different signaling filter 710C, and a single fixed filter 712A. Similarly, the signaling filter can include any of the following: adaptive loop filter, SAO filter, bilateral filter, cross-component SAO filter, or deblocking filter. Any type of fixed filter can be used.

[0165] The filters disclosed herein are not limited to the outputs of multiple different filters. In other examples, the outputs of the filters disclosed herein may be the inputs to multiple different other filters. Figure 4 illustrates various examples of this filter structure, but other combinations are also within the scope of this disclosure. The output of filter 800A may be the input to both signaling filter 810A and different signaling filters 810B. The output of filter 800B may be the input to both signaling filter 810A and fixed filter 812A. The output of filter 800C may be the input to both fixed filter 812A and different fixed filters 812B. Similarly, the signaling filters may include any of the following: adaptive loop filter, SAO filter, bilateral filter, cross-component SAO filter, or deblocking filter. Any type of fixed filter may be used.

[0166] Figure 5 is a block diagram illustrating an example video encoder 200 capable of performing the techniques of this disclosure. Figure 5 is provided for illustrative purposes and should not be construed as a limitation on the techniques extensively illustrated and described in this disclosure. For illustrative purposes, this disclosure describes the video encoder 200 in accordance with VVC and HEVC techniques. However, the techniques of this disclosure can be performed by video encoding devices configured for other video decoding standards and video decoding formats, such as AV1 and subsequent formats of AV1 video decoding.

[0167] In the example of Figure 5, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any or all of the video data memory 230, mode selection unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, DPB 218, and entropy coding unit 220 can be implemented in one or more processors or in processing circuitry. For example, the units of the video encoder 200 can be implemented as one or more circuit or logic elements as part of hardware circuitry, or as part of a processor, ASIC, or FPGA. Furthermore, the video encoder 200 may include additional or alternative processors or processing circuitry to perform these and other functions.

[0168] Video data memory 230 is an example of a memory system capable of storing video data to be encoded by components of video encoder 200. Video encoder 200 may receive video data stored in video data memory 230 from, for example, video source 104 (FIG. 1). DPB 218 is an example of a memory system capable of acting as a reference picture memory, storing reference video data for use when the video encoder 200 predicts subsequent video data. Video data memory 230 and DPB 218 may each be formed from any of a variety of one or more memory devices or memory cells, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. Video data memory 230 and DPB 218 may be provided by the same memory device or separate memory devices. In various examples, video data memory 230 may be on-chip with other components of video encoder 200 (as illustrated), or off-chip relative to those components.

[0169] In this disclosure, references to video data memory 230 should not be construed as limited to memory internal to video encoder 200 (unless specifically described) or memory external to video encoder 200 (unless specifically described). Rather, references to video data memory 230 should be understood as a reference memory storing video data received by video encoder 200 for encoding (e.g., video data for the current block to be encoded). Memory 106 of FIG1 may also provide temporary storage for outputs from various units of video encoder 200.

[0170] Various units in Figure 5 are illustrated to aid in understanding the operations performed by the video encoder 200. Units can be implemented as fixed-function circuits, programmable circuits, or combinations thereof. A fixed-function circuit is a circuit that provides specific functionality and is pre-defined for the operations that can be performed. A programmable circuit is a circuit that can be programmed to perform various tasks and provides flexible functionality for the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. A fixed-function circuit can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is generally immutable. In some examples, one or more units in the unit may be different circuit blocks (fixed-function or programmable), and in some examples, one or more units in the unit may be integrated circuits.

[0171] The video encoder 200 may include an arithmetic logic unit (ALU), an essential function unit (EFU), digital circuitry, analog circuitry, and / or a programmable core formed by programmable circuitry. In an example where the operation of the video encoder 200 is performed using software executed by programmable circuitry, memory 106 (FIG. 1) may store instructions (e.g., object code) of the software received and executed by the video encoder 200, or another memory (not shown) within the video encoder 200 may store such instructions.

[0172] The video data storage unit 230 is configured to store received video data. The video encoder 200 can retrieve images of the video data from the video data storage unit 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data storage unit 230 can be raw video data to be encoded.

[0173] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra-frame prediction unit 226. The mode selection unit 202 may include additional functional units for performing video prediction based on other prediction modes. As an example, the mode selection unit 202 may include a palette unit, an intra-frame block copying unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.

[0174] Mode selection unit 202 typically coordinates multiple coding channels to test combinations of coding parameters and the resulting rate-distortion values ​​for such combinations. Coding parameters may include the CTU-CU partitioning, the prediction mode for the CU, the transformation type of the residual data for the CU, the quantization parameters of the residual data for the CU, etc. Mode selection unit 202 can ultimately select a combination of coding parameters that has a better rate-distortion value compared to other tested combinations.

[0175] The video encoder 200 can divide an image retrieved from the video data storage 230 into a series of CTUs and encapsulate one or more CTUs within a slice. The mode selection unit 202 can divide the image's CTUs according to the tree structure described above (such as an MTT structure, QTBT structure, superblock structure, or the quadtree structure described above). As described above, the video encoder 200 can form one or more CUs by dividing the CTUs according to the tree structure. Such CUs are also commonly referred to as "video blocks" or "blocks".

[0176] Generally, mode selection unit 202 also controls its components (e.g., motion estimation unit 222, motion compensation unit 224, and intra-prediction unit 226) to generate prediction blocks for the current block (e.g., the current CU, or, in HEVC, the overlapping portion of PU and TU). To perform inter-frame prediction for the current block, motion estimation unit 222 may perform a motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously decoded pictures stored in DPB 218). Specifically, motion estimation unit 222 may calculate values ​​representing the similarity between a potential reference block and the current block, for example, based on sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), etc. Motion estimation unit 222 may typically perform these calculations using sample-by-sample differences between the current block and the reference blocks under consideration. Motion estimation unit 222 may identify reference blocks with the lowest values ​​produced by these calculations to indicate the reference block that best matches the current block.

[0177] Motion estimation unit 222 can generate one or more motion vectors (MVs) that define the location of a reference block in a reference image relative to the location of the current block in the current image. Motion estimation unit 222 can then provide the motion vectors to motion compensation unit 224. For example, for unidirectional inter-frame prediction, motion estimation unit 222 can provide a single motion vector, while for bidirectional inter-frame prediction, motion estimation unit 222 can provide two motion vectors. Motion compensation unit 224 can then use the motion vectors to generate a prediction block. For example, motion compensation unit 224 can use the motion vectors to retrieve data for the reference block. As another example, if the motion vectors have fractional sample accuracy, motion compensation unit 224 can interpolate the prediction block according to one or more interpolation filters. Furthermore, for bidirectional inter-frame prediction, motion compensation unit 224 can retrieve data for two reference blocks identified by corresponding motion vectors and combine the retrieved data, for example, by per-sample averaging or weighted averaging.

[0178] When operating according to the AV1 video decoding format, the motion estimation unit 222 and the motion compensation unit 224 can be configured to encode the decoded blocks of video data (e.g., both luma and chroma decoded blocks) using translational motion compensation, affine motion compensation, overlap block motion compensation (OBMC), and / or composite inter-intra-frame prediction.

[0179] As another example, for intra-prediction or intra-prediction decoding, intra-prediction unit 226 may generate a prediction block from samples adjacent to the current block. For example, for directional mode, intra-prediction unit 226 may typically mathematically combine the values ​​of adjacent samples and fill these calculated values ​​across the current block in a defined direction to produce a prediction block. As another example, for DC mode, intra-prediction unit 226 may calculate the average of the adjacent samples of the current block and generate a prediction block to include the resulting average for each sample of the prediction block.

[0180] When operating according to the AV1 video decoding format, the intra-frame prediction unit 226 can be configured to encode decoded blocks of video data (e.g., both luma and chroma decoded blocks) using directional intra-frame prediction, non-directional intra-frame prediction, recursive filter intra-frame prediction, luma-chroma (CFL) prediction, intra-block copying (IBC), and / or palette modes. The mode selection unit 202 may include additional functional units for performing video prediction based on other prediction modes.

[0181] Mode selection unit 202 provides a prediction block to residual generation unit 204. Residual generation unit 204 receives an uncoded raw version of the current block from video data memory 230 and a prediction block from mode selection unit 202. Residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines the residual block for the current block. In some examples, residual generation unit 204 may also determine the differences between sample values ​​in the residual block to generate the residual block using residual differential pulse decoding modulation (RDPCM). In some examples, residual generation unit 204 may be formed using one or more subtractor circuits performing binary subtraction.

[0182] In the example where mode selection unit 202 divides a CU into PUs, each PU can be associated with a luma prediction unit and a corresponding chroma prediction unit. Video encoder 200 and video decoder 300 can support PUs of various sizes. As noted above, the size of a CU can refer to the size of the luma decoding block of the CU, while the size of a PU can refer to the size of the luma prediction unit of the PU. Assuming a particular CU size is 2Nx2N, video encoder 200 can support PU sizes of 2Nx2N or NxN for intra-frame prediction, and symmetric PU sizes of 2Nx2N, 2NxN, Nx2N, NxN, or similar for inter-frame prediction. Video encoder 200 and video decoder 300 can also support asymmetric partitioning for PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter-frame prediction.

[0183] In an example where mode selection unit 202 does not further divide the CU into PUs, each CU can be associated with a luminance decoding block and a corresponding chrominance decoding block. As mentioned above, the size of the CU can refer to the size of the luminance decoding block of the CU. The video encoder 200 and the video decoder 300 can support CU sizes of 2Nx2N, 2NxN, or Nx2N.

[0184] For other video decoding techniques, such as intra-block copy mode decoding, affine mode decoding, and linear model (LM) mode decoding, as some examples, mode selection unit 202 generates a prediction block for the current block being encoded via a corresponding unit associated with the decoding technique. In some examples (such as palette mode decoding), mode selection unit 202 may not generate a prediction block, but instead generate syntax elements indicating how the block is reconstructed based on a selected palette. In such modes, mode selection unit 202 may provide these syntax elements to entropy coding unit 220 for encoding.

[0185] As described above, the residual generation unit 204 receives video data for the current block and the corresponding prediction block. Then, the residual generation unit 204 generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block.

[0186] Transform processing unit 206 applies one or more transformations to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 may apply various transformations to the residual block to form the transform coefficient block. For example, transform processing unit 206 may apply a discrete cosine transform (DCT), direction transform, Karhunen-Loeve transform (KLT), or conceptually similar transformations to the residual block. In some examples, transform processing unit 206 may perform multiple transformations on the residual block, such as primary and secondary transformations (e.g., rotation transformations). In some examples, transform processing unit 206 does not apply any transformations to the residual block.

[0187] When operating according to AV1, transform processing unit 206 may apply one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 may apply various transforms to the residual block to form the transform coefficient block. For example, transform processing unit 206 may apply a combination of horizontal / vertical transforms, which may include the Discrete Cosine Transform (DCT), the Asymmetric Discrete Sine Transform (ADST), the Reversed ADST (e.g., ADST in reverse order), and the Identity Transform (IDTX). When using the Identity Transform, the transform is skipped in either the vertical or horizontal direction. In some examples, transform processing may be skipped entirely.

[0188] Quantization unit 208 quantizes the transform coefficients in a transform coefficient block to produce a quantized transform coefficient block. Quantization unit 208 quantizes the transform coefficients of the transform coefficient block according to the quantization parameter (QP) value associated with the current block. Video encoder 200 (e.g., via mode selection unit 202) can adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may cause information loss, and therefore, the quantized transform coefficients may have lower accuracy compared to the original transform coefficients produced by transform processing unit 206.

[0189] The inverse quantization unit 210 and the inverse transform processing unit 212 can apply inverse quantization and inverse transform, respectively, to the quantized transform coefficient block to reconstruct the residual block based on the transform coefficient block. The reconstruction unit 214 can generate a reconstructed block corresponding to the current block (although potentially with some degree of distortion) based on the reconstructed residual block and the prediction block generated by the mode selection unit 202. For example, the reconstruction unit 214 can add samples of the reconstructed residual block to corresponding samples of the prediction block generated by the mode selection unit 202 to generate the reconstructed block.

[0190] Filter unit 216 may perform one or more filtering operations on the reconstructed block. For example, filter unit 216 may perform a deblocking operation to reduce block artifacts along the edges of the CU. In some examples, the operation of filter unit 216 may be skipped.

[0191] When operating according to AV1, filter unit 216 may perform one or more filtering operations on the reconstructed block. For example, filter unit 216 may perform a deblocking operation to reduce block artifacts along the edges of the CU. In other examples, filter unit 216 may apply a constrained direction enhancement filter (CDEF) after deblocking and may include the application of a non-separable, nonlinear, low-pass directional filter based on the estimated edge direction. Filter unit 216 may also include a loop recovery filter applied after CDEF and may include a separable symmetric normalized Wiener filter or a dual-guided filter.

[0192] The video encoder 200 stores reconstructed blocks in the DPB 218. For example, in an example where the filter unit 216 is not operated, the reconstruction unit 214 may store reconstructed blocks in the DPB 218. In an example where the filter unit 216 is operated, the filter unit 216 may store filtered reconstructed blocks in the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 may retrieve reference images formed by the reconstructed (and potentially filtered) blocks from the DPB 218 to perform inter-frame prediction for blocks of subsequent encoded images. Additionally, the intra-frame prediction unit 226 may use the reconstructed blocks of the current image in the DPB 218 to perform intra-frame prediction for other blocks in the current image.

[0193] The filter unit 216 may be further configured to perform any of the adaptive video filtering techniques discussed above with reference to Figures 2 through 4. For example, the filter unit 216 may be configured to receive video data and apply filters to samples of multiple types of the video data.

[0194] Generally, entropy coding unit 220 can entropy code syntax elements received from other functional components of video encoder 200. For example, entropy coding unit 220 can entropy code quantized transform coefficient blocks from quantization unit 208. As another example, entropy coding unit 220 can entropy code predictive syntax elements (e.g., motion information for inter-frame prediction or intra-frame mode information for intra-frame prediction) from mode selection unit 202. Entropy coding unit 220 can perform one or more entropy coding operations on syntax elements (another example of video data) to generate entropy-coded data. For example, entropy coding unit 220 can perform context-adaptive variable-length decoding (CAVLC), CABAC, variable-to-variable (V2V) length decoding, syntax-based context-adaptive binary arithmetic decoding (SBAC), probability interval partitioning entropy (PIPE) decoding, exponential Golomb coding, or another type of entropy coding operation on the data. In some examples, entropy coding unit 220 can operate in a bypass mode where syntax elements are not entropy-coded.

[0195] The video encoder 200 can output a bitstream that includes the entropy coding syntax elements required to reconstruct slices or blocks of images. Specifically, the entropy coding unit 220 can output a bitstream.

[0196] According to AV1, entropy coding unit 220 can be configured as a symbol-to-symbol adaptive multi-symbol arithmetic decoder. The syntax elements in AV1 consist of an N-element alphabet, and the context (e.g., a probability model) consists of a set of N probabilities. Entropy coding unit 220 can store the probabilities as an n-bit (e.g., 15-bit) cumulative distribution function (CDF). Entropy coding unit 220 can perform recursive scaling using an update factor based on the alphabet size to update the context.

[0197] The operations described above are relative to blocks. This description should be understood as operations applied to luma decoding blocks and / or chroma decoding blocks. As described above, in some examples, the luma decoding block and chroma decoding block are the luma and chroma components of the CU. In some examples, the luma decoding block and chroma decoding block are the luma and chroma components of the PU.

[0198] In some examples, it is not necessary to repeat the operations performed relative to the luma decoder for the chroma decoder block. As an example, the operations for identifying the motion vector (MV) and reference image of the luma decoder block do not need to repeat the MV and reference image used to identify the chroma block. Instead, the MV used for the luma decoder block can be scaled to determine the MV used for the chroma block, and the reference image can be the same. As another example, the intra-frame prediction process can be the same for both the luma and chroma decoders.

[0199] Video encoder 200 represents an example of a device configured to encode video data, the device including: a memory configured to store the video data; and one or more processing units implemented in circuitry and configured to apply filters to the input or output of one or more types of samples.

[0200] Figure 6 is a block diagram illustrating an example video decoder 300 capable of implementing the techniques of this disclosure. Figure 6 is provided for illustrative purposes and is not intended to limit the techniques extensively illustrated and described in this disclosure. For illustrative purposes, the video decoder 300 is described in accordance with VVC and HEVC techniques. However, the techniques of this disclosure can be implemented by video decoding devices configured for other video decoding standards.

[0201] In the example of Figure 6, the video decoder 300 includes a decoded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a DPB 314. Any or all of the CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filter unit 312, and DPB 314 can be implemented in one or more processors or in processing circuitry. For example, the units of the video decoder 300 can be implemented as one or more circuit or logic elements as part of hardware circuitry, or as part of a processor, ASIC, or FPGA. Furthermore, the video decoder 300 may include additional or alternative processors or processing circuitry to perform these and other functions.

[0202] The prediction processing unit 304 includes a motion compensation unit 316 and an intra-prediction unit 318. The prediction processing unit 304 may include additional units that perform predictions based on other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, an intra-block copying unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, the video decoder 300 may include more, fewer, or different functional components.

[0203] When operating according to AV1, motion compensation unit 316 can be configured to decode video data blocks (e.g., both luma and chroma blocks) using translational motion compensation, affine motion compensation, OBMC, and / or composite inter-intra-frame prediction, as described above. Intra-frame prediction unit 318 can be configured to decode video data blocks (e.g., both luma and chroma blocks) using directional intra-frame prediction, non-directional intra-frame prediction, recursive filter intra-frame prediction, CFL, IBC, and / or palette mode, as described above.

[0204] CPB memory 320 is an example of a memory system that can store video data (such as encoded video bitstreams) to be decoded by components of video decoder 300. For example, video data stored in CPB memory 320 can be obtained from computer-readable medium 110 (FIG. 1). CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Moreover, CPB memory 320 may store video data other than syntax elements of decoded images, such as temporary data representing the output from various units of video decoder 300. DPB 314 is an example of a memory system that typically stores decoded images, which video decoder 300 can output, and / or uses as reference video data when decoding subsequent data or images of the encoded video bitstream. CPB memory 320 and DPB 314 may each be formed from any of various memory devices or memory cells, such as DRAM (including SDRAM), MRAM, RRAM, or other types of memory devices. CPB memory 320 and DPB 314 may be provided by the same memory device or separate memory devices. In various examples, CPB memory 320 may be on-chip along with other components of video decoder 300, or off-chip relative to those components.

[0205] Additionally or alternatively, in some examples, the video decoder 300 may retrieve decoded video data from memory 120 (FIG. 1). That is, memory 120 may utilize CPB memory 320 to store data as discussed above. Similarly, when some or all of the functionality of the video decoder 300 is implemented in software to be executed by the processing circuitry of the video decoder 300, memory 120 may store instructions to be executed by the video decoder 300.

[0206] The various units shown in Figure 6 are illustrated to aid in understanding the operations performed by the video decoder 300. Units can be implemented as fixed-function circuits, programmable circuits, or combinations thereof. Similar to Figure 5, a fixed-function circuit refers to a circuit that provides specific functionality and is pre-defined for the operations that can be performed. A programmable circuit refers to a circuit that can be programmed to perform various tasks and provides flexible functionality for the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. A fixed-function circuit can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is generally immutable. In some examples, one or more units in the unit may be different circuit blocks (fixed-function or programmable), and in some examples, one or more units in the unit may be integrated circuits.

[0207] The video decoder 300 may include an ALU, EFU, digital circuitry, analog circuitry, and / or a programmable core formed by programmable circuitry. In an example where the operation of the video decoder 300 is performed by software executed on programmable circuitry, on-chip or off-chip memory may store instructions (e.g., object code) of the software received and executed by the video decoder 300.

[0208] The entropy decoding unit 302 can receive encoded video data from the CPB and perform entropy decoding on the video data to reproduce the syntax elements. The prediction processing unit 304, the inverse quantization unit 306, the inverse transform processing unit 308, the reconstruction unit 310, and the filter unit 312 can generate decoded video data based on the syntax elements extracted from the bitstream.

[0209] Generally speaking, the video decoder 300 reconstructs the image block by block. The video decoder 300 can perform the reconstruction operation on each block individually (where the block currently being reconstructed (i.e., decoded) can be referred to as the "current block").

[0210] Entropy decoding unit 302 can entropy decode the syntax elements of the quantized transform coefficients defining the quantized transform coefficient block, as well as transform information (such as quantization parameters (QP) and / or transform mode indications). Inverse quantization unit 306 can use the QP associated with the quantized transform coefficient block to determine the degree of quantization, and similarly, determine the degree of inverse quantization to be applied by inverse quantization unit 306. Inverse quantization unit 306 can, for example, perform a bit-by-bit left shift operation to inverse quantize the quantized transform coefficients. Inverse quantization unit 306 can thereby form a transform coefficient block including the transform coefficients.

[0211] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 may apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 may apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotation transform, an inverse direction transform, or another inverse transform to the transform coefficient block.

[0212] Furthermore, the prediction processing unit 304 generates a prediction block based on the prediction information syntax elements entropy-decoded by the entropy decoding unit 302. For example, if the prediction information syntax elements indicate that the current block is an inter-frame prediction, the motion compensation unit 316 may generate the prediction block. In this case, the prediction information syntax elements may indicate the reference picture from which the reference block is to be retrieved in the DPB 314, and a motion vector identifying the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion compensation unit 316 may generally perform the inter-frame prediction process in a manner substantially similar to that described with respect to the motion compensation unit 224 (FIG. 5).

[0213] As another example, when the prediction information syntax element indicates that the current block is intra-prediction, the intra-prediction unit 318 can generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Similarly, the intra-prediction unit 318 can generally perform the intra-prediction process in a manner substantially similar to that described with respect to intra-prediction unit 226 (FIG. 5). The intra-prediction unit 318 can retrieve data of neighboring samples of the current block from the DPB 314.

[0214] Reconstruction unit 310 can use the prediction block and the residual block to reconstruct the current block. For example, reconstruction unit 310 can add samples from the residual block to the corresponding samples from the prediction block to reconstruct the current block.

[0215] Filter unit 312 can perform one or more filtering operations on the reconstructed block. For example, filter unit 312 can perform a deblocking operation to reduce block artifacts along the edges of the reconstructed block. The operation of filter unit 312 is not necessarily performed in all examples.

[0216] The video decoder 300 can store reconstructed blocks in the DPB 314. For example, in an example where the operation of the filter unit 312 is not performed, the reconstruction unit 310 can store the reconstructed blocks in the DPB 314. In an example where the operation of the filter unit 312 is performed, the filter unit 312 can store the filtered reconstructed blocks in the DPB 314. As discussed above, the DPB 314 can provide reference information (such as samples of the current image for intra-frame prediction and samples of previously decoded images for subsequent motion compensation) to the prediction processing unit 304. Furthermore, the video decoder 300 can output decoded images (e.g., decoded video) from the DPB 314 for subsequent rendering on a display device such as the display device 118 of FIG. 1.

[0217] The filter unit 312 may be further configured to perform any of the adaptive video filtering techniques discussed above with reference to Figures 2 through 4. For example, the filter unit 312 may be configured to receive video data and apply filters to samples of multiple types of the video data.

[0218] In this way, video decoder 300 represents an example of a video decoding device, which includes: a memory configured to store video data; and one or more processing units implemented in a circuit and configured to apply filters to the input or output of one or more types of samples.

[0219] Figure 7 is a flowchart illustrating an example method for encoding a current block according to the technology of this disclosure. The current block may be or may include the current CU. Although described with respect to video encoder 200 (Figures 1 and 5), it should be understood that other devices may be configured to perform methods similar to those of Figure 7.

[0220] In this example, the video encoder 200 initially predicts the current block (400). For example, the video encoder 200 may form a prediction block for the current block. The video encoder 200 may then compute a residual block for the current block (402). To compute the residual block, the video encoder 200 may compute the difference between the unencoded original block for the current block and the prediction block. The video encoder 200 may then transform the residual block and quantize the transform coefficients of the residual block (404). Next, the video encoder 200 may scan the quantized transform coefficients of the residual block (406). During or after the scan, the video encoder 200 may entropy encode the transform coefficients (408). For example, the video encoder 200 may use CAVLC or CABAC to encode the transform coefficients. The video encoder 200 may then output the entropy-encoded data of the block (410).

[0221] Figure 8 is a flowchart illustrating an example method for decoding a current block of video data according to the technology of this disclosure. The current block may be or may include the current CU. Although described with respect to video decoder 300 (Figures 1 and 6), it should be understood that other devices may be configured to perform methods similar to those in Figure 8.

[0222] The video decoder 300 may receive entropy-coded data for the current block, such as entropy-coded prediction information and entropy-coded data for the transform coefficients of the residual block corresponding to the current block (500). The video decoder 300 may entropy decode the entropy-coded data to determine prediction information for the current block and reproduce the transform coefficients of the residual block (502). The video decoder 300 may predict the current block, for example, using an intra-frame prediction mode or inter-frame prediction mode indicated by the prediction information of the current block (504), to compute a prediction block for the current block. The video decoder 300 may then inverse scan the reproduced transform coefficients (506) to create a block of quantized transform coefficients. The video decoder 300 may then inverse quantize the transform coefficients and apply an inverse transform to the transform coefficients to produce a residual block (508). The video decoder 300 may finally decode the current block by combining the prediction block and the residual block (510).

[0223] Figure 9 is a flowchart illustrating an example method according to the technology of this disclosure. The technology of Figure 9 can be performed by one or more components (including filter unit 216 and filter unit 312) of the video encoder 200 and the video decoder 300, respectively.

[0224] In one example of this disclosure, video encoder 200 and video decoder 300 may be configured to receive video data (900) and apply filters (902) to samples of multiple types of video data. In one example, the multiple types of samples include two or more of the following: input to a fixed filter, output to a fixed filter, input to a signaling filter, output to a signaling filter, input to an adaptive loop filter, output to an adaptive loop filter, input to a sample adaptive offset (SAO) filter, output to a SAO filter, input to a bilateral filter, output to a bilateral filter, input to a cross-component SAO filter, output to a cross-component SAO filter, input to a deblocking filter, output to a deblocking filter, filtered reconstructed residual data, dequantization coefficients, filtered dequantization coefficients, predictors, or filtered predictors.

[0225] Figure 10 is a flowchart illustrating another example method according to the technology of this disclosure. More specifically, Figure 10 illustrates a particular example of applying filters to samples of video data of various kinds. For example, video encoder 200 and video decoder 300 may be configured to receive a first output (910) of a signaling filter, a second output (912) of a fixed filter, and apply filters (914) to the first and second outputs. In this example, the signaling filter includes one of an adaptive loop filter, a sample adaptive offset (SAO) filter, a bilateral filter, a cross-component SAO filter, or a deblocking filter.

[0226] In another example, to apply filters to samples of multiple types of video data, the video encoder 200 and video decoder 300 may be configured to receive first corresponding outputs of two or more signaled filters, receive second corresponding outputs of one or more fixed filters, and apply filters to the first and second corresponding outputs. In this example, the two or more signaled filters include one or more adaptive loop filters, one or more sample adaptive offset (SAO) filters, one or more bilateral filters, one or more cross-component SAO filters, and one or more deblocking filters.

[0227] In either of the examples in Figure 9 or Figure 10, the video encoder 200 and the video decoder 300 can be configured to apply a classifier to the output sample values ​​of samples of multiple classes of video data. In one example, the classifier is a classifier for a fixed filter or a classifier for a signaling filter.

[0228] The following numbered clauses illustrate one or more aspects of the devices and technologies described in this disclosure.

[0229] Aspect 1A - A method for decoding video data, the method comprising: applying a filter to the input or output of one or more types of samples.

[0230] Aspect 2A - According to the method of aspect 1A, applying the filter to the input or output of the one or more types of samples includes: applying the filter to the input or output of multiple types of samples.

[0231] Aspect 3A - The method according to Aspect 2A, wherein the plurality of samples comprises one or more of the following: an input to a fixed filter, an output to the fixed filter, an input to a signaling filter, an output to the signaling filter, an input to an adaptive loop filter, an output to the adaptive loop filter, an input to a sample adaptive offset (SAO) filter, an output to the SAO filter, an input to a bilateral filter, an output to the bilateral filter, an input to a cross-component SAO filter, an output to the cross-component SAO filter, an input to a deblocking filter, an output to the deblocking filter, filtered reconstructed residual data, dequantization coefficients, filtered dequantization coefficients, a predictor, or a filtered predictor.

[0232] Aspect 4A - The method according to aspect 1A further includes: applying a classifier to output sample values ​​of the one or more categories of samples.

[0233] Aspect 5A - The method according to aspect 4A, wherein the classifier is a classifier for a fixed filter or a classifier for a signaling filter.

[0234] Aspect 6A - A method according to any one of Aspects 1A to 5A, wherein decoding includes decoding.

[0235] Aspect 7A - The method according to any one of aspects 1A to 5A, wherein decoding includes encoding.

[0236] Aspect 8A - An apparatus for decoding video data, the apparatus comprising one or more components for performing a method according to any one of aspects 1A to 7A.

[0237] Aspect 9A - The device according to aspect 8A, wherein the one or more components include one or more processors implemented in a circuit.

[0238] Aspect 10A - The device according to any one of claims 8A and 9A, the device further comprising a memory for storing the video data.

[0239] Aspect 11A - The device according to any one of aspects 8A to 10A, the device further comprising a display configured to display decoded video data.

[0240] Aspect 12A - The device according to any one of aspects 8A to 11A, wherein the device includes one or more of a camera, computer, mobile device, broadcast receiver device or set-top box.

[0241] Aspect 13A - The apparatus according to any one of aspects 8A to 12A, wherein the apparatus includes a video decoder.

[0242] Aspect 14A - The device according to any one of aspects 8A to 13A, wherein the device includes a video encoder.

[0243] Aspect 15A - A computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to perform a method according to any one of aspects 1A to 7A.

[0244] Aspect 1B - A method for decoding video data, the method comprising: receiving video data; and applying filters to samples of the video data of multiple types.

[0245] Aspect 2B - According to the method of aspect 1B, wherein the plurality of samples comprises two or more of the following: an input to a fixed filter, an output to the fixed filter, an input to a signaling filter, an output to the signaling filter, an input to an adaptive loop filter, an output to the adaptive loop filter, an input to a sample adaptive offset (SAO) filter, an output to the SAO filter, an input to a bilateral filter, an output to the bilateral filter, an input to a cross-component SAO filter, an output to the cross-component SAO filter, an input to a deblocking filter, an output to the deblocking filter, filtered reconstructed residual data, dequantization coefficients, filtered dequantization coefficients, a predictor, or a filtered predictor.

[0246] Aspect 3B - The method according to any one of Aspects 1B to 2B, wherein applying the filter to the plurality of types of samples of the video data comprises: receiving a first output of a signaling filter; receiving a second output of a fixed filter; and applying the filter to the first output and the second output.

[0247] Aspect 4B - The method according to aspect 3B, wherein the signaling filter includes one of an adaptive loop filter, a sample adaptive offset (SAO) filter, a bilateral filter, a cross component SAO filter, or a deblocking filter.

[0248] Aspect 5B - The method according to any one of Aspects 1B to 2B, wherein applying the filter to samples of the plurality of types of the video data comprises: receiving a first corresponding output of two or more signaling filters; receiving a second corresponding output of one or more fixed filters; and applying the filter to the first corresponding output and the second corresponding output.

[0249] Aspect 6B - According to the method of aspect 5B, the two or more signaling filters include one or more adaptive loop filters, one or more sample adaptive offset (SAO) filters, one or more bilateral filters, one or more cross-component SAO filters, and one or more deblocking filters.

[0250] Aspect 7B - The method according to any one of Aspects 1B to 6B, the method further comprising: applying a classifier to the output sample values ​​of the plurality of types of samples of the video data.

[0251] Aspect 8B - The method according to aspect 7B, wherein the classifier is a classifier for a fixed filter or a classifier for a signaling filter.

[0252] Aspect 9B - The method according to any one of aspects 1B to 8B, wherein decoding includes decoding.

[0253] Aspect 10B - The method according to any one of aspects 1B to 8B, wherein decoding includes encoding.

[0254] Aspect 11B - An apparatus configured to decode video data, the apparatus comprising: a memory; and processing circuitry communicating with the memory, the processing circuitry being configured to: receive video data; and apply filters to samples of the video data of multiple types.

[0255] Aspect 12B - The apparatus according to aspect 11B, wherein the plurality of samples comprises two or more of the following: an input to a fixed filter, an output to the fixed filter, an input to a signaling filter, an output to the signaling filter, an input to an adaptive loop filter, an output to the adaptive loop filter, an input to a sample adaptive offset (SAO) filter, an output to the SAO filter, an input to a bilateral filter, an output to the bilateral filter, an input to a cross-component SAO filter, an output to the cross-component SAO filter, an input to a deblocking filter, an output to the deblocking filter, filtered reconstructed residual data, dequantization coefficients, filtered dequantization coefficients, a predictor, or a filtered predictor.

[0256] Aspect 13B - The apparatus according to any one of aspects 11B to 12B, wherein, in order to apply the filter to the plurality of types of samples of the video data, the processing circuit is further configured to: receive a first output of a signal-notified filter; receive a second output of a fixed filter; and apply the filter to the first output and the second output.

[0257] Aspect 14B - The apparatus according to aspect 13B, wherein the signaling filter comprises one of an adaptive loop filter, a sample adaptive offset (SAO) filter, a bilateral filter, a cross component SAO filter, or a deblocking filter.

[0258] Aspect 15B - An apparatus according to any one of aspects 11B to 12B, wherein, in order to apply the filter to samples of the plurality of types of the video data, the processing circuitry is further configured to: receive a first corresponding output of two or more signaled filters; receive a second corresponding output of one or more fixed filters; and apply the filter to the first corresponding output and the second corresponding output.

[0259] Aspect 16B - The apparatus according to aspect 15B, wherein the two or more signaling filters include one or more adaptive loop filters, one or more sample adaptive offset (SAO) filters, one or more bilateral filters, one or more cross-component SAO filters, and one or more deblocking filters.

[0260] Aspect 17B - An apparatus according to any one of aspects 11B to 16B, wherein the processing circuitry is further configured to apply a classifier to the output sample values ​​of the plurality of types of samples of the video data.

[0261] Aspect 18B - The apparatus according to aspect 17B, wherein the classifier is a classifier for a fixed filter or a classifier for a signaling filter.

[0262] Aspect 19B - An apparatus according to any one of aspects 11B to 18B, wherein the apparatus is configured to decode video data.

[0263] Aspect 20B - The apparatus according to any one of aspects 11B to 18B, wherein the apparatus is configured to encode video data.

[0264] It should be recognized that, based on the examples, certain actions or events of any technique described herein may be performed in a different sequence, and may be added, combined, or omitted entirely (e.g., not all actions or events described are necessary for implementing the technique). Furthermore, in some examples, actions or events may be performed concurrently (e.g., through multithreading, interrupt handling, or multiple processors) rather than sequentially.

[0265] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored as one or more instructions or code on a computer-readable medium or transmitted via a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include a computer-readable storage medium (which corresponds to a tangible medium such as a data storage medium) or a communication medium, including, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol. Thus, a computer-readable medium may generally correspond to (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products may include computer-readable media.

[0266] By way of example, and not limitation, such computer-readable storage media may include one or more of RAM, ROM, EEPROM, CD-ROM or other optical disc storage devices, magnetic disk storage devices or other magnetic storage devices, flash memory, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, and microwave), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology (such as infrared, radio, and microwave) is included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but instead refer to non-transient tangible storage media. As used herein, disks and optical discs include compact optical discs (CDs), laser optical discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs use lasers to reproduce data optically. Combinations of these should also be included within the scope of computer-readable media.

[0267] Instructions can be executed by one or more processors, such as one or more DSPs, general-purpose microprocessors, ASICs, FPGAs, or other equivalent integrated or discrete logic circuits. Therefore, the terms "processor" and "processing circuit" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Furthermore, these techniques can be fully implemented in one or more circuit or logic elements.

[0268] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or IC sets (e.g., chip sets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but implementation by different hardware units is not necessarily required. Specifically, as described above, various units may be combined in a codec hardware unit, or various units may be provided by a collection of interoperable hardware units (including one or more processors as described above) combined with appropriate software and / or firmware.

[0269] Various examples have been described. These and other examples are within the scope of the following claims.

Claims

1. A method for decoding video data, the method comprising: Receive video data; And filters are applied to samples of multiple types of the video data.

2. The method of claim 1, wherein the plurality of samples comprises two or more of the following: an input to a fixed filter, the output of the fixed filter, an input to a signaling filter, the output of the signaling filter, an input to an adaptive loop filter, the output of the adaptive loop filter, an input to a sample adaptive offset (SAO) filter, the output of the SAO filter, an input to a bilateral filter, the output of the bilateral filter, an input to a cross-component SAO filter, the output of the cross-component SAO filter, an input to a deblocking filter, the output of the deblocking filter, filtered reconstructed residual data, dequantization coefficients, filtered dequantization coefficients, a predictor, or a filtered predictor.

3. The method of claim 1, wherein applying the filter to the plurality of types of samples of the video data comprises: The first output of the filter that receives the signal notification; Receive the second output of the fixed filter; And apply the filter to the first output and the second output.

4. The method of claim 3, wherein the signaling filter comprises one of an adaptive loop filter, a sample adaptive offset (SAO) filter, a bilateral filter, a cross-component SAO filter, or a deblocking filter.

5. The method of claim 1, wherein applying the filter to the plurality of types of samples of the video data comprises: Receive the first corresponding output of two or more signaling filters; Receives the second corresponding output of one or more fixed filters; And apply the filter to the first corresponding output and the second corresponding output.

6. The method of claim 5, wherein the two or more signaling filters comprise one or more adaptive loop filters, one or more sample adaptive offset (SAO) filters, one or more bilateral filters, one or more cross-component SAO filters, and one or more deblocking filters.

7. The method according to claim 1, further comprising: A classifier is applied to the output sample values ​​of the various types of samples in the video data.

8. The method of claim 7, wherein the classifier is a classifier for a fixed filter or a classifier for a signaling filter.

9. The method of claim 1, wherein decoding includes decoding.

10. The method of claim 1, wherein decoding includes encoding.

11. An apparatus configured to decode video data, the apparatus comprising: Memory; and processing circuitry that communicates with the memory, the processing circuitry being configured to receive video data; And filters are applied to samples of multiple types of the video data.

12. The apparatus of claim 11, wherein the plurality of samples comprises two or more of the following: an input to a fixed filter, an output to the fixed filter, an input to a signaling filter, an output to the signaling filter, an input to an adaptive loop filter, an output to the adaptive loop filter, an input to a sample adaptive offset (SAO) filter, an output to the SAO filter, an input to a bilateral filter, an output to the bilateral filter, an input to a cross-component SAO filter, an output to the cross-component SAO filter, an input to a deblocking filter, an output to the deblocking filter, filtered reconstructed residual data, dequantization coefficients, filtered dequantization coefficients, a predictor, or a filtered predictor.

13. The apparatus of claim 11, wherein, in order to apply the filter to the plurality of types of samples of the video data, the processing circuit is further configured to: receive a first output of a signaling filter; receive a second output of a fixed filter; and apply the filter to the first output and the second output.

14. The apparatus of claim 13, wherein the signaling filter comprises one of an adaptive loop filter, a sample adaptive offset (SAO) filter, a bilateral filter, a cross-component SAO filter, or a deblocking filter.

15. The apparatus of claim 11, wherein, in order to apply the filter to the plurality of types of samples of the video data, the processing circuitry is further configured to: receive a first corresponding output of two or more signaled filters; receive a second corresponding output of one or more fixed filters; and apply the filter to the first corresponding output and the second corresponding output.

16. The apparatus of claim 15, wherein the two or more signaling filters comprise one or more adaptive loop filters, one or more sample adaptive offset (SAO) filters, one or more bilateral filters, one or more cross-component SAO filters, and one or more deblocking filters.

17. The apparatus of claim 11, wherein the processing circuitry is further configured to apply a classifier to the output sample values ​​of the plurality of types of samples of the video data.

18. The apparatus of claim 17, wherein the classifier is a classifier for a fixed filter or a classifier for a signaling filter.

19. The apparatus of claim 11, wherein the apparatus is configured to decode video data.

20. The apparatus of claim 11, wherein the apparatus is configured to encode video data.

Citation Information

Patent Citations

  • Adaptive video filter

    US20240357095A1