Adaptive video filter

By calculating the Laplacian activity value and category index of video blocks, the problem of bandwidth and processing power waste in adaptive video filtering is solved, thereby improving video quality and efficiency.

CN120982090APending Publication Date: 2025-11-18QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202480022360.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-03-14
Filing Date
2024-03-15
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

Existing adaptive video filtering techniques result in significant bandwidth consumption and wasted processing power, and require the calculation of differences that are known to be zero.

Method used

By determining the first value associated with the first window, calculating the sample difference within the second window, and decoding the video block based on the Laplacian activity value and category index, signaling and computational requirements are reduced.

Benefits of technology

Improve video quality, reduce bandwidth consumption and processing power usage, and simplify the video decoding process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120982090A_ABST
    Figure CN120982090A_ABST
Patent Text Reader

Abstract

An example device includes one or more memories and one or more processors coupled to the one or more memories. The one or more processors are configured to determine a first value associated with a first window containing a target block of video data. The one or more processors are configured to determine a respective difference between each sample value and the first value within a second window, the second window including the target block. The one or more processors are configured to determine a second value based on the respective difference. The one or more processors are configured to determine a Laplace activity value for the target block. The one or more processors are configured to determine a category index based on the second value and the Laplacian activity value, and to decode the target block based on the category index.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to the following applications: U.S. Patent Application No. 18 / 605,416, filed March 14, 2024; U.S. Provisional Patent Application No. 63 / 495,981, filed April 13, 2023; U.S. Provisional Patent Application No. 63 / 507,021, filed June 8, 2023; and U.S. Provisional Patent Application No. 63 / 508,765, filed June 16, 2023, the entire contents of each of which are incorporated herein by reference. U.S. Patent Application No. 18 / 605,416, filed March 14, 2024, claims the benefits of the following applications: U.S. Provisional Patent Application No. 63 / 495,981, filed April 13, 2023; U.S. Provisional Patent Application No. 63 / 507,021, filed June 8, 2023; and U.S. Provisional Patent Application No. 63 / 508,765, filed June 16, 2023. Technical Field

[0002] This disclosure relates to video encoding and video decoding. Background Technology

[0003] Digital video capabilities can be incorporated into a wide variety of devices, including digital televisions, digital live broadcast systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, tablets, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radio phones (so-called "smartphones"), video conferencing equipment, video streaming devices, and more. Digital video devices implement video decoding technologies such as those defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 (Part 10, Advanced Video Decoding (AVC)), ITU-T H.265 / High Efficiency Video Decoding (HEVC), ITU-T H.266 / Various Universal Video Decoding (VVC) and extensions to these standards, as well as proprietary video codecs / formats (such as those described in AOMediaVideo1 (AV1) developed by the Open Media Alliance). By implementing such video decoding technologies, video devices can send, receive, encode, decode, and / or store digital video information more efficiently.

[0004] Video decoding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in video sequences. For block-based video decoding, video slices (e.g., video pictures or portions of video pictures) can be divided into video blocks, which may also be referred to as decoding tree units (CTUs), decoding units (CUs), and / or decoding nodes. Video blocks in a slice of a picture that has been intra-decoded (I) are encoded using spatial prediction relative to reference samples in adjacent blocks within the same picture. Video blocks in a slice of a picture that has been inter-decoded (P or B) can use spatial prediction relative to reference samples in adjacent blocks within the same picture or temporal prediction relative to reference samples in other reference pictures. A picture may be called a frame, and a reference picture may be called a reference frame. Summary of the Invention

[0005] Generally, this disclosure describes techniques for adaptive video filtering. The techniques described herein can lead to improved video quality, reduced bandwidth consumption, reduced processing power usage, etc. Existing adaptive video filtering techniques can result in signaling of numerous syntax elements consuming a relatively large amount of bandwidth to indicate information that the video decoder may need to appropriately determine and apply adaptive video filtering to the video data. Furthermore, existing video filtering techniques may require the video encoder to calculate differences that should be known to be zero before the calculation, which can waste processing power. The techniques described herein can reduce the number of syntax elements signaled and / or remove or simplify calculations that would otherwise be performed by the video decoder.

[0006] In one example, a method includes: determining a first value associated with a first window, the first window comprising a target block of video data; determining a corresponding difference between each sample value within a second window and the first value, the second window comprising the target block, wherein the first window and the second window are the same window or different windows; determining a second value based on the corresponding differences; determining a Laplacian activity value for the target block; determining a category index based on the second value and the Laplacian activity value; and decoding the target block based on the category index.

[0007] In another example, a device includes: one or more memories configured to store video data; and one or more processors embedded in circuitry and coupled to the one or more memories, the processors being configured to: determine a first value associated with a first window, the first window including a target block of video data; determine a corresponding difference between each sample value within a second window and the first value, the second window including the target block, wherein the first window and the second window are the same window or different windows; determine a second value based on the corresponding differences; determine a Laplacian activity value for the target block; determine a category index based on the second value and the Laplacian activity value; and decode the target block based on the category index.

[0008] In another example, a computer-readable storage medium utilizes instructions to encode, when executed, instructions that cause one or more processors to: determine a first value associated with a first window, the first window comprising a target block of video data; determine a corresponding difference between each sample value within a second window and the first value, the second window comprising the target block, wherein the first window and the second window are the same window or different windows; determine a second value based on the corresponding differences; determine a Laplacian activity value for the target block, the second window comprising the target block, wherein the first window and the second window are the same window or different windows; determine a category index based on the second value and the Laplacian activity value; and decode the target block based on the category index.

[0009] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description, drawings, and claims. Attached Figure Description

[0010] Figure 1 This is a block diagram illustrating an example video encoding and decoding system that can perform the techniques described in this disclosure.

[0011] Figure 2 This is a block diagram illustrating an example video encoder that can perform the techniques described in this disclosure.

[0012] Figure 3 This is a block diagram illustrating an example video decoder that can perform the techniques described in this disclosure.

[0013] Figure 4 This is a flowchart illustrating an example adaptive video filter technique according to one or more aspects of this disclosure.

[0014] Figure 5 This is a flowchart illustrating an example method for encoding the current block according to the technology of this disclosure.

[0015] Figure 6This is a flowchart illustrating an example method for decoding the current block according to the technology of this disclosure. Detailed Implementation

[0016] Existing adaptive video filtering techniques can result in signaling for many syntax elements consuming relatively large amounts of bandwidth to represent information that the video decoder may need to appropriately determine and apply adaptive video filtering to the video data. Furthermore, existing video filtering techniques may require the video decoder to calculate differences that should be known to be zero before computation, potentially wasting processing power.

[0017] The techniques described in this article can improve adaptive video filtering. These techniques can lead to improved video quality, reduced bandwidth consumption, and reduced processing power usage.

[0018] Figure 1 This is a block diagram illustrating an example video encoding and decoding system 100 capable of performing the techniques of this disclosure. The techniques of this disclosure generally involve decoding (encoding and / or decoding) video data. Typically, video data includes any data used for processing video. Therefore, video data can include raw, unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata (such as signaling data).

[0019] like Figure 1 As shown in the example, system 100 includes source device 102, which provides encoded video data to be decoded and displayed by destination device 116. Specifically, source device 102 provides the video data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 can be or include any of a wide variety of devices, such as desktop computers, laptop computers, mobile devices, tablet computers, set-top boxes, mobile phones such as smartphones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, broadcast receivers, and so on. In some cases, source device 102 and destination device 116 may be equipped for wireless communication and therefore may be referred to as wireless communication devices.

[0020] exist Figure 1In the example, source device 102 includes a video source 104, memory 106, video encoder 200, and output interface 108. Destination device 116 includes an input interface 122, video decoder 300, memory 120, and display device 118. According to this disclosure, the video encoder 200 of source device 102 and the video decoder 300 of destination device 116 can be configured to apply techniques for adaptive video filtering. Therefore, source device 102 represents an example of a video encoding device, while destination device 116 represents an example of a video decoding device. In other examples, the source and destination devices may include other components or arrangements. For example, source device 102 may receive video data from an external video source such as an external camera. Similarly, destination device 116 may interface with an external display device, rather than including an integrated display device.

[0021] like Figure 1 The system 100 shown is merely an example. Typically, any digital video encoding and / or decoding device can perform techniques for adaptive video filtering. Source device 102 and destination device 116 are merely examples of such decoding devices, where source device 102 generates decoded video data for transmission to destination device 116. This disclosure refers to a “decoding” device as a device that performs the decoding (e.g., encoding and / or decoding) of data. Thus, video encoder 200 and video decoder 300 represent examples of decoding devices, specifically, a video encoder and a video decoder, respectively. In some examples, source device 102 and destination device 116 may operate in a substantially symmetrical manner, such that each of source device 102 and destination device 116 includes video encoding and decoding components. Therefore, system 100 can support one-way or two-way video transmission between source device 102 and destination device 116, for example, for video streaming, video playback, video broadcasting, or video telephony.

[0022] Typically, video source 104 represents a source of video data (i.e., raw, unencoded video data) and provides a sequential series of pictures (also referred to as “frames”) of the video data to video encoder 200, which encodes the data used for the pictures. Video source 104 of source device 102 may include video capture devices such as cameras, video archives containing previously captured raw video, and / or video feed interfaces for receiving video from video content providers. As a further alternative, video source 104 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In each case, video encoder 200 may encode the captured, pre-captured, or computer-generated video data. Video encoder 200 may rearrange the pictures from the receiving order (sometimes referred to as the “display order”) to a decoding order for decoding. Video encoder 200 may generate a bitstream comprising the encoded video data. Then, the source device 102 can output encoded video data to a computer-readable medium 110 via the output interface 108 for reception and / or retrieval by, for example, the input interface 122 of the destination device 116.

[0023] The memory 106 of source device 102 and the memory 120 of destination device 116 represent general-purpose memory. In some examples, memories 106 and 120 may store raw video data, such as raw video from video source 104 and raw decoded video data from video decoder 300. Alternatively or additionally, memories 106 and 120 may store software instructions executable by, for example, video encoder 200 and video decoder 300, respectively. Although memories 106 and 120 are shown separately from video encoder 200 and video decoder 300 in this example, it should be understood that video encoder 200 and video decoder 300 may also include internal memory for functionally similar or equivalent purposes. Furthermore, memories 106 and 120 may store, for example, encoded video data output from video encoder 200 and input to video decoder 300. In some examples, portions of memories 106 and 120 may be allocated as one or more video buffers, for example, to store raw, decoded, and / or encoded video data.

[0024] Computer-readable medium 110 can represent any type of medium or device capable of transmitting encoded video data from source device 102 to destination device 116. In one example, computer-readable medium 110 represents a communication medium that enables source device 102 to directly transmit encoded video data to destination device 116 in real time, for example, via a radio frequency network or a computer-based network. According to a communication standard such as a wireless communication protocol, output interface 108 can modulate the transmitted signal including the encoded video data, and input interface 122 can demodulate the received transmitted signal. The communication medium can include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include a router, switch, base station, or any other device that can be useful for facilitating communication from source device 102 to destination device 116.

[0025] In some examples, source device 102 can output encoded data from output interface 108 to storage device 112. Similarly, destination device 116 can access encoded data from storage device 112 via input interface 122. Storage device 112 may include any data storage medium of various distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.

[0026] In some examples, source device 102 can output encoded video data to file server 114 or another intermediate storage device that can store the encoded video data generated by source device 102. Destination device 116 can access the stored video data from file server 114 via streaming or download.

[0027] File server 114 can be any type of server device capable of storing encoded video data and sending such encoded video data to destination device 116. File server 114 can represent a web server (e.g., for a website), a server configured to provide file transfer protocol services (such as File Transfer Protocol (FTP) or File Transfer over One-Way Transmission (FLUTE) protocol), a Content Delivery Network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Service (MBMS) or Enhanced MBMS (eMBMS) server, and / or a Network Attached Storage (NAS) device. File server 114 may additionally or alternatively implement one or more HTTP streaming protocols, such as HTTP-based Dynamic Adaptive Streaming (DASH), HTTP Real-Time Streaming (HLS), Real-Time Streaming Protocol (RTSP), HTTP Dynamic Streaming, etc.

[0028] Destination device 116 can access encoded video data from file server 114 via any standard data connection, including an internet connection. This may include wireless channels (e.g., Wi-Fi connections), wired connections (e.g., digital subscriber line (DSL), cable modems, etc.), or combinations thereof, suitable for accessing encoded video data stored on file server 114. Input interface 122 can be configured to operate according to any one or more of the various protocols discussed above for retrieving or receiving media data from file server 114, or other such protocols for retrieving media data.

[0029] Output interface 108 and input interface 122 may represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 may be configured to transmit data (such as encoded video data) according to cellular communication standards (such as 4G, 4G-LTE (Long Term Evolution), improved LTE, 5G, etc.). In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 may be configured to operate according to other wireless standards (such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee)). TM Bluetooth TMThe source device 102 and / or destination device 116 may include corresponding system-on-chip (SoC) devices. For example, source device 102 may include an SoC device for performing functions belonging to video encoder 200 and / or output interface 108, and destination device 116 may include an SoC device for performing functions belonging to video decoder 300 and / or input interface 122.

[0030] The technology disclosed herein can be applied to video decoding to support any application in a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission (such as HTTP-based Dynamic Adaptive Streaming (DASH)), digital video encoded onto data storage media, decoding of digital video stored on data storage media, or other applications.

[0031] The input interface 122 of the destination device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., a communication medium, storage device 112, file server 114, etc.). The encoded video bitstream may include signaling information defined by the video encoder 200 and also used by the video decoder 300, such as syntax elements having values ​​describing the characteristics and / or processing of video blocks or other decoding units (e.g., slices, pictures, picture groups, sequences, etc.). The display device 118 displays decoded images of the decoded video data to the user. The display device 118 may represent any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.

[0032] Despite Figure 1 Not shown, but in some examples, the video encoder 200 and the video decoder 300 may each be integrated with the audio encoder and / or audio decoder, and may include appropriate MUX-DEMUX units or other hardware and / or software to process multiplexed streams that include both audio and video in a common data stream.

[0033] The video encoder 200 and video decoder 300 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is partially implemented in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the technology of this disclosure. Each of the video encoder 200 and video decoder 300 may be included in one or more encoders or decoders, and either one or more encoders or decoders may be integrated as part of a combined encoder / decoder (CODEC) in the respective device. Devices including the video encoder 200 and / or video decoder 300 may implement the video encoder 200 and / or video decoder 300 in processing circuitry such as integrated circuits and / or microprocessors. Such devices may be wireless communication devices (such as cellular phones) or any other type of device described herein.

[0034] Video encoder 200 and video decoder 300 may operate according to video decoding standards such as ITU-T H.265 (also known as High Efficiency Video Decoding (HEVC)) or extensions thereof such as MultiView and / or Scalable Video Decoding Extensions. Alternatively, video encoder 200 and video decoder 300 may operate according to other proprietary or industry standards such as the ITU-T H.266 standard (also known as Universal Video Decoding (VVC)). In other examples, video encoder 200 and video decoder 300 may operate according to proprietary video codecs / formats such as AOMedia Video 1 (AV1), extensions to AV1, and / or subsequent versions of AV1 (e.g., AV2). In other examples, video encoder 200 and video decoder 300 may operate according to other proprietary formats or industry standards. However, the techniques of this disclosure are not limited to any particular decoding standard or format. Typically, video encoder 200 and video decoder 300 may be configured to perform the techniques of this disclosure in conjunction with any video decoding technique using adaptive video filtering.

[0035] In general, the video encoder 200 and video decoder 300 can perform block-based decoding of images. The term "block" generally refers to a structure that includes data to be processed (e.g., to be encoded, decoded, or otherwise used in the encoding and / or decoding process). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. In general, the video encoder 200 and video decoder 300 can decode video data represented in YUV (e.g., Y, Cb, Cr) format. That is, instead of decoding the red, green, and blue (RGB) data used for images, the video encoder 200 and video decoder 300 can decode both the luminance and chrominance components, where the chrominance components may include both red hue and blue hue chrominance components. In some examples, the video encoder 200 converts the received RGB format data to a YUV representation before encoding, and the video decoder 300 converts the YUV representation to RGB format. Alternatively, preprocessing and postprocessing units (not shown) may perform these conversions.

[0036] In general, this disclosure can relate to the decoding (e.g., encoding and decoding) of images to include processes of encoding or decoding image data. Similarly, this disclosure can relate to the decoding of blocks of images to include processes of encoding or decoding data used for blocks, such as prediction and / or residual decoding. Encoded video bitstreams typically include a series of values ​​for syntax elements that represent decoding decisions (e.g., decoding modes) and the partitioning of images into blocks. Therefore, references to decoding images or blocks should generally be understood as decoding the values ​​of syntax elements used to form images or blocks.

[0037] HEVC defines various blocks, including decoding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video decoder (e.g., a video encoder 200) divides the decoding tree unit (CTU) into CUs according to a quadtree structure. That is, the video decoder divides the CTU and CU into four equal, non-overlapping squares, and each node of the quadtree has zero or four child nodes. Nodes without child nodes can be called "leaf nodes," and the CU of such leaf nodes can include one or more PUs and / or one or more TUs. The video decoder can further divide the PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents the partitioning of TUs. In HEVC, PUs represent inter-frame prediction data, while TUs represent residual data. CUs that are intra-frame predicted include intra-frame prediction information, such as intra-frame mode indication.

[0038] As another example, video encoder 200 and video decoder 300 can be configured to operate according to VVC. According to VVC, the video decoder (such as video encoder 200) partitions the image into multiple CTUs. Video encoder 200 can partition the CTUs according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure eliminates the concept of multiple partitioning types, such as the separation between CU, PU, ​​and TU in HEVC. The QTBT structure includes two levels: a first level based on quadtree partitioning and a second level based on binary tree partitioning. The root node of the QTBT structure corresponds to the CTU. The leaf nodes of the binary tree correspond to the CU.

[0039] In the MTT partitioning structure, blocks can be partitioned using quadtree (QT), binary tree (BT), and one or more types of ternary tree (TT) partitioning (also known as triplet tree (TT)). A ternary tree or triplet tree partition is a partition in which a block is divided into three sub-blocks. In some examples, a ternary tree or triplet tree partition divides a block into three sub-blocks without partitioning the original block by a center. The partitioning types in MTT (e.g., QT, BT, and TT) can be symmetric or asymmetric.

[0040] When operating according to the AV1 codec, the video encoder 200 and video decoder 300 can be configured to decode video data in blocks. In AV1, the largest decoded block that can be processed is called a superblock. In AV1, a superblock can be 128x128 luma samples or 64x64 luma samples. However, in subsequent video decoding formats (e.g., AV2), superblocks can be defined by different (e.g., larger) luma sample sizes. In some examples, the superblock is the highest level of the block quadtree. The video encoder 200 can further divide the superblock into smaller decoded blocks. The video encoder 200 can use square or non-square partitioning to divide the superblock and other decoded blocks into smaller blocks. Non-square blocks can include N / 2xN, NxN / 2, N / 4xN, and NxN / 4 blocks. The video encoder 200 and video decoder 300 can perform separate prediction and transform processes for each decoded block in the decoded block.

[0041] AV1 also defines video data tiles. A tile is a rectangular array of superblocks that can be decoded independently of other tiles. That is, the video encoder 200 and video decoder 300 can encode and decode the blocks within a tile separately, without using video data from other tiles. However, the video encoder 200 and video decoder 300 can perform filtering across tile boundaries. The tile size can be uniform or non-uniform. Tile-based decoding enables parallel processing and / or multithreading depending on the encoder and decoder implementation.

[0042] In some examples, the video encoder 200 and the video decoder 300 may use a single QTBT or MTT structure to represent each of the luma and chroma components, while in other examples, the video encoder 200 and the video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luma component and another QTBT / MTT structure for the two chroma components (or two QTBT / MTT structures for the respective chroma components).

[0043] The video encoder 200 and the video decoder 300 can be configured to use quadtree partitioning, QTBT partitioning, MTT partitioning, superblock partitioning or other partitioning structures.

[0044] In some examples, a CTU includes a decoded tree block (CTB) of luminance samples, two corresponding CTBs of chrominance samples of an image with three sample arrays, or a CTB of samples of a monochrome image or an image decoded using three separate color planes and a syntax structure for decoding the samples. A CTB can be an NxN sample block for some value of N, such that the splitting of components to CTBs is a partition. A component is an array or a single sample from one of the three arrays (luminance and two chrominance) that make up an image in a 4:2:0, 4:2:2, or 4:4:4 color format, or an array or a single sample from an array or array that makes up an image in monochrome format. In some examples, a decoded block is an MxN sample block for some values ​​of M and N, such that the splitting of CTBs to decoded blocks is a partition.

[0045] Blocks (e.g., CTUs or CUs) can be grouped in various ways within an image. As an example, a brick can refer to a rectangular area of ​​a row of CTUs within a specific tile in an image. A tile can be a rectangular area of ​​a CTU within a specific tile column and a specific tile row in an image. A tile column refers to a rectangular area of ​​a CTU that has a height equal to the height of the image and a width specified by syntax elements (e.g., as in an image parameter set). A tile row refers to a rectangular area of ​​a CTU that has a height specified by syntax elements (e.g., as in an image parameter set) and a width equal to the width of the image.

[0046] In some examples, a tile can be divided into multiple bricks, each brick potentially including one or more CTU rows within that tile. A tile that is not divided into multiple bricks can also be referred to as a brick. However, a brick that is a proper subset of a tile may not be referred to as a tile. Bricks in an image can also be arranged in slices. A slice can be an integer number of bricks in an image that can be uniquely contained within a single Network Abstraction Layer (NAL) unit. In some examples, a slice may include several complete tiles, or a continuous sequence of complete bricks representing only one tile.

[0047] This disclosure uses "NxN" and "N by N" interchangeably to refer to the sample dimensions of a block (such as a CU or other video block) in the vertical and horizontal dimensions, e.g., 16x16 samples or 16 by 16 samples. Typically, a 16x16 CU will have 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an N×N CU typically has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. Samples in a CU can be arranged in rows and columns. Furthermore, a CU does not necessarily need to have the same number of samples in the horizontal direction as it does in the vertical direction. For example, a CU can include NxM samples, where M is not necessarily equal to N.

[0048] The video encoder 200 encodes video data containing prediction information and / or residual information, as well as other information, for use in the control unit (CU). The prediction information indicates how the CU should be predicted to form a prediction block for the CU. The residual information typically represents the sample-by-sample difference between a sample of the CU before encoding and the prediction block.

[0049] To predict the Cubic Frame (CU), the video encoder 200 typically forms prediction blocks for the CU using either inter-frame prediction or intra-frame prediction. Inter-frame prediction generally refers to predicting the CU based on data from previously decoded images, while intra-frame prediction typically refers to predicting the CU based on data from previously decoded images of the same frame. To perform inter-frame prediction, the video encoder 200 can generate prediction blocks using one or more motion vectors. The video encoder 200 can typically perform motion search, for example, based on the difference between the CU and a reference block, to identify a reference block that closely matches the CU. The video encoder 200 can calculate a difference metric using the sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to determine whether the reference block closely matches the current CU. In some examples, the video encoder 200 can use unidirectional or bidirectional prediction to predict the current CU.

[0050] Some examples of VVC also provide affine motion compensation modes, which can be viewed as inter-frame prediction modes. In affine motion compensation mode, the video encoder 200 can determine two or more motion vectors representing non-translational motion (such as zooming in or out, rotation, perspective motion, or other irregular types of motion).

[0051] To perform intra-frame prediction, the video encoder 200 can select an intra-frame prediction mode to generate prediction blocks. Some examples of VVC provide 67 intra-frame prediction modes, including various orientation modes, as well as planar and DC modes. Typically, the video encoder 200 selects an intra-frame prediction mode that describes the neighboring samples of the current block (e.g., a block of a CU) according to which it predicts samples of the current block. Assuming the video encoder 200 decodes the CTU and CU in raster scan order (from left to right, from top to bottom), such samples are typically above, to the upper left, or to the left of the current block in the same image as the current block.

[0052] The video encoder 200 encodes data representing the prediction mode of the current block. For example, for inter-frame prediction modes, the video encoder 200 may encode data indicating which of the various available inter-frame prediction modes is used and the motion information used for the corresponding mode. For unidirectional or bidirectional inter-frame prediction, for example, the video encoder 200 may use Advanced Motion Vector Prediction (AMVP) or merging modes to encode motion vectors. The video encoder 200 may use similar modes to encode motion vectors used for affine motion compensation modes.

[0053] AV1 includes two common techniques for encoding and decoding blocks of video data. These two common techniques are intra-frame prediction (e.g., intra-frame prediction or spatial prediction) and inter-frame prediction (e.g., inter-frame prediction or temporal prediction). In the context of AV1, when using intra-frame prediction modes to predict blocks of video data for the current frame, the video encoder 200 and video decoder 300 do not use video data from other frames of the video data. For most intra-frame prediction modes, the video encoder 200 encodes blocks of the current frame based on the difference between sample values ​​in the current block and predicted values ​​generated from reference samples in the same frame. The video encoder 200 determines the predicted values ​​generated from the reference samples based on the intra-frame prediction mode.

[0054] Following a prediction, such as intra-frame or inter-frame prediction of a block, the video encoder 200 can compute residual data for that block. The residual data (such as a residual block) represents the sample-by-sample difference between the block and the prediction block formed using the corresponding prediction mode. The video encoder 200 can apply one or more transforms to the residual block to produce transformed data in the transform domain rather than the sample domain. For example, the video encoder 200 can apply a Discrete Cosine Transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. Additionally, the video encoder 200 can apply a secondary transform after the first transform, such as a Mode-dependent Inseparable Quadratic Transform (MDNSST), a Signal-dependent Transform, a Karhunen-Loeve Transform (KLT), etc. The video encoder 200 produces transform coefficients after applying one or more transforms.

[0055] As noted above, after any transformation to produce transform coefficients, the video encoder 200 can perform quantization of the transform coefficients. Quantization generally refers to the process of quantizing transform coefficients to reduce the amount of data used to represent them, thereby providing further compression. By performing the quantization process, the video encoder 200 can reduce the bit depth associated with some or all of the transform coefficients. For example, the video encoder 200 can round an n-bit value down to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder 200 can perform a bit-by-bit right shift of the value to be quantized.

[0056] After quantization, the video encoder 200 can scan the transform coefficients to generate a one-dimensional vector from a two-dimensional matrix containing the quantized transform coefficients. The scan can be designed to place higher-energy (and therefore lower-frequency) transform coefficients before the vector and lower-energy (and therefore higher-frequency) transform coefficients after the vector. In some examples, the video encoder 200 can utilize a predefined scan order to scan the quantized transform coefficients to produce a serialized vector, and then entropy-encode the quantized transform coefficients of the vector. In other examples, the video encoder 200 can perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 200 can entropy-encode the one-dimensional vector, for example, according to context-adaptive binary arithmetic decoding (CABAC). The video encoder 200 can also entropy-encode the values ​​of syntax elements, which describe metadata associated with the encoded video data for use by the video decoder 300 when decoding the video data.

[0057] To perform CABAC, the video encoder 200 can assign context within a context model to the symbols to be transmitted. The context may involve, for example, whether the neighboring values ​​of a symbol are zero. Probability determination can be based on the context assigned to the symbols.

[0058] The video encoder 200 may further generate, for example, grammar data (such as block-based grammar data, image-based grammar data, and sequence-based grammar data) or other grammar data (such as sequence parameter sets (SPS), image parameter sets (PPS), or video parameter sets (VPS)) destined for the video decoder 300 in image headers, block headers, and slice headers. The video decoder 300 may similarly decode such grammar data to determine how to decode the corresponding video data.

[0059] In this way, the video encoder 200 can generate a bitstream that includes encoded video data, such as syntax elements describing the partitioning of images into blocks (e.g., CUs) and prediction and / or residual information for the blocks. Finally, the video decoder 300 can receive this bitstream and decode the encoded video data.

[0060] Typically, the video decoder 300 performs the inverse of the process performed by the video encoder 200 to decode the encoded video data of the bitstream. For example, the video decoder 300 can use CABAC in a manner substantially similar to but inverse of the CABAC encoding process of the video encoder 200 to decode the values ​​of syntax elements used for the bitstream. Syntax elements can define partitioning information for dividing a picture into CTUs, and partitioning each CTU according to a corresponding partitioning structure (such as a QTBT structure) to define the CUs of the CTU. Syntax elements can further define prediction information and residual information for blocks (e.g., CUs) of the video data.

[0061] The residual information can be represented, for example, by quantized transform coefficients. The video decoder 300 can inversely quantize and inversely transform the quantized transform coefficients of the block to regenerate a residual block for that block. The video decoder 300 uses the prediction mode transmitted as a signal (intra-frame prediction or inter-frame prediction) and associated prediction information (e.g., motion information for inter-frame prediction) to form a prediction block for that block. The video decoder 300 can then combine the prediction block and the residual block (on a sample-by-sample basis) to regenerate the original block. The video decoder 300 can perform additional processing, such as performing a deblocking process to reduce visual artifacts along the block boundaries.

[0062] This disclosure can generally refer to the “signaling” of specific information (such as syntax elements). The term “signaling” can generally refer to the transmission of values ​​for syntax elements and / or other data for decoding encoded video data. That is, video encoder 200 can signal values ​​for syntax elements in the bitstream. Generally, signaling refers to generating values ​​in the bitstream. As noted above, source device 102 can transmit the bitstream to destination device 116 substantially in real time or not in real time (such as when syntax elements are stored in storage device 112 for later retrieval by destination device 116).

[0063] In a typical video encoder, frames of the original video sequence are divided into rectangular regions or blocks encoded in intra-frame mode (I-mode) or inter-frame mode. These blocks can be decoded using some form of transform decoding, such as DCT decoding. However, purely transform-based decoding only reduces inter-pixel correlation within a specific block, without considering inter-block correlation, and still produces a relatively high bit rate. Current digital image decoding standards also employ certain techniques to reduce the correlation of pixel values ​​between blocks.

[0064] Typically, blocks encoded in inter-frame mode are predicted based on several previously decoded and transmitted frames. The prediction information for the current block can be represented, for example, by two-dimensional (2D) motion vectors. For blocks encoded in I mode, a prediction block can be formed using spatial predictions from neighboring blocks already encoded within the same frame as the current block. The prediction error (e.g., the difference between the encoded current block and the prediction block) can be represented as a set of weighted basis functions of a discrete transform. The transform is typically performed on a block-by-block basis. The weights (e.g., transform coefficients) can then be quantized. Quantization introduces a loss of information, and therefore, the quantized transform coefficients can have lower precision than the original transform coefficients.

[0065] The quantized transform coefficients, along with the motion vector and some control information, can form a complete decoded sequence representation, which can be referred to as syntax elements. All syntax elements can be entropy-decoded before being transmitted from the video encoder to the video decoder to further reduce the number of bits required for their (e.g., syntax element) representation.

[0066] In a video decoder, the current block in the current frame is obtained by first constructing a prediction of the block in the same manner as in a video encoder, and then adding a compressed prediction error to that prediction. The compressed prediction error can be found by weighting the transform basis function using quantized transform coefficients. The difference between the reconstructed frame and the original frame can be called the reconstruction error.

[0067] In the field of video decoding, it is common to apply filtering to enhance the quality of the decoded video signal. Filters can be applied as post-filters (where the filtered frames are not used to predict future frames) or as in-loop filters (where the filtered frames are used to predict one or more future frames). For example, filters can be designed by minimizing the error between the original signal and the decoded filtered signal. Similar to transform coefficients, the coefficients of the filter h(k,l), k=-K,…,K,l=-K,…K can be quantized, for example, as follows:

[0068] c(k,l)=round(normFactor·h(k,l)),

[0069] It is decoded and sent to the video decoder. The normFactor is usually equal to 2. n A larger normFactor value results in more accurate quantization and better performance for the quantized filter coefficients c(k,l). On the other hand, a larger normFactor value produces coefficients c(k,l) that require more bits to transmit.

[0070] In the video decoder, the decoded filter coefficients c(k,l) are applied to the reconstructed image R(i,j) as follows:

[0071]

[0072] Where i and j are the pixel coordinates within a frame. The filter coefficients can also be applied to the difference f(k,l) between the sample R(i,j) to be filtered and its neighboring samples:

[0073] f(k,l)=R(i+k,j+l)-R(i,j).

[0074] In this case, the sample can be obtained by adding the sum to the reconstructed sample R(x,y). The difference f(k,l) can be modified, for example, by applying clipping.

[0075] An example ALF filter is set up in VVC with a block-based adaptive adaptive loop filter (ALF) (see M. Karczewicz et al., “VVC In-Loop Filters”, IEEE Trans. Circuit Systems. Video Technology, Vol. 31, No. 10, pp. 3907-3925, October 2021). Using such an ALF filter, sub-block or pixel-level filter adaptation can be applied. Based on the quantization values ​​of the block's directionality D and activity A, each M×M block can be classified into one of 25 classes:

[0076] C = 5D + A.

[0077] Each class can have its own assigned filter.

[0078] A Laplacian-based classifier can be used to derive the class C of samples within a target block. A window covering the target block can be used to classify that specific target block. Activity and directionality are derived using the values ​​of the horizontal, vertical, and two diagonal gradients calculated using 1-D Laplacian.

[0079] H k,l =|2R(k,l)-R(k-1,l)-R(k+1,l)|,

[0080] V k,l =|2R(k,l)-R(k,l-1)-R(k,l+1)|,

[0081] D1 k,l =|2R(k,l)-R(k-1,l-1)-R(k+1,l+1)|,

[0082] D2 k,l =|2R(k,l)-R(k-1,l+1)-R(k+1,l-1)|.

[0083] The sum of the horizontal, vertical, and two diagonal gradients within the window can be expressed as g. h g v g d1 and g d2 Directionality D can be determined by comparison.

[0084]

[0085] The activity A is determined by the set of thresholds. It can be calculated by g. h and g v It is derived by comparing the set of activity A with the set of thresholds.

[0086] Before filtering, certain geometric transformations (such as rotation, diagonalization, and / or vertical flipping) can be applied to pixels in the filter support region based on the orientation of the gradient of the filtered pixels (e.g., pixels multiplied by the filtered coefficients). For example, a video decoder 300 can apply such transformations. These transformations increase the similarity (e.g., their orientation) between different regions within the image. This can reduce the number of filters that must be sent to the video decoder 300, and thus reduce the number of bits required to represent the filters, or alternatively, reduce reconstruction errors. Applying transformations to the filter support region is equivalent to applying transformations directly to the filter coefficients.

[0087] To reduce the number of bits required to represent filter coefficients, different classes can be merged. Information about which classes to merge can be provided to the video decoder 300 by the video encoder 200 by sending an index i_C for each of the 25 classes. Classes with the same index i_C can share the same filters.

[0088] The ALF coefficients of a reference image can be stored, and the ALF coefficients of the reference image can be reused as the ALF coefficients of the current image. For the current image, the video encoder 200 can use the ALF coefficients stored for the reference image and bypass the ALF coefficients as signals to the video decoder. In this case, the video encoder 200 can simply send an index as a signal to one of the reference images, and the stored ALF coefficients of the indicated reference image can be simply inherited for the current image.

[0089] Y.-J. Chang, C.-C. Chen, J. Chen, J. Dong, Heegilmez, N. Hu, H. Huang, M. Karczewicz, J. Li, B. Ray, K. Reuze, V. Seregin, N. Shlyakhov, L. Pham Van, H. Wang, Y. Zhang, Z. Zhang, “A Compression Efficiency Method Beyond VVC” document JVET-U0100, 21st JVET meeting, January 2021 (hereinafter referred to as “Chang et al.”), discloses the use of three different classifiers (C0, C1, and C2) and three different filter sets (F0, F1, and F2). Sets F0 and F1 contain fixed filters with coefficients trained for classifiers C0 and C1. The coefficients of the filters in F2 are sent as signals. For example, video encoder 200 can send the coefficients of the filters in F2 as signals to video decoder 300. For a given sample, to extract the values ​​from set F... i Which filter to use depends on the classifier C. i The class C assigned to a given sample i The decision is made. All three classifiers are Laplace-based and differ from the classifiers used in VVC by using windows with different numbers of samples and several thresholds to determine activity and orientation.

[0090] In ECM-9.0, a third fixed filter, which can be called a Gaussian filter, can be applied to the samples before applying a deblocking filter. Signal filters can also be applied to the output of the Gaussian filter.

[0091] To further improve decoding efficiency, this disclosure proposes the following techniques. First, a difference-based classifier is described. Second, a multi-feature-based classifier is described. Such a classifier can be applied to both fixed filters and signal filters. Third, cascaded filtering is described. When cascaded filtering is applied, the output sample values ​​of one filter can be used as input sample values ​​for another filter. Fourth, a filter applied to sample values ​​in multiple reconstruction stages is described. When sample values ​​from one stage are unavailable and / or unused, sample values ​​from another stage can be used. Fifth, when a filter is applied to sample values ​​from only one stage, differential derivation of the input samples can be disabled. These described techniques can be applied individually or in any combination. The video encoder 200 or the video decoder 300 can employ such techniques.

[0092] The example of a difference-based classifier will now be described. For a target block in a reconstructed image, the video encoder 200 or video decoder 300 can compute the value m to derive the class index C.

[0093] For example, the value m could be the standard deviation, variance, median, or mean of a window p×q that could include the target block. In another example, the value m could be a sample value from the window.

[0094] The difference between each sample value and m can be calculated within a window that may include the target block. This window can be the same p×q window used to determine the value m or a different window. For example, the video encoder 200 or video decoder 300 can determine the difference between the first sample value and the value of m for the first sample within the window. The video encoder 200 or video decoder 300 can determine the difference between the second sample value and the value of m for the second sample within the window. The video encoder 200 or video decoder 300 can continue this difference determination for each sample within the window.

[0095] In one example, the window size may differ for different filters (such as different fixed filters). In some examples, the same sample window can be used for a difference-based classifier when applying a Laplacian classifier.

[0096] Another value v can be calculated based on the derived differences. The value v can be the sum of absolute differences, the sum of squared differences, the square root of the sum of squared differences, or any other combination based on the derived differences. For example, video encoder 200 or video decoder 300 can determine the value v.

[0097] The category index C can be derived from the value v. For example, a video encoder 200 or a video decoder 300 can determine the category index C based on the value v. In some examples, the value v can be further quantized by a scaling factor before deriving the category index. For example, video encoder 200 or video decoder 300 can determine category index C after applying a scaling factor.

[0098] In one example, the scaling factor can be derived based on the active value in the window, the window size, and / or the bit depth of the sample values. In one example, the active value A is derived as the sum of the values ​​of the horizontal and vertical gradients calculated using 1-D Laplacian. For example, video encoder 200 or video decoder 300 can determine the scaling factor and / or the active value A.

[0099] The quantified values ​​can be further... The value is cropped to within the allowed category index range. The cropped value can be used as the category index C. For example, a video encoder 200 or a video decoder 300 can crop the quantized value.

[0100] For example, (x, y) can represent the coordinates of a sample within a p×q window in a reconstructed image. The average value n of the p×q window can be calculated as...

[0101]

[0102] The average value m can be determined by the video encoder 200 or the video decoder 300.

[0103] The value v can be calculated as the square root of the sum of the squared differences between each sample and m in the p×q window.

[0104]

[0105] The video encoder 200 or the video decoder 300 can determine the value v.

[0106] Based on a Laplace-based classifier, the activity value of the current block can be A.

[0107] The category index C can be derived as

[0108]

[0109] The scaling factor array s[] = {2,2,4,4,8,8,8,8,16,16,16,16,32,32,32,32}, and the number of categories M = 8. The video encoder 200 or the video decoder 300 can determine the category index C.

[0110] This scaling factor can be further scaled based on the bit depth of the sample. In one example, it can be multiplied by 2. bitdepth-10 and category index

[0111]

[0112] To further modify s[A].

[0113] In one example

[0114]

[0115] This can be derived by accessing the sample values ​​in a single pass. Because

[0116]

[0117] For example, if r = p * q, then the variance can be calculated as follows:

[0118]

[0119] Therefore, when the value v is calculated as the square root of the sum of the squared differences between each sample and m in the p×q window, the sum of the sample values ​​in that window and the sum of the squared sample values ​​can be calculated to derive v and the further category index C. For example, a video encoder 200 or a video decoder 300 can derive v and the category index C.

[0120] When calculating the value v, in some examples of the technology disclosed herein, the video encoder 200 or video decoder 300 may incorporate an approximation of r (referred to as r′) into the calculation. The purpose of using r′ may be to allow for the software / hardware design of the technology without division. In other words, division ( / r) can be replaced by bit shifting. For example, the video encoder 200 or video decoder 300 may perform bit shifting instead of division ( / r).

[0121] As an example, consider an implementation using p = 10 and q = 10. The exact value of r = 10 * 10 = 100 is given by the equation. 2 It can be replaced with In this case, the equation becomes:

[0122]

[0123] When division can be performed using bit shifting:

[0124]

[0125] Furthermore, the arithmetic operations involved in the calculation of v can be divided into multiple stages to reduce the number of bits required for intermediate values. As an example, the approximation shown in the previous example can be derived from... Revised to:

[0126]

[0127] Now we will discuss an example of a classifier based on multiple features. In this technique, a final class index can be derived based on multiple (e.g., several) individually derived class indices. A mapping process can be introduced to derive the final class index from a combination of individual class indices. For example, a video encoder 200 or a video decoder 300 can determine the final class index based on multiple individually derived class indices.

[0128] In one example, the mapping process can be represented as follows. For example, C i (where i = 0…N-1) can represent the i-th individual classifier, and M i (where i = 0…N-1) can represent the total number of classes for the i-th individual classifier. For a given sample, C i C can represent the i-th individual classifier from i = 0 to N-1. i The derived category index. The category index C can be derived as follows:

[0129]

[0130] For example, video encoder 200 or video decoder 300 can determine category index C.

[0131] The derived category index C can be further mapped to the final category index. For example, two or more category index C values ​​can be mapped to the same final category index indicating the same filter to be used.

[0132] In another example of a classifier based on two features, C0 and C1 can represent the class indices derived from the previously described difference-based and Laplacian-based classifiers, respectively. The final class index can be derived as follows:

[0133] C = C0 * M1 + C1,

[0134] Where M1 is the total number of categories in classifier C1.

[0135] Cascaded filtering is now described. When multiple filters are applied, one filter can be applied to the output of one or more other filters. In some examples, a classifier can be applied to the output sample values ​​of another filter. For example, a video encoder 200 or a video decoder 300 can apply a filter and / or a classifier to the output of another filter.

[0136] For example, in Chang et al., both fixed filters can be applied to the input sample values ​​of the ALF. In one example, one fixed filter can be applied to the output sample values ​​of the other fixed filter. After applying both filters, the output can be used as the input for ALF filtering.

[0137] In another example, a fixed filter can be applied to several types of input. In one example, the output sample values ​​of one or more other fixed filters can be used as input, and sample values ​​prior to the ALF (e.g., samples prior to the deblocking filter, reconstructed residual samples, or predictors) can be used as another input. In some examples, geometric transpose can be applied to one type of input. In some examples, geometric transpose may not be applied to one type of input. In some examples, a geometric transpose of one type can be the same for all types of input. In some examples, the type of geometric transpose can be determined by one type of input, and the determined transpose can be applied to other (or all) types of input. For example, a video encoder 200 or a video decoder 300 can apply geometric transpose to input data.

[0138] The class index for each fixed filter can be derived from the reconstructed samples (e.g., input to ALF). In another example, the class index for a fixed filter can be derived from the output sample values ​​after another fixed filter has been applied.

[0139] Example filters applied to the residual sample values ​​during reconstruction are now described. Filters can be applied simultaneously to sample inputs obtained at different stages of the reconstruction process. In other words, the video encoder 200 or video decoder 300 can apply filters to samples at one stage of the reconstruction process and reapply the same filters to samples at another stage of the reconstruction process. For example, the video encoder 200 or video decoder 300 can apply filters to sample values ​​before and after some in-loop filters (such as deblocking filters, bilateral filters, sample adaptive offset (SAO) filters, cross-component SAO, etc.).

[0140] In another example, filtered sample values ​​can be obtained after intra-frame or inter-frame prediction. In yet another example, filtered sample values ​​can be obtained after the inverse transform; these sample values ​​are the reconstructed residuals.

[0141] If an input from a certain reconstruction stage is unavailable or unused, it can be replaced with another input for filtering purposes. For example, in a sense, such an input sample can be seen as being used twice in the filtering process. The clipping value or clipping index applied to the same sample can be averaged (first) before the filtering process. In another example, when an input is unavailable or unused, the portion of the filter corresponding to that input is not applied, or alternatively, that portion of the filter can be applied to zero input.

[0142] Now let's discuss disabling difference derivation for a single input. In some cases, filters can be applied to the difference between samples at different stages; however, when only one type of input is used (e.g., residual input only), multiple input stages are used to filter a single type of input instead, and when deriving the difference, it can produce zero difference when subtracting the center input from itself. Such computation can waste processing resources.

[0143] In this scenario, the video encoder 200 or video decoder 300 can use zero as the difference, for example, without performing any calculation. For instance, the video encoder 200 or video decoder 300 could discard determining the corresponding difference between the center sample value of the target block within the second window and the first value m, and set the corresponding difference for the center sample to be equal to 0. Alternatively, the difference can be left unapplied, and the residual input can be used without deriving the difference. In another example, the residual input can be multiplied by a factor (in one example, this factor could be equal to -1; in a previous sample (or in another example), this factor could be equal to 1).

[0144] Figure 2 This is a block diagram illustrating an example video encoder 200 that can perform the techniques described in this disclosure. Figure 2 This disclosure is provided for illustrative purposes and should not be construed as a limitation on the techniques extensively illustrated and described herein. For illustrative purposes, this disclosure describes the video encoder 200 in accordance with VVC and HEVC techniques. However, the techniques of this disclosure can be implemented by video encoding devices configured for other video decoding standards and video decoding formats, such as AV1 and subsequent versions of the AV1 video decoding format.

[0145] exist Figure 2In the example, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any or all of the units in the video data memory 230, mode selection unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, DPB 218, and entropy coding unit 220 can be implemented in one or more processors or in processing circuitry. For example, each unit of the video encoder 200 can be implemented as one or more circuit or logic elements as part of hardware circuitry, or as part of a processor, ASIC, or FPGA. Furthermore, the video encoder 200 may include additional or alternative processors or processing circuitry to perform these and other functions.

[0146] The video data storage device 230 can store video data to be encoded by the components of the video encoder 200. The video encoder 200 can obtain data from, for example, a video source 104 (…). Figure 1 The video data memory 230 receives video data stored in the video data memory 230. The DPB 218 can act as a reference picture memory, storing reference video data for use during prediction of subsequent video data by the video encoder 200. The video data memory 230 and DPB 218 can be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The video data memory 230 and DPB 218 can be provided by the same memory device or separate memory devices. In various examples, the video data memory 230 can be on-chip (as shown) with other components of the video encoder 200, or off-chip relative to those components.

[0147] In this disclosure, references to video data memory 230 should not be construed as limited to memory inside video encoder 200 (unless explicitly stated otherwise) or memory outside video encoder 200 (unless explicitly stated otherwise). Rather, references to video data memory 230 should be understood as reference memory storing video data received by video encoder 200 for encoding (e.g., video data for the current block to be encoded). Figure 1 The memory 106 can also provide temporary storage for the outputs from the various units of the video encoder 200.

[0148] Show Figure 2 The various units help to understand the operations performed by the video encoder 200. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide specific functionality and are pre-configured for the operations they can perform. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functionality in the operations they can perform. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by instructions from software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is generally immutable. In some examples, one or more units in the unit may be different circuit blocks (fixed-function or programmable), and in some examples, one or more units in the unit may be integrated circuits.

[0149] The video encoder 200 may include an arithmetic logic unit (ALU), an essential function unit (EFU), digital circuitry, analog circuitry, and / or a programmable core formed from programmable circuitry. In an example where the operation of the video encoder 200 is performed using software executed by programmable circuitry, memory 106 ( Figure 1 The video encoder 200 may store instructions (e.g., object code) of the software received and executed by the video encoder 200, or such instructions may be stored in another memory (not shown) within the video encoder 200.

[0150] The video data storage unit 230 is configured to store received video data. The video encoder 200 can retrieve images of the video data from the video data storage unit 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data storage unit 230 can be the raw video data to be encoded.

[0151] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra-frame prediction unit 226. The mode selection unit 202 may include additional functional units to perform video prediction based on other prediction modes. As an example, the mode selection unit 202 may include a palette unit, an intra-block copy unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.

[0152] Mode selection unit 202 typically coordinates multiple coding paths to test combinations of coding parameters and the resulting rate-distortion values ​​for such combinations. Coding parameters may include: CTU-CU partitioning, prediction modes for CUs, transformation types for residual data in CUs, quantization parameters for residual data in CUs, etc. Mode selection unit 202 can ultimately select a combination of coding parameters that yields a better rate-distortion value compared to other tested combinations.

[0153] The video encoder 200 can divide an image retrieved from the video data storage 230 into a series of CTUs and encapsulate one or more CTUs within a slice. The mode selection unit 202 can divide the image into CTUs according to a tree structure (such as the MTT structure, QTBT structure, superblock structure, or quadtree structure described above). As mentioned above, the video encoder 200 can form one or more CUs from the divided CTUs according to the tree structure. Such CUs can also be referred to as "video blocks" or "blocks".

[0154] Typically, mode selection unit 202 also controls its components (e.g., motion estimation unit 222, motion compensation unit 224, and intra-prediction unit 226) to generate predicted blocks for the current block (e.g., the current CU, or the overlapping portion of PU and TU in HEVC). For inter-frame prediction of the current block, motion estimation unit 222 may perform a motion search to identify one or more closely matching reference blocks in one or more reference pictures (e.g., one or more previously decoded pictures stored in DPB 218). Specifically, motion estimation unit 222 may calculate values ​​representing the similarity between the potential reference block and the current block, for example, based on sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared error (MSD), etc. Motion estimation unit 222 may typically perform these calculations using sample-by-sample differences between the current block and the reference blocks under consideration. Motion estimation unit 222 may identify the reference block with the lowest value obtained from these calculations, indicating the reference block that most closely matches the current block.

[0155] Motion estimation unit 222 can generate one or more motion vectors (MVs), each MV defining the position of a reference block in a reference image relative to the position of the current block in the current image. Motion estimation unit 222 can then provide these motion vectors to motion compensation unit 224. For example, for unidirectional inter-frame prediction, motion estimation unit 222 can provide a single motion vector, while for bidirectional inter-frame prediction, it can provide two motion vectors. Motion compensation unit 224 can then use the motion vectors to generate prediction blocks. For example, motion compensation unit 224 can use the motion vectors to retrieve data for the reference blocks. As another example, if the motion vectors have fractional-sample precision, motion compensation unit 224 can interpolate values ​​for the prediction blocks based on one or more interpolation filters. Furthermore, for bidirectional inter-frame prediction, motion compensation unit 224 can retrieve data for two reference blocks identified by corresponding motion vectors and combine the retrieved data, for example, by per-sample averaging or weighted averaging.

[0156] When operating according to the AV1 video decoding format, the motion estimation unit 222 and the motion compensation unit 224 can be configured to encode the decoded blocks of video data (e.g., both luma and chroma decoded blocks) using translational motion compensation, affine motion compensation, overlap block motion compensation (OBMC), and / or composite intra-frame prediction.

[0157] As another example, for intra-prediction or intra-prediction decoding, intra-prediction unit 226 can generate a prediction block from samples adjacent to the current block. For example, in directional mode, intra-prediction unit 226 can typically mathematically combine the values ​​of adjacent samples and fill these calculated values ​​across the current block in a defined direction to generate a prediction block. As another example, in DC mode, intra-prediction unit 226 can calculate the average of the adjacent samples of the current block and generate a prediction block to include the obtained average for each sample of the prediction block.

[0158] When operating according to the AV1 video decoding format, the intra-prediction unit 226 can be configured to encode decoded blocks of video data (e.g., both luma and chroma decoded blocks) using directional intra-prediction, non-directional intra-prediction, recursive filter intra-prediction, chroma from luma (CFL) prediction, intra-block copy (IBC), and / or palette modes. The mode selection unit 202 may include additional functional units to perform video prediction based on other prediction modes.

[0159] Mode selection unit 202 provides a prediction block to residual generation unit 204. Residual generation unit 204 receives the original uncoded version of the current block from video data memory 230 and the prediction block from mode selection unit 202. Residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines the residual block for the current block. In some examples, residual generation unit 204 may also determine the difference between sample values ​​in the residual block to generate the residual block using residual differential pulse decode-modulation (RDPCM). In some examples, one or more subtractor circuits performing binary subtraction may be used to form residual generation unit 204.

[0160] In the example where mode selection unit 202 divides a CU into PUs, each PU can be associated with a luma prediction unit and a corresponding chroma prediction unit. Video encoder 200 and video decoder 300 can support PUs of various sizes. As indicated above, the size of a CU can refer to the size of the luma decoding block of the CU, while the size of a PU can refer to the size of the luma prediction unit of the PU. Assuming a particular CU size is 2Nx2N, video encoder 200 can support PU sizes of 2Nx2N or NxN for intra-frame prediction, and symmetrical PU sizes of 2Nx2N, 2NxN, Nx2N, NxN, or similar sizes for inter-frame prediction. Video encoder 200 and video decoder 300 can also support asymmetric partitioning of PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter-frame prediction.

[0161] In the example where mode selection unit 202 does not further divide the CU into PUs, each CU can be associated with a luma decoding block and a corresponding chroma decoding block. As mentioned above, the size of the CU can refer to the size of the luma decoding block of the CU. Video encoder 200 and video decoder 300 can support CU sizes of 2Nx2N, 2NxN, or Nx2N.

[0162] For other video decoding techniques (such as intra-block copy mode decoding, affine mode decoding, and linear model (LM) mode decoding), as examples, mode selection unit 202 generates a prediction block for the current block being encoded via a corresponding unit associated with the decoding technique. In some examples, such as palette mode decoding, mode selection unit 202 may not generate a prediction block, but instead generate syntax elements indicating how the block will be reconstructed based on the selected palette. In such modes, mode selection unit 202 can provide these syntax elements to entropy coding unit 220 for encoding.

[0163] As described above, the residual generation unit 204 receives video data for the current block and the corresponding prediction block. Then, the residual generation unit 204 generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block.

[0164] Transform processing unit 206 applies one or more transformations to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 may apply various transformations to the residual block to form the transform coefficient block. For example, transform processing unit 206 may apply a Discrete Cosine Transform (DCT), a directional transformation, a Karhunen-Loeve Transform (KLT), or a conceptually similar transformation to the residual block. In some examples, transform processing unit 206 may perform multiple transformations on the residual block, such as a primary transformation and a secondary transformation (e.g., a rotation transformation). In some examples, transform processing unit 206 does not apply any transformations to the residual block.

[0165] When operating according to AV1, transform processing unit 206 may apply one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 may apply various transforms to the residual block to form the transform coefficient block. For example, transform processing unit 206 may apply a combination of horizontal / vertical transforms, which may include the Discrete Cosine Transform (DCT), the Asymmetric Discrete Sine Transform (ADST), the Reversed ADST (e.g., ADST in reverse order), and the Identity Transform (IDTX). When using the Identity Transform, the transform is skipped in either the vertical or horizontal direction. In some examples, transform processing may be skipped.

[0166] Quantization unit 208 can quantize the transform coefficients in the transform coefficient block to produce a quantized transform coefficient block. Quantization unit 208 can quantize the transform coefficients of the transform coefficient block based on the quantization parameter (QP) value associated with the current block. Video encoder 200 (e.g., via mode selection unit 202) can adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may introduce information loss, and therefore, the quantized transform coefficients may have lower accuracy compared to the original transform coefficients generated by transform processing unit 206.

[0167] The inverse quantization unit 210 and the inverse transform processing unit 212 can apply inverse quantization and inverse transform to the quantized transform coefficient block, respectively, to reconstruct the residual block from the transform coefficient block. The reconstruction unit 214 can generate a reconstructed block corresponding to the current block based on the reconstructed residual block and the prediction block generated by the mode selection unit 202 (although potentially with some degree of distortion). For example, the reconstruction unit 214 can add samples of the reconstructed residual block to corresponding samples of the prediction block generated by the mode selection unit 202 to generate the reconstructed block.

[0168] Filter unit 216 can perform one or more filter operations on the reconstructed block. For example, filter unit 216 can perform a deblocking operation to reduce blocky artifacts along the edges of the CU. In some examples, the operations of filter unit 216 can be skipped. In some examples, filter unit 216 can perform adaptive filtering techniques of this disclosure.

[0169] When operating according to AV1, filter unit 216 can perform one or more filter operations on the reconstructed block. For example, filter unit 216 can perform a deblocking operation to reduce blocky artifacts along the edges of the CU. In other examples, filter unit 216 can apply a constrained directional enhancement filter (CDEF) (which can be applied after deblocking) and can include an inseparable, nonlinear, low-pass directional filter applied based on the estimated edge direction. Filter unit 216 can also include a loop recovery filter applied after CDEF and can include a separable symmetric normalized Wiener filter or a dual-guided filter.

[0170] The video encoder 200 stores reconstructed blocks in the DPB 218. For example, in an example where the filter unit 216 is not operated, the reconstruction unit 214 may store the reconstructed blocks in the DPB 218. In an example where the filter unit 216 is operated, the filter unit 216 may store the filtered reconstructed blocks in the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 may retrieve a reference image formed from the reconstructed (and potentially filtered) blocks from the DPB 218 to perform inter-frame prediction of blocks in subsequently encoded images. Additionally, the intra-frame prediction unit 226 may use the reconstructed blocks of the current image in the DPB 218 to perform intra-frame prediction of other blocks in the current image.

[0171] Typically, entropy coding unit 220 can entropy-encode syntax elements received from other functional components of video encoder 200. For example, entropy coding unit 220 can entropy-encode quantized transform coefficient blocks from quantization unit 208. As another example, entropy coding unit 220 can entropy-encode predictive syntax elements (e.g., motion information for inter-frame prediction or intra-frame mode information for intra-frame prediction) from mode selection unit 202. Entropy coding unit 220 can perform one or more entropy coding operations on syntax elements (which is another example of video data) to generate entropy-encoded data. For example, entropy coding unit 220 can perform context-adaptive variable-length decoding (CAVLC), CABAC, variable-to-variable (V2V) length decoding, syntax-based context-adaptive binary arithmetic decoding (SBAC), probability interval partitioning entropy (PIPE) decoding, exponential Golomb coding, or another type of entropy coding operation on the data. In some examples, entropy coding unit 220 can operate in a bypass mode where syntax elements are not entropy-encoded.

[0172] The video encoder 200 can output a bitstream containing the entropy-encoded syntax elements required for reconstructing slices or blocks of images. In particular, the entropy coding unit 220 can output a bitstream.

[0173] According to AV1, entropy coding unit 220 can be configured as a symbol-to-symbol adaptive multi-symbol arithmetic decoder. The syntax elements in AV1 include an N-element alphabet, and the context (e.g., a probability model) includes a set of N probabilities. Entropy coding unit 220 can store the probabilities as an n-bit (e.g., 15-bit) cumulative distribution function (CDF). Entropy coding unit 220 can perform recursive scaling to update the context using an update factor based on the alphabet size.

[0174] The operations described above are relative to blocks. Such descriptions should be understood as operations applied to luma decoding blocks and / or chroma decoding blocks. As described above, in some examples, the luma decoding block and chroma decoding block are the luma and chroma components of the CU. In some examples, the luma decoding block and chroma decoding block are the luma and chroma components of the PU.

[0175] In some examples, it is not necessary to repeat the operations performed relative to the luma decoded block for the chroma decoded block. As an example, it is not necessary to repeat the operations used to identify the MV and reference image for the luma decoded block in order to identify the motion vector (MV) and reference image for the chroma block. Specifically, the MV for the luma decoded block can be scaled to determine the MV for the chroma block, and the reference image can be the same. As another example, the intra-frame prediction process can be the same for both the luma and chroma decoded blocks.

[0176] Video encoder 200 represents an example of a device configured to encode video data, the device including one or more memories configured to store the video data, and one or more processing units implemented in circuitry and configured to: determine a first value associated with a first window, the first window including a target block of video data; determine a corresponding difference between each sample value within a second window and the first value, the second window including the target block, wherein the first window and the second window are the same window or different windows; determine a second value based on the corresponding differences; determine a Laplacian activity value of the target block, the second window including the target block, wherein the first window and the second window are the same window or different windows; determine a category index based on the second value and the Laplacian activity value; and encode the target block based on the category index.

[0177] Video encoder 200 represents an example of a device configured to encode video data, the device including one or more memories configured to store the video data, and one or more processing units implemented in circuitry and configured to: determine the value of a first window, the first window including a target block of video data; determine a corresponding difference between each sample value of the target block within a second window and the value of the first window; determine a value based on the corresponding difference; determine a category index based on the value of the corresponding difference; and encode the target block based on the category index.

[0178] The video encoder 200 also represents an example of a device configured to encode video data, the device including one or more memories configured to store the video data, and one or more processing units implemented in a circuit and configured to perform the following operations: determining a first category index of the video data; determining a second category index of the video data; determining a final category index based on the first category index and the second category index; and encoding the video data based on the final category index.

[0179] Video encoder 200 represents an example of a device configured to encode video data, the device including one or more memories configured to store the video data, and one or more processing units implemented in circuitry and configured to: apply a first filter to the video data to generate a first output sample; apply at least one of a second filter or a classifier to the first output sample; and encode the video data based on the application of at least one of the second filter or classifier.

[0180] The video encoder 200 also represents an example of a device configured to encode video data, the device including one or more memories configured to store the video data, and one or more processing units implemented in circuitry and configured to: apply filters to the video data in a first stage of video reconstruction; apply filters to the video data in a second stage of video reconstruction; and encode the video data based on the application of the filters.

[0181] Video encoder 200 represents an example of a device configured to encode video data, the device including one or more memories configured to store the video data, and one or more processing units implemented in circuitry and configured to: determine that only one type of input is used in different reconstruction stages; set a difference to zero without performing subtraction or leaving the difference uncertain based on the one type of input used in different reconstruction stages; and encode the video data based on setting the difference to zero without performing subtraction or leaving the difference uncertain.

[0182] Figure 3 This is a block diagram illustrating an example video decoder 300 capable of performing the techniques described herein. Figure 3 This disclosure is provided for illustrative purposes and is not intended to limit the techniques as extensively illustrated and described herein. For illustrative purposes, this disclosure describes the video decoder 300 based on VVC and HEVC techniques. However, the techniques of this disclosure can be implemented by video decoding devices configured for other video decoding standards.

[0183] exist Figure 3 In the example, the video decoder 300 includes: a decoded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a DPB 314. Any or all of the CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filter unit 312, and DPB 314 can be implemented in one or more processors or in processing circuitry. For example, the units of the video decoder 300 can be implemented as one or more circuit or logic elements as part of hardware circuitry or as part of a processor, ASIC, or FPGA. Furthermore, the video decoder 300 may include additional or alternative processors or processing circuitry to perform these and other functions.

[0184] The prediction processing unit 304 includes a motion compensation unit 316 and an intra-frame prediction unit 318. The prediction processing unit 304 may include additional units for performing predictions based on other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, an intra-block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, the video decoder 300 may include more, fewer, or different functional components.

[0185] When operating according to AV1, motion compensation unit 316 can be configured to decode video data blocks (e.g., both luma and chroma blocks) using translational motion compensation, affine motion compensation, OBMC, and / or composite intra-inter-frame prediction, as described above. Intra-frame prediction unit 318 can be configured to decode video data blocks (e.g., both luma and chroma blocks) using directional intra-frame prediction, non-directional intra-frame prediction, recursive filter intra-frame prediction, CFL, IBC, and / or palette mode, as described above.

[0186] CPB memory 320 can store video data to be decoded by components of video decoder 300, such as encoded video bitstreams. For example, video data stored in CPB memory 320 can be obtained from computer-readable medium 110 (… Figure 1 The CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Additionally, the CPB memory 320 may store video data other than the syntax elements of the decoded picture, such as temporary data representing the output from various units of the video decoder 300. The DPB 314 typically stores decoded pictures that the video decoder 300 may output, and / or use as reference video data when decoding subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and DPB 314 may be formed from any of a variety of memory devices, such as DRAM (including SDRAM), MRAM, RRAM, or other types of memory devices. The CPB memory 320 and DPB 314 may be provided by the same memory device or separate memory devices. In various examples, the CPB memory 320 may be on-chip with other components of the video decoder 300, or off-chip relative to those components.

[0187] Alternatively, in some examples, the video decoder 300 can be derived from the memory 120 ( Figure 1The decoded video data is retrieved. That is, memory 120 can store data as discussed above with CPB memory 320. Similarly, when some or all of the functionality of the video decoder 300 is implemented in software to be executed by the processing circuitry of the video decoder 300, memory 120 can store instructions to be executed by the video decoder 300.

[0188] Show Figure 3 The various units shown aid in understanding the operations performed by the video decoder 300. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Similar to... Figure 2 Fixed-function circuits refer to circuits that provide specific functionality and are pre-programmed for the operations they can perform. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functionality in the operations they can perform. For example, a programmable circuit can execute software or firmware that causes it to operate in a manner defined by instructions from software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is typically immutable. In some examples, one or more units in the unit may be different circuit blocks (fixed-function or programmable), and in some examples, one or more units in the unit may be integrated circuits.

[0189] The video decoder 300 may include an ALU, EFU, digital circuitry, analog circuitry, and / or a programmable core formed from programmable circuitry. In an example where the operation of the video decoder 300 is performed by software executed on the programmable circuitry, on-chip or off-chip memory may store instructions (e.g., object code) of the software received and executed by the video decoder 300.

[0190] Entropy decoding unit 302 can receive encoded video data from CPB and perform entropy decoding on the video data to regenerate syntax elements. Prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, and filter unit 312 can generate decoded video data based on syntax elements extracted from the bitstream.

[0191] Typically, the video decoder 300 reconstructs the image on a block-by-block basis. The video decoder 300 can perform the reconstruction operation on each block individually (where the block currently being reconstructed (i.e., decoded) can be referred to as the "current block").

[0192] Entropy decoding unit 302 can perform entropy decoding on syntax elements defined as follows: quantized transform coefficients of a quantized transform coefficient block, and transform information such as quantization parameters (QP) and / or transform mode indications. Inverse quantization unit 306 can use the QP associated with the quantized transform coefficient block to determine the degree of quantization, and similarly, determine the degree of inverse quantization applied by inverse quantization unit 306. Inverse quantization unit 306 can, for example, perform a bit-left shift operation to inverse quantize the quantized transform coefficients. Inverse quantization unit 306 can thereby form a transform coefficient block including the transform coefficients.

[0193] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 can apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 can apply an inverse DCT, an inverse integer transform, an inverse Karhunen-Loeve transform (KLT), an inverse rotation transform, an inverse orientation transform, or another inverse transform to the transform coefficient block.

[0194] Furthermore, the prediction processing unit 304 generates a prediction block based on the prediction information syntax elements entropy decoded by the entropy decoding unit 302. For example, if the prediction information syntax elements indicate that the current block is inter-frame predicted, the motion compensation unit 316 can generate the prediction block. In this case, the prediction information syntax elements may indicate the reference picture from which the reference block is to be retrieved in the DPB 314, and a motion vector identifying the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion compensation unit 316 can typically be configured with respect to the motion compensation unit 224 ( Figure 2 The method described is essentially the same as the method used to perform the inter-frame prediction process.

[0195] As another example, if the prediction information syntax element indicates that the current block is intra-predictable, then intra-prediction unit 318 can generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Again, intra-prediction unit 318 can typically be configured with respect to intra-prediction unit 226 ( Figure 2 The intra-prediction process is performed in a manner substantially similar to that described above. The intra-prediction unit 318 can retrieve data from neighboring samples of the current block from the DPB 314.

[0196] Reconstruction unit 310 can use the prediction block and the residual block to reconstruct the current block. For example, reconstruction unit 310 can add samples from the residual block to the corresponding samples from the prediction block to reconstruct the current block.

[0197] Filter unit 312 can perform one or more filter operations on the reconstructed block. For example, filter unit 312 can perform a deblocking operation to reduce blocky artifacts along the edges of the reconstructed block. The operations of filter unit 312 are not necessarily performed in all examples. In some examples, filter unit 312 can perform the adaptive filtering techniques of this disclosure.

[0198] The video decoder 300 can store reconstructed blocks in the DPB 314. For example, in an example where the filter unit 312 is not operated, the reconstruction unit 310 can store the reconstructed blocks in the DPB 314. In an example where the filter unit 312 is operated, the filter unit 312 can store the filtered reconstructed blocks in the DPB 314. As discussed above, the DPB 314 can provide reference information to the prediction processing unit 304, such as samples of the current image for intra-frame prediction and previously decoded images for subsequent motion compensation. Furthermore, the video decoder 300 can output decoded images (e.g., decoded video) from the DPB 314 for display on a display device (such as...). Figure 1 The subsequent presentation on the display device 118).

[0199] In this manner, video decoder 300 represents an example of a video decoding device including a memory configured to store video data and one or more processing units implemented in circuitry and configured to: determine a first value associated with a first window, the first window comprising a target block of video data; determine a corresponding difference between each sample value within a second window and the first value, the second window comprising the target block, wherein the first window and the second window are the same window or different windows; determine a second value based on the corresponding differences; determine a Laplacian activity value for the target block, the second window comprising the target block, wherein the first window and the second window are the same window or different windows; determine a category index based on the second value and the Laplacian activity value; and decode the target block based on the category index.

[0200] In this manner, video decoder 300 represents an example of a video decoding device, which includes a memory configured to store video data and one or more processing units implemented in a circuit and configured to perform the following operations: determining the value of a first window, the first window including a target block of video data; determining a corresponding difference between each sample value of the target block within a second window and the value of the first window; determining a value based on the corresponding difference; determining a category index based on the value of the corresponding difference; and decoding the target block based on the category index.

[0201] The video decoder 300 also represents an example of a device configured to decode video data, the device including a memory configured to store the video data and one or more processing units implemented in a circuit and configured to: determine a first category index of the video data; determine a second category index of the video data; determine a final category index based on the first category index and the second category index; and decode the video data based on the final category index.

[0202] The video decoder 300 also represents an example of a device configured to decode video data, the device including a memory configured to store the video data and one or more processing units implemented in circuitry and configured to: apply a first filter to the video data to generate a first output sample; apply at least one of a second filter or a classifier to the first output sample; and decode the video data based on the application of at least one of the second filter or classifier.

[0203] The video decoder 300 also represents an example of a device configured to decode video data, the device including a memory configured to store video data and one or more processing units implemented in a circuit and configured to: apply filters to video data in a first stage of video reconstruction; apply filters to video data in a second stage of video reconstruction; and decode the video data based on the application of the filters.

[0204] The video decoder 300 also represents an example of a device configured to decode video data, the device including a memory configured to store the video data and one or more processing units implemented in a circuit and configured to: determine that only one type of input is used in different reconstruction stages; set the difference to zero without performing subtraction or uncertain difference based on the one type of input used in different reconstruction stages; and decode the video data based on setting the difference to zero without performing subtraction or uncertain difference.

[0205] Figure 4 This is a flowchart illustrating an example adaptive video filter technique according to one or more aspects of this disclosure. Video decoder 300 can determine a first value associated with a first window, the first window comprising a target block (400) of video data. For example, video decoder 300 can determine the first value m as the standard deviation, variance, median, or mean of samples in the first window p×q. In some examples, video decoder 300 can determine that the first value may be the mean of the values ​​of samples in the first window.

[0206] The video decoder 300 can determine the corresponding difference between each sample value within a second window and a first value, the second window comprising a target block, wherein the first window and the second window are the same window or different windows (402). For example, the video decoder 300 can determine the difference between the sample value of each sample in the second window and the first value m. In this way, the video decoder 300 can generate multiple differences.

[0207] The video decoder 300 can determine a second value (404) based on the corresponding differences. For example, the video decoder 300 can determine a second value v. The second value can be the sum of absolute differences, the sum of squared differences, or the square root of the sum of squared differences of the corresponding differences. In some examples, the video decoder 300 can determine the second value based on the square root of the sum of squared differences of the corresponding differences.

[0208] The video decoder 300 can determine the Laplacian activity value (406) of the target block. For example, the video decoder 300 can determine the value of activity A, as described above.

[0209] The video decoder 300 can determine the category index (408) based on the second value and the Laplacian activity value. In some examples, the video decoder 300 can determine the category index by scaling the second value using the Laplacian activity value. In some examples, the video decoder 300 can determine the category index by determining the first category index based on the second value, determining the second category index based on the Laplacian activity value, and determining the category index based on the first category index and the second category index.

[0210] The video decoder 300 can decode the target block based on the category index (410). For example, the video decoder 300 can use the category index to determine the adaptive filter to be applied to the target block, and apply the determined adaptive filter to the target block.

[0211] In some examples, the first value includes the mean of the sample values ​​within the first window. In some examples, the second value includes the square root of the sum of the squared differences of the corresponding differences. In some examples, as part of determining the second value, the video decoder 300 may apply a bit shift operation to the numerator of the square root of the sum of the squared differences of the corresponding differences to approximate a division operation.

[0212] In some examples, as part of determining the category index, the video decoder 300 may determine a scaling factor based on the Laplacian activity value. The video decoder 300 may then apply the scaling factor to a second value to generate a scaled second value. The video decoder 300 may then determine the category index based on this scaled second value.

[0213] In some examples, the category index is a third category index. In some examples, as part of determining the third category index, the video decoder 300 may determine the first category index based on a second value. The video decoder 300 may determine the second category index based on a Laplacian activity value. The video decoder 300 may determine the third category index based on both the first and second category indices.

[0214] In some examples, as part of determining the Laplacian activity value of the target block, the video decoder 300 can use a 1-D Laplacian transform to calculate the sum of the values ​​of the horizontal and vertical gradients.

[0215] In some examples, the video decoder 300 may discard the determination of the corresponding difference between the center sample value of the target block within the second window and the first value. In such examples, the video decoder 300 may set the corresponding difference for the center sample to be equal to 0.

[0216] In some examples, as part of decoding a target block based on a category index, the video decoder 300 may determine at least one of a first filter or a second filter based on the category index. The video decoder 300 may apply the first filter to samples of the target block to generate a first output sample. The video decoder 300 may apply the second filter to the first output sample to generate a second output sample. The video decoder 300 may decode the second output sample. In some examples, the video decoder 300 may determine that at least some of the multiple input samples used for one of the first or second filters are unavailable. Based on the unavailability of at least some of the multiple input samples used for one of the first or second filters, the video decoder 300 may replace the unavailable samples with reconstructed residual samples to generate modified multiple input samples. The video decoder 300 may apply one of the first or second filters to the modified multiple input samples.

[0217] In some examples, the video decoder 300 can average the cropping indices associated with multiple input samples. The video decoder 300 can crop the input samples based on the averaging of the cropping indices before applying a filter.

[0218] In some examples, the target block includes a first target block. In some examples, the video decoder 300 can determine that only one type of input (e.g., residual input) is used in different reconstruction stages of the second target block of video data. In some examples, the video decoder 300 can set the difference to zero without performing subtraction or uncertain difference based on the fact that only one type of input is used in different reconstruction stages of the second target block. In some examples, the video decoder 300 can decode the second target block based on setting the difference to zero without performing subtraction or uncertain difference.

[0219] Figure 5 This is a flowchart illustrating an example method for encoding a current block according to the technology of this disclosure. The current block may be or may include the current CU. Although this relates to a video encoder 200 ( Figure 1 and Figure 2 This is as described, but it should be understood that other devices can be configured to perform the same actions. Figure 5 Similar to the method.

[0220] In this example, the video encoder 200 first predicts the current block (350). For example, the video encoder 200 may form a predicted block for the current block. The video encoder 200 may then compute a residual block for the current block (352). To compute the residual block, the video encoder 200 may compute the difference between the original uncoded block and the predicted block for the current block. The video encoder 200 may then transform the residual block and quantize the transform coefficients of the residual block (354). Next, the video encoder 200 may scan the quantized transform coefficients of the residual block (356). During or after the scan, the video encoder 200 may entropy encode the transform coefficients (358). For example, the video encoder 200 may use CAVLC or CABAC to encode the transform coefficients. The video encoder 200 may then output the entropy-encoded data of the block (360).

[0221] Figure 6 This is a flowchart illustrating an example method for decoding a current block of video data according to the technology of this disclosure. The current block may be or may include the current CU. Although regarding the video decoder 300 ( Figure 1 and Figure 3 This is as described, but it should be understood that other devices can be configured to perform the same actions. Figure 6 Similar to the method.

[0222] The video decoder 300 can receive entropy-encoded data for the current block, such as entropy-encoded prediction information and entropy-encoded data for the transform coefficients of the residual block corresponding to the current block (370). The video decoder 300 can entropy decode the entropy-encoded data to determine the prediction information for the current block and regenerate the transform coefficients of the residual block (372). The video decoder 300 can predict the current block, for example, using an intra-frame or inter-frame prediction mode indicated by the prediction information for the current block (374), to compute a prediction block for the current block. The video decoder 300 can then inversely scan the regenerated transform coefficients (376) to create a block of quantized transform coefficients. The video decoder 300 can then inversely quantize the transform coefficients and apply the inverse transform to the transform coefficients to produce a residual block (378). The video decoder 300 can ultimately decode the current block by combining the prediction block and the residual block (380). The video decoder 300 can apply adaptive filtering techniques of this disclosure when decoding the current block.

[0223] The following numbered clauses describe one or more aspects of the devices and technologies described in this disclosure.

[0224] This disclosure includes the following non-restrictive terms.

[0225] Clause 1A: A method for decoding video data, the method comprising: determining the value of a first window, the first window including a target block of the video data; determining a corresponding difference between each sample value of the target block within a second window and the value of the first window; determining a value based on the corresponding difference; determining a category index based on the value of the corresponding difference; and decoding the target block based on the category index.

[0226] Clause 2A: The method according to Clause 1A, wherein the values ​​of the first window include the standard deviation, variance, median or mean of a plurality of sample values ​​within the first window or of a sample value within the first window.

[0227] Clause 3A: The method described in accordance with Clause 1A or Clause 2A, wherein the first window and the second window are identical.

[0228] Clause 4A: The method described in accordance with Clause 1A or Clause 2A, wherein the first window and the second window are different.

[0229] Clause 4.1A: The method according to any one of Clauses 1-4A, wherein at least one of the size of the first window or the size of the second window is based on the applied fixed filter.

[0230] Clause 5A: The method according to any one of Clauses 1A-4.1A, wherein the value based on the corresponding difference includes the sum of the differences of the differences, the sum of the squared differences of the differences, or the square root of the sum of the squared differences of the differences.

[0231] Clause 6A: The method according to any one of Clauses 1-5A, wherein determining the category index based on the value of the corresponding difference includes applying a scaling factor to the value based on the corresponding difference.

[0232] Clause 7A: The method according to Clause 6A further includes determining the scaling factor based on at least one of an activity value, the window size of the first window, or the window size of the second window.

[0233] Clause 8A: The method according to Clause 7A, wherein the activity value comprises the sum of the values ​​of the horizontal gradient and the vertical gradient calculated using 1-D Laplacian.

[0234] Clause 9A: The method according to any one of Clauses 6A-8A further includes cropping the scaling factor.

[0235] Clause 10A: A method for decoding video data, the method comprising: determining a first category index of the video data; determining a second category index of the video data; determining a final category index based on the first category index and the second category index; and decoding the video data based on the final category index.

[0236] Clause 11A: The method according to Clause 10A, wherein determining the final category index includes mapping the first category index and the second category index to the same final category index.

[0237] Clause 12A: A method for decoding video data, the method comprising: applying a first filter to the video data to generate a first output sample; applying at least one of a second filter or a classifier to the first output sample; and decoding the video data based on the application of the second filter or the at least one of the classifier.

[0238] Clause 13A: The method according to Clause 12A, wherein the first filter and the second filter are fixed filters.

[0239] Clause 14A: The method according to Clause 13A, wherein the second filter is applied to the first output sample to generate a second output sample, wherein the method further comprises applying an adaptive linear filter to the second output sample.

[0240] Clause 15A: The method according to any one of Clauses 12A-14A, wherein the input of the second filter includes the first output sample and a sample from the third filter.

[0241] Clause 16A: The method according to Clause 15A, wherein the third filter includes a deblocking filter.

[0242] Clause 16.1A: The method according to any one of Clauses 12A-16A further includes applying a geometric transpose to at least one of the inputs of the first filter, the second filter, or the third filter.

[0243] Clause 17A: The method according to Clause 13A further includes determining a category index for each fixed filter based on reconstructed samples or output samples of the first filter or the second filter.

[0244] Clause 18A: A method for decoding video data, the method comprising: applying a filter to video data in a first stage of video reconstruction; applying the filter to video data in a second stage of video reconstruction; and decoding the video data based on the application of the filter.

[0245] Clause 19A: The method according to Clause 18A, wherein the filter includes a first filter, and wherein applying the first filter to video data of the first stage of video reconstruction includes applying the first filter to sample values ​​before applying a second filter, and applying the first filter to video data of the second stage of video reconstruction includes applying the first filter to sample values ​​after applying the second filter.

[0246] Clause 20A: The method according to Clause 19A, wherein the second filter includes an in-loop filter.

[0247] Clause 21A: The method according to Clause 20A, wherein the in-loop filter comprises at least one of a deblocking filter, a bilateral filter, a SAO, or a cross component SAO.

[0248] Clause 22A: The method according to any one of Clauses 18A-21A, wherein the input to the filter includes sample values ​​after intra-frame prediction, inter-frame prediction, or inverse transform.

[0249] Clause 23A: The method according to any one of Clauses 18A-22A further includes: determining that at least some input samples of the filter are unavailable or unused; and based on the fact that the at least some input samples of the filter are unavailable or unused, not applying a portion of the filter corresponding to the at least some input samples or using zero input instead of the at least some input samples.

[0250] Clause 23.1A: The method according to any one of Clauses 18A-23A further includes averaging at least one of the clipping values ​​or clipping indices associated with the input sample before applying the filter.

[0251] Clause 24A: A method for decoding video data, the method comprising: determining that only residual inputs are used in different reconstruction stages; setting a difference to zero without performing subtraction or declaring the difference based on using only the residual inputs in the different reconstruction stages; and decoding the video data based on setting the difference to zero without performing the subtraction or declaring the difference.

[0252] Clause 25A: The method according to Clause 24A further includes multiplying the residual input by a factor.

[0253] Clause 26A: The method according to any one of Clauses 1A-25A, wherein the decoding includes decoding.

[0254] Clause 27A: The method according to any one of Clauses 1A-26A, wherein the decoding includes encoding.

[0255] Clause 28A: An apparatus for decoding video data, the apparatus comprising one or more units for performing the method according to any one of Clauses 1A-27A.

[0256] Clause 29A: The device according to Clause 28A, wherein the one or more units include one or more processors implemented in a circuit.

[0257] Clause 30A: The device described in Clause 28A or Clause 29A further includes a memory for storing the video data.

[0258] Clause 31A: The device according to any one of Clauses 28A-30A further includes: a display configured to display decoded video data.

[0259] Clause 32A: The device pursuant to any one of Clauses 28A-31A, wherein the device includes one or more of a camera, computer, mobile device, broadcast receiver device or set-top box.

[0260] Clause 33A: The device according to any one of Clauses 28A-32A, wherein the device includes a video decoder.

[0261] Clause 34A: The device according to any one of Clauses 28A-33A, wherein the device includes a video encoder.

[0262] Clause 35A: A computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to perform the method described in any one of Clauses 1A-27A.

[0263] Clause 36A: An apparatus for encoding video data, the apparatus comprising: a unit for performing a method according to any one of Clauses 1A-27A.

[0264] Clause 1B: A method for decoding video data, the method comprising: determining a first value associated with a first window, the first window including a target block of the video data; determining a corresponding difference between each sample value within a second window and the first value, the second window including the target block, wherein the first window and the second window are the same window or different windows; determining a second value based on the corresponding differences; determining a Laplacian activity value of the target block; determining a category index based on the second value and the Laplacian activity value; and decoding the target block based on the category index.

[0265] Clause 2B: The method according to Clause 1B, wherein the first value includes the mean of the sample values ​​within the first window.

[0266] Clause 3B: The method according to Clause 1B or Clause 2B, wherein the second value comprises the square root of the sum of the squared differences of the respective differences.

[0267] Clause 4B: The method according to Clause 3B, wherein determining the second value comprises applying a bit shift operation to the numerator of the square root of the sum of the squared differences of the respective differences in an approximate division operation.

[0268] Clause 5B: The method according to any one of Clauses 1B-4B, wherein determining the category index comprises: determining a scaling factor based on the Laplace activity value; applying the scaling factor to the second value to generate a scaled second value; and determining the category index based on the scaled second value.

[0269] Clause 6B: The method according to any one of Clauses 1B-5B, wherein the category index is a third category index, and wherein determining the third category index comprises: determining a first category index based on the second value; determining a second category index based on the Laplace activity value; and determining the third category index based on the first category index and the second category index.

[0270] Clause 7B: The method according to any one of Clauses 1B-6B, wherein determining the Laplace activity value of the target block comprises using a 1-D Laplace transform to calculate the sum of the values ​​of the horizontal and vertical gradients.

[0271] Clause 8B: The method according to any one of Clauses 1B-7B further includes: abandoning the determination of a corresponding difference between the center sample value of the target block within the second window and the first value; and setting the corresponding difference for the center sample to be equal to 0.

[0272] Clause 9B: The method according to any one of Clauses 1B-8B, wherein decoding the target block based on the category index comprises: determining at least one of a first filter or a second filter based on the category index; applying the first filter to a sample of the target block to generate a first output sample; applying the second filter to the first output sample to generate a second output sample; and decoding the second output sample.

[0273] Clause 10B: The method according to Clause 9B, wherein the method further comprises: determining that at least some of the plurality of input samples for one of the first filter or the second filter are unavailable; replacing the unavailable at least some of the plurality of input samples with reconstructed residual samples based on the unavailability of the at least some of the plurality of input samples for one of the first filter or the second filter to generate modified plurality of input sample values; and applying the first filter or the second filter to the modified plurality of input samples.

[0274] Clause 11B: The method according to Clause 10B further includes: averaging the cropping indices associated with the plurality of input samples; and cropping the input samples based on the averaging of the cropping indices before applying the filter.

[0275] Clause 12B: The method according to any one of Clauses 1B-9B, wherein the target block includes a first target block, the method further comprising: determining that only one type of input is used at different reconstruction stages of the second target block of the video data; setting a difference to zero without performing subtraction or declaring the difference based on the use of only said one type of input at the different reconstruction stages of the second target block; and decoding the second target block based on setting the difference to zero without performing the subtraction or declaring the difference.

[0276] Clause 13B: An apparatus for decoding video data, the apparatus comprising: one or more memories configured to store the video data; and one or more processors embedded in circuitry and coupled to the one or more memories, the one or more processors being configured to: determine a first value associated with a first window, the first window including a target block of the video data; determine a corresponding difference between each sample value within a second window and the first value, the second window including the target block, wherein the first window and the second window are the same window or different windows; determine a second value based on the corresponding differences; determine a Laplacian activity value of the target block; determine a category index based on the second value and the Laplacian activity value; and decode the target block based on the category index.

[0277] Clause 14B: The device as described in Clause 13B, wherein the first value includes the mean of the sample values ​​within the first window.

[0278] Clause 15B: The apparatus according to Clause 13B or Clause 14B, wherein the second value comprises the square root of the sum of the squared differences of the respective differences.

[0279] Clause 16B: The device according to Clause 15B, wherein, as part of determining the second value, the one or more processors are configured to apply a bit shift operation to the numerator of the square root of the sum of the squared differences of the respective differences in an approximate division operation.

[0280] Clause 17B: A device pursuant to any one of Clauses 13B-16B, wherein, as part of determining the category index, the one or more processors are configured to: determine a scaling factor based on the Laplace activity value; apply the scaling factor to the second value to generate a scaled second value; and determine the category index based on the scaled second value.

[0281] Clause 18B: A device pursuant to any one of Clauses 13B-17B, wherein the category index is a third category index, and wherein, as part of determining the third category index, the one or more processors are configured to: determine a first category index based on the second value; determine a second category index based on the Laplace activity value; and determine the third category index based on the first category index and the second category index.

[0282] Clause 19B: The device according to any one of Clauses 13B-18B further includes: a display configured to display decoded video data.

[0283] Clause 20B: A computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to: determine a first value associated with a first window, the first window comprising a target block of video data; determine a corresponding difference between each sample value within a second window and the first value, the second window comprising the target block, wherein the first window and the second window are the same window or different windows; determine a second value based on the corresponding differences; determine a Laplacian activity value of the target block, the second window comprising the target block, wherein the first window and the second window are the same window or different windows; determine a category index based on the second value and the Laplacian activity value; and decode the target block based on the category index.

[0284] It should be recognized that, depending on the examples, certain actions or events of any technique described herein may be performed in a different order, and may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for the practice of the technique). Furthermore, in certain examples, actions or events may be performed concurrently (e.g., through multithreading, interrupt handling, or multiple processors) rather than sequentially.

[0285] In one or more examples, the described functionality may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functionality may be stored or transmitted as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. A computer-readable medium may include: a computer storage medium, which corresponds to a tangible medium such as a data storage medium; or a communication medium, which includes, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol. In this way, a computer-readable medium may generally correspond to: (1) a non-transitory tangible computer-readable storage medium, or (2) a communication medium such as a signal or carrier wave. A data storage medium may be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for use in implementing the techniques described in this disclosure. Computer program products may include computer-readable media.

[0286] For example, rather than limiting, such computer-readable storage media may include one or more of the following: RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, flash memory, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of medium if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer to non-transient tangible storage media. As used herein, disks and optical discs include compact optical discs (CDs), laser optical discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, and optical discs typically reproduce data optically using lasers. The above combinations should also be included within the scope of computer-readable media.

[0287] Instructions can be executed by one or more processors (such as one or more DSPs, general-purpose microprocessors, ASICs, FPGAs, or other equivalent integrated or discrete logic circuits). Therefore, the terms "processor" and "processing circuit" as used herein can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Similarly, the techniques can be fully implemented in one or more circuit or logic elements.

[0288] The technologies disclosed herein can be implemented in various devices or apparatuses, including wireless handheld devices, integrated circuits (ICs), or collections of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed technologies, but they do not necessarily need to be implemented by different hardware units. Rather, as described above, various units can be combined in a codec hardware unit, or provided by a collection of interoperable hardware units (including one or more processors as described above) combined with suitable software and / or firmware.

[0289] Examples have been described. These and other examples are within the scope of the following claims.

Claims

1. A method for decoding video data, the method comprising: Determine a first value associated with a first window, the first window comprising a target block of the video data; Determine the corresponding difference between each sample value within a second window and the first value, the second window including the target block, wherein the first window and the second window are the same window or different windows; The second value is determined based on the corresponding difference; Determine the Laplace activity value of the target block; The category index is determined based on the second value and the Laplace activity value; and The target block is decoded based on the category index.

2. The method according to claim 1, wherein, The first value includes the mean of the sample values ​​within the first window.

3. The method according to claim 1, wherein, The second value includes the square root of the sum of the squared differences of the corresponding differences.

4. The method according to claim 3, wherein, Determining the second value involves applying a bit shift operation to the numerator of the square root of the sum of the squared differences of the corresponding differences in an approximate division operation.

5. The method according to claim 1, wherein, Determining the category index includes: The scaling factor is determined based on the Laplace activity value; Apply the scaling factor to the second value to generate a scaled second value; and The category index is determined based on the scaled second value.

6. The method according to claim 1, wherein, The category index is a third category index, and determining the third category index includes: The first category index is determined based on the second value; The second category index is determined based on the Laplace activity value; and The third category index is determined based on the first category index and the second category index.

7. The method according to claim 1, wherein, Determining the Laplace activity value of the target block involves using a 1-D Laplace transform to calculate the sum of the values ​​of the horizontal and vertical gradients.

8. The method according to claim 1, further comprising: Discard the determination of the corresponding difference between the center sample value of the target block within the second window and the first value; as well as Set the corresponding difference for the central sample to equal 0.

9. The method according to claim 1, wherein, Decoding the target block based on the category index includes: At least one of the first filter or the second filter is determined based on the category index; The first filter is applied to the samples of the target block to generate a first output sample; Apply the second filter to the first output sample to generate the second output sample; and The second output sample is decoded.

10. The method according to claim 9, wherein, The method further includes: Determine that at least some of the input samples among a plurality of input samples used for one of the first filter or the second filter are unavailable; Based on the fact that at least some of the input samples among the plurality of input samples used for one of the first filter or the second filter are unavailable, the unavailable at least some of the input samples among the plurality of input samples are replaced with reconstructed residual samples to generate modified plurality of input samples; and The first filter or the second filter is applied to the modified plurality of input samples.

11. The method of claim 10, further comprising: Average the cropping indices associated with the multiple input samples; as well as Before applying the filter, the input samples are cropped based on the average of the cropping index.

12. The method according to claim 1, wherein, The target block includes a first target block, and the method further includes: It is determined that only one type of input is used in different reconstruction stages for the second target block of the video data; Based on using only one type of input in the different reconstruction stages for the second target block, the difference is set to zero without performing subtraction or declaring the difference; and The second target block is decoded by setting the difference to zero without performing the subtraction or by leaving the difference uncertain.

13. An apparatus for decoding video data, the apparatus comprising: One or more memories configured to store the video data; as well as One or more processors are embedded in the circuitry and coupled to the one or more memories, the one or more processors being configured to: Determine a first value associated with a first window, the first window comprising a target block of the video data; Determine the corresponding difference between each sample value within a second window and the first value, the second window including the target block, wherein the first window and the second window are the same window or different windows; The second value is determined based on the corresponding difference; Determine the Laplace activity value of the target block; The category index is determined based on the second value and the Laplace activity value; and The target block is decoded based on the category index.

14. The device according to claim 13, wherein, The first value includes the mean of the sample values ​​within the first window.

15. The device according to claim 13, wherein, The second value includes the square root of the sum of the squared differences of the corresponding differences.

16. The device according to claim 15, wherein, As part of determining the second value, the one or more processors are configured to apply a bit shift operation to the numerator of the square root of the sum of the squared differences of the corresponding differences in an approximate division operation.

17. The device according to claim 13, wherein, As part of determining the category index, the one or more processors are configured to: The scaling factor is determined based on the Laplace activity value; The scaling factor is applied to the second value to generate a scaled second value; as well as The category index is determined based on the scaled second value.

18. The device according to claim 13, wherein, The category index is a third category index, and wherein, as part of determining the third category index, the one or more processors are configured to: The first category index is determined based on the second value; The second category index is determined based on the Laplace activity value; and The third category index is determined based on the first category index and the second category index.

19. The apparatus of claim 13, further comprising: A display configured to show decoded video data.

20. A computer-readable storage medium having instructions stored thereon, said instructions, when executed, causing one or more processors to perform the following operations: Determine a first value associated with a first window, the first window comprising a target block of video data; Determine the corresponding difference between each sample value within a second window and the first value, wherein the second window includes the target block, wherein... The first window and the second window may be the same window or different windows; The second value is determined based on the corresponding difference; Determine the Laplace activity value of the target block, wherein the second window includes the target block, and the first window and the second window are the same window or different windows; The category index is determined based on the second value and the Laplace activity value; and The target block is decoded based on the category index.