Joint clipping operation of filters for video coding
Patent Information
- Application Number
- CN202280040229.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2022-06-09
- Filing Date
- 2022-06-10
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2042-06-10
Smart Images

Figure CN117426097B_ABST
Abstract
Description
[0001] This application claims the benefit of U.S. Patent Application No. 17 / 806,192, filed June 9, 2022, and U.S. Provisional Application No. 63 / 210,438, filed June 14, 2021, the entire contents of each of which are incorporated herein by reference. U.S. Patent Application No. 17 / 806,192, filed June 9, 2022, claims the benefit of U.S. Provisional Application No. 63 / 210,438, filed June 14, 2021. Technical Field
[0002] This disclosure relates to video encoding and video decoding. Background Technology
[0003] Digital video capabilities can be incorporated into a wide variety of devices, including digital televisions, digital live broadcast systems, wireless broadcasting systems, personal digital assistants (PDAs), laptops or desktop computers, tablets, e-book readers, digital cameras, digital recording devices, digital media players, video game devices, video game consoles, cellular or satellite radio phones (so-called "smartphones"), video conferencing equipment, video streaming devices, and more. Digital video devices implement video decoding technologies (such as those defined by MPEG-2, MPEG-4, ITU-T H.263, ITU-T H.264 / MPEG-4 (Part 10, Advanced Video Decoding (AVC)), ITU-T H.265 / High Efficiency Video Decoding (HEVC), ITU-T H.266 / Various Universal Video Decoding (VVC) and extensions to such standards, as well as proprietary video codecs / formats (such as those described in AOMedia Video 1 (AV1) developed by the Open Media Alliance). By implementing such video decoding technology, video devices can more efficiently send, receive, encode, decode, and / or store digital video information.
[0004] Video decoding techniques include spatial (intra-picture) prediction and / or temporal (inter-picture) prediction to reduce or remove redundancy inherent in video sequences. For block-based video decoding, video slices (e.g., video pictures or portions of video pictures) can be segmented into video blocks, which may also be referred to as decoding tree units (CTUs), decoding units (CUs), and / or decoding nodes. Video blocks in an intra-coded (I) slice of a picture are encoded using spatial prediction relative to reference samples in adjacent blocks within the same picture. Video blocks in an inter-coded (P or B) slice of a picture can use spatial prediction relative to reference samples in adjacent blocks within the same picture or temporal prediction relative to reference samples in other reference pictures. A picture may be referred to as a frame, and a reference picture may be referred to as a reference frame. Summary of the Invention
[0005] In summary, this disclosure describes techniques for loop filtering in video decoding. Specifically, it describes techniques for joint truncation operations involving two or more loop filtering operations. Example loop filtering operations may include cross-component sample adaptive offset (CCSAO) filters, sample adaptive offset (SAO) filters, and / or bilateral filters (BIF). The techniques of this disclosure can reduce the implementation complexity of loop filter designs, including reducing the amount of memory required to implement loop filters. The techniques of this disclosure can further increase the throughput of loop filters in video decoding, thereby reducing latency during the decoding process.
[0006] In one example, this disclosure describes a method for decoding video data, the method comprising: reconstructing the video data to generate reconstructed video data; performing a plurality of loop filter operations in parallel on the reconstructed video data, wherein the plurality of loop filter operations includes a first filter operation that is not a bilateral filter operation or a sample adaptive offset (SAO) filter operation; and performing a joint truncation operation on a first output of the first filter operation and a second output of a second loop filter operation among the plurality of loop filter operations. In one example, the first filter operation is a CCSAO filter operation.
[0007] In another example, this disclosure describes an apparatus configured to decode video data, the apparatus comprising: a memory configured to store the video data; and one or more processors in communication with the memory, the one or more processors being configured to: reconstruct the video data to generate reconstructed video data; perform a plurality of loop filter operations on the reconstructed video data in parallel, wherein the plurality of loop filter operations includes a first filter operation that is not a bilateral filter operation or a sample adaptive offset (SAO) filter operation; and perform a joint truncation operation on a first output of the first filter operation and a second output of a second loop filter operation among the plurality of loop filter operations. In one example, the first filter operation is a CCSAO filter operation.
[0008] In another example, this disclosure describes an apparatus configured to decode video data, the apparatus comprising: a unit for reconstructing the video data to generate reconstructed video data; a unit for performing a plurality of loop filter operations in parallel on the reconstructed video data, wherein the plurality of loop filter operations includes a first filter operation that is not a bilateral filter operation or a sample adaptive offset (SAO) filter operation; and a unit for performing a joint truncation operation on a first output of the first filter operation and a second output of a second loop filter operation among the plurality of loop filter operations. In one example, the first filter operation is a CCSAO filter operation.
[0009] In another example, this disclosure describes a non-transitory computer-readable storage medium storing instructions that, when executed, cause one or more processors of a device configured to decode video data to: reconstruct the video data to generate reconstructed video data; perform a plurality of loop filter operations in parallel on the reconstructed video data, wherein the plurality of loop filter operations includes a first filter operation that is not a bilateral filter operation or a sample adaptive offset (SAO) filter operation; and perform a joint truncation operation on a first output of the first filter operation and a second output of a second loop filter operation among the plurality of loop filter operations. In one example, the first filter operation is a CCSAO filter operation.
[0010] Details of one or more examples are set forth in the accompanying drawings and the following description. Other features, objects, and advantages will be apparent from the description, drawings, and claims. Attached Figure Description
[0011] Figure 1 This is a block diagram illustrating an example video encoding and decoding system that can perform the techniques described in this disclosure.
[0012] Figure 2 This is a block diagram illustrating an example of a CCSAO process that can be used in conjunction with the techniques of this disclosure.
[0013] Figure 3 This is a conceptual diagram illustrating an example of candidate locations for a CCSAO classifier that can be used in conjunction with the techniques of this disclosure.
[0014] Figure 4 This is a flowchart illustrating an example of the processing of multiple loop filters in one example of the present disclosure.
[0015] Figure 5 This is a flowchart illustrating another example of a combined capture of SAO, BIF, and CCSAO, which demonstrates the contents of this disclosure.
[0016] Figure 6 This is a flowchart illustrating an example of offset truncation for SAO, BIF, and CCSAO, as another example of the content of this disclosure.
[0017] Figure 7 This is a flowchart illustrating an example of a cutoff applied to the output offset of a BIF, as another example of the content of this disclosure.
[0018] Figure 8 This is a flowchart illustrating an example of a joint cutoff of the output offsets of BIF and SAO, as shown in another example of the content of this disclosure.
[0019] Figure 9 This is a block diagram illustrating an example video encoder that can perform the techniques described in this disclosure.
[0020] Figure 10 This is a block diagram illustrating an example video decoder that can perform the techniques described in this disclosure.
[0021] Figure 11 This is a flowchart illustrating an example method for encoding the current block according to the technology of this disclosure.
[0022] Figure 12 This is a flowchart illustrating an example method for decoding the current block according to the technology of this disclosure.
[0023] Figure 13 This is a flowchart illustrating an example method for filtering the current block according to the techniques described in this disclosure. Detailed Implementation
[0024] In summary, this disclosure describes techniques for loop filtering in video decoding. Specifically, it describes techniques for joint truncation operations involving two or more loop filtering operations. Example loop filtering operations may include cross-component sample adaptive offset (CCSAO) filters, sample adaptive offset (SAO) filters, and / or bilateral filters (BIF). The techniques of this disclosure can reduce the implementation complexity of loop filter designs, including reducing the amount of memory required to implement loop filters. The techniques of this disclosure can further increase the throughput of loop filters in video decoding, thereby reducing latency during the decoding process.
[0025] Figure 1 This is a block diagram illustrating an example video encoding and decoding system 100 capable of performing the techniques of this disclosure. In summary, the techniques of this disclosure relate to decoding (encoding and / or decoding) video data. Typically, video data includes any data used for processing video. Therefore, video data can include raw, unencoded video, encoded video, decoded (e.g., reconstructed) video, and video metadata (e.g., signaling data).
[0026] like Figure 1 As shown in the example, system 100 includes source device 102, which provides encoded video data to be decoded and displayed by destination device 116. Specifically, source device 102 provides the video data to destination device 116 via computer-readable medium 110. Source device 102 and destination device 116 can include any of a wide variety of devices, including desktop computers, laptop computers, mobile devices, tablet computers, set-top boxes, mobile phones such as smartphones, televisions, cameras, display devices, digital media players, video game consoles, video streaming devices, broadcast receivers, etc. In some cases, source device 102 and destination device 116 may be equipped for wireless communication and may therefore be referred to as wireless communication devices.
[0027] exist Figure 1In the example, source device 102 includes a video source 104, memory 106, video encoder 200, and output interface 108. Destination device 116 includes an input interface 122, video decoder 300, memory 120, and display device 118. According to this disclosure, the video encoder 200 of source device 102 and the video decoder 300 of destination device 116 can be configured to apply techniques for performing a cut-off operation during filtering (e.g., loop filtering). Therefore, source device 102 represents an example of a video encoding device, while destination device 116 represents an example of a video decoding device. In other examples, the source and destination devices may include other components or arrangements. For example, source device 102 may receive video data from an external video source such as an external camera. Similarly, destination device 116 may interface with an external display device, rather than including an integrated display device.
[0028] like Figure 1 The system 100 shown is merely an example. Typically, any digital video encoding and / or decoding device can perform techniques for intercepting operations during filtering (e.g., loop filtering). Source device 102 and destination device 116 are merely examples of such decoding devices, where source device 102 generates decoded video data for transmission to destination device 116. In this disclosure, "decoding device" refers to a device that performs the decoding (encoding and / or decoding) of data. Therefore, video encoder 200 and video decoder 300 represent examples of decoding devices (specifically, video encoder and video decoder). In some examples, source device 102 and destination device 116 can operate in a substantially symmetrical manner, such that each of source device 102 and destination device 116 includes video encoding and decoding components. Thus, system 100 can support one-way or two-way video transmission between source device 102 and destination device 116, for example, for video streaming, video playback, video broadcasting, or video telephony.
[0029] Typically, video source 104 represents a source of video data (i.e., raw, unencoded video data) and provides a sequential series of pictures (also referred to as "frames") of the video data to video encoder 200, which encodes the data of the pictures. Video source 104 of source device 102 may include video capture devices such as cameras, video archives containing previously captured raw video, and / or video feed interfaces for receiving video from video content providers. Alternatively, video source 104 may generate computer graphics-based data as source video, or a combination of live video, archived video, and computer-generated video. In each case, video encoder 200 may encode the captured, pre-captured, or computer-generated video data. Video encoder 200 may rearrange the pictures from their received order (sometimes referred to as "display order") to a decoding order for decoding. Video encoder 200 may generate a bitstream comprising the encoded video data. Then, the source device 102 can output the encoded video data to the computer-readable medium 110 via the output interface 108 so that it can be received and / or retrieved by, for example, the input interface 122 of the destination device 116.
[0030] The memory 106 of source device 102 and the memory 120 of destination device 116 represent general-purpose memory. In some examples, memories 106 and 120 may store raw video data, such as raw video from video source 104 and raw decoded video data from video decoder 300. Alternatively, memories 106 and 120 may store software instructions executable by, for example, video encoder 200 and video decoder 300. Although memories 106 and 120 are shown separately from video encoder 200 and video decoder 300 in this example, it should be understood that video encoder 200 and video decoder 300 may also include internal memory for functionally similar or equivalent purposes. Furthermore, memories 106 and 120 may store, for example, encoded video data output from video encoder 200 and input to video decoder 300. In some examples, portions of memories 106 and 120 may be allocated as one or more video buffers, for example, to store raw decoded and / or encoded video data.
[0031] Computer-readable medium 110 can represent any type of medium or device capable of transmitting encoded video data from source device 102 to destination device 116. In one example, computer-readable medium 110 represents a communication medium that enables source device 102 to directly transmit encoded video data to destination device 116 in real time, for example, via a radio frequency network or a computer-based network. According to a communication standard such as a wireless communication protocol, output interface 108 can modulate the transmitted signal including the encoded video data, and input interface 122 can demodulate the received transmitted signal. The communication medium can include any wireless or wired communication medium, such as radio frequency (RF) spectrum or one or more physical transmission lines. The communication medium can form part of a packet-based network such as a local area network, a wide area network, or a global network such as the Internet. The communication medium can include a router, switch, base station, or any other device that may be useful for facilitating communication from source device 102 to destination device 116.
[0032] In some examples, source device 102 can output encoded data from output interface 108 to storage device 112. Similarly, destination device 116 can access encoded data from storage device 112 via input interface 122. Storage device 112 may include any of a variety of distributed or locally accessed data storage media, such as hard disk drives, Blu-ray discs, DVDs, CD-ROMs, flash memory, volatile or non-volatile memory, or any other suitable digital storage medium for storing encoded video data.
[0033] In some examples, source device 102 can output encoded video data to file server 114 or another intermediate storage device that can store the encoded video data generated by source device 102. Destination device 116 can access the stored video data from file server 114 via streaming or downloading.
[0034] File server 114 can be any type of server device capable of storing encoded video data and sending such encoded video data to destination device 116. File server 114 can represent a web server (e.g., for a website), a server configured to provide file transfer protocol services (such as File Transfer Protocol (FTP) or FLUTE-based file delivery protocol), a Content Delivery Network (CDN) device, a Hypertext Transfer Protocol (HTTP) server, a Multimedia Broadcast Multicast Service (MBMS) or enhanced MBMS (eMBMS) server, and / or a Network Attached Storage (NAS) device. File server 114 can additionally or alternatively implement one or more HTTP streaming protocols, such as HTTP-based Dynamic Adaptive Streaming (DASH), HTTP Real-Time Streaming (HLS), Real-Time Streaming Protocol (RTSP), HTTP Dynamic Streaming, etc.
[0035] Destination device 116 can access encoded video data from file server 114 via any standard data connection, including an internet connection. This may include a wireless channel (e.g., a Wi-Fi connection), a wired connection (e.g., a digital subscriber line (DSL), a cable modem, etc.), or a combination of both, suitable for accessing encoded video data stored on file server 114. Input interface 122 can be configured to operate according to any one or more of the various protocols discussed above for retrieving or receiving media data from file server 114, or other such protocols for retrieving media data.
[0036] Output interface 108 and input interface 122 can represent a wireless transmitter / receiver, a modem, a wired networking component (e.g., an Ethernet card), a wireless communication component operating according to any of the various IEEE 802.11 standards, or other physical components. In examples where output interface 108 and input interface 122 include wireless components, output interface 108 and input interface 122 can be configured to transmit data (such as encoded video data) according to cellular communication standards (such as 4G, 4G-LTE (Long Term Evolution), improved LTE, 5G, etc.). In some examples where output interface 108 includes a wireless transmitter, output interface 108 and input interface 122 can be configured to operate according to other wireless standards (such as the IEEE 802.11 specification, the IEEE 802.15 specification (e.g., ZigBee)). TM Bluetooth TMThe source device 102 and / or destination device 116 may include corresponding system-on-chip (SoC) devices. For example, source device 102 may include an SoC device for performing the functions assigned to video encoder 200 and / or output interface 108, and destination device 116 may include an SoC device for performing the functions assigned to video decoder 300 and / or input interface 122.
[0037] The technology disclosed herein can be applied to video decoding to support any of a variety of multimedia applications, such as over-the-air television broadcasting, cable television transmission, satellite television transmission, internet streaming video transmission (such as HTTP-based Dynamic Adaptive Streaming (DASH)), digital video encoded onto data storage media, decoding digital video stored on data storage media, or other applications.
[0038] The input interface 122 of the destination device 116 receives an encoded video bitstream from a computer-readable medium 110 (e.g., a communication medium, storage device 112, file server 114, etc.). The encoded video bitstream may include signaling information defined by the video encoder 200, such as syntax elements (also used by the video decoder 300), which have values describing the characteristics and / or processing of video blocks or other decoding units (e.g., slices, pictures, picture groups, sequences, etc.). The display device 118 displays a decoded image of the decoded video data to the user. The display device 118 may represent any of a variety of display devices, such as a liquid crystal display (LCD), a plasma display, an organic light-emitting diode (OLED) display, or another type of display device.
[0039] Despite Figure 1 Not shown, but in some examples, the video encoder 200 and the video decoder 300 may each be integrated with the audio encoder and / or audio decoder, and may include appropriate MUX-DEMUX units or other hardware and / or software to process multiplexed streams that include both audio and video in a common data stream.
[0040] The video encoder 200 and video decoder 300 can each be implemented as any of a variety of suitable encoder and / or decoder circuits, such as one or more microprocessors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), discrete logic, software, hardware, firmware, or any combination thereof. When the technology is implemented in part in software, the device may store instructions for the software in a suitable non-transitory computer-readable medium and use one or more processors to execute the instructions in hardware to perform the technology of this disclosure. Each of the video encoder 200 and video decoder 300 may be included in one or more encoders or decoders, and either encoder or decoder may be integrated as part of a combined encoder / decoder (CODEC) in the respective device. Devices including the video encoder 200 and / or video decoder 300 may include integrated circuits, microprocessors, and / or wireless communication devices (such as cellular phones).
[0041] The video encoder 200 and video decoder 300 may operate according to video decoding standards such as ITU-T H.265 (also known as the High Efficiency Video Decoding (HEVC) standard) or extensions thereof such as MultiView and / or Scalable Video Decoding Extensions. Alternatively, the video encoder 200 and video decoder 300 may operate according to other proprietary or industry standards such as ITU-T H.266 (also known as Universal Video Decoding (VVC)). In other examples, the video encoder 200 and video decoder 300 may operate according to proprietary video codecs / formats such as AOMedia Video 1 (AV1), extensions to AV1, and / or subsequent versions of AV1 (e.g., AV2)). In other examples, the video encoder 200 and video decoder 300 may operate according to other proprietary formats or industry standards. However, the technology of this disclosure is not limited to any particular decoding standard or format. Typically, the video encoder 200 and the video decoder 300 can be configured to perform the techniques of this disclosure by combining any video decoding technique that uses two or more filters (including two or more loop filters).
[0042] Typically, video encoder 200 and video decoder 300 can perform block-based decoding of images. The term "block" generally refers to a structure that includes data to be processed (e.g., encoded, decoded, or otherwise used during encoding and / or decoding). For example, a block may include a two-dimensional matrix of samples of luminance and / or chrominance data. Typically, video encoder 200 and video decoder 300 can decode video data represented in YUV (e.g., Y, Cb, Cr) format. That is, video encoder 200 and video decoder 300 can decode luminance and chrominance components, rather than decoding red, green, and blue (RGB) data samples used for images, where chrominance components may include both red hue and blue hue chrominance components. In some examples, video encoder 200 converts the received RGB-formatted data to a YUV representation before encoding, and video decoder 300 converts the YUV representation to RGB format. Alternatively, preprocessing and post-processing units (not shown) can perform these conversions.
[0043] In summary, this disclosure may relate to the decoding (e.g., encoding and decoding) of images to include the process of encoding or decoding the data of an image. Similarly, this disclosure may relate to the decoding of blocks of an image to include the process of encoding or decoding the data used for the blocks (e.g., prediction and / or residual decoding). Encoded video bitstreams typically include a series of values representing decoding decisions (e.g., decoding modes) and syntax elements that segment the image into blocks. Therefore, references to decoding images or blocks should generally be understood as decoding the values of the syntax elements that form the image or block.
[0044] HEVC defines various blocks, including decoding units (CUs), prediction units (PUs), and transform units (TUs). According to HEVC, a video decoder (such as a video encoder 200) partitions a decoding tree unit (CTU) into CUs based on a quadtree structure. That is, the video decoder partitions the CTU and CU into four equal, non-overlapping squares, and each node of the quadtree has zero or four child nodes. Nodes without child nodes can be called "leaf nodes," and the CU of such leaf nodes can include one or more PUs and / or one or more TUs. The video decoder can further partition PUs and TUs. For example, in HEVC, a residual quadtree (RQT) represents a partition of the TU. In HEVC, PUs represent inter-frame prediction data, while TUs represent residual data. CUs with intra-frame prediction include intra-frame prediction information, such as intra-frame mode indication.
[0045] As another example, video encoder 200 and video decoder 300 can be configured to operate according to VVC. According to VVC, the video decoder (such as video encoder 200) segments the image into multiple decoding tree units (CTUs). Video encoder 200 can segment CTUs according to a tree structure (such as a quadtree-binary tree (QTBT) structure or a multi-type tree (MTT) structure). The QTBT structure eliminates the concept of multiple segmentation types, such as segmentation between CUs, PUs, and TUs in HEVC. The QTBT structure includes two levels: a first level segmented according to quadtree segmentation and a second level segmented according to binary tree segmentation. The root node of the QTBT structure corresponds to a CTU. The leaf nodes of the binary tree correspond to decoding units (CUs).
[0046] In the MTT partitioning structure, blocks can be partitioned using quadtree (QT) partitioning, binary tree (BT) partitioning, and one or more types of ternary tree (TT) partitioning (also known as triplet tree (TT)) partitioning. A ternary tree or triplet tree partitioning is a partition in which a block is divided into three sub-blocks. In some examples, a ternary tree or triplet tree partitioning divides a block into three sub-blocks without partitioning the original block through a center. The partitioning types in MTT (e.g., QT, BT, and TT) can be symmetric or asymmetric.
[0047] When operating according to the AV1 codec, the video encoder 200 and video decoder 300 can be configured to decode video data in blocks. In AV1, the largest decoded block that can be processed is called a superblock. In AV1, a superblock can be 128x128 luma samples or 64x64 luma samples. However, in subsequent video decoding formats (e.g., AV2), superblocks can be defined by different (e.g., larger) luma sample sizes. In some examples, the superblock is the top level of a block quadtree. The video encoder 200 can further divide the superblock into smaller decoded blocks. The video encoder 200 can use square or non-square partitioning to divide the superblock and other decoded blocks into smaller blocks. Non-square blocks can include N / 2xN, NxN / 2, N / 4xN, and NxN / 4 blocks. The video encoder 200 and video decoder 300 can perform separate prediction and transform processes for each decoded block.
[0048] AV1 also defines video data tiles. A tile is a rectangular array of superblocks that can be decoded independently of other tiles. That is, the video encoder 200 and video decoder 300 can encode and decode the blocks within a tile separately, without using video data from other tiles. However, the video encoder 200 and video decoder 300 can perform filtering across tile boundaries. The tile size can be uniform or non-uniform. Tile-based decoding enables parallel processing and / or multithreading for both the encoder and decoder.
[0049] In some examples, the video encoder 200 and the video decoder 300 may use a single QTBT or MTT structure to represent each of the luma and chroma components, while in other examples, the video encoder 200 and the video decoder 300 may use two or more QTBT or MTT structures, such as one QTBT / MTT structure for the luma component and another QTBT / MTT structure for the two chroma components (or two QTBT / MTT structures for the respective chroma components).
[0050] The video encoder 200 and the video decoder 300 can be configured to use quadtree segmentation, QTBT segmentation, MTT segmentation, superblock segmentation or other segmentation structures.
[0051] In some examples, a CTU includes a decoded tree block (CTB) of luminance samples, two corresponding CTBs of chrominance samples of an image with three sample arrays, or a CTB of samples of a monochrome image or an image decoded using three separate color planes and a syntax structure for decoding the samples. A CTB can be an N×N block of samples (for some value of N) such that dividing a component into a CTB is a partition. A component is an array or a single sample of one of the three arrays (one luminance and two chrominance) that make up an image in a 4:2:0, 4:2:2, or 4:4:4 color format, or an array or a single sample of an array or array that makes up an image in monochrome format. In some examples, a decoded block is an M×N block of samples (for some values of M and N) such that dividing a CTB into a decoded block is a partition.
[0052] Blocks (e.g., CTUs or CUs) can be grouped in various ways within an image. As an example, a brick can refer to a rectangular area of a CTU row within a specific tile in an image. A tile can be a rectangular area of a CTU within a specific tile column and a specific tile row in an image. A tile column refers to a rectangular area of a CTU with a height equal to the height of the image and a width specified by a syntax element (e.g., in an image parameter set). A tile row refers to a rectangular area of a CTU with a height specified by a syntax element (e.g., in an image parameter set) and a width equal to the width of the image.
[0053] In some examples, a tile can be divided into multiple bricks, where each brick can include one or more CTU rows within the tile. A tile that is not divided into multiple bricks can also be referred to as a brick. However, bricks that are a true subset of a tile may not be referred to as tiles. Bricks in an image can also be arranged in slices. A slice can be an integer number of bricks in the image, which can be uniquely contained within a single Network Abstraction Layer (NAL) unit. In some examples, a slice includes multiple complete tiles or a continuous sequence of complete bricks that include only one tile.
[0054] This disclosure uses “NxN” and “N by N” interchangeably to refer to the sample dimensions of a block (such as a CU or other video block) in the vertical and horizontal dimensions, for example, 16x16 samples or 16 by 16 samples. Typically, a 16x16 CU will have 16 samples in the vertical direction (y = 16) and 16 samples in the horizontal direction (x = 16). Similarly, an NxN CU typically has N samples in the vertical direction and N samples in the horizontal direction, where N represents a non-negative integer value. Samples in a CU can be arranged in rows and columns. Furthermore, a CU does not necessarily need to have the same number of samples in the horizontal direction as it does in the vertical direction. For example, a CU can include NxM samples, where M is not necessarily equal to N.
[0055] The video encoder 200 encodes video data for use in predicting and / or residual information, as well as other information, for the CU. The prediction information indicates how the CU will be predicted to form a prediction block for the CU. The residual information typically represents the sample-by-sample difference between a sample of the CU before encoding and the prediction block.
[0056] To predict the Cubic Frame (CU), the video encoder 200 typically forms prediction blocks for the CU using either inter-frame prediction or intra-frame prediction. Inter-frame prediction generally refers to predicting the CU based on data from a previously decoded image, while intra-frame prediction generally refers to predicting the CU based on data from a previously decoded image within the same frame. To perform inter-frame prediction, the video encoder 200 can generate prediction blocks using one or more motion vectors. The video encoder 200 can typically perform a motion search to identify, for example, a reference block that closely matches the CU in terms of the difference between the CU and a reference block. The video encoder 200 can calculate a difference metric using the sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared difference (MSD), or other such difference calculations to determine whether the reference block closely matches the current CU. In some examples, the video encoder 200 can use unidirectional or bidirectional prediction to predict the current CU.
[0057] Some examples of VVC also provide an affine motion compensation mode, which can be considered an inter-frame prediction mode. In affine motion compensation mode, the video encoder 200 can determine two or more motion vectors representing non-translational motion (such as zooming in or out, rotation, perspective motion, or other irregular types of motion).
[0058] To perform intra-frame prediction, the video encoder 200 can select an intra-frame prediction mode to generate prediction blocks. Some examples of VVC provide sixty-seven intra-frame prediction modes, including various directional modes, as well as planar and DC modes. Typically, the video encoder 200 selects an intra-frame prediction mode that describes the samples of the current block (e.g., a block of a CU) to be predicted based on, which are the neighboring samples of the current block. Assuming the video encoder 200 decodes the CTU and CU in raster scan order (from left to right, from top to bottom), such samples can typically be located above, to the upper left, or to the left of the current block within the same image.
[0059] The video encoder 200 encodes data representing the prediction mode used for the current block. For example, for inter-frame prediction modes, the video encoder 200 may encode data indicating which of the various available inter-frame prediction modes is used, as well as motion information for the corresponding mode. For unidirectional or bidirectional inter-frame prediction, for example, the video encoder 200 may use Advanced Motion Vector Prediction (AMVP) or merging modes to encode motion vectors. The video encoder 200 may use similar modes to encode motion vectors used for affine motion compensation modes.
[0060] AV1 includes two common techniques for encoding and decoding blocks of video data. These two common techniques are intra-frame prediction (e.g., intra-frame prediction or spatial prediction) and inter-frame prediction (e.g., inter-frame prediction or temporal prediction). In the context of AV1, when using intra-frame prediction modes to predict blocks of the current frame of video data, the video encoder 200 and video decoder 300 do not use video data from other frames of the video data. For most intra-frame prediction modes, the video encoder 200 encodes the blocks of the current frame based on the difference between sample values in the current block and predicted values generated from reference samples in the same frame. The video encoder 200 determines the predicted values generated from the reference samples based on the intra-frame prediction mode.
[0061] Following a prediction, such as intra-frame or inter-frame prediction of a block, the video encoder 200 can compute residual data for that block. The residual data (such as a residual block) represents the sample-by-sample difference between the block and the prediction block used to form the block, which is formed using the corresponding prediction mode. The video encoder 200 can apply one or more transforms to the residual block to produce transformed data in the transform domain rather than the sample domain. For example, the video encoder 200 can apply a Discrete Cosine Transform (DCT), an integer transform, a wavelet transform, or a conceptually similar transform to the residual video data. Additionally, the video encoder 200 can apply a secondary transform after the first transform, such as a Mode-dependent Inseparable Quadratic Transform (MDNSST), a Signal-dependent Transform, a Karhunen-Loeve Transform (KLT), etc. The video encoder 200 produces transform coefficients after applying one or more transforms.
[0062] As mentioned above, after any transformation to produce transform coefficients, the video encoder 200 can perform quantization of the transform coefficients. Quantization generally refers to the process of quantizing the transform coefficients to potentially reduce the amount of data used to represent them, thereby providing further compression. By performing the quantization process, the video encoder 200 can reduce the bit depth associated with some or all of the transform coefficients. For example, the video encoder 200 can round an n-bit value down to an m-bit value during quantization, where n is greater than m. In some examples, to perform quantization, the video encoder 200 can perform a bitwise right shift of the values to be quantized.
[0063] After quantization, the video encoder 200 can scan the transform coefficients to generate a one-dimensional vector from a two-dimensional matrix including the quantized transform coefficients. The scan can be designed to place higher-energy (and therefore lower-frequency) transform coefficients before the vector and lower-energy (and therefore higher-frequency) transform coefficients after the vector. In some examples, the video encoder 200 can utilize a predefined scan order to scan the quantized transform coefficients to produce a serialized vector, and then entropy-encode the quantized transform coefficients of the vector. In other examples, the video encoder 200 can perform an adaptive scan. After scanning the quantized transform coefficients to form a one-dimensional vector, the video encoder 200 can entropy-encode the one-dimensional vector, for example, according to context-adaptive binary arithmetic decoding (CABAC). The video encoder 200 can also entropy-encode the values of syntax elements used to describe metadata associated with the encoded video data for use by the video decoder 300 when decoding the video data.
[0064] To perform CABAC, the video encoder 200 can assign context from within a context model to the symbols to be transmitted. Context may involve, for example, whether the neighboring values of a symbol are zero. Probability determination can be based on the context assigned to the symbol.
[0065] The video encoder 200 can also generate syntax data (such as block-based syntax data, image-based syntax data, and sequence-based syntax data) or other syntax data (such as sequence parameter sets (SPS), image parameter sets (PPS), or video parameter sets (VPS)) for the video decoder 300, for example, in image headers, block headers, or slice headers. Similarly, the video decoder 300 can decode such syntax data to determine how to decode the corresponding video data.
[0066] In this way, the video encoder 200 can generate a bitstream that includes encoded video data, such as syntax elements describing the segmentation of an image into blocks (e.g., CUs) and prediction and / or residual information for those blocks. Finally, the video decoder 300 can receive the bitstream and decode the encoded video data.
[0067] Typically, the video decoder 300 performs the reverse process of the video encoder 200 to decode the encoded video data of the bitstream. For example, the video decoder 300 can use CABAC to decode the values of syntax elements used for the bitstream in a manner substantially similar to, but reversed, the CABAC encoding process of the video encoder 200. Syntax elements can define segmentation information for segmenting images into CTUs, and for segmenting each CTU according to a corresponding segmentation structure (such as a QTBT structure) to define the CUs of the CTU. Syntax elements can also define prediction and residual information for blocks (e.g., CUs) of the video data.
[0068] The residual information can be represented, for example, by quantized transform coefficients. The video decoder 300 can inversely quantize and inversely transform the quantized transform coefficients of the block to reconstruct the residual block used for that block. The video decoder 300 uses a signal-informed prediction mode (intra-frame prediction or inter-frame prediction) and associated prediction information (e.g., motion information for inter-frame prediction) to form a prediction block for that block. The video decoder 300 can then combine the prediction block and the residual block (on a sample-by-sample basis) to reconstruct the original block. The video decoder 300 can perform additional processing, such as performing a deblocking process to reduce visual artifacts along the block boundaries.
[0069] In summary, this disclosure may involve "signaling" certain information (such as syntax elements). The term "signaling" can generally refer to the transmission of values for syntax elements and / or other data for decoding encoded video data. That is, video encoder 200 can signal values for syntax elements in the bitstream. Generally, signaling refers to generating values in the bitstream. As mentioned above, source device 102 can transmit the bitstream to destination device 116 substantially in real time or not in real time (such as when syntax elements are stored in storage device 112 for later retrieval by destination device 116).
[0070] According to the technology of this disclosure, as will be explained in more detail below, the video encoder 200 and the video decoder 300 can be configured to: reconstruct video data to generate reconstructed video data; perform multiple loop filter operations on the reconstructed video data in parallel, wherein the multiple loop filter operations include a first filter operation that is not a bilateral filter operation or a sample adaptive offset (SAO) filter operation; and perform a joint truncation operation on a first output of the first filter operation and a second output of a second loop filter operation among the multiple loop filter operations. In one example, the first filter operation is a CCSAO filter operation.
[0071] In the enhanced compression model (ECM 1.0) (in ECM-1.0: https: / / vcgit.hhi.fraunhofer.de / ecm / VVCSoftware_VTM / - / tree / ECM-1.0 Available at [location], two filters, known as Sample Adaptive Offset (SAO) and Bilateral Filter (BIF), are configured to operate in parallel. In ECM 1.0, both SAO and BIF are configured to use deblocked samples (e.g., samples after applying the deblocking filter) as input, and the outputs of both the bilateral filter and the SAO filter are offset together and jointly truncated. These truncated output samples are further fed to an Adaptive Loop Filter (ALF) as input samples.
[0072] Cross-Component Sample Adaptive Offset (CCSAO) is a loop filter currently being investigated in exploratory experiments with ECM. Like the SAO and BIF filters, the cross-component SAO also takes deblocked samples from the deblocking filter as input. In one example, the CCSAO filter operates in parallel with the SAO and BIF. However, the offset generated by the CCSAO is only added to the truncated outputs of the SAO and BIF. Given this implementation, the example CCSAO design has several drawbacks. One drawback is the need for an additional truncating operation for each sample, which increases the computational complexity of both the video encoder and the video decoder. Another drawback is the need to generate the truncated outputs of the SAO and BIF, and then store the truncated outputs in memory before adding the CCSAO offset. In some examples, this implementation may increase memory requirements or potentially add additional latency to sample processing.
[0073] Similar to SAO, the CCSAO filter is configured to classify reconstructed samples into distinct categories. The CCSAO filter can then derive an offset for each category and add that offset to the reconstructed samples within that category. However, unlike SAO, which uses only one luminance / chrominance component of the current sample as input, the CCSAO filter utilizes all three components (e.g., YUV or YCbCr) to classify the current sample into distinct categories. To facilitate parallel processing, the output samples from the deblocking filter are used as input to the CCSAO filter. Figure 2 This is a block diagram illustrating an example of the CCSAO process 400.
[0074] exist Figure 2 In this process, a deblocking filter (DBF), a SAO filter, and a CCSAO filter are used to filter the luminance (Y) and chrominance (U, V) samples. DBF Y 402A is configured to perform deblocking filtering on the luminance (Y) samples of the reconstructed video data. DBF U 402B is configured to perform deblocking filtering on the first chrominance (U) sample of the reconstructed video data. DBF V 402B is configured to perform deblocking filtering on the second chrominance (V) sample of the reconstructed video data.
[0075] SAO Y 404A performs SAO filtering on the output of DBF Y 402A. SAO U 404B performs SAO filtering on the output of DBF U 402B. SAO V 404C performs SAO filtering on the output of DBF V 402C. Each of CCSAO Y 406A, CCSAO U 406B, and CCSAO V 406C is configured to perform CCSAO filtering on the outputs of DBF Y 402A, DBF U 402B, and DBF V 402C. Adder 408A sums the sample values of the outputs of CCSAO Y 406A and SAO Y 404A. Adder 408B sums the sample values of the outputs of CCSAO U 406B and SAO U 404B. Adder 408C sums the sample values of the outputs of CCSAO V 406C and SAOV 404C.
[0076] In one example of the CCSAO design, to achieve a better complexity / performance tradeoff, only band bias (BO) is used to enhance the quality of the reconstructed samples. For a given luma / chroma sample, three candidate samples are selected to classify the given sample into different categories: a co-located Y sample, a co-located U sample, and a co-located V sample. The video encoder 200 and video decoder 300 can then classify the sample values of these three selected samples into three different band biases. Y band U band V}, and the class of a given sample can be indicated using the joint index i. Determining an offset and adding it to the reconstructed samples falling into that class can be formulated as:
[0077]
[0078] In equation (1), {Y col U col V col} represents three co-located samples selected to classify the current sample; {N} Y N U N V} are respectively applied to {Y col U col V col The total number of equal-width segments; BD is the internal decoding bit depth; C ree and C′ rec These are the reconstructed samples before and after applying the CCSAO filter; σ CCSAO[i] is the value of the CCSAO offset applied to the i-th BO category. Clip1 is the clipping function. In one example of CCSAO, co-located luminance samples can be selected from nine (9) candidate locations, while the co-located chrominance sample locations are fixed, as shown in Figure 3 As depicted in the text. Figure 3 This is a conceptual diagram showing an example of candidate locations for the CCSAO classifier. Figure 3 Nine candidates 410 (labeled 0-8) for co-located luminance samples are shown, as well as a co-located U chromaticity sample 412 at position 4 and a co-located V chromaticity sample 414 at position 4.
[0079] Similar to SAO, different classifiers can be applied to different local regions to further enhance the overall image quality. At the frame level, signals are used to inform the parameters (e.g., Y) used for each classifier. col N Y N U N V The position and offset), and explicitly signal and switch which classifier to use at the decode tree block (CTB) level. In one example, for each classifier, {N Y N U N V The maximum value of} is set to {16, 4, 4}, and the offset is constrained to the range [-15, 15]. In one example, although the maximum classifier per frame is constrained to 4, other constraints can be used.
[0080] Figure 4 This is a flowchart illustrating an example of joint capture using SAO, BIF, and CCSAO. Figure 4 In the example CCSAO design, all three filters (SAO 420, BIF 422, and CCSAO 424) take the deblocked samples from deblocking filter 426 as input, and each filter generates an output offset. The output offsets of SAO 420 and BIF 422 are added to the samples from deblocking filter 426 by adder 428, and the summed output of adder 428 is truncated. Adder 430 then sums the output offset of CCSAO 424 with the summed truncated output samples of SAO 420 and BIF 422 (e.g., the output of adder 428). The output of adder 430 is then sent to an adaptive loop filter (ALF). In other examples, the output of adder 430 is stored in memory (such as a decoded image buffer). Figure 4 The example CCSAO design may exhibit several drawbacks:
[0081] 1) The first drawback is that each sample requires an additional truncation operation, which adds computational complexity. The truncation operation essentially ensures that the filtered sample values do not underflow or overflow, and that the final filtered values are within a specified dynamic range. This dynamic range can vary based on the components (luminance or chrominance) and depends on the bit depth of the input video.
[0082] 2) The second drawback is that before adding the output offset of CCSAO 424, the truncated outputs of SAO 420 and BIF 422 need to be generated and stored in memory, which may increase memory requirements in some implementations or add additional latency to sample processing.
[0083] In view of the above-mentioned drawbacks, this disclosure describes the following techniques for joint interception when two or more filters (e.g., loop filters) are applied. The techniques of this disclosure can be used in any combination or with any combination of filters. The techniques of this disclosure are described below with reference to CCSAO filters, but the techniques of this disclosure are not limited thereto. The techniques of this disclosure can be applied to joint interception of any filter operation that is not a bilateral filter operation or a sample adaptive offset (SAO) filter operation (when such filter operations are performed in parallel with other filter operations). Filter operations are performed in parallel when they operate on the same input and not necessarily simultaneously. For example, refer to... Figure 4 SAO 420, BIF 422 and CCSAO 424 operate in parallel because they each take the output of the deblocking filter 426 as input.
[0084] Typically, the video encoder 200 and video decoder 300 can be configured to: reconstruct video data to generate reconstructed video data; perform multiple loop filter operations on the reconstructed video data in parallel, wherein the multiple loop filter operations include a first filter operation that is not a bilateral filter operation or a sample adaptive offset (SAO) filter operation; and perform a joint truncation operation on a first output of the first filter operation and a second output of a second loop filter operation among the multiple loop filter operations. In one example, the first filter operation is a CCSAO filter operation. However, the first filter can be another filter (e.g., other than BIF or SAO) performed in parallel with other filter operations. In one example, the multiple loop filter operations also include at least one of a bilateral filter (BIF) operation or a sample adaptive offset (SAO) filter operation. In one example, the video encoder 200 and video decoder 300 can be configured to: perform a deblocking filter operation on the reconstructed video data before performing the multiple loop filter operations. Alternatively, the video encoder 200 and video decoder 300 can also be configured to: perform an adaptive loop filter operation after performing two or more loop filter operations.
[0085] In the first example of this disclosure, such as Figure 5 As shown, a single joint truncation operation is performed at adder 500. Deblocking filter 426 generates a deblocked output sample. The output sample generated by deblocking filter 426 is used as input to SAO 420, BIF 422, and CCSAO 424. Each of SAO 420, BIF 422, and CCSAO generates an output offset. Adder 500 adds the output offsets calculated by SAO 420, BIF 422, and CCSAO 424 and adds them to the deblocked sample generated from deblocking filter 426. After adding the sum, adder 500 performs a joint truncation operation on the sum. Figure 5 The technology mitigates the aforementioned drawbacks while maintaining compatibility with... Figure 4 The CCSAO technique proposed in this paper has the same objective and subjective quality.
[0086] exist Figure 5 In the example, in order to perform a joint truncation operation on the first output of CCSAO 424 and at least the second output of the second loop filter in a plurality of loop filter operations (e.g., deblocking filter 426, SAO 420, or BIF 422), the video encoder 200 and / or the video decoder 300 are configured to: add the first output of CCSAO 424 to the respective outputs of SAO 420, BIF 422, and deblocking filter 426 to generate a first sum; and perform a joint truncation operation on the first sum.
[0087] In the second example of this disclosure, to avoid overflow of the offsets of the SAO filter, BIF, and CCSAO filter, the video encoder 200 and video decoder 300 can apply a truncation operation to the output offset of each filter, such as... Figure 6 As shown in the image.
[0088] Again, in Figure 6 In the example, deblocking filter 426 generates deblocked output samples. The output samples generated by deblocking filter 426 are used as inputs to SAO 420, BIF 422, and CCSAO 424. Each of SAO 420, BIF 422, and CCSAO generates an output offset. A truncation function 600 is used to truncate the output offset of SAO 420. A truncation function 602 is used to truncate the output offset of BIF 422. A truncation function 604 is used to truncate the output offset of CCSAO 424. Adder 606 adds the truncated outputs of each of SAO 420, BIF 422, and CCSAO 424 to the deblocked samples generated by deblocking filter 426 to generate a sum. Adder 606 then truncates this sum.
[0089] exist Figure 6 In the example, to perform a joint truncation operation on the first output of CCSAO 424 and at least the second output of the second loop filter of multiple loop filter operations (e.g., deblocking filter 426, SAO 420, or BIF 422), the video encoder 200 and / or the video decoder 300 are configured to truncate the corresponding output of each of SAO 420, BIF 422, and CCSAO 424 (e.g., using truncation functions 600, 602, and 604) to generate a corresponding truncated output. The video encoder 200 and the video decoder 300 may add the corresponding truncated outputs of SAO 420, BIF 422, and CCSAO 424 together with samples from deblocking filter 426 (e.g., using adder 606) to generate a first sum, and perform a joint truncation operation on the first sum (e.g., at the output of adder 606).
[0090] In the third example of this disclosure, since the offsets of the SAO and CCSAO filters can be signaled in the bitstream, bitstream constraints can be applied to avoid overflow of the SAO and CCSAO offsets. Therefore, as... Figure 7 As shown, the truncation can be applied only to the output offset of the BIF.
[0091] Again, in Figure 7In the example, deblocking filter 426 generates deblocked output samples. The output samples generated by deblocking filter 426 are used as inputs to SAO 420, BIF 422, and CCSAO 424. Each of SAO 420, BIF 422, and CCSAO generates an output offset. A truncation function 700 is used to truncate the output offset of BIF 422. Adder 702 adds the truncated output of truncation function 700 to the output offsets of each of SAO and CCSAO 424, along with the deblocked samples generated by deblocking filter 426, to generate a sum. Adder 702 then truncates this sum.
[0092] exist Figure 7 In the example, to perform a joint truncation operation on the first output of CCSAO 424 and at least the second output of the second loop filter of multiple loop filter operations (e.g., deblocking filter 426, SAO 420, or BIF 422), video encoder 200 and / or video decoder 300 are configured to truncate the third output of BIF 422 (e.g., using truncation function 700) to generate a truncated output of BIF. Video encoder 200 and video decoder 300 may add the truncated output of BIF 422 to the first output of CCSAO 424 and the corresponding outputs of SAO 420 and deblocking filter 426 to generate a first sum. Video encoder 200 and video decoder 300 may then perform a joint truncation operation on the first sum (e.g., at the output of adder 702).
[0093] In the fourth example of this disclosure, because SAO is applied in the previous codec, it is possible to achieve a certain level of performance relative to... Figure 4 Reuse SAO's control logic, such as Figure 8 As shown in the image. Figure 8 This is a flowchart illustrating an example of joint capture, showing the output offset of BIF 422 and the sum of the outputs of SAO 420 and deblocking filter 426.
[0094] Again, in Figure 8In the example, deblocking filter 426 generates deblocked output samples. The output samples generated by deblocking filter 426 are used as inputs to SAO 420, BIF 422, and CCSAO 424. Each of SAO 420, BIF 422, and CCSAO generates an output offset. Adder 800 adds the output offset of SAO 420 to the output of deblocking filter 426 and generates a sum (e.g., a SAO sum). This SAO sum (e.g., the result of adding the SAO offset to the deblocked samples) can be truncated by truncation function 804. The output offset of BIF 422 can be truncated by truncation function 802. In some examples, the output offset of CCSAO 424 can also be truncated. Then, the output offsets of BIF 422 (e.g., truncated by truncation function 802) and CCSAO 424 can be added to the truncated SAO sum by adder 806. Then, adder 806 can truncate the resulting sum.
[0095] exist Figure 8 In the example, to perform a joint truncation operation on the first output of CCSAO 424 and at least the second output of the second loop filter of multiple loop filter operations (e.g., deblocking filter 426, SAO 420, or BIF 422), video encoder 200 and / or video decoder 300 are configured to: truncate the third output of BIF 422 (e.g., using truncation function 802) to generate a first truncated output. Video encoder 200 and / or video decoder 300 may sum the corresponding outputs of SAO 420 and deblocking filter 426 to create a first sum, and may truncate the first sum to form a second truncated output. The video encoder 200 and / or video decoder 300 may further sum the first truncated output, the second truncated output, and the first output of the CCSAO 424 (e.g., using adder 806) to create a second sum, and then a joint truncation operation may be performed on the second sum (e.g., at the output of adder 806).
[0096] Figure 9 This is a block diagram illustrating an example video encoder 200 that can perform the techniques described in this disclosure. Figure 9 This disclosure is provided for illustrative purposes and should not be construed as limiting the scope of the techniques illustrated and described in this disclosure. For illustrative purposes, this disclosure describes a video encoder 200 based on VVC (ITU-T H.266 under development) and HEVC (ITU-T H.265) technologies. However, the technologies of this disclosure can be implemented by video encoding devices configured for other video decoding standards and video decoding formats, such as AV1 and subsequent versions of the AV1 video decoding format.
[0097] exist Figure 9 In the example, the video encoder 200 includes a video data memory 230, a mode selection unit 202, a residual generation unit 204, a transform processing unit 206, a quantization unit 208, an inverse quantization unit 210, an inverse transform processing unit 212, a reconstruction unit 214, a filter unit 216, a decoded picture buffer (DPB) 218, and an entropy coding unit 220. Any or all of the video data memory 230, mode selection unit 202, residual generation unit 204, transform processing unit 206, quantization unit 208, inverse quantization unit 210, inverse transform processing unit 212, reconstruction unit 214, filter unit 216, DPB 218, and entropy coding unit 220 can be implemented in one or more processors or in processing circuitry. For example, the units of the video encoder 200 can be implemented as one or more circuit or logic elements as part of hardware circuitry, or as part of a processor, ASIC, or FPGA. Furthermore, the video encoder 200 may include additional or alternative processors or processing circuitry to perform these and other functions.
[0098] The video data storage device 230 can store video data to be encoded by the components of the video encoder 200. The video encoder 200 can obtain data from, for example, a video source 104 (…). Figure 1 The video data memory 230 receives video data stored in the video data memory 230. The DPB 218 can act as a reference picture memory, storing reference video data for use when the video encoder 200 predicts subsequent video data. The video data memory 230 and DPB 218 can be formed from any of a variety of memory devices, such as dynamic random access memory (DRAM) (including synchronous DRAM (SDRAM)), magnetoresistive RAM (MRAM), resistive RAM (RRAM), or other types of memory devices. The video data memory 230 and DPB 218 can be provided by the same memory device or separate memory devices. In various examples, the video data memory 230 can be on-chip (as shown) with other components of the video encoder 200, or off-chip relative to those components.
[0099] In this disclosure, references to video data memory 230 should not be construed as limited to memory internal to video encoder 200 (unless so specifically described) or memory external to video encoder 200 (unless so specifically described). Rather, references to video data memory 230 should be understood as reference memory storing video data received by video encoder 200 for encoding (e.g., video data for the current block to be encoded). Figure 1The memory 106 can also provide temporary storage for the outputs from the various units of the video encoder 200.
[0100] It shows Figure 9 The various units help to understand the operations performed by the video encoder 200. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Fixed-function circuits refer to circuits that provide a specific function and are pre-configured regarding the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functionality in terms of the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is generally immutable. In some examples, one or more of these units can be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units can be integrated circuits.
[0101] The video encoder 200 may include an arithmetic logic unit (ALU), an essential function unit (EFU), digital circuitry, analog circuitry, and / or a programmable core, all formed by programmable circuitry. In an example where software executed by programmable circuitry is used to perform the operation of the video encoder 200, memory 106 ( Figure 1 The video encoder 200 may store instructions (e.g., object code) of the software received and executed by the video encoder 200, or another memory (not shown) within the video encoder 200 may store such instructions.
[0102] The video data storage unit 230 is configured to store the received video data. The video encoder 200 can retrieve images of the video data from the video data storage unit 230 and provide the video data to the residual generation unit 204 and the mode selection unit 202. The video data in the video data storage unit 230 can be the raw video data to be encoded.
[0103] The mode selection unit 202 includes a motion estimation unit 222, a motion compensation unit 224, and an intra-frame prediction unit 226. The mode selection unit 202 may include additional functional units that perform video prediction based on other prediction modes. As an example, the mode selection unit 202 may include a palette unit, an intra-block copy unit (which may be part of the motion estimation unit 222 and / or the motion compensation unit 224), an affine unit, a linear model (LM) unit, etc.
[0104] Mode selection unit 202 typically coordinates multiple coding passes to test combinations of coding parameters and the rate-distortion values obtained for such combinations. Coding parameters may include segmenting the CTU into CUs, the prediction mode for the CUs, the transformation type of the residual data for the CUs, and the quantization parameters of the residual data for the CUs. Mode selection unit 202 can ultimately select a combination of coding parameters that yields a better rate-distortion value than other tested combinations.
[0105] The video encoder 200 can segment images retrieved from the video data storage 230 into a series of CTUs and encapsulate one or more CTUs within a slice. The mode selection unit 202 can segment the CTUs of the image according to a tree structure (such as the MTT structure, QTBT structure, superblock structure, or quadtree structure described above). As mentioned above, the video encoder 200 can form one or more CUs by segmenting CTUs according to a tree structure. Such CUs can also generally be referred to as "video blocks" or "blocks".
[0106] Typically, mode selection unit 202 also controls its components (e.g., motion estimation unit 222, motion compensation unit 224, and intra-prediction unit 226) to generate prediction blocks for the current block (e.g., the current CU, or the overlapping portion of PU and TU in HEVC). To perform inter-frame prediction for the current block, motion estimation unit 222 may perform a motion search to identify one or more closely matching reference blocks in one or more reference images (e.g., one or more previously decoded images stored in DPB 218). Specifically, motion estimation unit 222 may calculate values representing the similarity between a potential reference block and the current block, for example, based on sum of absolute differences (SAD), sum of squared differences (SSD), mean absolute difference (MAD), mean squared error (MSD), etc. Motion estimation unit 222 may typically perform these calculations using sample-by-sample differences between the current block and the reference blocks being considered. Motion estimation unit 222 may identify the reference block with the lowest value obtained from these calculations, indicating the reference block that most closely matches the current block.
[0107] Motion estimation unit 222 can generate one or more motion vectors (MVs) that define the position of a reference block in a reference image relative to the position of the current block in the current image. Motion estimation unit 222 can then provide the motion vectors to motion compensation unit 224. For example, for unidirectional inter-frame prediction, motion estimation unit 222 can provide a single motion vector, while for bidirectional inter-frame prediction, motion estimation unit 222 can provide two motion vectors. Motion compensation unit 224 can then use the motion vectors to generate prediction blocks. For example, motion compensation unit 224 can use the motion vectors to retrieve data for the reference blocks. As another example, if the motion vectors have fractional-sample precision, motion compensation unit 224 can interpolate the values used for the prediction blocks according to one or more interpolation filters. Furthermore, for bidirectional inter-frame prediction, motion compensation unit 224 can retrieve data for two reference blocks identified by corresponding motion vectors and combine the retrieved data, for example, by per-sample averaging or weighted averaging.
[0108] When operating according to the AV1 video decoding format, the motion estimation unit 222 and the motion compensation unit 224 can be configured to encode the decoded blocks of video data (e.g., both luma and chroma decoded blocks) using translational motion compensation, affine motion compensation, overlap block motion compensation (OBMC), and / or composite inter-intra-frame prediction.
[0109] As another example, for intra-prediction or intra-prediction decoding, intra-prediction unit 226 can generate a prediction block based on samples adjacent to the current block. For example, in directional mode, intra-prediction unit 226 can typically mathematically combine the values of adjacent samples and fill these calculated values across the current block in a defined direction to generate a prediction block. As another example, in DC mode, intra-prediction unit 226 can calculate the average of the adjacent samples of the current block and generate a prediction block to include the obtained average for each sample in the prediction block.
[0110] When operating according to the AV1 video decoding format, the intra-frame prediction unit 226 can be configured to encode decoded blocks of video data (e.g., both luma-decoded blocks and chroma-decoded blocks) using directional intra-frame prediction, non-directional intra-frame prediction, recursive filter intra-frame prediction, chroma-based (CFL) prediction, intra-block copy (IBC), and / or palette modes. The mode selection unit 202 may include additional functional units for performing video prediction based on other prediction modes.
[0111] Mode selection unit 202 provides a prediction block to residual generation unit 204. Residual generation unit 204 receives the original, uncoded version of the current block from video data memory 230 and the prediction block from mode selection unit 202. Residual generation unit 204 calculates the sample-by-sample difference between the current block and the prediction block. The resulting sample-by-sample difference defines the residual block for the current block. In some examples, residual generation unit 204 may also determine the difference between sample values in the residual block to generate the residual block using residual differential pulse decode-modulation (RDPCM). In some examples, one or more subtractor circuits performing binary subtraction may be used to form residual generation unit 204.
[0112] In the example where mode selection unit 202 divides a CU into PUs, each PU can be associated with a luma prediction unit and a corresponding chroma prediction unit. Video encoder 200 and video decoder 300 can support PUs of various sizes. As noted above, the size of a CU can refer to the size of the luma decoding block of the CU, while the size of a PU can refer to the size of the luma prediction unit of the PU. Assuming a particular CU size is 2Nx2N, video encoder 200 can support PU sizes of 2Nx2N or NxN for intra-frame prediction, and symmetrical PU sizes of 2Nx2N, 2NxN, Nx2N, NxN, or similar sizes for inter-frame prediction. Video encoder 200 and video decoder 300 can also support asymmetric segmentation for PU sizes of 2NxnU, 2NxnD, nLx2N, and nRx2N for inter-frame prediction.
[0113] In an example where mode selection unit 202 does not further divide the CU into PUs, each CU can be associated with a luminance decoding block and a corresponding chrominance decoding block. As mentioned above, the size of the CU can refer to the size of the luminance decoding block of the CU. The video encoder 200 and the video decoder 300 can support CU sizes of 2Nx2N, 2NxN, or Nx2N.
[0114] For other video decoding techniques (such as block-copying mode decoding, affine mode decoding, and linear model (LM) mode decoding), mode selection unit 202 generates a prediction block for the current block being encoded via a corresponding unit associated with the decoding technique. In some examples (such as palette mode decoding), mode selection unit 202 may not generate a prediction block, but instead generate syntax elements indicating how the block should be reconstructed based on the selected palette. In such a mode, mode selection unit 202 can provide these syntax elements to entropy coding unit 220 for encoding.
[0115] As described above, the residual generation unit 204 receives video data for the current block and the corresponding prediction block. Then, the residual generation unit 204 generates a residual block for the current block. To generate the residual block, the residual generation unit 204 calculates the sample-by-sample difference between the prediction block and the current block.
[0116] Transform processing unit 206 applies one or more transformations to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 can apply various transformations to the residual block to form the transform coefficient block. For example, transform processing unit 206 can apply a discrete cosine transform (DCT), direction transform, Karhunen-Loeve transform (KLT), or conceptually similar transformations to the residual block. In some examples, transform processing unit 206 can perform multiple transformations on the residual block, such as primary and secondary transformations (e.g., rotation transformations). In some examples, transform processing unit 206 does not apply any transformations to the residual block.
[0117] When operating according to AV1, transform processing unit 206 may apply one or more transforms to the residual block to generate a block of transform coefficients (referred to herein as a "transform coefficient block"). Transform processing unit 206 may apply various transforms to the residual block to form the transform coefficient block. For example, transform processing unit 206 may apply a combination of horizontal / vertical transforms, which may include the Discrete Cosine Transform (DCT), the Asymmetric Discrete Sine Transform (ADST), the Inverted ADST (e.g., the Reverse ADST), and the Identity Transform (IDTX). When using the Identity Transform, the transform is skipped in either the vertical or horizontal direction. In some examples, transform processing may be skipped.
[0118] Quantization unit 208 can quantize the transform coefficients in the transform coefficient block to produce a quantized transform coefficient block. Quantization unit 208 can quantize the transform coefficients of the transform coefficient block based on the quantization parameter (QP) value associated with the current block. Video encoder 200 (e.g., via mode selection unit 202) can adjust the degree of quantization applied to the transform coefficient block associated with the current block by adjusting the QP value associated with the CU. Quantization may cause information loss, and therefore, the quantized transform coefficients may have lower accuracy compared to the original transform coefficients produced by transform processing unit 206.
[0119] The inverse quantization unit 210 and the inverse transform processing unit 212 can apply inverse quantization and inverse transform, respectively, to the quantized transform coefficient block to reconstruct the residual block from the transform coefficient block. The reconstruction unit 214 can generate a reconstructed block corresponding to the current block based on the reconstructed residual block and the prediction block generated by the mode selection unit 202 (although potentially with some degree of distortion). For example, the reconstruction unit 214 can add samples from the reconstructed residual block to corresponding samples from the prediction block generated by the mode selection unit 202 to generate the reconstructed block.
[0120] Filter unit 216 can perform one or more filtering operations on the reconstructed block. For example, filter unit 216 can perform a deblocking operation to reduce block artifacts along the edges of the CU. In some examples, the operations of filter unit 216 can be skipped. In some examples, filter unit 216 can be configured to perform two or more loop filtering operations, including CCSAO, SAO, and BIF. Filter unit 216 can also be configured to perform one or more of the joint truncation operations described above.
[0121] When operating according to AV1, filter unit 216 can perform one or more filter operations on the reconstructed block. For example, filter unit 216 can perform a deblocking operation to reduce blocky artifacts along the edges of the CU. In other examples, filter unit 216 can apply a constrained directional enhancement filter (CDEF), which can be applied after deblocking, and can include an inseparable, nonlinear, low-pass directional filter applied based on the estimated edge direction. Filter unit 216 can also include a loop recovery filter applied after CDEF, and can include a separable, symmetric normalized Wiener filter or a dual-self-guided filter.
[0122] The video encoder 200 stores the reconstructed blocks in the DPB 218. For example, in an example where the filter unit 216 is not operated, the reconstruction unit 214 may store the reconstructed blocks in the DPB 218. In an example where the filter unit 216 is operated, the filter unit 216 may store the filtered reconstructed blocks in the DPB 218. The motion estimation unit 222 and the motion compensation unit 224 may retrieve a reference image formed by the reconstructed (and potentially filtered) blocks from the DPB 218 to perform inter-frame prediction of blocks in subsequent encoded images. Additionally, the intra-frame prediction unit 226 may use the reconstructed blocks of the current image in the DPB 218 to perform intra-frame prediction of other blocks in the current image.
[0123] Typically, entropy coding unit 220 can entropy-encode syntax elements received from other functional components of video encoder 200. For example, entropy coding unit 220 can entropy-encode quantized transform coefficient blocks from quantization unit 208. As another example, entropy coding unit 220 can entropy-encode predictive syntax elements (e.g., motion information for inter-frame prediction or intra-frame mode information for intra-frame prediction) from mode selection unit 202. Entropy coding unit 220 can perform one or more entropy coding operations on syntax elements, another example of video data, to generate entropy-encoded data. For example, entropy coding unit 220 can perform context-adaptive variable-length decoding (CAVLC), CABAC, variable-to-variable (V2V) length decoding, syntax-based context-adaptive binary arithmetic decoding (SBAC), probabilistic interval partitioned entropy (PIPE) decoding, exponential Golomb coding, or another type of entropy coding operation on the data. In some examples, entropy coding unit 220 can operate in a bypass mode where syntax elements are not entropy-encoded.
[0124] The video encoder 200 can output a bitstream that includes the entropy-encoded syntax elements required to reconstruct slices or blocks of images. Specifically, the entropy coding unit 220 can output a bitstream.
[0125] According to AV1, entropy coding unit 220 can be configured as a symbol-to-symbol adaptive multi-symbol arithmetic decoder. The syntax elements in AV1 include an N-element alphabet, and the context (e.g., a probability model) includes a set of N probabilities. Entropy coding unit 220 can store the probabilities as an n-bit (e.g., 15-bit) cumulative distribution function (CDF). Entropy coding unit 220 can perform recursive scaling to update the context using an update factor based on letter size.
[0126] The above operations are described in relation to the blocks. Such a description should be understood as referring to the operations used for the luma decoding block and / or the chroma decoding block. As mentioned above, in some examples, the luma decoding block and the chroma decoding block are the luma and chroma components of the CU. In some examples, the luma decoding block and the chroma decoding block are the luma and chroma components of the PU.
[0127] In some examples, it is not necessary to repeat the operations performed for the luma decoding block for the chroma decoding block. As an example, it is not necessary to repeat the operations used to identify the motion vector (MV) and reference image for the luma decoding block to identify the MV and reference image for the chroma block. Specifically, the MV for the luma decoding block can be scaled to determine the MV for the chroma block, and the reference image can be the same. As another example, the intra-frame prediction process can be the same for both the luma and chroma decoding blocks.
[0128] Video encoder 200 represents an example of a device configured to encode video data, the device including: a memory configured to store video data; and one or more processing units implemented in circuitry and configured to: reconstruct the video data to generate reconstructed video data; perform multiple loop filter operations on the reconstructed video data in parallel, wherein the multiple loop filter operations include a first filter operation that is not a bilateral filter operation or a sample adaptive offset (SAO) filter operation; and perform a joint truncation operation on a first output of the first filter operation and a second output of a second loop filter operation among the multiple loop filter operations.
[0129] Figure 10 This is a block diagram illustrating an example video decoder 300 capable of performing the techniques described herein. Figure 10 This disclosure is provided for illustrative purposes and does not limit the techniques illustrated and described in this disclosure in a general manner. For illustrative purposes, this disclosure describes a video decoder 300 based on VVC (ITU-T H.266 under development) and HEVC (ITU-T H.265) technologies. However, the technologies of this disclosure can be implemented by video decoding devices configured for other video decoding standards.
[0130] exist Figure 10 In the example, the video decoder 300 includes a decoded picture buffer (CPB) memory 320, an entropy decoding unit 302, a prediction processing unit 304, an inverse quantization unit 306, an inverse transform processing unit 308, a reconstruction unit 310, a filter unit 312, and a decoded picture buffer (DPB) 134. Any or all of the CPB memory 320, entropy decoding unit 302, prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, filter unit 312, and DPB 134 can be implemented in one or more processors or in processing circuitry. For example, units of the video decoder 300 can be implemented as one or more circuit or logic elements as part of hardware circuitry, or as part of a processor, ASIC, or FPGA. Furthermore, the video decoder 300 may include additional or alternative processors or processing circuitry to perform these and other functions.
[0131] The prediction processing unit 304 includes a motion compensation unit 316 and an intra-frame prediction unit 318. The prediction processing unit 304 may include an addition unit that performs predictions based on other prediction modes. As an example, the prediction processing unit 304 may include a palette unit, an intra-block copy unit (which may form part of the motion compensation unit 316), an affine unit, a linear model (LM) unit, etc. In other examples, the video decoder 300 may include more, fewer, or different functional components.
[0132] When operating according to AV1, compensation unit 316 can be configured to decode video data decoding blocks (e.g., both luma and chroma decoding blocks) using translational motion compensation, affine motion compensation, OBMC, and / or composite inter-intra-frame prediction, as described above. Intra-frame prediction unit 318 can be configured to decode video data decoding blocks (e.g., both luma and chroma decoding blocks) using directional intra-frame prediction, non-directional intra-frame prediction, recursive filter intra-frame prediction, CFL prediction, intra-block copy (IBC), and / or palette mode, as described above.
[0133] CPB memory 320 can store video data to be decoded by components of video decoder 300, such as encoded video bitstreams. For example, it can be stored from computer-readable medium 110 ( Figure 1 The video data stored in the CPB memory 320 is obtained. The CPB memory 320 may include a CPB that stores encoded video data (e.g., syntax elements) from the encoded video bitstream. Additionally, the CPB memory 320 may store video data other than the syntax elements of the decoded picture, such as temporary data representing the output from various units of the video decoder 300. The DPB 314 typically stores decoded pictures that the video decoder 300 may output and / or use as reference video data when decoding subsequent data or pictures of the encoded video bitstream. The CPB memory 320 and DPB 314 may be formed of any of a variety of memory devices, such as DRAM, including SDRAM, MRAM, RRAM, or other types of memory devices. The CPB memory 320 and DPB 314 may be provided by the same memory device or separate memory devices. In various examples, the CPB memory 320 may be on-chip with other components of the video decoder 300, or off-chip relative to those components.
[0134] Alternatively or concurrently, in some examples, the video decoder 300 can be derived from the memory 120 ( Figure 1The decoded video data is retrieved. In other words, memory 120 can utilize CPB memory 320 to store data, as discussed above. Similarly, when some or all of the functions of video decoder 300 are implemented using software to be executed by the processing circuitry of video decoder 300, memory 120 can store instructions to be executed by video decoder 300.
[0135] It shows Figure 10 The various units shown help to understand the operations performed by the video decoder 300. These units can be implemented as fixed-function circuits, programmable circuits, or a combination thereof. Similar to... Figure 9 Fixed-function circuits refer to circuits that provide a specific function and are pre-configured regarding the operations that can be performed. Programmable circuits refer to circuits that can be programmed to perform various tasks and provide flexible functionality in terms of the operations that can be performed. For example, a programmable circuit can execute software or firmware that causes the programmable circuit to operate in a manner defined by the instructions of the software or firmware. Fixed-function circuits can execute software instructions (e.g., to receive or output parameters), but the type of operation performed by a fixed-function circuit is typically immutable. In some examples, one or more of these units may be different circuit blocks (fixed-function or programmable), and in some examples, one or more of these units may be integrated circuits.
[0136] The video decoder 300 may include an ALU, EFU, digital circuitry, analog circuitry, and / or a programmable core formed by programmable circuitry. In an example where the operation of the video decoder 300 is performed by software executing on the programmable circuitry, on-chip or off-chip memory may store instructions (e.g., object code) of the software received and executed by the video decoder 300.
[0137] Entropy decoding unit 302 can receive encoded video data from the CPB and perform entropy decoding on the video data to reconstruct the syntax elements. Prediction processing unit 304, inverse quantization unit 306, inverse transform processing unit 308, reconstruction unit 310, and filter unit 312 can generate decoded video data based on the syntax elements extracted from the bitstream.
[0138] Typically, the video decoder 300 reconstructs the image block by block. The video decoder 300 can perform the reconstruction operation on each block individually (where the block currently being reconstructed (i.e., decoded) can be referred to as the "current block").
[0139] Entropy decoding unit 302 can entropy decode the syntax elements of the quantized transform coefficients that define the quantized transform coefficient block, as well as transform information such as quantization parameters (QPs) and / or transform mode indications. Inverse quantization unit 306 can use the QPs associated with the quantized transform coefficient block to determine the degree of quantization, and similarly, determine the degree of inverse quantization to be applied by inverse quantization unit 306. Inverse quantization unit 306 can, for example, perform a bitwise left shift operation to inverse quantize the quantized transform coefficients. Inverse quantization unit 306 can thus form a transform coefficient block including the transform coefficients.
[0140] After the inverse quantization unit 306 forms the transform coefficient block, the inverse transform processing unit 308 can apply one or more inverse transforms to the transform coefficient block to generate a residual block associated with the current block. For example, the inverse transform processing unit 308 can apply the inverse DCT, inverse integer transform, inverse Karhunen-Loeve transform (KLT), inverse rotation transform, inverse direction transform, or another inverse transform to the transform coefficient block.
[0141] Furthermore, the prediction processing unit 304 generates a prediction block based on the prediction information syntax elements entropy-decoded by the entropy decoding unit 302. For example, if the prediction information syntax elements indicate that the current block is inter-frame predicted, the motion compensation unit 316 can generate the prediction block. In this case, the prediction information syntax elements may indicate the reference picture from which the reference block is to be retrieved in the DPB 314, and a motion vector identifying the position of the reference block in the reference picture relative to the position of the current block in the current picture. The motion compensation unit 316 can typically be configured with respect to the motion compensation unit 224 ( Figure 9 The method described is basically similar to the way the inter-frame prediction process is performed.
[0142] As another example, if the prediction information syntax element indicates that the current block is intra-predicted, then intra-prediction unit 318 can generate a prediction block according to the intra-prediction mode indicated by the prediction information syntax element. Again, intra-prediction unit 318 can typically be configured with respect to intra-prediction unit 226 ( Figure 9 The intra-prediction process is performed in a manner substantially similar to that described above. The intra-prediction unit 318 can retrieve data from neighboring samples of the current block from the DPB 314.
[0143] Reconstruction unit 310 can reconstruct the current block using the prediction block and the residual block. For example, reconstruction unit 310 can reconstruct the current block by adding the samples of the residual block to the corresponding samples of the prediction block.
[0144] Filter unit 312 can perform one or more filtering operations on the reconstructed block. For example, filter unit 312 can perform a deblocking operation to reduce block artifacts along the edges of the reconstructed block. The operation of filter unit 312 is not necessarily performed in all examples. In some examples, filter unit 312 can be configured to perform two or more loop filtering operations, including CCSAO, SAO, and BIF. Filter unit 316 can also be configured to perform one or more of the joint truncation operations described above.
[0145] The video decoder 300 can store the reconstructed blocks in the DPB 314. For example, in an example where the operation of the filter unit 312 is not performed, the reconstruction unit 310 can store the reconstructed blocks in the DPB 314. In an example where the operation of the filter unit 312 is performed, the filter unit 312 can store the filtered reconstructed blocks in the DPB 314. As discussed above, the DPB 314 can provide reference information (such as samples of the current image for intra-frame prediction and samples of previously decoded images for subsequent motion compensation) to the prediction processing unit 304. Furthermore, the video decoder 300 can output decoded images (e.g., decoded video) from the DPB 314 for use in applications such as... Figure 1 The subsequent presentation on display devices such as display device 118.
[0146] In this manner, video decoder 300 represents an example of a video decoding device, which includes: a memory configured to store video data; and one or more processing units implemented in circuitry and configured to: reconstruct the video data to generate reconstructed video data; perform multiple loop filter operations on the reconstructed video data in parallel, wherein the multiple loop filter operations include a first filter operation that is not a bilateral filter operation or a sample adaptive offset (SAO) filter operation; and perform a joint truncation operation on a first output of the first filter operation and a second output of a second loop filter operation among the multiple loop filter operations.
[0147] Figure 11 This is a flowchart illustrating an example method for encoding a current block according to the technology of this disclosure. The current block may include the current CU. Although regarding video encoder 200 ( Figure 1 and 9 The description is provided, but it should be understood that other devices can be configured to perform the same actions. Figure 11 Similar to the method.
[0148] In this example, the video encoder 200 initially predicts the current block (350). For example, the video encoder 200 may form a prediction block for the current block. Then, the video encoder 200 may compute a residual block for the current block (352). To compute the residual block, the video encoder 200 may compute the difference between the original unencoded block and the prediction block for the current block. Then, the video encoder 200 may transform the residual block and quantize the transform coefficients of the residual block (354). Next, the video encoder 200 may scan the quantized transform coefficients of the residual block (356). During or after the scan, the video encoder 200 may entropy encode the transform coefficients (358). For example, the video encoder 200 may use CAVLC or CABAC to encode the transform coefficients. Then, the video encoder 200 may output the entropy-encoded data of the block (360).
[0149] Figure 12 This is a flowchart illustrating an example method for decoding a current block of video data according to the technology of this disclosure. The current block may include the current CU. Although regarding the video decoder 300 ( Figure 1 and 10 The description is provided, but it should be understood that other devices can be configured to perform the same actions. Figure 12 Similar to the method.
[0150] The video decoder 300 can receive entropy-encoded data for the current block (such as entropy-encoded prediction information and entropy-encoded data for the transform coefficients of the residual block corresponding to the current block) (370). The video decoder 300 can entropy decode the entropy-encoded data to determine the prediction information for the current block and reproduce the transform coefficients of the residual block (372). The video decoder 300 can predict the current block, for example, using an intra-frame or inter-frame prediction mode indicated by the prediction information for the current block (374), to compute a prediction block for the current block. The video decoder 300 can then perform an inverse scan on the reproduced transform coefficients (376) to create a block of quantized transform coefficients. The video decoder 300 can then inverse quantize the transform coefficients and apply the inverse transform to the transform coefficients to produce a residual block (378). Finally, the video decoder 300 can decode the current block by combining the prediction block and the residual block (380).
[0151] Figure 13 This is a flowchart illustrating an example method for filtering the current block according to the techniques described in this disclosure. Figure 13 The technology can be derived from one or more structural components of the video encoder 200 and the video decoder 300 (including the filtering unit 216 of the video encoder 200 (see...)). Figure 9) and / or the filtering unit 312 of the video decoder 300 (see Figure 10 ))implement. Figure 13 The technique can be performed in the reconstruction loop of the video encoder 200, which is typically in... Figure 11 The transformation and quantization residual block (354) process is executed after that. Figure 13 The technology can also be performed by the video decoder 300, and is typically available in [the following]. Figure 12 This occurs after the combined prediction block and residual block (380) process.
[0152] In one example of this disclosure, video encoder 200 and video decoder 300 may be configured to reconstruct video data to generate reconstructed video data (1200). In the example of video encoding, video encoder 200 may be configured to reconstruct video data in a reconstruction loop to generate reconstructed video data. In the example of video decoding, video decoder 300 may be configured to decode video data to generate reconstructed video data.
[0153] The video encoder 200 and video decoder 300 can also be configured to perform multiple loop filter operations on the reconstructed video data in parallel, wherein the multiple loop filter operations include a first filter operation (1202) that is not a bilateral filter operation or a Sample Adaptive Shift (SAO) filter operation. The video encoder 200 and video decoder 300 can also be configured to perform a joint truncation operation (1204) on a first output of the first filter operation and a second output of a second loop filter operation among the multiple loop filter operations. In one example, the first filter operation is a CCSAO filter operation. In another example, the multiple loop filter operations also include at least one of a bilateral filter operation or a SAO filter operation.
[0154] In another example of this disclosure, the video encoder 200 and video decoder 300 may be configured to perform deblocking filtering on the reconstructed video data before performing multiple loop filter operations. The video encoder 200 and video decoder 300 may also be configured to perform adaptive loop filtering after performing multiple loop filter operations.
[0155] In one example, to perform a joint truncation operation on a first output of a CCSAO filter operation and at least a second output of a second loop filter operation among a plurality of loop filter operations, the video encoder 200 and the video decoder 300 are configured to: add the first output of the CCSAO filter operation to the corresponding outputs of the SAO filter operation, the bilateral filter operation, and the deblocking filter operation to generate a first sum; and perform a joint truncation operation on the first sum.
[0156] In another example, to perform a joint truncation operation on a first output of a CCSAO filter operation and at least a second output of a second loop filter operation among a plurality of loop filter operations, the video encoder 200 and the video decoder 300 are configured to: truncate the corresponding output of each of the SAO filter operation, the bilateral filter operation, and the CCSAO filter operation to generate a corresponding truncated output; add the corresponding truncated outputs of the SAO filter operation, the bilateral filter operation, and the CCSAO filter operation together with samples from the deblocking filter operation to generate a first sum; and perform a joint truncation operation on the first sum.
[0157] In another example, to perform a joint truncation operation on a first output of a CCSAO filter operation and at least a second output of a second loop filter operation among a plurality of loop filter operations, the video encoder 200 and the video decoder 300 are configured to: truncate a third output of a bilateral filter operation to generate a truncated output of the bilateral filter operation; add the truncated output of the bilateral filter operation to the first output of the CCSAO filter operation and the corresponding outputs of the SAO filter operation and the deblocking filter operation to generate a first sum; and perform a joint truncation operation on the first sum.
[0158] In another example, to perform a joint truncation operation on a first output of a CCSAO filter operation and at least a second output of a second loop filter operation among multiple loop filter operations, the video encoder 200 and the video decoder 300 are configured to: truncate a third output of a bilateral filter operation to generate a first truncated output; add the corresponding outputs of the SAO filter operation and the deblocking filter operation to create a first sum; truncate the first sum to form a second truncated output; add the first truncated output, the second truncated output, and the first output of the CCSAO filter operation to create a second sum; and perform a joint truncation operation on the second sum.
[0159] Additional aspects of this disclosure are described below.
[0160] Aspect 1A - A method for decoding video data, the method comprising: reconstructing the video data; and performing two or more loop filter operations on the reconstructed video data, including performing a joint truncation operation on the output of the two or more loop filter operations.
[0161] Aspect 2A - The method according to aspect 1A, wherein the two or more loop filter operations include a bilateral filter (BIF), a sample adaptive offset (SAO) filter, and a cross-component SAO (CCSAO) filter.
[0162] Aspect 3A - The method according to any one of Aspects 1A-2A further includes: performing a deblocking filter operation on the reconstructed video data before performing the two or more loop filter operations.
[0163] Aspect 4A - The method according to any one of aspects 1A-3A further includes: performing an adaptive loop filtering operation after performing the two or more loop filter operations.
[0164] Aspect 5A - The method according to any one of Aspects 1A-4A, wherein performing the two or more loop filter operations comprises: adding the outputs of the SAO filter, BIF and CCSAO filter together with samples from the deblocking filter; and performing a truncation operation on the sum.
[0165] Aspect 6A - The method according to any one of Aspects 1A-4A, wherein performing the two or more loop filter operations comprises: truncating the output of each of the SAO filter, BIF and CCSAO filter; adding the truncated outputs of the SAO filter, the BIF and CCSAO filter together with samples from the deblocking filter; and performing a truncating operation on the sum.
[0166] Aspect 7A - The method according to any one of Aspects 1A-4A, wherein performing the two or more loop filter operations comprises: truncating the output of the BIF; adding the truncated output of the BIF together with the outputs of the SAO filter and the CCSAO filter and a sample from the deblocking filter; and performing a truncating operation on the sum.
[0167] Aspect 8A - A method according to any one of Aspects 1A-4A, wherein performing the two or more loop filter operations comprises: truncating the output of the BIF; adding the output of the SAO filter to a sample of the deblocking filter to create a first sum; truncating the first sum to form a first truncated output; adding the truncated output of the BIF to the first truncated output and the output of the CCSAO filter to create a second truncated output; and performing a truncating operation on the second truncated output.
[0168] Aspect 9A - The method according to any one of Aspects 1A-8A, wherein decoding includes decoding.
[0169] Aspect 10A - The method according to any one of aspects 1A-8A, wherein decoding includes encoding.
[0170] Aspect 11A - An apparatus for decoding video data, the apparatus comprising one or more units for performing the method according to any one of aspects 1A-10A.
[0171] Aspect 12A - The device according to aspect 11A, wherein the one or more units include one or more processors implemented in a circuit.
[0172] Aspect 13A - The device according to any one of aspects 11A and 12A further includes: a memory for storing the video data.
[0173] The device according to any one of aspects 11A-13A further includes: a display configured to display decoded video data.
[0174] Aspect 15A - The device according to any one of Aspects 11A-14A, wherein the device includes one or more of a camera, a computer, a mobile device, a broadcast receiver device, or a set-top box.
[0175] Aspect 16A - The device according to any one of aspects 11A-15A, wherein the device includes a video decoder.
[0176] Aspect 17A - The device according to any one of aspects 11A-16A, wherein the device includes a video encoder.
[0177] Aspect 18A - A computer-readable storage medium having instructions stored thereon, which, when executed, cause one or more processors to perform the method according to any one of aspects 1A-10A.
[0178] Aspect 1B - A method for decoding video data, the method comprising: reconstructing the video data to generate reconstructed video data; performing a plurality of loop filter operations in parallel on the reconstructed video data, wherein the plurality of loop filter operations includes a first filter operation that is not a bilateral filter operation or a sample adaptive offset (SAO) filter operation; and performing a joint truncation operation on a first output of the first filter operation and a second output of a second loop filter operation among the plurality of loop filter operations.
[0179] Aspect 2B - The method according to aspect 1B, wherein the plurality of loop filter operations further includes at least one of the bilateral filter operation or the SAO filter operation.
[0180] Aspect 3B - The method according to aspect 1B further includes: performing a deblocking filter operation on the reconstructed video data before performing the plurality of loop filter operations.
[0181] Aspect 4B - The method according to aspect 1B further includes: performing an adaptive loop filtering operation after performing the plurality of loop filter operations.
[0182] Aspect 5B - According to the method of aspect 1B, wherein the first filter operation is a cross-component sample adaptive offset (CCSAO) filter operation.
[0183] Aspect 6B - According to the method of aspect 5B, wherein performing the joint truncation operation on the first output of the CCSAO filter operation and at least the second output of the second loop filter operation among the plurality of loop filter operations comprises: adding the first output of the CCSAO filter operation to the corresponding outputs of the SAO filter operation, the bilateral filter operation, and the deblocking filter operation to generate a first sum; and performing the joint truncation operation on the first sum.
[0184] Aspect 7B - According to the method of aspect 5B, wherein performing the joint truncation operation on the first output of the CCSAO filter operation and at least the second output of the second loop filter operation of the plurality of loop filter operations comprises: truncating the corresponding output of each of the SAO filter operation, the bilateral filter operation, and the CCSAO filter operation to generate a corresponding truncated output; adding the corresponding truncated outputs of the SAO filter operation, the bilateral filter operation, and the CCSAO filter operation together with samples from the deblocking filter operation to generate a first sum; and performing the joint truncation operation on the first sum.
[0185] Aspect 8B - According to the method of aspect 5B, wherein performing the joint truncation operation on the first output of the CCSAO filter operation and at least the second output of the second loop filter operation among the plurality of loop filter operations comprises: truncating the third output of the bilateral filter operation to generate a truncated output of the bilateral filter operation; adding the truncated output of the bilateral filter operation to the first output of the CCSAO filter operation and the corresponding outputs of the SAO filter operation and the deblocking filter operation to generate a first sum; and performing the joint truncation operation on the first sum.
[0186] Aspect 9B - According to the method of aspect 5B, wherein performing the joint truncation operation on the first output of the CCSAO filter operation and at least the second output of the second loop filter operation of the plurality of loop filter operations comprises: truncating the third output of the bilateral filter operation to generate a first truncated output; adding the corresponding outputs of the SAO filter operation and the deblocking filter operation to create a first sum; truncating the first sum to form a second truncated output; adding the first truncated output, the second truncated output, and the first output of the CCSAO filter operation to create a second sum; and performing the joint truncation operation on the second sum.
[0187] Aspect 10B - The method according to aspect 1B, wherein decoding includes encoding, and wherein reconstructing the video data to generate the reconstructed video data includes: reconstructing the video data in a reconstruction loop of a video encoder to generate the reconstructed video data.
[0188] Aspect 11B - The method according to aspect 1B, wherein decoding includes decoding, and wherein reconstructing the video data to generate the reconstructed video data includes: decoding the video data to generate the reconstructed video data.
[0189] Aspect 12B - An apparatus configured to decode video data, the apparatus comprising: a memory configured to store video data; and one or more processors in communication with the memory, the one or more processors configured to: reconstruct the video data to generate reconstructed video data; perform a plurality of loop filter operations on the reconstructed video data in parallel, wherein the plurality of loop filter operations includes a first filter operation that is not a bilateral filter operation or a sample adaptive offset (SAO) filter operation; and perform a joint truncation operation on a first output of the first filter operation and a second output of a second loop filter operation among the plurality of loop filter operations.
[0190] Aspect 13B - The apparatus according to aspect 12B, wherein the plurality of loop filter operations further include at least one of the bilateral filter operation or the sample adaptive offset (SAO) filter operation.
[0191] Aspect 14B - The apparatus according to aspect 12B, wherein the one or more processors are further configured to perform a deblocking filter operation on the reconstructed video data prior to performing the plurality of loop filter operations.
[0192] Aspect 15B - The apparatus according to aspect 12B, wherein the one or more processors are further configured to perform an adaptive loop filtering operation after performing the plurality of loop filter operations.
[0193] Aspect 16B - The apparatus according to aspect 12B, wherein the first filter operation is a cross-component sample adaptive offset (CCSAO) filter operation.
[0194] Aspect 17B - The apparatus according to aspect 16B, wherein, in order to perform the joint truncation operation on the first output of the CCSAO filter operation and at least the second output of the second loop filter operation of the plurality of loop filter operations, the one or more processors are further configured to: add the first output of the CCSAO filter operation to the corresponding outputs of the SAO filter operation, the bilateral filter operation, and the deblocking filter operation to generate a first sum; and perform the joint truncation operation on the first sum.
[0195] Aspect 18B - The apparatus according to aspect 16B, wherein, in order to perform the joint truncation operation on the first output of the CCSAO filter operation and at least the second output of the second loop filter operation of the plurality of loop filter operations, the one or more processors are further configured to: truncate the corresponding output of each of the SAO filter operation, the bilateral filter operation, and the CCSAO filter operation to generate a corresponding truncated output; add the corresponding truncated output of the SAO filter operation, the bilateral filter operation, and the CCSAO filter operation to a sample from the deblocking filter operation to generate a first sum; and perform the joint truncation operation on the first sum.
[0196] Aspect 19B - The apparatus according to aspect 16B, wherein, in order to perform the joint truncation operation on the first output of the CCSAO filter operation and at least the second output of the second loop filter operation of the plurality of loop filter operations, the one or more processors are further configured to: truncate the third output of the bilateral filter operation to generate a truncated output of the bilateral filter operation; add the truncated output of the bilateral filter operation to the first output of the CCSAO filter operation and the corresponding outputs of the SAO filter operation and the deblocking filter operation to generate a first sum; and perform the joint truncation operation on the first sum.
[0197] Aspect 20B - The apparatus according to aspect 16B, wherein, in order to perform the joint truncation operation on the first output of the CCSAO filter operation and at least the second output of the second loop filter operation of the plurality of loop filter operations, the one or more processors are further configured to: truncate the third output of the bilateral filter operation to generate a first truncated output; add the corresponding outputs of the SAO filter operation and the deblocking filter operation to create a first sum; truncate the first sum to form a second truncated output; add the first truncated output, the second truncated output, and the first output of the CCSAO filter operation to create a second sum; and perform the joint truncation operation on the second sum.
[0198] Aspect 21B - The apparatus according to aspect 12B, wherein the apparatus is a video encoder, and wherein, in order to reconstruct the video data to generate the reconstructed video data, the one or more processors are further configured to: reconstruct the video data in a reconstruction loop of the video encoder to generate the reconstructed video data.
[0199] Aspect 22B - The apparatus according to aspect 12B, wherein the apparatus is a video decoder, and wherein, in order to reconstruct the video data to generate the reconstructed video data, the one or more processors are further configured to: decode the video data to generate the reconstructed video data.
[0200] Aspect 23B - An apparatus configured to decode video data, the apparatus comprising: a unit for reconstructing the video data to generate reconstructed video data; a unit for performing a plurality of loop filter operations in parallel on the reconstructed video data, wherein the plurality of loop filter operations includes a first filter operation that is not a bilateral filter operation or a sample adaptive offset (SAO) filter operation; and a unit for performing a joint truncation operation on a first output of the first filter operation and a second output of a second loop filter operation among the plurality of loop filter operations.
[0201] Aspect 24B - A non-transitory computer-readable storage medium storing instructions that, when executed, cause one or more processors of a device configured to decode video data to: reconstruct the video data to generate reconstructed video data; perform a plurality of loop filter operations on the reconstructed video data in parallel, wherein the plurality of loop filter operations includes a first filter operation that is not a bilateral filter operation or a sample adaptive offset (SAO) filter operation; and perform a joint truncation operation on a first output of the first filter operation and a second output of a second loop filter operation among the plurality of loop filter operations.
[0202] Aspect 1C - A method for decoding video data, the method comprising: reconstructing the video data to generate reconstructed video data; performing a plurality of loop filter operations in parallel on the reconstructed video data, wherein the plurality of loop filter operations includes a first filter operation that is not a bilateral filter operation or a sample adaptive offset (SAO) filter operation; and performing a joint truncation operation on a first output of the first filter operation and a second output of a second loop filter operation among the plurality of loop filter operations.
[0203] Aspect 2C - The method according to aspect 1C, wherein the plurality of loop filter operations further includes at least one of the bilateral filter operation or the SAO filter operation.
[0204] Aspect 3C - The method according to any one of Aspects 1C-2C further includes: performing a deblocking filter operation on the reconstructed video data before performing the plurality of loop filter operations.
[0205] Aspect 4C - The method according to any one of aspects 1C-3C further includes: performing an adaptive loop filtering operation after performing the plurality of loop filter operations.
[0206] Aspect 5C - The method according to any one of Aspects 1C-4C, wherein the first filter operation is a cross-component sample adaptive offset (CCSAO) filter operation.
[0207] Aspect 6C - According to the method of aspect 5C, wherein performing the joint truncation operation on the first output of the CCSAO filter operation and at least the second output of the second loop filter operation among the plurality of loop filter operations comprises: adding the first output of the CCSAO filter operation to the corresponding outputs of the SAO filter operation, the bilateral filter operation, and the deblocking filter operation to generate a first sum; and performing the joint truncation operation on the first sum.
[0208] Aspect 7C - According to the method of aspect 5C, wherein performing the joint truncation operation on the first output of the CCSAO filter operation and at least the second output of the second loop filter operation of the plurality of loop filter operations comprises: truncating the corresponding output of each of the SAO filter operation, the bilateral filter operation, and the CCSAO filter operation to generate a corresponding truncated output; adding the corresponding truncated outputs of the SAO filter operation, the bilateral filter operation, and the CCSAO filter operation together with samples from the deblocking filter operation to generate a first sum; and performing the joint truncation operation on the first sum.
[0209] Aspect 8C - The method according to aspect 5C, wherein performing the joint truncation operation on the first output of the CCSAO filter operation and at least the second output of the second loop filter operation of the plurality of loop filter operations comprises: truncating the third output of the bilateral filter operation to generate a truncated output of the bilateral filter operation; adding the truncated output of the bilateral filter operation to the first output of the CCSAO filter operation and the corresponding outputs of the SAO filter operation and the deblocking filter operation to generate a first sum; and performing the joint truncation operation on the first sum.
[0210] Aspect 9C - According to the method of aspect 5C, wherein performing the joint truncation operation on the first output of the CCSAO filter operation and at least the second output of the second loop filter operation of the plurality of loop filter operations comprises: truncating the third output of the bilateral filter operation to generate a first truncated output; adding the corresponding outputs of the SAO filter operation and the deblocking filter operation to create a first sum; truncating the first sum to form a second truncated output; adding the first truncated output, the second truncated output, and the first output of the CCSAO filter operation to create a second sum; and performing the joint truncation operation on the second sum.
[0211] Aspect 10C - The method according to any one of Aspects 1C-9C, wherein decoding includes encoding, and wherein reconstructing the video data to generate the reconstructed video data includes: reconstructing the video data in a reconstruction loop of a video encoder to generate the reconstructed video data.
[0212] Aspect 11C - The method according to any one of Aspects 1C-9C, wherein decoding includes decoding, and wherein reconstructing the video data to generate the reconstructed video data includes: decoding the video data to generate the reconstructed video data.
[0213] Aspect 12C - An apparatus configured to decode video data, the apparatus comprising: a memory configured to store video data; and one or more processors in communication with the memory, the one or more processors configured to: reconstruct the video data to generate reconstructed video data; perform a plurality of loop filter operations on the reconstructed video data in parallel, wherein the plurality of loop filter operations includes a first filter operation that is not a bilateral filter operation or a sample adaptive offset (SAO) filter operation; and perform a joint truncation operation on a first output of the first filter operation and a second output of a second loop filter operation among the plurality of loop filter operations.
[0214] Aspect 13C - The apparatus according to aspect 12C, wherein the plurality of loop filter operations further include at least one of the bilateral filter operation or the sample adaptive offset (SAO) filter operation.
[0215] Aspect 14C - The apparatus according to any one of aspects 12C-13C, wherein the one or more processors are further configured to perform a deblocking filter operation on the reconstructed video data prior to performing the plurality of loop filter operations.
[0216] Aspect 15C - The apparatus according to any one of aspects 12C-14C, wherein the one or more processors are further configured to perform an adaptive loop filtering operation after performing the plurality of loop filter operations.
[0217] Aspect 16C - The apparatus according to any one of aspects 12C-15C, wherein the first filter operation is a cross-component sample adaptive offset (CCSAO) filter operation.
[0218] Aspect 17C - The apparatus according to aspect 16C, wherein, in order to perform the joint truncation operation on the first output of the CCSAO filter operation and at least the second output of the second loop filter operation of the plurality of loop filter operations, the one or more processors are further configured to: add the first output of the CCSAO filter operation to the corresponding outputs of the SAO filter operation, the bilateral filter operation, and the deblocking filter operation to generate a first sum; and perform the joint truncation operation on the first sum.
[0219] Aspect 18C - The apparatus according to aspect 16C, wherein, in order to perform the joint truncation operation on the first output of the CCSAO filter operation and at least the second output of the second loop filter operation of the plurality of loop filter operations, the one or more processors are further configured to: truncate the corresponding output of each of the SAO filter operation, the bilateral filter operation, and the CCSAO filter operation to generate a corresponding truncated output; add the corresponding truncated output of the SAO filter operation, the bilateral filter operation, and the CCSAO filter operation to a sample from the deblocking filter operation to generate a first sum; and perform the joint truncation operation on the first sum.
[0220] Aspect 19C - The apparatus according to aspect 16C, wherein, in order to perform the joint truncation operation on the first output of the CCSAO filter operation and at least the second output of the second loop filter operation of the plurality of loop filter operations, the one or more processors are further configured to: truncate the third output of the bilateral filter operation to generate a truncated output of the bilateral filter operation; add the truncated output of the bilateral filter operation to the first output of the CCSAO filter operation and the corresponding outputs of the SAO filter operation and the deblocking filter operation to generate a first sum; and perform the joint truncation operation on the first sum.
[0221] Aspect 20C - The apparatus according to aspect 16C, wherein, in order to perform the joint truncation operation on the first output of the CCSAO filter operation and at least the second output of the second loop filter operation of the plurality of loop filter operations, the one or more processors are further configured to: truncate the third output of the bilateral filter operation to generate a first truncated output; add the corresponding outputs of the SAO filter operation and the deblocking filter operation to create a first sum; truncate the first sum to form a second truncated output; add the first truncated output, the second truncated output, and the first output of the CCSAO filter operation to create a second sum; and perform the joint truncation operation on the second sum.
[0222] Aspect 21C - An apparatus according to any one of Aspects 12C-20C, wherein the apparatus is a video encoder, and wherein, in order to reconstruct the video data to generate the reconstructed video data, the one or more processors are further configured to: reconstruct the video data in a reconstruction loop of the video encoder to generate the reconstructed video data.
[0223] Aspect 22C - An apparatus according to any one of Aspects 12C-20C, wherein the apparatus is a video decoder, and wherein, in order to reconstruct the video data to generate the reconstructed video data, the one or more processors are further configured to: decode the video data to generate the reconstructed video data.
[0224] It should be recognized that, based on the examples, certain actions or events of any technique described herein may be performed in a different order, and may be added, combined, or omitted entirely (e.g., not all described actions or events are necessary for implementing the technique). Furthermore, in some examples, actions or events may be performed concurrently rather than sequentially, for example, through multithreaded processing, interrupt handling, or multiple processors.
[0225] In one or more examples, the described functionality can be implemented using hardware, software, firmware, or any combination thereof. If implemented in software, the functionality can be stored on or transmitted through a computer-readable medium as one or more instructions or code and executed by a hardware-based processing unit. A computer-readable medium can include a computer-readable storage medium, which corresponds to a tangible medium such as a data storage medium or a communication medium, including, for example, any medium that facilitates the transfer of a computer program from one place to another according to a communication protocol. In this way, a computer-readable medium can generally correspond to (1) a non-transitory tangible computer-readable storage medium or (2) a communication medium such as a signal or carrier wave. A data storage medium can be any available medium that can be accessed by one or more computers or one or more processors to retrieve instructions, code, and / or data structures for implementing the techniques described in this disclosure. Computer program products can include computer-readable media.
[0226] For example, rather than limiting, such computer-readable storage media may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, disk storage or other magnetic storage devices, flash memory, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible by a computer. Furthermore, any connection is appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies (such as infrared, radio, and microwave), then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies (such as infrared, radio, and microwave) are included in the definition of medium. However, it should be understood that computer-readable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but rather refer instead to non-transient tangible storage media. As used herein, disks and optical discs include compact optical discs (CDs), laser optical discs, optical discs, digital versatile optical discs (DVDs), floppy disks, and Blu-ray discs, wherein disks typically magnetically copy data, while optical discs utilize lasers to optically copy data. Combinations of the above items should also be included within the scope of computer-readable media.
[0227] Instructions can be executed by one or more processors, such as one or more DSPs, general-purpose microprocessors, ASICs, FPGAs, or other equivalent integrated or discrete logic circuits. Therefore, as used herein, the terms "processor" and "processing circuitry" can refer to any of the foregoing structures or any other structure suitable for implementing the techniques described herein. Additionally, in some aspects, the functionality described herein can be provided within dedicated hardware and / or software modules configured for encoding and decoding, or incorporated into combined codecs. Furthermore, the techniques can be implemented entirely within one or more circuit or logic elements.
[0228] The techniques disclosed herein can be implemented in a wide variety of devices or apparatuses, including wireless mobile phones, integrated circuits (ICs), or a set of ICs (e.g., chipsets). Various components, modules, or units are described in this disclosure to emphasize functional aspects of a device configured to perform the disclosed techniques, but are not necessarily required to be implemented through different hardware units. Specifically, as described above, various units may be combined in a codec hardware unit, or provided by a collection of interoperable hardware units (including one or more processors as described above) combined with appropriate software and / or firmware.
[0229] Various examples have been described. These and other examples are within the scope of the appended claims.
Claims
1. A method for decoding video data, the method comprising: Reconstruct video data to generate reconstructed video data; Multiple loop filter operations are performed in parallel on the reconstructed video data, wherein the multiple loop filter operations include a first loop filter operation, a second loop filter operation, and a third loop filter operation, wherein the first loop filter operation is a cross-component sample adaptive offset (CCSAO) filter operation, the second loop filter operation is a bilateral filter operation, and the third loop filter operation is a sample adaptive offset (SAO) filter operation. The corresponding outputs of the CCSAO filter operation, the SAO filter operation, the bilateral filter operation, and the deblocking filter operation are added together to generate a first sum; and Perform a joint interception operation on the first sum.
2. The method according to claim 1, further comprising: The deblocking filter operation is performed on the reconstructed video data before the execution of the plurality of loop filter operations.
3. The method according to claim 1, further comprising: After performing the multiple loop filter operations, an adaptive loop filter operation is performed.
4. The method according to claim 1, wherein, Decoding includes encoding, and wherein reconstructing the video data to generate the reconstructed video data includes: The video data is reconstructed in the reconstruction loop of the video encoder to generate the reconstructed video data.
5. The method according to claim 1, wherein, Decoding includes decoding, and wherein reconstructing the video data to generate the reconstructed video data includes: The video data is decoded to generate the reconstructed video data.
6. An apparatus configured to decode video data, the apparatus comprising: The memory is configured to store video data; as well as One or more processors communicating with the memory, the one or more processors being configured to: The video data is reconstructed to generate reconstructed video data; Multiple loop filter operations are performed in parallel on the reconstructed video data, wherein the multiple loop filter operations include a first loop filter operation, a second loop filter operation, and a third loop filter operation, wherein the first loop filter operation is a cross-component sample adaptive offset (CCSAO) filter operation, the second loop filter operation is a bilateral filter operation, and the third loop filter operation is a sample adaptive offset (SAO) filter operation. The corresponding outputs of the CCSAO filter operation, the SAO filter operation, the bilateral filter operation, and the deblocking filter operation are added together to generate a first sum; and Perform a joint interception operation on the first sum.
7. The apparatus according to claim 6, wherein, The one or more processors are further configured to: The deblocking filter operation is performed on the reconstructed video data before the plurality of loop filter operations are executed.
8. The apparatus according to claim 6, wherein, The one or more processors are further configured to: After performing the multiple loop filter operations, an adaptive loop filter operation is performed.
9. The apparatus according to claim 6, wherein, The apparatus is a video encoder, and wherein, in order to reconstruct the video data to generate the reconstructed video data, the one or more processors are further configured to: The video data is reconstructed in the reconstruction loop of the video encoder to generate the reconstructed video data.
10. The apparatus according to claim 6, wherein, The apparatus is a video decoder, and wherein, in order to reconstruct the video data to generate the reconstructed video data, the one or more processors are further configured to: The video data is decoded to generate the reconstructed video data.
11. The apparatus according to claim 6, wherein, The device is a wireless communication device.
12. A non-transitory computer-readable storage medium storing instructions, which, when executed, cause one or more processors of a device configured to decode video data to perform the following operations: The video data is reconstructed to generate reconstructed video data; Multiple loop filter operations are performed in parallel on the reconstructed video data, wherein... The plurality of loop filter operations include a first loop filter operation, a second loop filter operation, and a third loop filter operation, wherein the first loop filter operation is a cross-component sample adaptive offset (CCSAO) filter operation, the second loop filter operation is a bilateral filter operation, and the third loop filter operation is a sample adaptive offset (SAO) filter operation. The corresponding outputs of the CCSAO filter operation, the SAO filter operation, the bilateral filter operation, and the deblocking filter operation are added together to generate a first sum; and Perform a joint interception operation on the first sum.
Citation Information
Patent Citations
Nonlinear extensions of adaptive loop filtering for video coding
US20200404335A1