System and method for dual query transformer-context adaptive neural network in-loop filter
The DQT-CALF network addresses the limitations of existing video coding standards by learning dual features and using spatial attention to enhance video quality and reduce artifacts, achieving superior coding efficiency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
- Filing Date
- 2024-10-23
- Publication Date
- 2026-04-30
AI Technical Summary
Existing video coding standards like VVC suffer from limitations in effectively handling complex and rich textures of video content, leading to undesirable compression artifacts, particularly at high compression rates, despite advancements in neural network-based in-loop filters.
A dual-query transformer-content adaptive neural network (DQT-CALF) is introduced, which learns low-frequency global and high-frequency local features through a dual-query mechanism, utilizing spatial attention mechanisms and auxiliary inputs to generate rich feature representations, enhancing the reconstruction of video content.
DQT-CALF significantly reduces compression artifacts and improves video quality by effectively integrating diverse information types, outperforming traditional methods in coding efficiency and artifact reduction.
Smart Images

Figure CN2024126859_30042026_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR DUAL QUERY TRANSFORMER-CONTEXT ADAPTIVE NEURAL NETWORK IN-LOOP FILTERBACKGROUND
[0001] Embodiments of the present disclosure relate to video coding.
[0002] Digital video has become mainstream and is being used in a wide range of applications including digital television, video telephony, and teleconferencing. These digital video applications are feasible because of the advances in computing and communication technologies as well as efficient video coding techniques. Various video coding techniques may be used to compress video data, such that coding on the video data can be performed using one or more video coding standards. Exemplary video coding standards may include, but not limited to, versatile video coding (H. 266 / VVC) , high-efficiency video coding (H. 265 / HEVC) , advanced video coding (H. 264 / AVC) , moving picture expert group (MPEG) coding, to name a few.SUMMARY
[0003] According to one aspect of the present disclosure, a method of video decoding is provided. The method may include obtaining, by a processor, a first reference picture from a decoded picture buffer. The method may include generating, by the processor, a first feature map based on first information associated with the first reference picture. The first information may include a reconstructed picture. The method may include generating, by the processor, a spatial attention map based on second information associated with the first reference picture. The method may include generating, by the processor, a second feature map based on the first feature map and the spatial attention map by performing at least one multi-type feature fusion block (MFFB) operation. The method may include generating, by the processor, a second reference picture based on the reconstructed picture and the second feature map. The method may include decoding, by the processor, a current picture region based on the second reference picture.
[0004] According to another aspect of the present disclosure, a video decoder is provided. The video decoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to obtain a first reference picture from a decoded picture buffer. The memory storing instructions, which when executed by the processor, may cause the processor to generate a first feature map based on first information associated with the first reference picture. The first information may include a reconstructed picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a spatial attention map based on second information associated with the first reference picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second reference picture based on the reconstructed picture and the second feature map. The memory storing instructions, which when executed by the processor, may cause the processor to decode a current picture region based on the second reference picture.
[0005] According to a further aspect of the present disclosure, an apparatus for video decoding is provided. The video decoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to obtain a first reference picture from a decoded picture buffer. The memory storing instructions, which when executed by the processor, may cause the processor to generate a first feature map based on first information associated with the first reference picture. The first information may include a reconstructed picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a spatial attention map based on second information associated with the first reference picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second reference picture based on the reconstructed picture and the second feature map. The memory storing instructions, which when executed by the processor, may cause the processor to decode a current picture region based on the second reference picture.
[0006] According to yet another aspect of the present disclosure, a non-transitory computer-readable medium storing instructions for a video decoder is provided. The instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to obtain a first reference picture from a decoded picture buffer. The instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to generate a first feature map based on first information associated with the first reference picture. The first information may include a reconstructed picture. The instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to generate a spatial attention map based on second information associated with the first reference picture. The instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to generate a second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation. The instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to generate a second reference picture based on the reconstructed picture and the second feature map. The instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to decode a current picture region based on the second reference picture.
[0007] According to one aspect of the present disclosure, a method of video encoding is provided. The method may include obtaining, by a processor, a first reference picture from a decoded picture buffer. The method may include generating, by the processor, a first feature map based on first information associated with the first reference picture. The first information may include a reconstructed picture. The method may include generating, by the processor, a spatial attention map based on second information associated with the first reference picture. The method may include generating, by the processor, a second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation. The method may include generating, by the processor, a second reference picture based on the reconstructed picture and the second feature map. The method may include encoding, by the processor, a current picture region based on the second reference picture.
[0008] According to another aspect of the present disclosure, a video encoder is provided. The video encoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to obtain a first reference picture from a decoded picture buffer. The memory storing instructions, which when executed by the processor, may cause the processor to generate a first feature map based on first information associated with the first reference picture. The first information may include a reconstructed picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a spatial attention map based on second information associated with the first reference picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second reference picture based on the reconstructed picture and the second feature map. The memory storing instructions, which when executed by the processor, may cause the processor to encode a current picture region based on the second reference picture.
[0009] According to a further aspect of the present disclosure, an apparatus for video encoding is provided. The video encoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to obtain a first reference picture from a decoded picture buffer. The memory storing instructions, which when executed by the processor, may cause the processor to generate a first feature map based on first information associated with the first reference picture. The first information may include a reconstructed picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a spatial attention map based on second information associated with the first reference picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second reference picture based on the reconstructed picture and the second feature map. The memory storing instructions, which when executed by the processor, may cause the processor to encode a current picture region based on the second reference picture.
[0010] According to yet another aspect of the present disclosure, a non-transitory computer-readable medium storing instructions for a video encoder is provided. The instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to obtain a first reference picture from a decoded picture buffer. The instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to generate a first feature map based on first information associated with the first reference picture. The first information may include a reconstructed picture. The instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to generate a spatial attention map based on second information associated with the first reference picture. The instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to generate a second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation. The instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to generate a second reference picture based on the reconstructed picture and the second feature map. The instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to encode a current picture region based on the second reference picture.
[0011] According to yet a further aspect of the present disclosure, a non-transitory computer-readable medium storing a bitstream is provided. The bitstream may be generated based on one or more of the operations described herein.
[0012] These illustrative embodiments are mentioned not to limit or define the present disclosure, but to provide examples to aid understanding thereof. Additional embodiments are described in the Detailed Description, and further description is provided there.BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate embodiments of the present disclosure and, together with the description, further serve to explain the principles of the present disclosure and to enable a person skilled in the pertinent art to make and use the present disclosure.
[0014] FIG. 1 illustrates a block diagram of an exemplary video codec system, according to some embodiments of the present disclosure.
[0015] FIG. 2A illustrates a block diagram of an exemplary encoding apparatus, according to some embodiments of the present disclosure.
[0016] FIG. 2B illustrates a block diagram of an exemplary decoding apparatus, according to some embodiments of the present disclosure.
[0017] FIG. 2C illustrates a detailed block diagram of an exemplary video coder that includes a DQT-CALF network, according to some embodiments of the present disclosure.
[0018] FIG. 3A illustrates a detailed block diagram of an exemplary dual query transformer (DQT) -Content Adaptive Neural Network-Based In-Loop Filter (CALF) network for luma components, according to some embodiments of the present disclosure.
[0019] FIG. 3B illustrates a detailed block diagram of an exemplary DQT-CALF network for chroma components, according to some embodiments of the present disclosure.
[0020] FIG. 4 illustrates a detailed block diagram of a Multi-Type Feature Fusion Block (MFFB) of the DQT-CALF network depicted FIGs. 3A and 3B, according to some embodiments of the present disclosure.
[0021] FIG. 5A illustrates a detailed block diagram of a high frequency-local feature generation (HLFG) block of the MFFB depicted in FIG. 4, according to some embodiments of the present disclosure.
[0022] FIG. 5B illustrates a detailed block diagram of a low frequency-global feature generation (LGFG) block of the MFFB depicted in FIG. 4, according to some embodiments of the present disclosure.
[0023] FIG. 5C illustrates a detailed block diagram of a DQT of the MFFB depicted in FIG. 4, according to some embodiments of the present disclosure.
[0024] FIG. 6 illustrates an exemplary progressive learning strategy for a DQT-CALF network based on quantization parameter (QP) distance, according to some embodiments of the present disclosure.
[0025] FIG. 7A illustrates exemplary test results of a first stage of the exemplary progressive learning strategy depicted in FIG. 6 implemented with the AI configuration, according to some embodiments of the present disclosure.
[0026] FIG. 7B illustrates exemplary test results of a second stage of the exemplary progressive learning strategy depicted in FIG. 6 implemented with the AI configuration, according to some embodiments of the present disclosure.
[0027] FIG. 7C illustrates exemplary test results of a third stage of the exemplary progressive learning strategy depicted in FIG. 6 implemented with the AI configuration, according to some embodiments of the present disclosure.
[0028] FIG. 7D illustrates exemplary final test results of a fourth stage of the exemplary progressive learning strategy depicted in FIG. 6 implemented with the AI configuration, according to some embodiments of the present disclosure.
[0029] FIG. 7E illustrates exemplary test results of the DQT-CALF network obtained under a random access (RA) configuration, according to some embodiments of the present disclosure.
[0030] FIG. 7F illustrates an overall BD-rate comparison with latest in-loop filters in JVET under AI configuration, according to some embodiments of the present disclosure.
[0031] FIG. 7G illustrates the results of an ablation study on the spatial attention (SA) in DQT-CALF, according to some embodiments of the present disclosure.
[0032] FIG. 7H illustrates the results of an ablation study on MFFB in DQT-CALF, according to some embodiments of the present disclosure.
[0033] FIG. 7I illustrates the results of an ablation study on the effect of the reconstructed luma component (Rec_Y) on the chroma component in the chroma model, according to some embodiments of the present disclosure.
[0034] FIG. 8A is a first table illustrating average rate distortion (RD) results on all test sequences under all intra (AI) configuration in terms of PSNR for a first dataset, according to some embodiments of the present disclosure.
[0035] FIG. 8B is a second table illustrating second average RD results on all test sequences under AI configuration in terms of PSNR for a second dataset, according to some embodiments of the present disclosure.
[0036] FIG. 8C is a third table illustrating third average RD results on all test sequences under AI configuration in terms of PSNR for a third dataset, according to some embodiments of the present disclosure.
[0037] FIG. 8D is a fourth table illustrating fourth average RD results on all test sequences under AI configuration in terms of PSNR for a fourth dataset, according to some embodiments of the present disclosure.
[0038] FIG. 8E is a fifth table illustrating fifth average RD results on all test sequences under AI configuration in terms of PSNR for a fifth dataset, according to some embodiments of the present disclosure.
[0039] FIG. 8F is a sixth table illustrating sixth average RD results on all test sequences under AI configuration in terms of PSNR for a sixth dataset, according to some embodiments of the present disclosure.
[0040] FIG. 9 illustrates a flow chart of an exemplary method of video decoding, according to some embodiments of the present disclosure.
[0041] FIG. 10 illustrates a flow chart of an exemplary method of video encoding, according to some embodiments of the present disclosure.
[0042] Embodiments of the present disclosure will be described with reference to the accompanying drawings.DETAILED DESCRIPTION
[0043] Although some configurations and arrangements are discussed, it should be understood that this is done for illustrative purposes only. A person skilled in the pertinent art will recognize that other configurations and arrangements can be used without departing from the spirit and scope of the present disclosure. It will be apparent to a person skilled in the pertinent art that the present disclosure can also be employed in a variety of other applications.
[0044] It is noted that references in the specification to “one embodiment, ” “an embodiment, ” “an example embodiment, ” “some embodiments, ” “certain embodiments, ” etc., indicate that the embodiment described may include a particular feature, structure, or characteristic, but every embodiment may not necessarily include the particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. Further, when a particular feature, structure, or characteristic is described in connection with an embodiment, it would be within the knowledge of a person skilled in the pertinent art to affect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
[0045] In general, terminology may be understood at least in part from usage in context. For example, the term “one or more” as used herein, depending at least in part upon context, may be used to describe any feature, structure, or characteristic in a singular sense or may be used to describe combinations of features, structures or characteristics in a plural sense. Similarly, terms, such as “a, ” “an, ” or “the, ” again, may be understood to convey a singular usage or to convey a plural usage, depending at least in part upon context. In addition, the term “based on” may be understood as not necessarily intended to convey an exclusive set of factors and may, instead, allow for existence of additional factors not necessarily expressly described, again, depending at least in part on context.
[0046] Various aspects of video coding systems will now be described with reference to various apparatus and methods. These apparatus and methods will be described in the following detailed description and illustrated in the accompanying drawings by various modules, components, circuits, steps, operations, processes, algorithms, etc. (collectively referred to as “elements” ) . These elements may be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether such elements are implemented as hardware, firmware, or software depends upon the particular application and design constraints imposed on the overall system.
[0047] The techniques described herein may be used for various video coding applications. As described herein, video coding includes both encoding and decoding a video. Encoding and decoding of a video can be performed by the unit of block. For example, an encoding / decoding process such as transform, quantization, prediction, in-loop filtering, reconstruction, or the like may be performed on a coding block, a transform block, or a prediction block. As described herein, a block to be encoded / decoded will be referred to as a “current block. ” For example, the current block may represent a coding block, a transform block, or a prediction block according to a current encoding / decoding process. In addition, it is understood that the term “unit” used in the present disclosure indicates a basic unit for performing a specific encoding / decoding process, and the term “block” indicates a sample array of a predetermined size. Unless otherwise stated, the “block, ” “unit, ” and “component” may be used interchangeably.
[0048] Video coding aims at effectively compressing video data while maintaining video quality to meet various transmission and storage needs. From the early MPEG-1 and MPEG-2 standards to the widely used H. 264 / AVC and H. 265 / HEVC, video coding has achieved tremendous advancements in recent decades. H. 264 / AVC, with its efficient compression capabilities and broad compatibility, became the mainstream video coding standard for many streaming services and blu-ray discs. H. 265 / HEVC further enhanced compression efficiency although its hardware support and decoding complexity are relatively higher. With the proliferation of 5G networks and the increasing demand for ultra-high-definition (UHD) video content, new video encoding standards such as the versatile video coding (VVC) and AV1 are being developed and promoted. These standards aim to provide higher compression efficiency and better video quality, while supporting a wider range of applications, including virtual reality (VR) , augmented reality (AR) , and 8K videos. With the integration of artificial intelligence technology, intelligent encoding techniques are gradually developing, enabling dynamic adjustment of encoding strategies based on the complexity of the video content to achieve optimal compression efficiency. Overall, video encoding technology is progressing towards higher efficiency, greater intelligence, and better adaptation to future digital media needs. Since VVC is a block-based hybrid coding framework, it inevitably produces undesirable compression artifacts, especially at a high compression rate. To address this issue, VVC employs a series of advanced filtering techniques to eliminate or reduce compression artifacts. The in-loop filters include an Adaptive Loop Filter (ALF) , a Deblocking Filter (DBF) , and a Sample Adaptive Offset (SAO) , and optimize the video signal during both encoding and decoding processes. ALF dynamically adjusts the filter strength based on the local characteristics of the video content, thereby reducing blocking artifacts while maintaining edge sharpness. DBF improves overall video quality by applying filters at the boundaries of coding blocks to reduce visual discontinuities. SAO reduces high-frequency noise and artifacts by adjusting pixel values, particularly addressing detail loss caused by quantization during the encoding process. In VVC, the handcrafted in-loop filters often suffer from handling the complexity and rich textures of video content. In contrast, neural network-based in-loop filters (NNLFs) demonstrate outstanding capabilities of recovering video content. These advanced filters not only excel in significantly reducing compression artifacts but also remarkably enhancing video quality via learning.
[0049] A convolutional neural network (CNN) in-loop filter based on quantization parameters (QP) has also been proposed. The QP Attention Module (QPAM) effectively adapted to different QP settings, thus avoiding the need to train and deploy multiple models for different networks. Another in-loop filter, referred to as an “attention-based dual-scale CNN” (ADCNN) , focuses on reducing compression artifacts in I-frames and B-frames. ADCNN enhanced the quality of video encoding by fully utilizing information priors such as QP and partition information. ADCNN effectively reduced compression artifacts by combining attention mechanisms and dual-scale processing. Some neural network-based in-loop filters are based on residual block (ResBlock) and Transformer in VVC, referred to as a “residual transformer neural network” (RTNN) . RTNN uses ResBlocks with attention to extract shallow features, and then utilizes the Transformer blocks to extract and process deep features. At a recent JVET conference, two unified NNLFs were released: the low complexity operating point (LOP) and the high performance operating point (HOP) . To further enhance the performance of HOP, a three-stage training strategy was also proposed at the conference. Subsequently, researchers conducted a series of experimental studies on the network architecture and hyper parameters to enhance coding efficiency while reducing complexity. However, further advancements in in-loop filters are still needed.
[0050] To overcome the limitations of existing in-loop filtering techniques, the present disclosure provides a content-adaptive neural network loop filter (CALF) based on a dual-query transformer (DQT) , referred to herein as a “DQT-CALF” for VVC. DQT learns low-frequency global features and high-frequency local features through a dual-query mechanism, effectively integrating various types of information to generate rich feature representations. To remove compression artifacts in the reconstructed pictures, DQT-CALF utilizes auxiliary inputs such as the predicted picture, partition map, and QP map. Unlike traditional methods that simply concatenate auxiliary information with the reconstructed picture, DQT-CALF employs a spatial attention mechanism to generate spatial attention maps from the auxiliary inputs, enhancing the reconstructed picture. In the network backbone, the input features are decomposed into four types: low-frequency features, high-frequency features, local features, and global features. Low-frequency features correspond to smooth areas in the image with slow changes, while high-frequency features correspond to rapidly changing areas, such as textures and edges. Local features represent structure, textures, and shapes, while global features represent the overall characteristics of the image. The combination of low-frequency and global features into low-frequency global information helps represent the overall scene, while the combination of high-frequency and local features into high-frequency local information significantly enhances the network's ability to capture details and local variations. Each feature set is extracted through residual blocks and processed through modules that generate high-frequency local features and low-frequency global auxiliary information, with final fusion achieved through DQT, convolution, and concatenation operations.
[0051] To train DQT-CALF, we design a fourth-stage progressive learning strategy based on QP distance. The training sets for four stages are obtained by the original VTM compression with QP settings of 7, 12, 17, 22, 27, 32, 37, and 42. When setting the QP distance to 5, the input QP values for the training set are set to {22, 27, 32, 37, 42} , while the corresponding label QP values are set to {17, 22, 27, 32, 37} . Subsequently, in the next training stage, the model from the previous stage is loaded, and the QP distance for the training set is increased. In the first three stages, the QP distance increases by 5, while in the final stage, the label QP values are set to 7.
[0052] FIG. 1 is a block diagram of a video codec system, according to some embodiments of the present disclosure. The video codec system, according to an embodiment, may include an encoding apparatus 10 and a decoding apparatus 20. The encoding apparatus 10 may deliver encoded video and / or picture information or data to the decoding apparatus 20 in the form of a file or streaming via a digital storage medium or network.
[0053] The encoding apparatus 10, according to an embodiment, may include a video source generator 11, an encoding unit 12, and a transmitter 13. The decoding apparatus 20, according to an embodiment, may include a receiver 21, a decoding unit 22, and a renderer 23. The encoding unit 12 may be called a video / picture encoding unit, and the decoding unit 22 may be called a video / picture decoding unit. The transmitter 13 may be included in the encoding unit 12. The receiver 21 may be included in the decoding unit 22. The renderer 23 may include a display, and the display may be configured as a separate device or an external component.
[0054] The video source generator 11 may acquire a video / picture through a process of capturing, synthesizing, or generating the video / picture. The video source generator 11 may include a video / picture capture device and / or a video / picture generating device. The video / picture capture device may include, for example, one or more cameras, video / picture archives including previously captured video / pictures, and the like. The video / picture-generating device may include, for example, computers, tablets, and smartphones, and may (electronically) generate video / pictures. For example, a virtual video / picture may be generated through a computer or the like. In this case, the video / picture capturing process may be replaced by a process of generating related data.
[0055] The encoding unit 12 may encode an input video / picture. The encoding unit 12 may perform a series of procedures such as prediction, transform, and quantization for compression and coding efficiency. The encoding unit 12 may output encoded data (encoded video / picture information) in the form of a bitstream.
[0056] The transmitter 13 may transmit the encoded video / picture information or data output in the form of a bitstream to the receiver 21 of the decoding apparatus 20 through a digital storage medium or a network in the form of a file or streaming. The digital storage medium may include various storage mediums such as universal serial bus (USB) , secure digital (SD) , compact disc (CD) , digital video disc (DVD) , Blu-ray, hard disk drive (HDD) , solid-state drive (SSD) , and the like. The transmitter 13 may include an element for generating a media file through a predetermined file format and may include an element for transmission through a broadcast / communication network. The receiver 21 may extract / receive the bitstream from the storage medium or network and transmit the bitstream to the decoding unit 22.
[0057] The decoding unit 22 may decode the video / picture by performing a series of procedures such as dequantization, inverse transform, and prediction corresponding to the operation of the encoding unit 12.
[0058] The renderer 23 may render the decoded video / picture. The rendered video / picture may be displayed through the display.
[0059] FIG. 2A is a schematic block diagram of an encoding apparatus, in accordance with some aspects of the present disclosure. Referring to FIG. 2A, the encoding apparatus 200 includes a picture partitioner 210, a predictor 220, a residual processor 230, an entropy encoder 240, an adder 251, a filter 261, and a memory 271. The predictor 220 may include an inter predictor 221 and an intra predictor 222. The residual processor 230 may include a transformer 232, a quantizer 233, a dequantizer 234, and an inverse transformer 235. The residual processor 230 may further include a subtractor 231. The adder 251 may be called a reconstructor or a reconstructed block generator. The picture partitioner 210, the predictor 220, the residual processor 230, the entropy encoder 240, the adder 251, and the filter 261 may be configured by at least one hardware component (e.g., an encoder chipset or processor) , according to an embodiment. In addition, the memory 271 may include a decoded picture buffer (DPB) or may be configured by a digital storage medium. The hardware component may further include the memory 271 as an internal / external component.
[0060] The picture partitioner 210 may partition an input picture (or a picture or a frame) input to the encoding apparatus 200 into one or more processors. For example, the processor may be called a coding unit (CU) . In this case, the coding unit may be recursively partitioned according to a quad-tree binary-tree ternary-tree (QTBTTT) structure from a coding tree unit (CTU) or a largest coding unit (LCU) . For example, one coding unit may be partitioned into a plurality of coding units of a deeper depth based on a quad tree structure, a binary tree structure, and / or a ternary structure. In this case, for example, the quad tree structure may be applied first, and the binary tree structure and / or ternary structure may be applied later. Alternatively, the binary tree structure may be applied first. The coding procedure according to this invention may be performed based on the final coding unit that is no longer partitioned. In this case, the largest coding unit may be used as the final coding unit based on coding efficiency according to picture characteristics, or if necessary, the coding unit may be recursively partitioned into coding units of deeper depth, and a coding unit having an optimal size may be used as the final coding unit. Here, the coding procedure may include a procedure of prediction, transform, and reconstruction, which will be described later. As another example, the processor may further include a prediction unit (PU) or a transform unit (TU) . In this case, the prediction unit and the transform unit may be split or partitioned from the aforementioned final coding unit. The prediction unit may be a unit of sample prediction, and the transform unit may be a unit for deriving a transform coefficient and / or a unit for deriving a residual signal from the transform coefficient.
[0061] The unit may be used interchangeably with terms such as block or area in some cases. In a general case, an M×N block may represent a set of samples or transform coefficients composed of M columns and N rows. A sample may generally represent a pixel or a value of a pixel, may represent only a pixel / pixel value of a luma component or represent only a pixel / pixel value of a chroma component. A sample may be used as a term corresponding to one picture (or picture) for a pixel or a pel.
[0062] In the encoding apparatus 200, a prediction signal (predicted block, prediction sample array) output from the inter predictor 221 or the intra predictor 222 is subtracted from an input picture signal (original block, original sample array) to generate a residual signal residual block, residual sample array) , and the generated residual signal is transmitted to the transformer 232. In this case, as shown, a unit for subtracting a prediction signal (predicted block, prediction sample array) from the input picture signal (original block, original sample array) in the encoding apparatus 200 may be called a subtractor 231. The predictor may perform prediction on a block to be processed (hereinafter, referred to as a current block) and generate a predicted block including prediction samples for the current block. The predictor may determine whether intra prediction or inter prediction is applied on a current block or CU basis. As described later in the description of each prediction mode, the predictor may generate various information related to prediction, such as prediction mode information, and transmit the generated information to the entropy encoder 240. The information on the prediction may be encoded in the entropy encoder 240 and output in the form of a bitstream.
[0063] The intra predictor 222 may predict the current block by referring to the samples in the current picture. The referred samples may be located in the neighborhood of the current block or may be located apart according to the prediction mode. In the intra prediction, prediction modes may include a plurality of non-directional modes and a plurality of directional modes. The non-directional mode may include, for example, a DC mode and a planar mode. The directional mode may include, for example, 33 directional prediction modes or 65 directional prediction modes according to the degree of detail of the prediction direction. However, this is merely an example, and more or less directional prediction modes may be used depending on the setting. The intra predictor 222 may determine the prediction mode applied to the current block by using a prediction mode applied to a neighboring block.
[0064] The inter predictor 221 may derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. Here, in order to reduce the amount of motion information transmitted in the inter prediction mode, the motion information may be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between the neighboring block and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc. ) information. In the case of inter prediction, the neighboring block may include a spatial neighboring block present in the current picture and a temporal neighboring block present in the reference picture. The reference picture including the reference block and the reference picture including the temporal neighboring block may be the same or different. The temporal neighboring block may be called a collocated reference block, a co-located CU (colCU) , and the like, and the reference picture including the temporal neighboring block may be called a collocated picture (colPic) . For example, the inter predictor 221 may configure a motion information candidate list based on neighboring blocks and generate information indicating which candidate is used to derive a motion vector and / or a reference picture index of the current block. Inter prediction may be performed based on various prediction modes. For example, in the case of a skip mode and a merge mode, the inter predictor 221 may use motion information of the neighboring block as motion information of the current block. In the skip mode, unlike the merge mode, the residual signal may not be transmitted. In the case of the motion vector prediction (MVP) mode, the motion vector of the neighboring block may be used as a motion vector predictor, and the motion vector of the current block may be indicated by signaling a motion vector difference.
[0065] The predictor 220 may generate a prediction signal based on various prediction methods described below. For example, the predictor may not only apply intra prediction or inter prediction to predict one block but also simultaneously apply both intra prediction and inter prediction. This may be called combined inter and intra prediction (CIIP) . In addition, the predictor may be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content picture / video coding of a game or the like, for example, screen content coding (SCC) . The IBC basically performs prediction in the current picture but may be performed similarly to inter prediction in that a reference block is derived in the current picture. That is, the IBC may use at least one of the inter prediction techniques described in the present disclosure. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, a sample value within a picture may be signaled based on information on the palette table and the palette index.
[0066] The prediction signal generated by the predictor (including the inter predictor 221 and / or the intra predictor 222) may be used to generate a reconstructed signal or to generate a residual signal. The transformer 232 may generate transform coefficients by applying a transform technique to the residual signal. For example, the transform technique may include at least one of a discrete cosine transform (DCT) , a discrete sine transform (DST) , a karhunen-loève transform (KLT) , a graph-based transform (GBT) , or a conditionally non-linear transform (CNT) . Here, the GBT means transform obtained from a graph when relationship information between pixels is represented by the graph. The CNT refers to the transform generated based on a prediction signal generated using all previously reconstructed pixels. In addition, the transform process may be applied to square pixel blocks having the same size or may be applied to blocks having a variable size rather than a square.
[0067] The quantizer 233 may quantize the transform coefficients and transmit them to the entropy encoder 240, and the entropy encoder 240 may encode the quantized signal (information on the quantized transform coefficients) and output a bitstream. The information on the quantized transform coefficients may be referred to as residual information. The quantizer 233 may rearrange block type quantized transform coefficients into a one-dimensional vector form based on a coefficient scanning order and generate information on the quantized transform coefficients based on the quantized transform coefficients in the one-dimensional vector form. Information on transform coefficients may be generated. The entropy encoder 240 may perform various encoding methods such as, for example, exponential Golomb, context-adaptive variable length coding (CAVLC) , context-adaptive binary arithmetic coding (CABAC) , and the like. The entropy encoder 240 may encode information necessary for video / picture reconstruction other than quantized transform coefficients (e.g., values of syntax elements, etc. ) together or separately. Encoded information (e.g., encoded video / picture information) may be transmitted or stored in units of NALs (network abstraction layer) in the form of a bitstream. The video / picture information may further include information on various parameter sets, such as an adaptation parameter set (APS) , a picture parameter set (PPS) , a sequence parameter set (SPS) , or a video parameter set (VPS) . In addition, the video / picture information may further include general constraint information. In the present disclosure, information and / or syntax elements transmitted / signaled from the encoding apparatus to the decoding apparatus may be included in video / picture information. The video / picture information may be encoded through the above-described encoding procedure and included in the bitstream. The bitstream may be transmitted over a network or may be stored in a digital storage medium. The network may include a broadcasting network and / or a communication network, and the digital storage medium may include various storage media such as USB, SD, CD, DVD, Blu-ray, HDD, SSD, and the like. A transmitter (not shown) transmitting a signal output from the entropy encoder 240 and / or a storage unit (not shown) storing the signal may be included as internal / external element of the encoding apparatus 200, and alternatively, the transmitter may be included in the entropy encoder 240.
[0068] The quantized transform coefficients output from the quantizer 233 may be used to generate a prediction signal. For example, the residual signal (residual block or residual samples) may be reconstructed by applying dequantization and inverse transform to the quantized transform coefficients through the dequantizer 234 and the inverse transformer 235. The adder 251 adds the reconstructed residual signal to the prediction signal output from the inter predictor 221 or the intra predictor 222 to generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) . If there is no residual for the block to be processed, such as in a case where the skip mode is applied, the predicted block may be used as the reconstructed block. The adder 251 may be called a reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture and may be used for inter prediction of the next picture through filtering as described below.
[0069] Meanwhile, luma mapping with chroma scaling (LMCS) may be applied during picture encoding and / or reconstruction.
[0070] The filter 261 may improve subjective / objective picture quality by applying filtering to the reconstructed signal. For example, the filter 261 may generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 271, specifically, a DPB of the memory 271. The various filtering methods may include, for example, deblocking filtering, a sample adaptive offset, an adaptive loop filter, a bilateral filter, and the like. The filter 261 may generate various information related to the filtering and transmit the generated information to the entropy encoder 240, as described later in the description of each filtering method. The information related to the filtering may be encoded by the entropy encoder 240 and output in the form of a bitstream.
[0071] The modified reconstructed picture transmitted to the memory 271 may be used as the reference picture in the inter predictor 221. When the inter prediction is applied through the encoding apparatus, prediction mismatch between the encoding apparatus 200 and the decoding apparatus 250 may be avoided, and encoding efficiency may be improved.
[0072] The DPB of the memory 271 DPB may store the modified reconstructed picture for use as a reference picture in the inter predictor 221. The memory 271 may store the motion information of the block from which the motion information in the current picture is derived (or encoded) and / or the motion information of the blocks in the picture that have already been reconstructed. The stored motion information may be transmitted to the inter predictor 221 and used as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 271 may store reconstructed samples of reconstructed blocks in the current picture and may transfer the reconstructed samples to the intra predictor 222.
[0073] FIG. 2B is a schematic block diagram of a decoding, in accordance with some embodiments of the present disclosure. Referring to FIG. 2B, the decoding apparatus 250 may include an entropy decoder 270, a residual processor 252, a predictor 258, an adder 264, a filter 266, and a memory 268. The predictor 258 may include an inter predictor 260 and an intra predictor 262. The residual processor 252 may include a dequantizer 254 and an inverse transformer 256. The entropy decoder 270, the residual processor 252, the predictor 258, the adder 264, and the filter 266 may be configured by a hardware component (e.g., a decoder chipset or a processor) according to an embodiment. In addition, the memory 268 may include a decoded picture buffer (DPB) or may be configured by digital storage.
[0074] When a bitstream including video / picture information is input, the decoding apparatus 250 may reconstruct a picture corresponding to a process in which the video / picture information is processed in the encoding apparatus of FIG. 2A. For example, the decoding apparatus 250 may derive units / blocks based on block partition-related information obtained from the bitstream. The decoding apparatus 250 may perform decoding using a processor applied in the encoding apparatus. Thus, the processor of decoding may be a coding unit, for example, and the coding unit may be partitioned according to a quad tree structure, binary tree structure and / or ternary tree structure from the coding tree unit or the largest coding unit. One or more transform units may be derived from the coding unit. The input picture data signal decoded and output through the decoding apparatus 250 may be reproduced through a reproducing apparatus.
[0075] The decoding apparatus 250 may receive a signal output from the encoding apparatus of FIG. 2A in the form of a bitstream, and the received signal may be decoded through the entropy decoder 270. For example, the entropy decoder 270 may parse the bitstream to derive information (e.g., video / picture information) necessary for picture reconstruction (or picture reconstruction) . The video / picture information may further include information on various parameter sets, such as an adaptation parameter set (APS) , a picture parameter set (PPS) , a sequence parameter set (SPS) , or a video parameter set (VPS) . In addition, the video / picture information may further include general constraint information. The decoding apparatus may further decode picture based on the information on the parameter set and / or the general constraint information. Signaled / received information and / or syntax elements described later in the present disclosure may be decoded by the decoding procedure and obtained from the bitstream. For example, the entropy decoder 270 decodes the information in the bitstream based on a coding method such as exponential Golomb coding, CAVLC, or CABAC, and output syntax elements required for picture reconstruction and quantized values of transform coefficients for residual. More specifically, the CABAC entropy decoding method may receive a bin corresponding to each syntax element in the bitstream, determine a context model using a decoding target syntax element information, decoding information of a decoding target block or information of a symbol / bin decoded in a previous stage, and perform an arithmetic decoding on the bin by predicting a probability of occurrence of a bin according to the determined context model, and generate a symbol corresponding to the value of each syntax element. In this case, the CABAC entropy decoding method may update the context model by using the information of the decoded symbol / bin for a context model of the next symbol / bin after determining the context model. The information related to the prediction among the information decoded by the entropy decoder 270 may be provided to the predictor (the inter predictor 260 and the intra predictor 262) , and the residual value on which the entropy decoding was performed in the entropy decoder 270, that is, the quantized transform coefficients and related parameter information, may be input to the residual processor 252. The residual processor 252 may derive the residual signal (the residual block, the residual samples, and the residual sample array) . In addition, information on filtering among information decoded by the entropy decoder 270 may be provided to the filter 266. Meanwhile, a receiver (not shown) for receiving a signal output from the encoding apparatus may be further configured as an internal / external element of the decoding apparatus 250, or the receiver may be a component of the entropy decoder 270. Meanwhile, the decoding apparatus in the present disclosure may be referred to as a video / picture / picture decoding apparatus, and the decoding apparatus may be classified into an information decoder (video / picture / picture information decoder) and a sample decoder (video / picture / picture sample decoder) . The information decoder may include the entropy decoder 270, and the sample decoder may include at least one of the dequantizer 254, the inverse transformer 256, the adder 264, the filter 266, the memory 268, the inter predictor 260, and the intra predictor 262.
[0076] The dequantizer 254 may dequantize the quantized transform coefficients and output the transform coefficients. The dequantizer 254 may rearrange the quantized transform coefficients in the form of a two-dimensional block form. In this case, the rearrangement may be performed based on the coefficient scanning order performed in the encoding apparatus. The dequantizer 254 may perform dequantization on the quantized transform coefficients by using a quantization parameter (e.g., quantization step size information) and obtain transform coefficients.
[0077] The inverse transformer 256 inversely transforms the transform coefficients to obtain a residual signal (residual block, residual sample array) .
[0078] The predictor may perform prediction on the current block and generate a predicted block including prediction samples for the current block. The predictor may determine whether intra prediction or inter prediction is applied to the current block based on the information on the prediction output from the entropy decoder 270 and may determine a specific intra / inter prediction mode.
[0079] The predictor 258 may generate a prediction signal based on various prediction methods described below. For example, the predictor may not only apply intra prediction or inter prediction to predict one block but also simultaneously apply intra prediction and inter prediction. This may be called combined inter and intra prediction (CIIP) . In addition, the predictor may be based on an intra block copy (IBC) prediction mode or a palette mode for prediction of a block. The IBC prediction mode or palette mode may be used for content picture / video coding of a game or the like, for example, screen content coding (SCC) . The IBC basically performs prediction in the current picture but may be performed similarly to inter prediction in that a reference block is derived in the current picture. That is, the IBC may use at least one of the inter prediction techniques described in this document. The palette mode may be considered an example of intra coding or intra prediction. When the palette mode is applied, a sample value within a picture may be signaled based on information on the palette table and the palette index.
[0080] The intra predictor 262 may predict the current block by referring to the samples in the current picture. The referred samples may be located in the neighborhood of the current block or may be located apart according to the prediction mode. In the intra prediction, prediction modes may include a plurality of nondirectional modes and a plurality of directional modes. The intra predictor 262 may determine the prediction mode applied to the current block by using a prediction mode applied to a neighboring block.
[0081] The intra predictor 262 may predict the current block by referring to the samples in the current picture. The referenced samples may be located in the neighborhood of the current block or may be located apart according to the prediction mode. In intra prediction, prediction modes may include a plurality of nondirectional modes and a plurality of directional modes. The intra predictor 262 may determine the prediction mode applied to the current block by using the prediction mode applied to the neighboring block.
[0082] The inter predictor 260 may derive a predicted block for the current block based on a reference block (reference sample array) specified by a motion vector on a reference picture. In this case, in order to reduce the amount of motion information transmitted in the inter prediction mode, motion information may be predicted in units of blocks, subblocks, or samples based on the correlation of motion information between the neighboring block and the current block. The motion information may include a motion vector and a reference picture index. The motion information may further include inter prediction direction (L0 prediction, L1 prediction, Bi prediction, etc. ) information. In the case of inter prediction, the neighboring block may include a spatial neighboring block present in the current picture and a temporal neighboring block present in the reference picture. For example, the inter predictor 260 may configure a motion information candidate list based on neighboring blocks and derive a motion vector of the current block and / or a reference picture index based on the received candidate selection information. Inter prediction may be performed based on various prediction modes, and the information on the prediction may include information indicating a mode of inter prediction for the current block.
[0083] The adder 264 may generate a reconstructed signal (reconstructed picture, reconstructed block, reconstructed sample array) by adding the obtained residual signal to the prediction signal (predicted block, predicted sample array) output from the predictor (including the inter predictor 260 and / or the intra predictor 262) . If there is no residual for the block to be processed, such as when the skip mode is applied, the predicted block may be used as the reconstructed block.
[0084] The adder 264 may be called reconstructor or a reconstructed block generator. The generated reconstructed signal may be used for intra prediction of the next block to be processed in the current picture, may be output through filtering as described below, or may be used for inter prediction of the next picture.
[0085] Meanwhile, luma mapping with chroma scaling (LMCS) may be applied in the picture decoding process.
[0086] The filter 266 may improve subjective / objective picture quality by applying filtering to the reconstructed signal. For example, the filter 266 may generate a modified reconstructed picture by applying various filtering methods to the reconstructed picture and store the modified reconstructed picture in the memory 268, specifically, a DPB of the memory 268. The various filtering methods may include, for example, deblocking filtering, a sample adaptive offset, an adaptive loop filter, a bilateral filter, and the like.
[0087] The (modified) reconstructed picture stored in the DPB of the memory 268 may be used as a reference picture in the inter predictor 260. The memory 268 may store the motion information of the block from which the motion information in the current picture is derived (or decoded) and / or the motion information of the blocks in the picture that have already been reconstructed. The stored motion information may be transmitted to the inter predictor 221 so as to be utilized as the motion information of the spatial neighboring block or the motion information of the temporal neighboring block. The memory 268 may store reconstructed samples of reconstructed blocks in the current picture and transfer the reconstructed samples to the intra predictor 262.
[0088] In the present disclosure, the embodiments described in the filter 261, the inter predictor 221, and the intra predictor 222 of the encoding apparatus 200 may be the same as or respectively applied to correspond to the filter 266, the inter predictor 260, and the intra predictor 262of the decoding apparatus 250. The same may also apply to the inter predictor 260 and the intra predictor 262.
[0089] In-loop filtering techniques for traditional video codecs are one of the representative components of NN-based coding models or tools. Since VVC is a block-based hybrid coding framework, it inevitably produces undesirable compression artifacts, especially at a high compression rate. To address this issue, VVC employs a series of advanced filtering techniques to eliminate or reduce compression artifacts. As mentioned above, the present disclosure provides an exemplary DQT-CALF in-loop filter, as described below in connection to FIGs. 2C, 3A, and 3B.
[0090] FIG. 2C illustrates a detailed block diagram of an exemplary video coder 255 that includes a DQT-CALF network 288, according to some embodiments of the present disclosure.
[0091] As shown in FIG. 2C, video coder 255 may include a transform block 272, quantization block 274, entropy coding block 276, inverse quantization block 278, inverse transform block 280, in-loop filter 273, decoded picture buffer 294, inter prediction unit 296, and intra prediction unit 298.
[0092] Still referring to FIG. 2C, in-loop filter 273 may include, e.g., an LMCS unit 282, a DBF unit 284, an SAO unit 286, DQT-CALF network 288, a rate-distortion optimization (RDO) unit 290, and an ALF unit 292. LMCS unit 282 may utilize signal ranges to enhance coding efficiency; DBF unit 284 may reduce block artifacts and serve as the primary filter; SAO unit 286 may minimize ringing artifacts by finely adjusting captured intensity variations; DQT-CALF may implement a dual query mechanism to fuse global and local features to filter input picture data; RDO unit 290 may optimize the amount of distortion against the amount of data required to encode the video; and ALF unit 292 may correct signal values based on linear filter samples. Examples of DQT-CALF network 288 are provided below in connection with FIGs. 3A and 3B.
[0093] FIG. 3A illustrates a detailed block diagram of an exemplary DQT-CALF network 300 for luma components, according to some embodiments of the present disclosure. FIG. 3B illustrates a detailed block diagram of an exemplary DQT-CALF network 301 for chroma components, according to some embodiments of the present disclosure. FIGs. 3A and 3B will be described together.
[0094] FIGs. 3A and 3B illustrate the proposed DQT-CALF network architecture for the luma and chroma components, respectively. Both the luma and chroma models consist of three main parts: feature extraction (head) , feature enhancement (body) , and reconstruction (tail) . Although there are subtle differences between the luma and chroma models in feature extraction and reconstruction, they share the same feature enhancement structure, namely the multi-type feature fusion block (MFFB) 306, but with a different number of basic blocks (e.g., 6 MFFB in the luma model and 4 MFFB in the chroma model) . The inputs for the luma model include the luma reconstructed picture (Rec_Y) 302a, the luma predicted picture (Pre_Y) 302b, the luma partition map (Par_Y) 302c, and the QP map (QP_map) 302d. On the other hand, the inputs for the chroma model include the luma reconstructed picture (Rec_Y) 302a, the chroma reconstructed picture (Rec_UV) 302e, the chroma prediction picture (Pre_UV) 302f, the chroma partition map (Par_UV) 302g, and the QP_map 302h. The QP maps 302d and 302h of the luma model and the chroma model are inconsistent in size, and size of the chroma QP map (302h) is one-quarter the size of the luma QP map (302d) .
[0095] As the main input of DQT-CALF, the reconstructed picture (s) contain (s) the most information and is obtained by adding the predicted picture and the residual map before in-loop filtering. The predicted picture, as an auxiliary input, is composed of prediction blocks generated through intra-frame prediction (e.g., performed by intra predictor 222, 262) and inter-frame prediction (e.g., performed by inter predictor 221, 260) . In the AI configuration, intra-frame prediction is used, where decoded blocks are utilized to predict subsequent blocks. The partition map reflects block partitioning information for the current picture. The QP map differentiates blocks based on varying QP values, enabling DQT-CALF to adapt to different QPs.
[0096] In the feature extraction module (head) , the partition map, prediction picture and QP map are used as the auxiliary information. At the recent JVET conferences, most neural network-based in-loop filters, such as JVET-Z0091, JVET-AA0111, JVETAB0090, JVET-AC0118, JVET-AD0380, JVET-AE0191, JVET-AF0041, and JVET-AF0296, adopt the approach of directly concatenating auxiliary information with the reconstructed picture. However, due to significant content difference between the reconstructed picture and the auxiliary information, direct concatenation not only fails to effectively utilize the auxiliary information but may also cause interference that affects feature extraction from the reconstructed picture. Therefore, for the auxiliary information (luma model: Pre_Y 302b, Par_Y 302c, QP_map 302d; chroma model: Pre_UV 302f, Par_UV 302g, QP_map 302h) , we first employ a spatial attention (SA) mechanism to generate spatial attention maps 304. The spatial attention maps 304 are then combined with the feature maps associated with the reconstructed picture to achieve precise feature fusion.
[0097] FIG. 4 illustrates a detailed block diagram of a MFFB 306 of the DQT-CALF network 300, 301 depicted FIGs. 3A and 3B, according to some embodiments of the present disclosure. As shown in FIG. 4, the feature enhancement module (body) in DQT-CALF network 300, 301 employs MFFB 306 as the basic building blocks, with the luma model comprising 6 MFFBs and the chroma model consisting of 4. Each MFFB 306 has two inputs (e.g., L-GF and H_LF) and two outputs (e.g., low-frequency global feature maps 412 and high-frequency local feature maps 414) . Each MFFB 306 includes 4 residual blocks (RBs) 402, two High-frequency Local Feature Generation (HLFG) blocks 404, two Low-frequency Global Feature Generation (LGFG) blocks 406, a Double Query Transformer (DQT) 408, and three 1x1 convolution layers 410. The MFFB 306 is divided into three parts: Low-frequency Global Features (L-GF) , High-frequency Local Features (H-LF) , and the transformation between L-GF and H-LF, which results in low-frequency global feature maps 412 and high-frequency local feature maps 414.
[0098] FIG. 5A illustrates a detailed block diagram of HLFG block 404 of the MFFB 306 depicted in FIG. 4, according to some embodiments of the present disclosure. FIG. 5B illustrates a detailed block diagram of LGFG block 406 of the MFFB 306 depicted in FIG. 4, according to some embodiments of the present disclosure. FIGs. 5A and 5B will be described together.
[0099] The two core modules for DQT-CALF, e.g., HLFG block 404 and LGFG block 406, are shown in FIGs. 5A and 5B, respectively. These two modules aim to transform and integrate low-frequency global features with high-frequency local features.
[0100] Referring to FIG. 5A, HLFG block 404 employs a max pooling layer 502 and an edge detection block 504 to extract high-frequency feature from the input. Max pooling layer 502 effectively preserves texture details of an image by extracting significant information from the feature maps while removing redundant information. Edge detection block 504 identifies the contours of objects and features in an image, which are essential components of high-frequency information. During image restoration, edge information guides the reconstruction of details in the image. Additionally, the HLFG block 404 utilizes a 3x3 convolutional layer 508 (with stride 2) and a 1x1 convolutional layer 506 to extract local features from the input.
[0101] Referring to FIG. 5B, LGFG block 406, on the other hand, utilizes a convolutional transpose layer 510 and 1x1 convolutional layer 506 to convert the enhanced high-frequency local features into the auxiliary information for low-frequency global features.
[0102] FIG. 5C illustrates a detailed block diagram of DQT 408 of the MFFB 306 depicted in FIG. 4, according to some embodiments of the present disclosure.
[0103] As another core module for DQT-CALF, the present disclosure employs a DQT 408, as shown in FIG. 5C. DQT 408 adopts a dual query mechanism, one of which is used to capture low-frequency global features and the other is used to capture high-frequency local features. The attention map over these two different queries (Q1 and Q2) and keys (K) is computed, which are then multiplied and summed by the values (V) to obtain the weighted result. Finally, these results are concatenated in the channel dimension and dimensionality reduction is performed. The dual query key mechanism effectively fuses different types of information, thereby generating enriched feature representations. This mechanism not only enables the DQT-CALF network to comprehensively capture different characteristics of the input data, but also enhances the representation ability, allowing it to better adapt to the complexity of the input data. At the same time, by using dual queries, the risk of information loss can be reduced, especially when the input data has multiple important characteristics. Therefore, the dual query mechanism enhances model performance and robustness of DQT-CALF while generating accurate output. To implement the dual query mechanism, DQT 408 employs, e.g., 1x1 convolutional layers 506, 3x3 convolutional layers 508, softmax activation functions 512, a reshape (R) function 516, a GELU (G) activation function 518, matrix multiplication (X) function 520, addition function (+) 522, and element-wise (·) function 524.
[0104] FIG. 6 illustrates an exemplary progressive learning strategy 600 for a DQT-CALF network based on QP distance, according to some embodiments of the present disclosure.
[0105] The present disclosure uses a four-stage progressive learning strategy based on QP distance to train the DQT-CALF network. As shown in FIG. 6, the training sets for four stages are obtained by the original VTM compression with QP settings of 7, 12, 17, 22, 27, 32, 37, and 42. When setting the QP distance to 5, the input QP values for the training set are set to {22, 27, 32, 37, 42} , while the corresponding label QP values are set to {17, 22, 27, 32, 37} . Subsequently, in the next training stage, the model from the previous stage is loaded, and the QP distance for the training set is increased. In the first three stages, the QP distance increases by 5, while in the final stage, the label QP values are set to 7.
[0106] Two models were trained using the exemplary DQT-CALF network: one for processing luma components and another for processing chroma components. To validate the effectiveness of DQT-CALF, the trained models were embedded into VTM 11.0_NNVC-2.0 and conducted evaluations under the Common Test Conditions (CTC) specified by JVET. FIGs. 7A, 7B, and 7C present the test results for the first three training stages under the AI configuration. In these stages, the BD-rates for the Y, U, and V channels are {-8.24%, -16.43%, -17.23%} , {-8.54%, -20.27%, -21.18%} , and {-8.72%, -21.33%, -22.24%} , respectively, showing significantly improved gains in the later stages compared to the earlier ones.
[0107] FIG. 7A illustrates exemplary test results 700 of a first stage of the exemplary progressive learning strategy depicted in FIG. 6 implemented with the AI configuration, according to some embodiments of the present disclosure.
[0108] Referring to FIG. 7A, the first-stage trained models were embedded into VTM-11.0 NNVC-2.0 and the testing results were obtained under the AI configuration with a QP distance set to 5. The input QP values of the training set were {22, 27, 32, 37, 42} , corresponding to label QP values of {17, 22, 27, 32, 37} .
[0109] FIG. 7B illustrates exemplary test results 705 of a second stage of the exemplary progressive learning strategy depicted in FIG. 6 implemented with the AI configuration, according to some embodiments of the present disclosure.
[0110] Referring to FIG. 7B, the second-stage trained models were embedded into VTM-11.0 NNVC-2.0 and the testing results were obtained under the AI configuration with a QP distance set to 10. The input QP values of the training set were {22, 27, 32, 37, 42} , corresponding to label QP values of {12, 17, 22, 27, 32} .
[0111] FIG. 7C illustrates exemplary test results 710 of a third stage of the exemplary progressive learning strategy depicted in FIG. 6 implemented with the AI configuration, according to some embodiments of the present disclosure.
[0112] Referring to FIG. 7C, the third-stage trained models were embedded into VTM-11.0 NNVC-2.0 and the testing results were obtained under the AI configuration with a QP distance set to 15. The input QP values of the training set were {22, 27, 32, 37, 42} , corresponding to label QP values of {7, 12, 17, 22, 27} .
[0113] FIG. 7D illustrates exemplary final test results 715 of a fourth stage of the exemplary progressive learning strategy depicted in FIG. 6 implemented with the AI configuration, according to some embodiments of the present disclosure.
[0114] Referring to FIG. 7D, the final results from the fourth stage are shown. To further verify the effectiveness of DQT-CALF, we selected one test sequence from Classes A1, A2, B, C, D, and E and plotted RD curves for the Y channel, as shown in FIGs. 8A-8F. The results demonstrate that DQT-CALF achieves outstanding performance.
[0115] Still referring to FIG. 7D, the models trained in the fourth stage were embedded into VTM-11.0 NNVC-2.0, and testing was conducted under the AI configuration. The input QP values for the training set were {22, 27, 32, 37, 42} , with corresponding labeled QP values all set to {7} .
[0116] FIG. 7E illustrates exemplary test results 720 of the DQT-CALF network obtained under an RA configuration, according to some embodiments of the present disclosure.
[0117] Referring to FIG. 7E, the BD rates under the RA configuration are shown. To train a model suitable for the RA configuration, a pre-trained model from the AI configuration is utilized as the initialization and trained it solely using the dataset with QP_Distance = 10. The trained models were embedded into VTM-11.0 NNVC-2.0 and the testing results were obtained under the RA configuration with a QP_distance set to 10. The input QP values of the training set were {22, 27, 32, 37, 42} , corresponding to label QP values of {12, 17, 22, 27, 32} .
[0118] FIG. 7F illustrates an overall BD-rate comparison 725 with latest in-loop filters in JVET under AI configuration, according to some embodiments of the present disclosure.
[0119] Referring to FIG. 7F, additional results for comparison with other NNLFs from the JVET conference are shown. The results indicate that DQT-CALF demonstrates exceptional performance across all test sequences, achieving optimal results. The results indicate that DQT-CALF network of the present disclosure demonstrates improved performance across all test sequences, while optimizing performance.
[0120] FIG. 7G illustrates the results of an ablation study 730 on the spatial attention (SA) in DQT-CALF, according to some embodiments of the present disclosure.
[0121] Referring to FIG. 7G, the results of ablation study on the SA in DQT-CALF network is shown. The BD-rates are obtained by directly concatenating the auxiliary information with the reconstructed picture under the AI configuration. In the feature extraction stage (head) , to better utilize auxiliary information, we first generate spatial attention maps through a spatial attention mechanism and then apply them to the reconstructed pictures. To validate this mechanism’s effectiveness, we removed the spatial attention mechanism from the feature extraction stage and instead directly concatenated the reconstructed pictures with the auxiliary information, while keeping the rest of the network, training, and testing conditions unchanged. Table 7 shows the corresponding results. From the table, it can be observed that without using the spatial attention mechanism in the feature extraction stage, the performance gain decreased by {0.37% (Y) , 0.74% (U) , 0.84% (V) } .
[0122] FIG. 7H illustrates the results of an ablation study 735 on MFFB in DQT-CALF, according to some embodiments of the present disclosure.
[0123] Referring to FIG. 7H, the results of an ablation study on MFFB in DQT-CALF network is shown. The BD-rates are obtained without the feature decomposition and fusion in MFFB on the input feature under the AI configuration.
[0124] In MFFB, the input features are divided into two groups: one containing low-frequency and global information, and the other containing high-frequency and local information. To validate this approach, we conducted an ablation experiment where we did not process the input features in groups but directly used residual blocks (RB) and Transform to process the input information. FIG. 7H shows the results. From Fig. 7H, it can be seen that compared to the grouped processing and complementary enhancement of input features, the performance gain decreased by {0.23% (Y) , 2.85% (U) , 3.0% (V) } .
[0125] FIG. 7I illustrates the results of an ablation study 740 on the effect of the reconstructed luma component (Rec_Y) on the chroma component in the chroma model, according to some embodiments of the present disclosure.
[0126] Referring to FIG. 7I, the results of the ablation study on the effect of the reconstructed luma component (Rec_Y) on the chroma component in the chroma model are shown.
[0127] Still referring to FIG. 7I, in the chroma model, using the reconstructed luma component (Rec_Y) as a supplement further enhances the chroma component. FIG. 7I provides a detailed demonstration of the effect of Rec_Y on the chroma component. The table in FIG. 7I shows that, compared to the case when Rec_Y is used, the gain of the U and V channels decreased from {-22.09%, -22.99%} to {-9.33%, -11.40%} when Rec_Y was not used as input. Rec_Y provides essential structural and detailed information about the image, which is crucial for enhancing the chroma component. It offers the necessary context for the chroma component, allowing for a more comprehensive consideration of the overall structure and details of the image, thus effectively improving performance. Without Rec_Y, DQT-CALF may struggle to fully utilize the limited chroma information, leading to restricted performance.
[0128] FIG. 8A is a first table illustrating average rate distortion (RD) results 800 on all test sequences under AI configuration in terms of PSNR for a first dataset (Class A1) , according to some embodiments of the present disclosure. FIG. 8B is a second table illustrating second average RD results 805 on all test sequences under AI configuration in terms of PSNR for a second dataset (Class A2) , according to some embodiments of the present disclosure. FIG. 8C is a third table illustrating third average RD results 810 on all test sequences under AI configuration in terms of PSNR for a third dataset (Class B) , according to some embodiments of the present disclosure. FIG. 8D is a fourth table illustrating fourth average RD results 815 on all test sequences under AI configuration in terms of PSNR for a fourth dataset (Class C) , according to some embodiments of the present disclosure. FIG. 8E is a fifth table illustrating fifth average RD results 820 on all test sequences under AI configuration in terms of PSNR for a fifth dataset (Class D) , according to some embodiments of the present disclosure. FIG. 8F is a sixth table illustrating sixth average RD results 825 on all test sequences under AI configuration in terms of PSNR for a sixth dataset (Class E) , according to some embodiments of the present disclosure.
[0129] FIG. 9 illustrates a flow chart of an exemplary method 900 of video decoding, according to some embodiments of the present disclosure. Method 900 may be performed by an apparatus, e.g., such as decoding apparatus 20, 250 decoding unit 22, in-loop filter 273, DQT-CALF network 288, 300, 301, etc. Method 900 may include operations 902-914 as described below. It is understood that some of the operations may be optional, and some of the operations may be performed simultaneously, or in a different order other than shown in FIG. 9.
[0130] At 902, the apparatus may obtain a first reference picture from a decoded picture buffer. Referring to FIG. 2C, a first picture may be obtained from decoded picture buffer 294, and one or more reconstructed pictures, prediction pictures, and / or partition maps may be generated by inter prediction unit 296 and / or intra prediction unit 298.
[0131] At 904, the apparatus may generate a first feature map based on first information associated with the first reference picture. In some implementations, the first information may include a reconstructed picture. In some implementations, the first information may include a luma reconstructed picture and a luma prediction picture associated with the first reference picture. In some implementations, the first information may include a luma reconstructed picture and a chroma reconstructed picture associated with the first reference picture. For example, the first feature map may be generated based on one or more operations described above in connection with FIGs. 3A and / or 3B.
[0132] At 906, the apparatus may generate a spatial attention map based on second information associated with the first reference picture. In some implementations, the second information may include a luma partition map and a QP map. In some implementations, the second information may include a chroma prediction picture, a chroma partition map, and a QP map. For example, the spatial attention map may be generated based on one or more operations described above in connection with FIGs. 3A and / or 3B.
[0133] At 908, the apparatus may generate a second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation. In some implementations, the at least one MFFB operation may include more than one MFFB operation. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the apparatus may generate low-frequency information and high-frequency information based on the first feature map and the spatial attention map. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the apparatus may generate a low-frequency global feature map based on the low-frequency information and the high-frequency information by performing a first residual block operation, a first HFLG block operation, a first LGFG block operation, and a DQT operation. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the apparatus may generate a high-frequency local feature map based on the low-frequency information and the high-frequency information by performing a second residual block operation, a second HFLG block operation, a second LGFG block operation, and the DQT operation. In some implementations, the second feature map including the low-frequency global feature map and the high-frequency local feature map. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the apparatus may perform a matrix multiplication of the first feature map and the spatial attention map to generate an intermediate feature map. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the apparatus may generate the low-frequency information based on the intermediate feature map by performing a first convolutional layer operation. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the apparatus may generate the high-frequency information based on the low-frequency information by performing a max-pooling layer operation and a second convolutional layer operation. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the apparatus may generate an attention map by multiplying a first query, a second query, and a key by a value and performing a summation as part of the DQT operation. In some implementations, the attention map may be used to generate the low-frequency global feature map and the high-frequency local feature map.
[0134] At 910, the apparatus may generate a second reference picture based on the reconstructed picture and the second feature map. In some implementations, to generate the second reference picture based on the reconstructed picture and the second feature map, the apparatus may generate a final feature map based on the second feature map by performing a convolutional layer operation and a pixel-shuffle layer operation. In some implementations, to generate the second reference picture based on the reconstructed picture and the second feature map, the apparatus may combine the reconstructed picture associated with the first reference picture and the final feature map to generate the second reference picture.
[0135] At 912, the apparatus may add the second reference picture into the decoded picture buffer. For example, referring to FIG. 2C, in-loop filter 273 may add the second reference picture to decoded picture buffer 294.
[0136] At 914, the apparatus may decode a current picture region based on the second reference picture. Referring to FIG. 2C, video coder 255 may decode a current picture region based on the second reference picture added to decoded picture buffer 294.
[0137] FIG. 10 illustrates a flow chart of an exemplary method 1000 of video encoding, according to some embodiments of the present disclosure. Method 1000 may be performed by an apparatus, e.g., such as encoding apparatus 10, 200, encoding unit 12, in-loop filter 273, DQT-CALF network 288, 300, 301, etc. Method 1000 may include operations 1002-1014 as described below. It is understood that some of the operations may be optional, and some of the operations may be performed simultaneously, or in a different order other than shown in FIG. 10.
[0138] At 1002, the apparatus may obtain a first reference picture from a decoded picture buffer. Referring to FIG. 2C, a first picture may be obtained from decoded picture buffer 294, and one or more reconstructed pictures, prediction pictures, and / or partition maps may be generated by inter prediction unit 296 and / or intra prediction unit 298.
[0139] At 1004, the apparatus may generate a first feature map based on first information associated with the first reference picture. In some implementations, the first information may include a reconstructed picture. In some implementations, the first information may include a luma reconstructed picture and a luma prediction picture associated with the first reference picture. In some implementations, the first information may include a luma reconstructed picture and a chroma reconstructed picture associated with the first reference picture. For example, the first feature map may be generated based on one or more operations described above in connection with FIGs. 3A and / or 3B.
[0140] At 1006, the apparatus may generate a spatial attention map based on second information associated with the first reference picture. In some implementations, the second information may include a luma partition map and a QP map. In some implementations, the second information may include a chroma prediction picture, a chroma partition map, and a QP map. For example, the spatial attention map may be generated based on one or more operations described above in connection with FIGs. 3A and / or 3B.
[0141] At 1008, the apparatus may generate a second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation. In some implementations, the at least one MFFB operation may include more than one MFFB operation. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the apparatus may generate low-frequency information and high-frequency information based on the first feature map and the spatial attention map. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the apparatus may generate a low-frequency global feature map based on the low-frequency information and the high-frequency information by performing a first residual block operation, a first HFLG block operation, a first LGFG block operation, and a DQT operation. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the apparatus may generate a high-frequency local feature map based on the low-frequency information and the high-frequency information by performing a second residual block operation, a second HFLG block operation, a second LGFG block operation, and the DQT operation. In some implementations, the second feature map including the low-frequency global feature map and the high-frequency local feature map. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the apparatus may perform a matrix multiplication of the first feature map and the spatial attention map to generate an intermediate feature map. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the apparatus may generate the low-frequency information based on the intermediate feature map by performing a first convolutional layer operation. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the apparatus may generate the high-frequency information based on the low-frequency information by performing a max-pooling layer operation and a second convolutional layer operation. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the apparatus may generate an attention map by multiplying a first query, a second query, and a key by a value and performing a summation as part of the DQT operation. In some implementations, the attention map may be used to generate the low-frequency global feature map and the high-frequency local feature map.
[0142] At 1010, the apparatus may generate a second reference picture based on the reconstructed picture and the second feature map. In some implementations, to generate the second reference picture based on the reconstructed picture and the second feature map, the apparatus may generate a final feature map based on the second feature map by performing a convolutional layer operation and a pixel-shuffle layer operation. In some implementations, to generate the second reference picture based on the reconstructed picture and the second feature map, the apparatus may combine the reconstructed picture associated with the first reference picture and the final feature map to generate the second reference picture.
[0143] At 1012, the apparatus may add the second reference picture into the decoded picture buffer. For example, referring to FIG. 2C, in-loop filter 273 may add the second reference picture to decoded picture buffer 294.
[0144] At 1014, the apparatus may encode a current picture region based on the second reference picture. Referring to FIG. 2C, video coder 255 may encode a current picture region based on the second reference picture added to decoded picture buffer 294.
[0145] In various aspects of the present disclosure, the functions described herein may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as instructions on a non-transitory computer-readable medium. Computer-readable media includes computer storage media. Storage media may be any available media that can be accessed by a processor, such as a processor in encoding unit 12 or decoding unit 22 in FIG. 1. By way of example, and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, HDD, such as magnetic disk storage or other magnetic storage devices, Flash drive, SSD, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a processing system, such as a mobile device or a computer. Disk and disc, as used herein, include CD, laser disc, optical disc, digital video disc (DVD) , and floppy disk where disks usually reproduce data magnetically, while discs reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0146] According to one aspect of the present disclosure, a method of video decoding is provided. The method may include obtaining, by a processor, a first reference picture from a decoded picture buffer. The method may include generating, by the processor, a first feature map based on first information associated with the first reference picture. The first information may include a reconstructed picture. The method may include generating, by the processor, a spatial attention map based on second information associated with the first reference picture. The method may include generating, by the processor, a second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation. The method may include generating, by the processor, a second reference picture based on the reconstructed picture and the second feature map. The method may include decoding, by the processor, a current picture region based on the second reference picture.
[0147] In some implementations, the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation may include generating, by the processor, low-frequency information and high-frequency information based on the first feature map and the spatial attention map. In some implementations, the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation may include generating, by the processor, a low-frequency global feature map based on the low-frequency information and the high-frequency information by performing a first residual block operation, a first HFLG block operation, a first LGFG block operation, and a DQT operation. In some implementations, the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation may include generating, by the processor, a high-frequency local feature map based on the low-frequency information and the high-frequency information by performing a second residual block operation, a second HFLG block operation, a second LGFG block operation, and the DQT operation, the second feature map including the low-frequency global feature map and the high-frequency local feature map.
[0148] In some implementations, the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation may include performing, by the processor, a matrix multiplication of the first feature map and the spatial attention map to generate an intermediate feature map. In some implementations, the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation may include generating, by the processor, the low-frequency information based on the intermediate feature map by performing a first convolutional layer operation. In some implementations, the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation may include generating, by the processor, the high-frequency information based on the low-frequency information by performing a max-pooling layer operation and a second convolutional layer operation.
[0149] In some implementations, the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation may include generating, by the processor, an attention map by multiplying a first query, a second query, and a key by a value and performing a summation as part of the DQT operation. In some implementations, the attention map may be used to generate the low-frequency global feature map and the high-frequency local feature map.
[0150] In some implementations the generating, by the processor, the second reference picture based on the reconstructed picture and the second feature map may include generating, by the processor, a final feature map based on the second feature map by performing a convolutional layer operation and a pixel-shuffle layer operation. In some implementations the generating, by the processor, the second reference picture based on the reconstructed picture and the second feature map may include combining, by the processor, the reconstructed picture associated with the first reference picture and the final feature map to generate the second reference picture.
[0151] In some implementations, the first information may include a luma reconstructed picture and a luma prediction picture associated with the first reference picture. In some implementations, the second information includes a luma partition map and a QP map.
[0152] In some implementations, the first information may include a luma reconstructed picture and a chroma reconstructed picture associated with the first reference picture. In some implementations, the second information may include a chroma prediction picture, a chroma partition map, and a QP map.
[0153] In some implementations, the at least one MFFB operation includes more than one MFFB operation.
[0154] In some implementations, the method may further include adding, by the processor, the second reference picture into the decoded picture buffer.
[0155] According to another aspect of the present disclosure, a video decoder is provided. The video decoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to obtain a first reference picture from a decoded picture buffer. The memory storing instructions, which when executed by the processor, may cause the processor to generate a first feature map based on first information associated with the first reference picture. The first information may include a reconstructed picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a spatial attention map based on second information associated with the first reference picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second reference picture based on the reconstructed picture and the second feature map. The memory storing instructions, which when executed by the processor, may cause the processor to decode a current picture region based on the second reference picture.
[0156] In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, may cause the processor to generate low-frequency information and high-frequency information based on the first feature map and the spatial attention map. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, may cause the processor to generate a low-frequency global feature map based on the low-frequency information and the high-frequency information by performing a first residual block operation, a first HFLG block operation, a first low-frequency LGFG block operation, and a DQT operation. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, may cause the processor to generate a high-frequency local feature map based on the low-frequency information and the high-frequency information by performing a second residual block operation, a second HFLG block operation, a second LGFG block operation, and the DQT operation, the second feature map including the low-frequency global feature map and the high-frequency local feature map.
[0157] In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, may cause the processor to perform a matrix multiplication of the first feature map and the spatial attention map to generate an intermediate feature map. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, may cause the processor to generate the low-frequency information based on the intermediate feature map by performing a first convolutional layer operation. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, may cause the processor to generate the high-frequency information based on the low-frequency information by performing a max-pooling layer operation and a second convolutional layer operation.
[0158] In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, may cause the processor to generate an attention map by multiplying a first query, a second query, and a key by a value and performing a summation as part of the DQT operation. In some implementations, the attention map may be used to generate the low-frequency global feature map and the high-frequency local feature map.
[0159] In some implementations, to generate the second reference picture based on the reconstructed picture and the second feature map, the memory storing instructions, which when executed by the processor, may cause the processor to generate a final feature map based on the second feature map by performing a convolutional layer operation and a pixel-shuffle layer operation. In some implementations, to generate the second reference picture based on the reconstructed picture and the second feature map, the memory storing instructions, which when executed by the processor, may cause the processor to combine the reconstructed picture associated with the first reference picture and the final feature map to generate the second reference picture.
[0160] In some implementations, the first information may include a luma reconstructed picture and a luma prediction picture associated with the first reference picture. In some implementations, the second information includes a luma partition map and a QP map.
[0161] In some implementations, the first information may include a luma reconstructed picture and a chroma reconstructed picture associated with the first reference picture. In some implementations, the second information includes a chroma prediction picture, a chroma partition map, and a QP map.
[0162] In some implementations, the at least one MFFB operation may include more than one MFFB operation.
[0163] In some implementations, the memory storing instructions, which when executed by the processor, may cause the processor to add the second reference picture into the decoded picture buffer.
[0164] According to a further aspect of the present disclosure, an apparatus for video decoding is provided. The video decoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to obtain a first reference picture from a decoded picture buffer. The memory storing instructions, which when executed by the processor, may cause the processor to generate a first feature map based on first information associated with the first reference picture. The first information may include a reconstructed picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a spatial attention map based on second information associated with the first reference picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second reference picture based on the reconstructed picture and the second feature map. The memory storing instructions, which when executed by the processor, may cause the processor to decode a current picture region based on the second reference picture.
[0165] According to yet another aspect of the present disclosure, a non-transitory computer-readable medium storing instructions for a video decoder is provided. The instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to obtain a first reference picture from a decoded picture buffer. The instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to generate a first feature map based on first information associated with the first reference picture. The first information may include a reconstructed picture. The instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to generate a spatial attention map based on second information associated with the first reference picture. The instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to generate a second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation. The instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to generate a second reference picture based on the reconstructed picture and the second feature map. The instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to decode a current picture region based on the second reference picture.
[0166] In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to generate low-frequency information and high-frequency information based on the first feature map and the spatial attention map. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to generate a low-frequency global feature map based on the low-frequency information and the high-frequency information by performing a first residual block operation, a first HFLG block operation, a first low-frequency LGFG block operation, and a DQT operation. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to generate a high-frequency local feature map based on the low-frequency information and the high-frequency information by performing a second residual block operation, a second HFLG block operation, a second LGFG block operation, and the DQT operation, the second feature map including the low-frequency global feature map and the high-frequency local feature map.
[0167] In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to perform a matrix multiplication of the first feature map and the spatial attention map to generate an intermediate feature map. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to generate the low-frequency information based on the intermediate feature map by performing a first convolutional layer operation. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to generate the high-frequency information based on the low-frequency information by performing a max-pooling layer operation and a second convolutional layer operation.
[0168] In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to generate an attention map by multiplying a first query, a second query, and a key by a value and performing a summation as part of the DQT operation. In some implementations, the attention map may be used to generate the low-frequency global feature map and the high-frequency local feature map.
[0169] In some implementations, to generate the second reference picture based on the reconstructed picture and the second feature map, the instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to generate a final feature map based on the second feature map by performing a convolutional layer operation and a pixel-shuffle layer operation. In some implementations, to generate the second reference picture based on the reconstructed picture and the second feature map, the instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to combine the reconstructed picture associated with the first reference picture and the final feature map to generate the second reference picture.
[0170] In some implementations, the first information may include a luma reconstructed picture and a luma prediction picture associated with the first reference picture. In some implementations, the second information includes a luma partition map and a QP map.
[0171] In some implementations, the first information may include a luma reconstructed picture and a chroma reconstructed picture associated with the first reference picture. In some implementations, the second information includes a chroma prediction picture, a chroma partition map, and a QP map.
[0172] In some implementations, the at least one MFFB operation may include more than one MFFB operation.
[0173] In some implementations, the instructions, which when executed by the processor of the video decoder, may cause the processor of the video decoder to add the second reference picture into the decoded picture buffer.
[0174] According to one aspect of the present disclosure, a method of video encoding is provided. The method may include obtaining, by a processor, a first reference picture from a decoded picture buffer. The method may include generating, by the processor, a first feature map based on first information associated with the first reference picture. The first information may include a reconstructed picture. The method may include generating, by the processor, a spatial attention map based on second information associated with the first reference picture. The method may include generating, by the processor, a second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation. The method may include generating, by the processor, a second reference picture based on the reconstructed picture and the second feature map. The method may include encoding, by the processor, a current picture region based on the second reference picture.
[0175] In some implementations, the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation may include generating, by the processor, low-frequency information and high-frequency information based on the first feature map and the spatial attention map. In some implementations, the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation may include generating, by the processor, a low-frequency global feature map based on the low-frequency information and the high-frequency information by performing a first residual block operation, a first HFLG block operation, a first LGFG block operation, and a DQT operation. In some implementations, the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation may include generating, by the processor, a high-frequency local feature map based on the low-frequency information and the high-frequency information by performing a second residual block operation, a second HFLG block operation, a second LGFG block operation, and the DQT operation, the second feature map including the low-frequency global feature map and the high-frequency local feature map.
[0176] In some implementations, the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation may include performing, by the processor, a matrix multiplication of the first feature map and the spatial attention map to generate an intermediate feature map. In some implementations, the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation may include generating, by the processor, the low-frequency information based on the intermediate feature map by performing a first convolutional layer operation. In some implementations, the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation may include generating, by the processor, the high-frequency information based on the low-frequency information by performing a max-pooling layer operation and a second convolutional layer operation.
[0177] In some implementations, the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation may include generating, by the processor, an attention map by multiplying a first query, a second query, and a key by a value and performing a summation as part of the DQT operation. In some implementations, the attention map may be used to generate the low-frequency global feature map and the high-frequency local feature map.
[0178] In some implementations the generating, by the processor, the second reference picture based on the reconstructed picture and the second feature map may include generating, by the processor, a final feature map based on the second feature map by performing a convolutional layer operation and a pixel-shuffle layer operation. In some implementations the generating, by the processor, the second reference picture based on the reconstructed picture and the second feature map may include combining, by the processor, the reconstructed picture associated with the first reference picture and the final feature map to generate the second reference picture.
[0179] In some implementations, the first information may include a luma reconstructed picture and a luma prediction picture associated with the first reference picture. In some implementations, the second information includes a luma partition map and a QP map.
[0180] In some implementations, the first information may include a luma reconstructed picture and a chroma reconstructed picture associated with the first reference picture. In some implementations, the second information may include a chroma prediction picture, a chroma partition map, and a QP map.
[0181] In some implementations, the at least one MFFB operation includes more than one MFFB operation.
[0182] In some implementations, the method may further include adding, by the processor, the second reference picture into the decoded picture buffer.
[0183] According to another aspect of the present disclosure, a video encoder is provided. The video encoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to obtain a first reference picture from a decoded picture buffer. The memory storing instructions, which when executed by the processor, may cause the processor to generate a first feature map based on first information associated with the first reference picture. The first information may include a reconstructed picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a spatial attention map based on second information associated with the first reference picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second reference picture based on the reconstructed picture and the second feature map. The memory storing instructions, which when executed by the processor, may cause the processor to encode a current picture region based on the second reference picture.
[0184] In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, may cause the processor to generate low-frequency information and high-frequency information based on the first feature map and the spatial attention map. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, may cause the processor to generate a low-frequency global feature map based on the low-frequency information and the high-frequency information by performing a first residual block operation, a first HFLG block operation, a first low-frequency LGFG block operation, and a DQT operation. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, may cause the processor to generate a high-frequency local feature map based on the low-frequency information and the high-frequency information by performing a second residual block operation, a second HFLG block operation, a second LGFG block operation, and the DQT operation, the second feature map including the low-frequency global feature map and the high-frequency local feature map.
[0185] In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, may cause the processor to perform a matrix multiplication of the first feature map and the spatial attention map to generate an intermediate feature map. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, may cause the processor to generate the low-frequency information based on the intermediate feature map by performing a first convolutional layer operation. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, may cause the processor to generate the high-frequency information based on the low-frequency information by performing a max-pooling layer operation and a second convolutional layer operation.
[0186] In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, may cause the processor to generate an attention map by multiplying a first query, a second query, and a key by a value and performing a summation as part of the DQT operation. In some implementations, the attention map may be used to generate the low-frequency global feature map and the high-frequency local feature map.
[0187] In some implementations, to generate the second reference picture based on the reconstructed picture and the second feature map, the memory storing instructions, which when executed by the processor, may cause the processor to generate a final feature map based on the second feature map by performing a convolutional layer operation and a pixel-shuffle layer operation. In some implementations, to generate the second reference picture based on the reconstructed picture and the second feature map, the memory storing instructions, which when executed by the processor, may cause the processor to combine the reconstructed picture associated with the first reference picture and the final feature map to generate the second reference picture.
[0188] In some implementations, the first information may include a luma reconstructed picture and a luma prediction picture associated with the first reference picture. In some implementations, the second information includes a luma partition map and a QP map.
[0189] In some implementations, the first information may include a luma reconstructed picture and a chroma reconstructed picture associated with the first reference picture. In some implementations, the second information includes a chroma prediction picture, a chroma partition map, and a QP map.
[0190] In some implementations, the at least one MFFB operation may include more than one MFFB operation.
[0191] In some implementations, the memory storing instructions, which when executed by the processor, may cause the processor to add the second reference picture into the decoded picture buffer.
[0192] According to a further aspect of the present disclosure, an apparatus for video encoding is provided. The video encoder may include a processor and memory storing instructions. The memory storing instructions, which when executed by the processor, may cause the processor to obtain a first reference picture from a decoded picture buffer. The memory storing instructions, which when executed by the processor, may cause the processor to generate a first feature map based on first information associated with the first reference picture. The first information may include a reconstructed picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a spatial attention map based on second information associated with the first reference picture. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation. The memory storing instructions, which when executed by the processor, may cause the processor to generate a second reference picture based on the reconstructed picture and the second feature map. The memory storing instructions, which when executed by the processor, may cause the processor to encode a current picture region based on the second reference picture.
[0193] According to yet another aspect of the present disclosure, a non-transitory computer-readable medium storing instructions for a video encoder is provided. The instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to obtain a first reference picture from a decoded picture buffer. The instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to generate a first feature map based on first information associated with the first reference picture. The first information may include a reconstructed picture. The instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to generate a spatial attention map based on second information associated with the first reference picture. The instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to generate a second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation. The instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to generate a second reference picture based on the reconstructed picture and the second feature map. The instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to encode a current picture region based on the second reference picture.
[0194] In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to generate low-frequency information and high-frequency information based on the first feature map and the spatial attention map. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to generate a low-frequency global feature map based on the low-frequency information and the high-frequency information by performing a first residual block operation, a first HFLG block operation, a first low-frequency LGFG block operation, and a DQT operation. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to generate a high-frequency local feature map based on the low-frequency information and the high-frequency information by performing a second residual block operation, a second HFLG block operation, a second LGFG block operation, and the DQT operation, the second feature map including the low-frequency global feature map and the high-frequency local feature map.
[0195] In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to perform a matrix multiplication of the first feature map and the spatial attention map to generate an intermediate feature map. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to generate the low-frequency information based on the intermediate feature map by performing a first convolutional layer operation. In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to generate the high-frequency information based on the low-frequency information by performing a max-pooling layer operation and a second convolutional layer operation.
[0196] In some implementations, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to generate an attention map by multiplying a first query, a second query, and a key by a value and performing a summation as part of the DQT operation. In some implementations, the attention map may be used to generate the low-frequency global feature map and the high-frequency local feature map.
[0197] In some implementations, to generate the second reference picture based on the reconstructed picture and the second feature map, the instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to generate a final feature map based on the second feature map by performing a convolutional layer operation and a pixel-shuffle layer operation. In some implementations, to generate the second reference picture based on the reconstructed picture and the second feature map, the instructions, which when executed by the processor of the video encoder, may cause the processor of the video encoder to combine the reconstructed picture associated with the first reference picture and the final feature map to generate the second reference picture.
[0198] In some implementations, the first information may include a luma reconstructed picture and a luma prediction picture associated with the first reference picture. In some implementations, the second information includes a luma partition map and a QP map.
[0199] In some implementations, the first information may include a luma reconstructed picture and a chroma reconstructed picture associated with the first reference picture. In some implementations, the second information includes a chroma prediction picture, a chroma partition map, and a QP map.
[0200] In some implementations, the at least one MFFB operation may include more than one MFFB operation.
[0201] In some implementations, the instructions, which when executed by the processor of the video encoder, may cause the processor to add the second reference picture to the decoded picture buffer.
[0202] According to yet a further aspect of the present disclosure, a non-transitory computer-readable medium storing a bitstream is provided. The bitstream may be generated based on one or more of the operations described herein.
[0203] The foregoing description of the embodiments will so reveal the general nature of the present disclosure that others can, by applying knowledge within the skill of the art, readily modify and / or adapt for various applications such embodiments, without undue experimentation, without departing from the general concept of the present disclosure. Therefore, such adaptations and modifications are intended to be within the meaning and range of equivalents of the disclosed embodiments, based on the teaching and guidance presented herein. It is to be understood that the phraseology or terminology herein is for the purpose of description and not of limitation, such that the terminology or phraseology of the present specification is to be interpreted by the skilled artisan in light of the teachings and guidance.
[0204] Embodiments of the present disclosure have been described above with the aid of functional building blocks illustrating the implementation of specified functions and relationships thereof. The boundaries of these functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternate boundaries can be defined so long as the specified functions and relationships thereof are appropriately performed.
[0205] The Summary and Abstract sections may set forth one or more but not all exemplary embodiments of the present disclosure as contemplated by the inventor (s) , and thus, are not intended to limit the present disclosure and the appended claims in any way.
[0206] Various functional blocks, modules, and steps are disclosed above. The arrangements provided are illustrative and without limitation. Accordingly, the functional blocks, modules, and steps may be reordered or combined in different ways than in the examples provided above. Likewise, some embodiments include only a subset of the functional blocks, modules, and steps, and any such subset is permitted.
[0207] The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.
Claims
1.A method of video decoding, comprising:obtaining, by a processor, a first reference picture from a decoded picture buffer;generating, by the processor, a first feature map based on first information associated with the first reference picture, the first information including a reconstructed picture;generating, by the processor, a spatial attention map based on second information associated with the first reference picture;generating, by the processor, a second feature map based on the first feature map and the spatial attention map by performing at least one multi-type feature fusion block (MFFB) operation;generating, by the processor, a second reference picture based on the reconstructed picture and the second feature map; anddecoding, by the processor, a current picture region based on the second reference picture.2.The method of claim 1, wherein the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation comprises:generating, by the processor, low-frequency information and high-frequency information based on the first feature map and the spatial attention map;generating, by the processor, a low-frequency global feature map based on the low-frequency information and the high-frequency information by performing a first residual block operation, a first high-frequency local feature generation (HFLG) block operation, a first low-frequency global-feature generation (LGFG) block operation, and a dual query transformer (DQT) operation; andgenerating, by the processor, a high-frequency local feature map based on the low-frequency information and the high-frequency information by performing a second residual block operation, a second HFLG block operation, a second LGFG block operation, and the DQT operation, the second feature map including the low-frequency global feature map and the high-frequency local feature map.3.The method of claim 2, wherein the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation further comprises:performing, by the processor, a matrix multiplication of the first feature map and the spatial attention map to generate an intermediate feature map;generating, by the processor, the low-frequency information based on the intermediate feature map by performing a first convolutional layer operation; andgenerating, by the processor, the high-frequency information based on the low-frequency information by performing a max-pooling layer operation and a second convolutional layer operation.4.The method of claim 3, wherein the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation further comprises:generating, by the processor, an attention map by multiplying a first query, a second query, and a key by a value and performing a summation as part of the DQT operation, the attention map being used to generate the low-frequency global feature map and the high-frequency local feature map.5.The method of claim 1, wherein the generating, by the processor, the second reference picture based on the reconstructed picture and the second feature map comprises:generating, by the processor, a final feature map based on the second feature map by performing a convolutional layer operation and a pixel-shuffle layer operation; andcombining, by the processor, the reconstructed picture associated with the first reference picture and the final feature map to generate the second reference picture.6.The method of claim 1, wherein:the first information includes a luma reconstructed picture and a luma prediction picture associated with the first reference picture, andthe second information includes a luma partition map and a quantization parameters (QP) map.7.The method of claim 1, wherein:the first information includes a luma reconstructed picture and a chroma reconstructed picture associated with the first reference picture, andthe second information includes a chroma prediction picture, a chroma partition map, and a quantization parameters (QP) map.8.The method of claim 1, wherein the at least one MFFB operation includes more than one MFFB operation.9.The method of claim 1, further comprising:adding, by the processor, the second reference picture into the decoded picture buffer.10.A video decoder, comprising:a processor; andmemory storing instructions, which when executed by the processor, cause the processor to:obtain a first reference picture from a decoded picture buffer;generate a first feature map based on first information associated with the first reference picture, the first information including a reconstructed picture;generate a spatial attention map based on second information associated with the first reference picture;generate a second feature map based on the first feature map and the spatial attention map by performing at least one multi-type feature fusion block (MFFB) operation;generate a second reference picture based on the reconstructed picture and the second feature map; anddecode a current picture region based on the second reference picture.11.The video decoder of claim 10, wherein, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, cause the processor to:generate low-frequency information and high-frequency information based on the first feature map and the spatial attention map;generate a low-frequency global feature map based on the low-frequency information and the high-frequency information by performing a first residual block operation, a first high-frequency local feature generation (HFLG) block operation, a first low-frequency global-feature generation (LGFG) block operation, and a dual query transformer (DQT) operation; andgenerate a high-frequency local feature map based on the low-frequency information and the high-frequency information by performing a second residual block operation, a second HFLG block operation, a second LGFG block operation, and the DQT operation, the second feature map including the low-frequency global feature map and the high-frequency local feature map.12.The video decoder of claim 11, wherein, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, cause the processor to:perform a matrix multiplication of the first feature map and the spatial attention map to generate an intermediate feature map;generate the low-frequency information based on the intermediate feature map by performing a first convolutional layer operation; andgenerate the high-frequency information based on the low-frequency information by performing a max-pooling layer operation and a second convolutional layer operation.13.The video decoder of claim 12, wherein, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, cause the processor to:generate an attention map by multiplying a first query, a second query, and a key by a value and performing a summation as part of the DQT operation, the attention map being used to generate the low-frequency global feature map and the high-frequency local feature map.14.The video decoder of claim 10, wherein, to generate the second reference picture based on the reconstructed picture and the second feature map, the memory storing instructions, which when executed by the processor, cause the processor to:generate a final feature map based on the second feature map by performing a convolutional layer operation and a pixel-shuffle layer operation; andcombine the reconstructed picture associated with the first reference picture and the final feature map to generate the second reference picture.15.The video decoder of claim 10, wherein:the first information includes a luma reconstructed picture and a luma prediction picture associated with the first reference picture, andthe second information includes a luma partition map and a quantization parameters (QP) map.16.The video decoder of claim 10, wherein:the first information includes a luma reconstructed picture and a chroma reconstructed picture associated with the first reference picture, andthe second information includes a chroma prediction picture, a chroma partition map, and a quantization parameters (QP) map.17.The video decoder of claim 10, wherein the at least one MFFB operation includes more than one MFFB operation.18.The video decoder of claim 10, wherein the memory storing instructions, which when executed by the processor, cause the processor to:add the second reference picture into the decoded picture buffer.19.An apparatus for video decoding, comprising:a processor; andmemory storing instructions, which when executed by the processor, cause the processor to:obtain a first reference picture from a decoded picture buffer;generate a first feature map based on first information associated with the first reference picture, the first information including a reconstructed picture;generate a spatial attention map based on second information associated with the first reference picture;generate a second feature map based on the first feature map and the spatial attention map by performing at least one multi-type feature fusion block (MFFB) operation;generate a second reference picture based on the reconstructed picture and the second feature map; anddecode a current picture region based on the second reference picture.20.A non-transitory computer-readable medium storing instructions, which when executed by a processor of a video decoder, cause the processor of the video decoder to:obtain a first reference picture from a decoded picture buffer;generate a first feature map based on first information associated with the first reference picture, the first information including a reconstructed picture;generate a spatial attention map based on second information associated with the first reference picture;generate a second feature map based on the first feature map and the spatial attention map by performing at least one multi-type feature fusion block (MFFB) operation;generate a second reference picture based on the reconstructed picture and the second feature map; anddecode a current picture region based on the second reference picture.21.The non-transitory computer-readable medium of claim 20, wherein, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video decoder, cause the processor of the video decoder to:generate low-frequency information and high-frequency information based on the first feature map and the spatial attention map;generate a low-frequency global feature map based on the low-frequency information and the high-frequency information by performing a first residual block operation, a first high-frequency local feature generation (HFLG) block operation, a first low-frequency global-feature generation (LGFG) block operation, and a dual query transformer (DQT) operation; andgenerate a high-frequency local feature map based on the low-frequency information and the high-frequency information by performing a second residual block operation, a second HFLG block operation, a second LGFG block operation, and the DQT operation, the second feature map including the low-frequency global feature map and the high-frequency local feature map.22.The non-transitory computer-readable medium of claim 21, wherein, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video decoder, cause the processor of the video decoder to:perform a matrix multiplication of the first feature map and the spatial attention map to generate an intermediate feature map;generate the low-frequency information based on the intermediate feature map by performing a first convolutional layer operation; andgenerate the high-frequency information based on the low-frequency information by performing a max-pooling layer operation and a second convolutional layer operation.23.The non-transitory computer-readable medium of claim 22, wherein, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video decoder, cause the processor of the video decoder to:generate an attention map by multiplying a first query, a second query, and a key by a value and performing a summation as part of the DQT operation, the attention map being used to generate the low-frequency global feature map and the high-frequency local feature map.24.The non-transitory computer-readable medium of claim 20, wherein, to generate the second reference picture based on the reconstructed picture and the second feature map, the instructions, which when executed by the processor of the video decoder, cause the processor of the video decoder to:generate a final feature map based on the second feature map by performing a convolutional layer operation and a pixel-shuffle layer operation; andcombine the reconstructed picture associated with the first reference picture and the final feature map to generate the second reference picture.25.The non-transitory computer-readable medium of claim 20, wherein:the first information includes a luma reconstructed picture and a luma prediction picture associated with the first reference picture, andthe second information includes a luma partition map and a quantization parameters (QP) map.26.The non-transitory computer-readable medium of claim 20, wherein:the first information includes a luma reconstructed picture and a chroma reconstructed picture associated with the first reference picture, andthe second information includes a chroma prediction picture, a chroma partition map, and a quantization parameters (QP) map.27.The non-transitory computer-readable medium of claim 20, wherein the at least one MFFB operation includes more than one MFFB operation.28.The non-transitory computer-readable medium of claim 20, wherein the instructions, which when executed by the processor of the video decoder, cause the processor of the video decoder to:add the second reference picture into the decoded picture buffer.29.A method of video encoding, comprising:obtaining, by a processor, a first reference picture from a decoded picture buffer;generating, by the processor, a first feature map based on first information associated with the first reference picture, the first information including a reconstructed picture;generating, by the processor, a spatial attention map based on second information associated with the first reference picture;generating, by the processor, a second feature map based on the first feature map and the spatial attention map by performing at least one multi-type feature fusion block (MFFB) operation;generating, by the processor, a second reference picture based on the reconstructed picture and the second feature map; andencoding, by the processor, a current picture region based on the second reference picture.30.The method of claim 29, wherein the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation comprises:generating, by the processor, low-frequency information and high-frequency information based on the first feature map and the spatial attention map;generating, by the processor, a low-frequency global feature map based on the low-frequency information and the high-frequency information by performing a first residual block operation, a first high-frequency local feature generation (HFLG) block operation, a first low-frequency global-feature generation (LGFG) block operation, and a dual query transformer (DQT) operation; andgenerating, by the processor, a high-frequency local feature map based on the low-frequency information and the high-frequency information by performing a second residual block operation, a second HFLG block operation, a second LGFG block operation, and the DQT operation, the second feature map including the low-frequency global feature map and the high-frequency local feature map.31.The method of claim 30, wherein the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation further comprises:performing, by the processor, a matrix multiplication of the first feature map and the spatial attention map to generate an intermediate feature map;generating, by the processor, the low-frequency information based on the intermediate feature map by performing a first convolutional layer operation; andgenerating, by the processor, the high-frequency information based on the low-frequency information by performing a max-pooling layer operation and a second convolutional layer operation.32.The method of claim 31, wherein the generating, by the processor, the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation further comprises:generating, by the processor, an attention map by multiplying a first query, a second query, and a key by a value and performing a summation as part of the DQT operation, the attention map being used to generate the low-frequency global feature map and the high-frequency local feature map.33.The method of claim 29, wherein the generating, by the processor, the second reference picture based on the reconstructed picture and the second feature map comprises:generating, by the processor, a final feature map based on the second feature map by performing a convolutional layer operation and a pixel-shuffle layer operation; andcombining, by the processor, the reconstructed picture associated with the first reference picture and the final feature map to generate the second reference picture.34.The method of claim 29, wherein:the first information includes a luma reconstructed picture and a luma prediction picture associated with the first reference picture, andthe second information includes a luma partition map and a quantization parameters (QP) map.35.The method of claim 29, wherein:the first information includes a luma reconstructed picture and a chroma reconstructed picture associated with the first reference picture, andthe second information includes a chroma prediction picture, a chroma partition map, and a quantization parameters (QP) map.36.The method of claim 29, wherein the at least one MFFB operation includes more than one MFFB operation.37.The method of claim 29, further comprising:adding, by the processor, the second reference picture into the decoded picture buffer.38.A video encoder, comprising:a processor; andmemory storing instructions, which when executed by the processor, cause the processor to:obtain a first reference picture from a decoded picture buffer;generate a first feature map based on first information associated with the first reference picture, the first information including a reconstructed picture;generate a spatial attention map based on second information associated with the first reference picture;generate a second feature map based on the first feature map and the spatial attention map by performing at least one multi-type feature fusion block (MFFB) operation;generate a second reference picture based on the reconstructed picture and the second feature map; andencode a current picture region based on the second reference picture.39.The video encoder of claim 38, wherein, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, cause the processor to:generate low-frequency information and high-frequency information based on the first feature map and the spatial attention map;generate a low-frequency global feature map based on the low-frequency information and the high-frequency information by performing a first residual block operation, a first high-frequency local feature generation (HFLG) block operation, a first low-frequency global-feature generation (LGFG) block operation, and a dual query transformer (DQT) operation; andgenerate a high-frequency local feature map based on the low-frequency information and the high-frequency information by performing a second residual block operation, a second HFLG block operation, a second LGFG block operation, and the DQT operation, the second feature map including the low-frequency global feature map and the high-frequency local feature map.40.The video encoder of claim 39, wherein, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, cause the processor to:perform a matrix multiplication of the first feature map and the spatial attention map to generate an intermediate feature map;generate the low-frequency information based on the intermediate feature map by performing a first convolutional layer operation; andgenerate the high-frequency information based on the low-frequency information by performing a max-pooling layer operation and a second convolutional layer operation.41.The video encoder of claim 40, wherein, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the memory storing instructions, which when executed by the processor, cause the processor to:generate an attention map by multiplying a first query, a second query, and a key by a value and performing a summation as part of the DQT operation, the attention map being used to generate the low-frequency global feature map and the high-frequency local feature map.42.The video encoder of claim 38, wherein, to generate the second reference picture based on the reconstructed picture and the second feature map, the memory storing instructions, which when executed by the processor, cause the processor to:generate a final feature map based on the second feature map by performing a convolutional layer operation and a pixel-shuffle layer operation; andcombine the reconstructed picture associated with the first reference picture and the final feature map to generate the second reference picture.43.The video encoder of claim 38, wherein:the first information includes a luma reconstructed picture and a luma prediction picture associated with the first reference picture, andthe second information includes a luma partition map and a quantization parameters (QP) map.44.The video encoder of claim 38, wherein:the first information includes a luma reconstructed picture and a chroma reconstructed picture associated with the first reference picture, andthe second information includes a chroma prediction picture, a chroma partition map, and a quantization parameters (QP) map.45.The video encoder of claim 38, wherein the at least one MFFB operation includes more than one MFFB operation.46.The video encoder of claim 38, wherein the memory storing instructions, which when executed by the processor, cause the processor to:add the second reference picture into the decoded picture buffer.47.An apparatus for video encoding, comprising:a processor; andmemory storing instructions, which when executed by the processor, cause the processor to:obtain a first reference picture from a decoded picture buffer;generate a first feature map based on first information associated with the first reference picture, the first information including a reconstructed picture;generate a spatial attention map based on second information associated with the first reference picture;generate a second feature map based on the first feature map and the spatial attention map by performing at least one multi-type feature fusion block (MFFB) operation;generate a second reference picture based on the reconstructed picture and the second feature map; andencode a current picture region based on the second reference picture.48.A non-transitory computer-readable medium storing instructions, which when executed by a processor of a video encoder, cause the processor of the video encoder to:obtain a first reference picture from a decoded picture buffer;generate a first feature map based on first information associated with the first reference picture, the first information including a reconstructed picture;generate a spatial attention map based on second information associated with the first reference picture;generate a second feature map based on the first feature map and the spatial attention map by performing at least one multi-type feature fusion block (MFFB) operation;generate a second reference picture based on the reconstructed picture and the second feature map; andencode a current picture region based on the second reference picture.49.The non-transitory computer-readable medium of claim 48, wherein, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video encoder, cause the processor of the video encoder to:generate low-frequency information and high-frequency information based on the first feature map and the spatial attention map;generate a low-frequency global feature map based on the low-frequency information and the high-frequency information by performing a first residual block operation, a first high-frequency local feature generation (HFLG) block operation, a first low-frequency global-feature generation (LGFG) block operation, and a dual query transformer (DQT) operation; andgenerate a high-frequency local feature map based on the low-frequency information and the high-frequency information by performing a second residual block operation, a second HFLG block operation, a second LGFG block operation, and the DQT operation, the second feature map including the low-frequency global feature map and the high-frequency local feature map.50.The non-transitory computer-readable medium of claim 49, wherein, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video encoder, cause the processor of the video encoder to:perform a matrix multiplication of the first feature map and the spatial attention map to generate an intermediate feature map;generate the low-frequency information based on the intermediate feature map by performing a first convolutional layer operation; andgenerate the high-frequency information based on the low-frequency information by performing a max-pooling layer operation and a second convolutional layer operation.51.The non-transitory computer-readable medium of claim 50, wherein, to generate the second feature map based on the first feature map and the spatial attention map by performing at least one MFFB operation, the instructions, which when executed by the processor of the video encoder, cause the processor of the video encoder to:generate an attention map by multiplying a first query, a second query, and a key by a value and performing a summation as part of the DQT operation, the attention map being used to generate the low-frequency global feature map and the high-frequency local feature map.52.The non-transitory computer-readable medium of claim 48, wherein, to generate the second reference picture based on the reconstructed picture and the second feature map, the instructions, which when executed by the processor of the video encoder, cause the processor of the video encoder to:generate a final feature map based on the second feature map by performing a convolutional layer operation and a pixel-shuffle layer operation; andcombine the reconstructed picture associated with the first reference picture and the final feature map to generate the second reference picture.53.The non-transitory computer-readable medium of claim 48, wherein:the first information includes a luma reconstructed picture and a luma prediction picture associated with the first reference picture, andthe second information includes a luma partition map and a quantization parameters (QP) map.54.The non-transitory computer-readable medium of claim 48, wherein:the first information includes a luma reconstructed picture and a chroma reconstructed picture associated with the first reference picture, andthe second information includes a chroma prediction picture, a chroma partition map, and a quantization parameters (QP) map.55.The non-transitory computer-readable medium of claim 48, wherein the at least one MFFB operation includes more than one MFFB operation.56.The non-transitory computer-readable medium of claim 48, wherein the instructions, which when executed by the processor of the video encoder, cause the processor of the video encoder to:add the second reference picture into the decoded picture buffer.57.A non-transitory computer-readable medium storing a bitstream, the bitstream being generated based on one or more of claims 29-37.
Citation Information
Patent Citations
LDCT artifact suppression method based on fusion of edge prior and CNN-Transform
CN117036198A
Chroma quantization parameter (QP) derivation for video coding
US20210058620A1
Quantization parameter signaling in video processing
US20210092460A1
Method and apparatus for determining chroma quantization parameters when using separate coding trees for luma and chroma
US20220038704A1
Method and apparatus for signaling of mapping function of chroma quantization parameter
US20220070460A1