In-loop filter based on neural network
By introducing a neural network-based in-loop filter in video encoding, combining residual attention blocks and converter blocks, using auxiliary information adaptive selection features, the problems of compression artifacts and information loss in video encoding are solved, and higher quality video reconstruction is achieved.
Patent Information
- Application Number
- CN202380089820.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-01-03
- Filing Date
- 2023-03-03
- Publication Date
- 2025-08-12
Smart Images

Figure CN120476420A_ABST
Abstract
Description
Background Art
[0001] Embodiments of the present disclosure relate to video encoding.
[0002] Digital video has become mainstream and is widely used in various applications including digital television, video telephony, and teleconferencing. These digital video applications are feasible due to advances in computing and communication technology and efficient video coding techniques. Various video coding techniques can be used to compress video data so that encoding of the video data can be performed using one or more video coding standards. For example, exemplary video coding standards may include, but are not limited to, Versatile Video Coding (H.266 / VVC), High Efficiency Video Coding (H.265 / HEVC), Advanced Video Coding (H.264 / AVC), and Moving Picture Experts Group (MPEG) coding. Summary of the Invention
[0003] According to one aspect of the present disclosure, a method for enhancing the quality of a frame is provided. A processor receives the frame and auxiliary information associated with the frame. A neural network (NN)-based in-loop filter is applied to the frame based on the auxiliary information to enhance the quality of the frame. The NN-based in-loop filter includes three parts: feature extraction, backbone, and reconstruction. The backbone part includes a residual-attention block (RAB) and a transformer block, and an attention block is used in the RAB to better refine the features by introducing auxiliary information.
[0004] According to another aspect of the present disclosure, a system for enhancing frame quality includes a memory configured to store instructions and a processor coupled to the memory. The processor is configured to, when executing the instructions, receive a frame and auxiliary information associated with the frame and apply a NN-based in-loop filter to the frame based on the auxiliary information to enhance the frame quality. The NN-based in-loop filter includes a backbone comprising a RAB and a transformer block, and employs an attention block within the RAB to better refine features by incorporating the auxiliary information.
[0005] According to another aspect of the present disclosure, a tangible computer-readable device is provided. The tangible computer-readable device has instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform a plurality of operations. The plurality of operations include receiving a frame and auxiliary information associated with the frame, and applying a NN-based in-loop filter to the frame based on the auxiliary information to enhance the quality of the frame. The NN-based in-loop filter includes a backbone portion comprising an RAB and a transformer block, and employing an attention block in the RAB to better refine features by introducing auxiliary information.
[0006] These illustrative embodiments are mentioned not to limit or define the present disclosure, but to provide examples to aid understanding of the present disclosure. Additional embodiments are described in the detailed description, and further description is provided in the detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] The accompanying drawings, which are incorporated herein and form a part of this specification, illustrate embodiments of the present disclosure and, together with the description, further serve to explain the principles of the present disclosure and enable one skilled in the relevant art to make and use the present disclosure.
[0008] Figure 1 A block diagram of a video encoder with an in-loop filter according to some embodiments of the present disclosure is shown.
[0009] Figure 2 A block diagram of an exemplary encoding system according to some embodiments of the present disclosure is shown.
[0010] Figure 3 A block diagram of an exemplary decoding system according to some embodiments of the present disclosure is shown.
[0011] Figure 4 Some embodiments of the present disclosure are shown Figure 1 Block diagram of an exemplary NN-based in-loop filter for a video encoder.
[0012] Figure 5A Some embodiments of the present disclosure are shown Figure 4 Detailed block diagram of an exemplary feature extraction portion of an NN-based in-loop filter for the luminance component.
[0013] Figure 5B Some embodiments of the present disclosure are shown Figure 4 Detailed block diagram of another exemplary feature extraction portion of the NN-based in-loop filter for chroma components.
[0014] Figure 6A Some embodiments of the present disclosure are shown Figure 4Detailed block diagram of an exemplary reconstruction portion of a NN-based in-loop filter for the luma component.
[0015] Figure 6B Some embodiments of the present disclosure are shown Figure 4 Detailed block diagram of another exemplary reconstruction portion of the NN-based in-loop filter for the chroma component.
[0016] Figure 7A Some embodiments of the present disclosure are shown Figure 4 Detailed block diagram of an exemplary backbone portion of a NN-based in-loop filter.
[0017] Figure 7B Some embodiments of the present disclosure are shown Figure 4 Detailed block diagram of another exemplary backbone portion of a NN-based in-loop filter.
[0018] Figure 8 Some embodiments of the present disclosure are shown Figure 7A and Figure 7B Detailed block diagram of an exemplary residual attention block (RAB) for the backbone part of .
[0019] Figure 9 Some embodiments of the present disclosure are shown Figure 8 Detailed block diagram of an exemplary residual block of RAB.
[0020] Figure 10 Some embodiments of the present disclosure are shown Figure 8 Detailed block diagram of an exemplary attention block of RAB.
[0021] Figure 11 Some embodiments of the present disclosure are shown Figure 10 Detailed block diagram of an exemplary channel attention block of the attention block.
[0022] Figure 12A Some embodiments of the present disclosure are shown Figure 11 Detailed block diagram of an exemplary strength channel attention block.
[0023] Figure 12B Some embodiments of the present disclosure are shown Figure 11 Detailed block diagram of an exemplary contrast channel attention block for .
[0024] Figure 13 Some embodiments of the present disclosure are shown Figure 10 Detailed block diagram of an exemplary spatial attention block of the attention block.
[0025] Figure 14Some embodiments of the present disclosure are shown Figure 7A and Figure 7B Detailed block diagram of an exemplary transformer block (TB) of the backbone portion of FIG.
[0026] Figure 15 A flow chart illustrating an exemplary method of enhancing the quality of a frame according to some embodiments of the present disclosure is shown.
[0027] Figure 16 A block diagram of an exemplary in-loop filter for training according to some embodiments of the present disclosure is shown.
[0028] Figure 17A A block diagram of the in-loop filter training strategy is shown.
[0029] Figure 17B A block diagram illustrating an exemplary in-loop filter training strategy according to some embodiments of the present disclosure is shown.
[0030] Figure 18 Shown Figure 4 Comparison of visual quality of NN-based in-loop filters on an exemplary test dataset.
[0031] Figures 19A-19C The use of some embodiments according to the present disclosure is shown Figure 4 The NN-based in-loop filter is used Figure 18 The test results of the relationship between peak signal-to-noise ratio (PSNR) and bit rate for video encoding using a test dataset.
[0032] Figure 20 The training according to some embodiments of the present disclosure is shown Figure 4 Flowchart of an exemplary method for a NN-based in-loop filter.
[0033] Figure 21 Compression artifacts with different quantization parameters (QP) are shown.
[0034] Embodiments of the present disclosure will be described with reference to the accompanying drawings. DETAILED DESCRIPTION
[0035] Although some configurations and arrangements have been discussed, it should be understood that this is for illustrative purposes only. Those skilled in the relevant art will recognize that other configurations and arrangements may be used without departing from the spirit and scope of the present disclosure. It will be apparent to those skilled in the relevant art that the present disclosure may also be used for various other applications.
[0036] Note that references in the specification to "one embodiment," "an embodiment," "example embodiment," "some embodiments," "certain embodiments," etc., indicate that the described embodiments may include a particular feature, structure, or characteristic, but not every embodiment necessarily includes that particular feature, structure, or characteristic. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in conjunction with an embodiment, it is within the knowledge of persons skilled in the relevant art to implement such feature, structure, or characteristic in conjunction with other embodiments, whether or not explicitly described.
[0037] In general, terms can be understood, at least in part, from their use in context. For example, the term "one or more" as used herein may be used to describe any feature, structure, or characteristic in the singular, or may be used to describe a combination of features, structures, or characteristics in the plural, depending, at least in part, on the context. Similarly, terms such as "a," "an," or "the" may also be understood to convey singular usage or to convey plural usage, depending, at least in part, on the context. Furthermore, the term "based on" may be understood to not necessarily be intended to convey an exclusive set of factors, but rather may allow for the presence of additional factors that are not necessarily explicitly described, again depending, at least in part, on the context.
[0038] Various aspects of the video coding system will now be described with reference to various apparatuses and methods. These apparatuses and methods are described in the detailed description below and illustrated in the accompanying drawings by various modules, components, circuits, steps, operations, processes, algorithms, etc. (collectively, "elements"). These elements can be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether these elements are implemented as hardware, firmware, or software depends on the specific application and the design constraints imposed on the overall system.
[0039] The techniques described herein can be used in various video coding applications. As described herein, video coding includes both encoding and decoding of video. Encoding and decoding of video can be performed in units of blocks. For example, encoding / decoding processes such as transform, quantization, prediction, in-loop filtering, or reconstruction can be performed on a coding block, a transform block, or a prediction block. As described herein, a block to be encoded / decoded is referred to as a "current block." For example, depending on the current encoding / decoding process, the current block can represent a coding block, a transform block, or a prediction block. In addition, it should be understood that the term "unit" used in this disclosure refers to a basic unit for performing a specific encoding / decoding process, and the term "block" refers to a sample array of a predetermined size. Unless otherwise specified, "block", "unit", "portion", and "component" can be used interchangeably.
[0040] For existing video compression methods (e.g., HEVC and VVC), blocking and quantization are performed during encoding, resulting in irreversible information loss and various compression artifacts such as blocking artifacts, blurring artifacts, and banding artifacts. This phenomenon is particularly noticeable when the compression ratio, such as represented by the quantization parameter (QP), is high. For example, in Figure 21 , compression artifacts with different QPs are shown, for example, (a) QP = 22, (b) QP = 27, (c) QP = 32, (d) QP = 37, and (e) QP = 42. As QP increases, the artifacts become more severe and the image quality deteriorates.
[0041] Currently, there are many deep learning-based methods for improving the quality of compressed images and videos, primarily for reducing blocking artifacts, banding artifacts, and noise. In one example, Versatile Video Coding (VVC) employs in-loop filters in the encoder to suppress compression artifacts and reduce distortion. To name a few, these in-loop filters can include deblocking filters (DBF), sample adaptive offset (SAO), and adaptive loop filters (ALF). DBF and SAO are two filters designed to reduce artifacts caused by the encoding process. DBF focuses on visual artifacts at block boundaries, while SAO supplementally reduces artifacts that may be caused by quantization of transform coefficients within a block. ALF can enhance adaptive filtering of the reconstructed signal, using a Wiener-based adaptive filter to reduce the mean square error (MSE) between the original and reconstructed samples. Although these filters significantly reduce compression artifacts, they are handcrafted and developed based on signal processing theory that assumes signal stability. Since natural video sequences are often non-stationary, their performance is limited. Therefore, there is still much room for improvement in the loop filter of VVC.
[0042] With the development of deep learning, various image and video quality enhancement methods based on convolutional neural networks (CNNs) have emerged. Recently, some video encoders have been designed with neural network (NN)-based in-loop filters, which include trained CNN filters embedded in the VVC loop. This can be achieved by inserting loop filter components or replacing some loop filter components.
[0043] Figure 1A block diagram of a video encoder 100 is shown having an in-loop filter 114 with a NN-based in-loop filter (CNN ILF) 122, according to some embodiments of the present disclosure. To name a few, the video encoder 100 may include, for example, a video sequence component 102, a transform component 104, a quantization component 106, an inverse quantization component 108, an inverse transform component 110, an encoding component 112, an in-loop filter 114, a decoded picture buffer 126, an inter-frame prediction component 128, and an intra-frame prediction component 130. The in-loop filter 114 may include, for example, a lummapping with chroma scaling (LMCS) 116, a DBF 118, a SAO 120, a CNN ILF 122, and an ALF 124.
[0044] Some video encoders use in-loop filters based on QP-variable NNs for VVC intra-frame coding. To avoid training and deployment in multiple networks, these encoders use a QP attention module (QPAM), which can capture the compression noise level of different QPs and highlight meaningful features along the channel dimension. The QPAM can be embedded in the residual block as part of the network architecture, which is designed for controllability of different QPs. To fine-tune the network, these video encoders can use a focal mean squared error (MSE) loss function.
[0045] However, these video encoders do not use additional auxiliary information (also known as side information) as input; therefore, the performance of this network is limited. At the same time, directly fusing shallow features into the backbone part of the network (also known as the backend) limits the performance of the model; therefore, the performance of this network further degrades.
[0046] In other video encoders, in-loop filters based on dense residual convolutional neural networks (DRNs) can be used for VVC. These video encoders use residual learning components, dense shortcuts, and bottleneck layers to solve the gradient vanishing problem, promote feature reuse, and reduce computational resources. Unfortunately, the performance of these video encoders cannot achieve a good balance between complexity and performance.
[0047] In some other existing video encoders, NN-based filters can be used to enhance the quality of VVC intra-frame coded frames by taking auxiliary information such as segmentation and prediction information as input. For chrominance, the auxiliary information also includes luma samples. Although the filter achieves sufficient performance on the Y channel, the performance on other channels (e.g., the U and V channels of luma) is relatively low, and has undesirable coding delay and high complexity.
[0048] To overcome these and other challenges, the present disclosure provides an improved NN-based in-loop filter (e.g., Figure 1 The CNN ILF 122 in FIG1 is a CNN network that is constructed from a neural network. The in-loop filter is based on the residual block (resblock) and the transformer block in the backbone portion of the CNN network. Compared to other NN-based in-loop filters, the NN-based in-loop filter 122 can achieve improved performance with lower computational complexity.
[0049] According to some aspects of the present invention, the NN-based in-loop filter 122 is applicable to various image sample and slice types, such as two sample types (luminance and chrominance) and two slice types (I slices and B slices). Depending on the different types of samples and slices, different feature extraction networks, backbone networks, and reconstruction networks can be designed to enhance the applicability of the NN-based in-loop filter 122.
[0050] According to some aspects of the present disclosure, the backbone of the NN-based in-loop filter 122 uses residual blocks and transformer blocks to extract and process features. For example, a residual attention block can be used to first extract shallow features of the input and collect correlations between local features. Then, a transformer block can be used to collect long-range correlations between features, thereby enabling the NN-based in-loop filter disclosed herein to obtain useful residual features.
[0051] According to some aspects of the present disclosure, auxiliary information such as partition maps, prediction maps, and QP maps of the reconstructed frame is used to guide the NN-based in-loop filter 122 to adaptively select and refine important features. For example, for the spatial attention block, the partition map and QP map can be introduced to better locate areas with blocking artifacts and distortion. For the channel attention block, the QP map can be introduced to combine intensity attention and contrast attention.
[0052] Consistent with the scope of the present disclosure, the NN-based in-loop filter 122 is based on a residual block and a transformer block, which are used to work with the DBF and SAO within the in-loop filter 114 to improve the objective quality of the reconstructed frame. In some embodiments, four models for different types of inputs are introduced: a luminance model for I slices, a chrominance model for I slices, a luminance model for B slices, and a chrominance model for B slices, and these four models can be processed by different designs of the NN-based in-loop filter 122. In some embodiments, in terms of network architecture, the residual block and the transformer block are combined as a backbone network to achieve better performance with acceptable complexity. In some embodiments, more auxiliary information (e.g., partition map and QP map) is introduced into the attention block to achieve effective feature refinement. In some embodiments, multi-stage progressive training and iterative training are used to maximize the learning ability of the proposed network. Therefore, the NN-based in-loop filter 122 can achieve a good balance between complexity and performance. The following is combined with Figures 2 to 20 More details are provided on an exemplary training strategy for the NN-based in-loop filter 122 and its model.
[0053] It should be understood that although the NN-based in-loop filter 122 is primarily used to enhance the quality of compressed images, the CNN used in the NN-based in-loop filter 122 can also be used as post-processing after decoding to enhance the quality of decoded frames. In addition, if the downsampling convolution in feature extraction is removed, the CNN used in the NN-based in-loop filter 122 can also be used for image super-resolution.
[0054] Figure 2 A block diagram of an exemplary encoding system 200 is shown, according to some embodiments of the present disclosure. Figure 3 A block diagram of an exemplary decoding system 300 according to some embodiments of the present disclosure is shown. Each system 200 or 300 can be applied to or integrated into various systems and devices capable of data processing, such as computers and wireless communication devices. For example, the system 200 or 300 can be the whole or part of the following items: a mobile phone; a desktop computer; a laptop computer; a tablet computer, a car computer; a game console; a printer; a positioning device; a wearable electronic device; a smart sensor; a virtual reality (VR) device, an augmented reality (AR) device; or any other suitable electronic device with data processing capabilities. Figure 2 and Figure 3As shown, system 200 or 300 may include processor 202, memory 204, and interface 206. These components are shown as being interconnected via a bus, but other connection types are also permitted. It should be understood that system 200 or 300 may include any other suitable components for performing the functions described herein.
[0055] The processor 202 may include a microprocessor, such as a graphics processing unit (GPU), an image signal processor (ISP), a central processing unit (CPU), a digital signal processor (DSP), a tensor processing unit (TPU), a vision processing unit (VPU), a neural processing unit (NPU), a synergistic processing unit (SPU) or a physics processing unit (PPU), a microcontroller unit (MCU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device (PLD), a state machine, gating logic, discrete hardware circuits, and other suitable hardware configured to perform the various functions described throughout the present disclosure. Although Figure 2 and Figure 3 Only one processor is shown, but it should be understood that multiple processors may be included. Processor 202 may be a hardware device having one or more processing cores. Processor 202 may execute software. Software should be broadly interpreted as instructions, instruction sets, codes, code segments, program codes, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executable programs, execution threads, processes, functions, etc. Whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise, software may include computer instructions written in an interpreted language, compiled language, or machine code. Under the broad category of software, other technologies for indicating hardware are also permitted.
[0056] The memory 204 may broadly include memory (also referred to as main memory / system memory) and storage devices (also referred to as secondary memory). For example, the memory 204 may include random-access memory (RAM), read-only memory (ROM), static RAM (SRAM), dynamic RAM (DRAM), ferroelectric RAM (FRAM), electrically erasable programmable ROM (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage devices, hard disk drives (HDD) such as magnetic disk storage devices or other magnetic storage devices, flash drives, solid-state drives (SSD), or any other medium that can be used to carry or store the required program code in the form of instructions that can be accessed and executed by the processor 202. Broadly speaking, the memory 204 may be embodied by any computer-readable medium, such as a non-transitory computer-readable medium. Although in Figure 2 and Figure 3 Only one memory is shown in the figure, but it should be understood that multiple memories may be included.
[0057] The interface 206 may broadly include a data interface and a communication interface configured to receive and send signals in the process of receiving and sending information with other external network elements. For example, the interface 206 may include input / output (I / O) devices and wired or wireless transceivers. Figure 2 and Figure 3 Only one memory is shown, but it should be understood that multiple interfaces may be included.
[0058] The processor 202, memory 204, and interface 206 can be implemented in various forms in the system 200 or 300 to perform video encoding functions. In some embodiments, the processor 202, memory 204, and interface 206 of the system 200 or 300 are implemented on one or more system-on-chip (SOC) (e.g., integrated on one or more SOCs). In one example, the processor 202, memory 204, and interface 206 can be integrated on an application processor (AP) SoC, which manipulates application processing in an operating system (OS) environment, including running video encoding and decoding applications. In another example, the processor 202, memory 204, and interface 206 can be integrated on a dedicated processor chip for video encoding, such as a GPU or ISP chip dedicated to image and video processing in a real-time operating system (RTOS).
[0059] like Figure 2 As shown, in the encoding system 200, the processor 202 may include one or more modules, such as an encoder 201 (also referred to as a pre-processing network). Figure 2 The encoder 201 is shown within a processor 202, but it should be understood that the encoder 201 may include one or more submodules that can be implemented on different processors that are close to or far away from each other. The encoder 201 (and any corresponding submodules or subunits) can be a hardware unit (e.g., part of an integrated circuit) of the processor 202, which is designed to be used with other components or software units implemented by the processor 202 by executing at least a portion of a program (i.e., instructions). The instructions of the program can be stored on a computer-readable medium such as the memory 204, and when these instructions are executed by the processor 202, they can perform a process having one or more functions related to video coding, such as picture segmentation, inter-frame prediction, intra-frame prediction, transform, quantization, filtering, entropy coding, etc., as described in detail below.
[0060] Similarly, if Figure 3 As shown, in the decoding system 300, the processor 202 may include one or more modules, such as a decoder 301 (also referred to as a post-processing network). Figure 3The decoder 301 is shown within a processor 202, but it should be understood that the decoder 301 may include one or more submodules that can be implemented on different processors that are close to or far away from each other. The decoder 301 (and any corresponding submodules or subunits) can be a hardware unit of the processor 202 (e.g., part of an integrated circuit) that is designed to be used with other components or software units implemented by the processor 202 by executing at least a portion of a program (i.e., instructions). The instructions of the program can be stored on a computer-readable medium such as the memory 204, and when these instructions are executed by the processor 202, a process having one or more functions related to video decoding can be performed, such as entropy decoding, inverse quantization, inverse transform, inter-frame prediction, intra-frame prediction, filtering, as described in detail below.
[0061] Return Reference Figure 2 , the encoder 201 may be a video encoder 100, the video encoder 100 including the following combination Figures 4 to 20 The described NN-based in-loop CNN filter 122. Figure 4 Some embodiments of the present disclosure are shown Figure 1 1. Block diagram of the NN-based in-loop filter 122 of the video encoder 100.
[0062] refer to Figure 4 , the NN-based in-loop filter 122 may include a feature extraction portion 407, a backbone portion 403, and a reconstruction portion 405. Inputs to the NN-based in-loop filter 122 may include, for example, a reconstructed frame (rec) 401, a prediction map (pred) 402a associated with the reconstructed frame 401, a partition map (par) 402b associated with the reconstructed frame 401, and a QP map (qp) 402c associated with the reconstructed frame 401, each of which is generated by any suitable component of the video encoder 100. The feature extraction portion 407 is configured to extract features from the reconstructed frame 401, the prediction map 402a, the partition map 402b, and the QP map 402c. The backbone portion 403 is configured to process these features based on the RAB 416 and the transformer block 417 to generate global features of the input features 502. The reconstruction part 405 is configured to reconstruct a residual map based on the global features and add the residual map to the reconstructed frame 401 to generate an enhanced reconstructed frame (rec) 418 .
[0063] Reconstructed frame 401 is a reconstruction of the current video frame by video encoder 100 for quality enhancement. Information other than reconstructed frame 401 (e.g., prediction map 402a, partition map 402b, and QP map 402c) is considered auxiliary information in this disclosure. This auxiliary information can help NN-based in-loop filter 122 better enhance the quality of reconstructed frame 401. Prediction map 402a contains prediction information for reconstructed frame 401 generated by video encoder 100. Since prediction map 402a is generated from adjacent frames, using prediction map 402a is equivalent to introducing temporal information into NN-based in-loop filter 122, thereby helping NN-based in-loop filter 122 better locate and process artifacts in moving areas. Partition map 402b represents block information of reconstructed frame 401 and can therefore effectively help NN-based in-loop filter 122 estimate blocky areas, thereby removing blocky artifacts in reconstructed frame 401. QP map 402c is used to indicate the QP value used for reconstructing frame 401. The QP map 402 c is related to the distortion of the reconstructed frame 401 and can improve the quality of the reconstructed frame at different QPs at the reconstructed portion 405 .
[0064] The feature extraction portion 407 may be implemented in various embodiments based on different types of input reconstructed frames 401 (e.g., having different image samples and / or slice types). In one example where the reconstructed frame 401 is a reconstructed frame of luminance samples (e.g., in the Y channel), Figure 5A Some embodiments of the present invention are shown Figure 4 Detailed block diagram of the feature extraction portion 407 of the NN-based in-loop filter 122.
[0065] for Figure 5A The input auxiliary information includes a prediction map 402a, a partition map 402b and a QP map 402c, all of which are associated with the reconstructed frame 401 of luma samples. Figure 5A The feature extraction portion 407 in FIG. 4 is configured to extract features 502 (eg, feature maps) from the reconstructed frame 401 of luma samples, the prediction map 402a, the partition map 402b, and the QP map 402c. In some embodiments, Figure 5AThe reconstruction part 405 in includes a plurality of parallel standard convolutional layers (Conv(x,y)), and each parallel convolutional layer is used to integrate and extract shallow features of its corresponding input. The (x,y) of the convolutional layer represents the number of input channels and output channels of each convolutional layer, and can vary for different inputs, for example, based on the features extracted from each input. Each parallel convolutional layer can be followed by a corresponding parameter rectified linear unit (PReLU). Afterwards, the shallow features from different PReLUs can be concatenated using a concatenation layer (Concat) and fused using a convolutional layer with 64 output channels (Conv(64)) to obtain fused shallow features. The fused shallow features can then be downsampled using a convolutional layer with a stride of 2 (Conv(s=2)) to reduce computational complexity and output features 502 (also called intermediate features).
[0066] In another example where the reconstructed frame 401 is a reconstructed frame of chroma samples (eg, in the U channel and the V channel), Figure 5B Some embodiments of the present disclosure are shown Figure 4 Detailed block diagram of the feature extraction portion 407 of the NN-based in-loop filter 122.
[0067] for Figure 5A The input auxiliary information includes a prediction map 402a, a partition map 402b and a QP map 402c, all of which are associated with the reconstructed frame 401 of chroma samples. Figure 5B The input to the feature extraction section 407 in
[15] also includes the reconstructed frame 503 of luma samples (reconstructed frame Y↓ (rec Y↓)). In other words, the Y channel luma information of the same frame can be used to help extract features from chroma samples. Because luma samples contain more information, they can be used to provide more accurate structural and texture information for feature extraction of chroma samples, thereby further improving the quality of chroma samples. It should be understood that, in general, the useful information in luma samples is related to the QP value. For example, for a higher QP, the feature extraction section 407 may require more structural information, while for a lower QP, more detailed texture information may be more desirable.
[0068] Figure 5B The feature extraction part 407 in FIG. 4 is configured to extract features 502 from the reconstructed frame 401 of chroma samples based on the prediction map 402a, the partition map 402b, the QP map 402c, and the reconstructed frame 503 of luma samples. Figure 5A The examples in are different because the types of input information are different. Figure 5BA progressive fusion method is used to obtain more accurate fused features. In some embodiments, since chroma samples have strong correlation, they should be fused first. For example, the shallow features of the reconstructed frame 401, the prediction map 402a, and the partition map 402b from the chroma samples can be fused first to obtain chroma-related fused shallow features. Thereafter, the shallow features of the reconstructed frame 503 from the luma samples and the QP map 402c can be fused with the chroma-related fused shallow features using concatenation and convolution to obtain the final fused shallow features. Finally, the final fused shallow features can be downsampled using convolutional layers and PReLU to reduce computational complexity and output features 502 (also called intermediate features).
[0069] Figure 5A and Figure 5B The above example can be used for I slice input. It should be understood that in some examples, for B slice input, the partition map 402b may not be used as input auxiliary information for the feature extraction part 407.
[0070] The reconstruction portion 405 may also be implemented in various embodiments based on different types of input reconstruction frames 401 (e.g., having different image samples and / or slice types). In one example where the reconstruction frame 401 is a reconstruction frame of luminance samples (e.g., in the Y channel), Figure 6A Some embodiments of the present disclosure are shown Figure 4 Detailed block diagram of the reconstruction portion 405 of the NN-based in-loop filter 122. Figure 6A The reconstruction part 405 in the image can receive the global features (gl fea) 602 from the backbone part 403 and use a 64×4 convolution layer (Conv(64,4)) to reduce the channel dimension of the global features 602. Then, the reconstruction part 405 can use a pixel shuffle (PS) layer to upsample the reduced dimensionality features to obtain a residual map. Finally, the residual map can be obtained by the summation layer ( Figure 4 As shown), the obtained residual map is added to the reconstructed frame 401, thereby generating an enhanced reconstructed frame 418.
[0071] In another example where the reconstructed frame 401 is a reconstructed frame of chroma samples (eg, in the U channel and the V channel), Figure 6B Some embodiments of the present disclosure are shown Figure 4 Detailed block diagram of the reconstruction portion 405 of the NN-based in-loop filter 122. Figure 6B The reconstruction part 405 in the image can receive the global feature (gl fea) 602 from the backbone part 403 and use a 64×2 convolution layer (Conv(64,2)) to reduce the channel dimension of the global feature 602 to obtain a residual map. Finally, the residual map can be obtained by the summation layer ( Figure 4As shown), the obtained residual map is added to the reconstructed frame 401, thereby generating an enhanced reconstructed frame 418.
[0072] Return Reference Figure 4 , consistent with the scope of the present disclosure, the backbone portion 403 of the NN-based in-loop filter 122 includes at least one RAB 416 and at least one TB 417. The number of RABs 416 and TBs in the NN-based in-loop filter 122 may vary in different embodiments, for example, based on different types of input reconstructed frames 401 (e.g., having different image samples and / or slice types). In one example, where the reconstructed frame 401 is a reconstructed frame of luma samples (e.g., in the Y channel), the backbone portion 403 may include three RABs (RAB×3) 416 and six TBs (TB×6) 417, as shown in FIG. Figure 7A As shown in the figure, three RABs 416 and six TBs 417 can be configured to handle Figure 5A The feature extraction part 407 in the embodiment re-extracts the features 502 from the reconstructed frame of luma samples to obtain the global features 602. In another example where the reconstructed frame 401 is a reconstructed frame of chroma samples (e.g., in the U channel and the V channel), the backbone part 403 may include one RAB (RAB×1) 416 and three TBs (TB×3) 417, as shown in FIG. Figure 7B One RAB 416 and three TBs 417 can be configured to handle Figure 5B The feature extraction portion 407 in FIG. 5 extracts features 502 from the reconstructed frame of chroma samples to obtain global features 602. It should be understood that in other examples, the number of RABs 416 and TBs 417 may vary.
[0073] RAB 416 and TB 417 are configured to extract and process intermediate features. In some embodiments, RAB 416 is configured to process features 502 to obtain local correlations between features 502. For example, RAB 416 can be used to extract shallow features of input information and collect correlations between local features. However, considering only local correlations between features may not be enough. In a video frame, similar patterns often repeat. It is important to calculate the response of a single position by weighting all non-locally correlated positions. Therefore, in some embodiments, TB 417 is configured to process features 502 to obtain long-range correlations between features 502. For example, TB 417 can be used to collect long-range correlations between features, thereby helping the NN-based in-loop filter 122 to obtain more effective residual features.
[0074] Figure 14 Some embodiments of the present disclosure are shown Figure 7A and Figure 7BDetailed block diagram of TB 417 of the backbone portion 403. Figure 14 As shown, TB 417 includes a normalization layer (LayerNorm), a multi-head attention network (self-attention) that performs spatially rich feature interactions across channels, and a feed-forward network (Feed-froward) for controlled feature transformation, that is, the feed-forward network allows useful information to be further propagated so that it can capture long-range pixel interactions while still remaining applicable to large images.
[0075] Figure 8 Some embodiments of the present disclosure are shown Figure 7A and Figure 7B Detailed block diagram of RAB 416 of backbone portion 403. RAB 416 includes at least one residual block (ResBlock) 802 (e.g., Figure 8 ) and at least one attention block (AttBlock) 804 (e.g., in Figure 8 (one in the middle). Figure 9 Some embodiments of the present disclosure are shown Figure 8 Detailed block diagram of the residual block 802 of the RAB 416. Figure 9 As shown, the residual block 802 may include a convolutional layer (Conv(3×3)), a PReLU, and then another convolutional layer (Conv(3×3)). The result after the second convolutional layer may be added to the input of the first convolutional layer as the output of the residual block 802.
[0076] Figure 10 Some embodiments of the present disclosure are shown Figure 8 Detailed block diagram of the attention block 804 of the RAB 416. Figure 10 As shown, in addition to the features 502, the attention block 804 also receives at least some auxiliary information 402 (e.g., partition map 402b and QP map 402c), which can guide the attention block 804 to adaptively select and refine important features. The attention block 804 includes a spatial attention block (SA) 1004 and a channel attention block (CA) 1002, and the outputs of the spatial attention block 1004 and the channel attention block 1002 are combined into the output of the attention block 804 by summing. Through these two branches, with the help of the auxiliary information 402, the features 502 retain important feature information in the channel dimension and the spatial dimension respectively, and then the two parts of information are fused to obtain the final output.
[0077] Figure 11 Some embodiments of the present disclosure are shown Figure 10Detailed block diagram of the channel attention block 1002 of the attention block 804 of FIG. The channel attention block 1002 is configured to combine the intensity attention and contrast attention of the reconstructed frame 401 based on the QP map 402c. The intensity attention and contrast attention can be processed by the intensity channel attention block (Intensity) 1102 and the contrast channel attention block (Contrast) 1104, respectively. Figure 12A and Figure 12B As further shown. Figure 12A As shown, the intensity channel attention block 1102 extracts the weight of each channel through global average pooling (Avgpool) and channel compression and expansion using multiple convolutional layers (Conv(x,y)), ReLus layers, and sigmoid layers. The extracted weight is then multiplied (Mul) with the input feature 502 to obtain a channel attention map.
[0078] like Figure 12B As shown, the main difference between the contrast channel attention block 1104 and the intensity channel attention block 1102 is that the input is the sum of the mean and variance of the feature 502, rather than the result of the global average pooling. The calculation process of the mean and variance is shown in the following formula:
[0079] as well as
[0080]
[0081] In high-level domains, the importance of feature maps depends on the high-value areas of activation, because global average pooling is used to collect global information. Although average pooling can indeed improve the peak signal-to-noise ratio (PSNR) value, it lacks information about structure, texture, and edges that is conducive to enhancing image details (related to the structural similarity index measure (SSIM)). Therefore, according to some embodiments, the contrast channel attention block 1104 supplements the intensity channel attention block 1102 by replacing the global average pooling sum of the standard deviation and the mean (to evaluate the contrast of the feature map).
[0082] Return Reference Figure 11 , a QP map 402c is introduced to fuse the results of the intensity channel attention block 1102 and the contrast channel attention block 1104 using weighted summation. This is mainly due to the observation that the NN-based in-loop filter 122 tends to focus on structural features of larger QP inputs, while the NN-based in-loop filter 122 tends to focus on texture features of smaller QP inputs.
[0083] Figure 13 Some embodiments of the present disclosure are shown Figure 10 Detailed block diagram of the spatial attention block 1004 of the attention block 804. The spatial attention block 1004 is configured to locate regions with blockiness and distortion in the reconstructed frame 401 based on the partition map 402b and the QP map 402c. In the spatial attention block 1004, the input auxiliary information 402 mainly includes the partition map 402b and the QP map 402c, which can better locate the regions where blockiness and distortion are located in space.
[0084] At the same time, feature 502 is processed by the average pooling (avgpool) layer and the maximum pooling (maxpool) layer to obtain the average pooling map and the maximum pooling map. The maximum pooling map and the average pooling map can be used to merge and obtain important spatial features in the current feature map. First, the partition map 402b, the QP map 402c, the maximum pooling map and the average pooling map are combined into a set of inputs through a series operation, and then the features are extracted through two convolution operations, and the spatial attention map is obtained through the sigmoid activation function. Finally, the attention map is used to highlight important spatial features through point-by-point multiplication.
[0085] Figure 15 15. A flowchart of an exemplary method 1500 for enhancing the quality of a frame according to some embodiments of the present disclosure is shown. The method 1500 may be performed by, for example, an apparatus such as the encoder 201 or any other suitable video encoding and / or compression system. The method 1500 may include operations 1502 to 1508 as described below. It should be understood that some of the operations may be optional, and some of the operations may be performed simultaneously or in different order. Figure 15 Execute in the order shown.
[0086] refer to Figure 15 At 1502, a frame and auxiliary information associated with the frame are received. In some embodiments, the auxiliary information includes a prediction map, a partition map, and a QP map, all associated with the frame. In some embodiments, the frame is a reconstructed frame. For example, the frame is a reconstructed frame of chroma samples or luma samples. In one example, the frame and auxiliary information are received by processor 202. For example, Figure 4 As shown, the NN-based loop filter 122 receives a reconstructed frame 401 and side information 402 , where the side information 402 includes a prediction map 402 a , a partition map 402 b , and a QP map 402 c .
[0087] Then, a NN-based in-loop filter is applied to the frame based on the auxiliary information to enhance the quality of the frame. The NN-based in-loop filter includes a backbone portion, the backbone portion including at least one transformer block and at least one attention block, the at least one attention block receiving at least a portion of the auxiliary information. In one example, the processor 202 applies the NN-based in-loop filter to the frame. For example, Figure 4 As shown, the NN-based in-loop filter 122 is applied to the reconstructed frame 401 based on the auxiliary information 402 including the prediction map 402a, the partition map 402b and the QP map 402c. Figure 4 and Figure 8 As shown, the NN-based in-loop filter 122 includes a backbone portion 403, which includes at least one TB 417 and at least one RAB 416. The RAB 416 includes at least one attention block 804. Figure 11 and Figure 13 As shown, the attention block 804 receives at least some auxiliary information 402, which includes a partition map 402b and a QP map 402c.
[0088] To apply the NN-based in-loop filter to the frame, at 1504, features are extracted from the frame based on the auxiliary information, for example, using a feature extraction portion of the NN-based in-loop filter. In some embodiments, to extract features from the reconstructed frame of chroma samples, features are extracted from the reconstructed frame of chroma samples based on the auxiliary information and the reconstructed frame of luma samples using a feature extraction portion. For example, Figure 5A and Figure 5B As shown, features 502 are extracted from the reconstructed frame 401 based on the side information 402 using the feature extraction portion 407 of the NN-based in-loop filter 122 .
[0089] To apply the NN-based in-loop filter to the frame, at 1506, features are processed based on at least a portion of the side information using the backbone portion to generate global features for the frame. Figure 7A and Figure 7B As shown, the backbone portion 403 includes the RAB 416 and the TB 417 , and the features 502 are processed by the RAB 416 and the TB 417 and the backbone portion 403 to generate the global features 602 of the reconstructed frame 401 .
[0090] In some embodiments, at least one attention block is used to process features based on the partition map and the QP map. In one example, a spatial attention block is used to locate regions of the frame with blockiness and distortion based on the partition map and the QP map, and a channel attention block is used to combine intensity attention and contrast attention of the frame based on the QP map. For example, Figure 11 and Figure 13As shown, the spatial attention block 1004 locates areas of the reconstructed frame 401 with blockiness and distortion based on the partition map 402b and the QP map 402c, and the channel attention block 1002 combines the intensity attention and contrast attention of the reconstructed frame 401 based on the QP map 402c.
[0091] In some embodiments, residual blocks and attention blocks are used to process features to obtain local correlations between features. Figure 7A 、 Figure 7B and Figures 8 to 13 As shown, RAB 416 processes features 502 to obtain local correlations between features 502. In some embodiments, a transformer block is used to process features to obtain long-range correlations between features. Figure 7A 、 Figure 7B and Figure 14 As shown, TB 417 processes features 502 to obtain long-range correlations between features 502 .
[0092] At 1508, the frame is reconstructed based on the global features using the reconstructed portion to generate an enhanced frame. In one example, the processor 202 reconstructs the frame based on the global features to generate an enhanced frame. For example, Figure 4 、 Figure 6A and Figure 6B As shown, the reconstruction portion 405 of the NN-based in-loop filter 122 reconstructs the reconstructed frame 401 based on the global features 602 to generate an enhanced reconstructed frame 418 .
[0093] Figure 16 FIG. 1 shows a block diagram of an in-loop filter 114 for training according to some embodiments of the present disclosure. Figure 16 , the in-loop filter 114 may include LMCS 116, DBF 118, SAO 120, and ALF 124. A compressed dataset (e.g., reconstructed frame, prediction map, and partition map) may be obtained after LMCS 116, and a label set may be obtained after ALF 124. The compressed dataset and the label set may be used to train the NN-based in-loop filter 122 described herein. To this end, the NN-based in-loop filter 122 may be trained using a loss function f(x) using weighted L1 loss and L2 loss as shown below:
[0094] f(x)=8×Loss y +Loss u +Loss v ,
[0095] Among them, Loss represents the L1 loss or L2 loss in the Y channel, U channel, and V channel. For example, the loss function of the brightness (luma) sample and the chroma (chroma) sample is expressed as follows:
[0096] L luma =Loss(rec Y -label Y ),as well as
[0097] L chroma =Loss(rec U -label U )+Loss(rec V -label V ).
[0098] In some examples, L1 loss can be used in the early and middle stages of training, while L2 loss can be used in the late stages of training.
[0099] Figure 17A A block diagram of an in-loop filter training strategy for training a traditional neural network (NN) filter 1704 is shown. Figure 17B A block diagram illustrating an exemplary in-loop filter training strategy for training the in-loop filter 114 according to some embodiments of the present disclosure is shown. Figure 17A and Figure 17B Will be described together.
[0100] refer to Figure 17A , network training typically uses uncompressed images as labels 1706a. However, because input 1702a is compressed with different QPs, the distance between each compressed image and each label is different, making it difficult for the network to learn from all training data. Furthermore, this distance severely limits the overall performance of the network.
[0101] refer to Figure 17B , an exemplary in-loop filter training strategy implemented by the in-loop filter 114 uses inputs 1702b (e.g., compressed images) with different QPs to maximize the learning ability of the network. For inputs 1702b with the current QP, lower QP compressed images are used instead of uncompressed images as labels to train the in-loop filter 114. In this way, the distance between the input 1702b and the label 1706b remains consistent, thereby improving the stability of network performance and addressing the shortcomings of traditional training methods.
[0102] Still refer to Figure 17B, the training strategy can set a parameter qp_dis, which represents the QP difference between input 1702a and label 1706b. Because a smaller QP represents a higher quality, the QP value of label 1706b is lower than the QP value of input 1702b. First, the training strategy can use a smaller qp_dis to train the in-loop filter 114 until convergence. Then, the training strategy can gradually increase qp_dis and continue to train the in-loop filter 114. Since the loss function is a multi-stage loss, the training strategy combines the loss function with the training strategy to implement a multi-stage training strategy. First, the training strategy sets qp_dis = 5 and continuously trains the in-loop filter 114 using the L1 and L2 loss functions. Then, the training strategy increases qp_dis by 10 and trains the in-loop filter 114 again using the L1 and L2 loss functions. Thereafter, the training strategy increases qp_dis again and trains the in-loop filter 114 using the L1 and L2 loss functions. In this way, the exemplary in-loop filter training strategy achieves network convergence and maximizes the learning ability of the training strategy.
[0103] Figure 18 Shows the Figure 4 Comparison of visual quality of an exemplary test dataset of NN-based in-loop filters, where (a) shows a video frame of category C-BasketballDrill, (b) shows the result of VTM-11.0_NNVC-2.0, (c) shows the result of NN-based in-loop filter 122, and (d) shows the ground truth. Figure 18 It can be observed in , that the NN-based in-loop filter 122 can effectively remove artifacts and blur in the compressed image (e.g., basketball) while making the edges of the object clearer.
[0104] Figures 19A-19C It is shown that according to some embodiments of the present disclosure, Figure 4 The NN-based in-loop filter 122 is used Figure 18 The test results of the relationship between PSNR and bit rate for video encoding using the test dataset are as follows. Specifically, Figure 19A The rate-distortion (RD) curve of the Y channel is shown. Figure 19B shows the RD curve of the U channel, Figure 19C The RD curve of the V channel is shown in FIG. Figures 19A-19C As shown, the NN-based in-loop filter 122 achieves significant gains.
[0105] Figure 201 is a flow chart illustrating an exemplary method 2000 for training an NN-based in-loop filter 122 according to some embodiments of the present disclosure. The method 2000 may be performed by any suitable computing system. The method 2000 may include operations 2002 to 2012 as described below. It should be understood that some of the multiple operations may be optional, and some of the multiple operations may be performed simultaneously or in different order. Figure 20 Execute in the order shown.
[0106] refer to Figure 20 In 2002, the apparatus may obtain a training data set through a processor, the training data set including a reconstructed frame, a prediction map, and a partition map at each QP. Figure 17B , the training strategy implemented by the in-loop filter 114 uses inputs 1702b with different QPs to maximize the learning ability of the network.
[0107] In 2003, the apparatus may apply a filter, such as DBF, SAO, and ALF, to the training data set via a processor. Figure 16 , the in-loop filter 114 may include LMCS 116, DBF 118, SAO 120, and ALF 124. A compressed data set (e.g., reconstructed frame, prediction map, and partition map) may be obtained after LMCS 116, and a label set may be obtained after ALF 124. DBF 118, SAO 120, and ALF 124 may be applied to the compressed data set to obtain the label set.
[0108] In some embodiments, the network can be implemented in the PyTorch platform. The total number of epochs can be 800, the batch size can be set to 32, and the learning rate can be set to 1e-4. The proposed filter can be trained using the DIV2K and BVI-DVC datasets. All images can be compressed using VTM 11.0-NNVC in all intra (AI) and random access (RA) configurations. Specifically, all images in DIV2K and all I frames of class A in BVI-DVC can be compressed in the AI configuration to train the I-slice model, while all videos of classes B, C, and D in the BVI-DVC dataset can be compressed in the RA configuration. Compressed images can be extracted from all compressed videos at intervals of 20 frames to train the B-slice model. It should be noted that DBF and SAO may be disabled when compressing training data for the I-slice model in the AI configuration. The compressed images can be randomly cropped into multiple 144×144 blocks and data augmentation can be performed with random horizontal and vertical flips.
[0109] At 2003, the apparatus may obtain, via a processor, a label set associated with an enhanced reconstructed frame as an output of the ALF and associated with a second set of QPs that is smaller than the first set of QPs. Figure 16 , a compressed dataset (e.g., reconstructed frame, prediction map, and partition map) may be obtained after LMCS 116, and a label set may be obtained after ALF 124. DBF 118, SAO 120, and ALF 124 may be applied to the compressed dataset to obtain the label set.
[0110] At 2304, the apparatus may define, via the processor, a parameter qp_dis as the difference between the input (e.g., reconstructed frames, prediction maps, and partition maps compressed at the input QP) and the label (e.g., reconstructed frames of the ALF output compressed at the output QP). Figure 17B The training strategy can set a parameter qp_dis, which represents the QP difference between input 1702b and label 1706b. Because a smaller QP indicates higher quality, the QP value of label 1706b is lower than that of input 1702b.
[0111] At 2006, the device may train the in-loop filter 114 under the current qp_dis. For example, referring to Figure 17B , this training strategy can use a smaller qp_dis to train the in-loop filter 114.
[0112] In 2008, the device may increase the current qp_dis after the network converges. For example, referring to Figure 17B , the training strategy may use a smaller qp_dis for the loop filter 114 until convergence. Then, the training strategy may gradually increase qp_dis and continue to train the loop filter 114. Since the loss function is a multi-stage loss, the training strategy combines the loss function with the training strategy to implement a multi-stage training strategy. First, the training strategy sets qp_dis=5 and continuously trains the in-loop filter 114 with the L1 and L2 loss functions. Then, the training strategy increases qp_dis by 10 and trains the in-loop filter 114 again with the L1 and L2 loss functions. Thereafter, the training strategy increases qp_dis again and trains the in-loop filter 114 with the L1 and L2 loss functions.
[0113] In 2010, the device can determine whether network performance is stagnant. For example, referring to Figure 17B , the training strategy can use the increased qp_dis to determine whether the performance of the network increases with subsequent training. If not, the operation can return to 2006; if so, the operation can move to 2012.
[0114] In 2012, the device can fix the network parameters and end the training. For example, refer to Figure 17B Once the network parameters are fixed, the NN-based in-loop filter 122 is generated for use by the in-loop filter 114. In this way, the in-loop filter training strategy achieves network convergence and maximizes learning ability.
[0115] In various aspects of the present disclosure, the functions described herein may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as instructions on a non-transitory computer-readable medium. Computer-readable media include computer storage media. The storage medium may be a computer program that can be executed by a processor (e.g., Figure 2 and Figure 3 Any available medium accessible by the processor 202 in the computer readable medium. By way of example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage devices, HDD (e.g., magnetic disk storage or other magnetic storage devices), flash drives, SSDs, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a processing system (e.g., a mobile device or a computer). Disks and optical disks, as used herein, include compact disks (CDs), laser disks, optical disks, digital video disks (DVDs), and floppy disks, where disks typically reproduce data magnetically, while optical disks reproduce data optically with lasers. Combinations of the above should also be included within the scope of computer-readable media.
[0116] According to one aspect of the present disclosure, a method for enhancing the quality of a frame is provided. A processor receives a frame and auxiliary information associated with the frame. A NN-based in-loop filter is applied to the frame based on the auxiliary information to enhance the quality of the frame. The NN-based in-loop filter includes a backbone portion, the backbone portion including at least one transformer block and at least one RAB. The at least one RAB includes at least one attention block, the at least one attention block receiving at least a portion of the auxiliary information.
[0117] In some embodiments, the auxiliary information includes a prediction map, a partition map, and a QP map, all associated with the frame.
[0118] In some embodiments, the NN-based in-loop filter further comprises a feature extraction portion. In some embodiments, in order to apply the NN-based in-loop filter, the feature extraction portion is used to extract features from the frame based on the auxiliary information.
[0119] In some embodiments, the frame is a reconstructed frame of chroma samples. In some embodiments, to extract features from the frame, a feature extraction portion is used to extract features from the reconstructed frame of chroma samples based on the auxiliary information and the reconstructed frame of luma samples.
[0120] In some embodiments, the at least a portion of the auxiliary information includes a partition map and a QP map, both associated with the frame. In some embodiments, to apply the NN-based in-loop filter, features are processed based on the partition map and the QP map using at least one attention block.
[0121] In some embodiments, the at least one attention block includes a spatial attention block and a channel attention block. In some embodiments, to process features, the spatial attention block is used to locate regions of the frame with blockiness and distortion based on the partition map and the QP map, and the channel attention block is used to combine intensity attention and contrast attention of the frame based on the QP map.
[0122] In some embodiments, the RAB further comprises at least one residual block.In some embodiments, to apply the NN-based in-loop filter, the features are processed using the RAB to obtain local correlations between features, and the features are processed using the transformer block to obtain long-range correlations between features.
[0123] In some embodiments, the frame is a reconstructed frame of luma samples, and the backbone portion includes three attention blocks and six transformer blocks.
[0124] In some embodiments, the frame is a reconstructed frame of chroma samples, and the backbone portion includes one attention block and three transformer blocks.
[0125] In some embodiments, the NN-based in-loop filter further comprises a reconstruction portion. In some embodiments, to apply the NN-based in-loop filter, a backbone portion is used to process features based on at least a portion of the auxiliary information to generate global features of the frame, and the reconstruction portion is used to reconstruct the frame based on the global features to generate an enhanced frame.
[0126] According to another aspect of the present disclosure, a system for enhancing frame quality includes a memory configured to store instructions and a processor coupled to the memory. The processor is configured, when executing the instructions, to receive a frame and auxiliary information associated with the frame and apply a NN-based in-loop filter to the frame based on the auxiliary information to enhance the frame quality. The NN-based in-loop filter includes a backbone comprising at least one transformer block and at least one RAB. The at least one RAB includes at least one attention block that receives at least a portion of the auxiliary information.
[0127] In some embodiments, the auxiliary information includes a prediction map, a partition map, and a QP map, all associated with the frame.
[0128] In some embodiments, the NN-based in-loop filter further comprises a feature extraction portion. In some embodiments, to apply the NN-based in-loop filter, the processor is further configured to extract features from the frame based on the auxiliary information using the feature extraction portion.
[0129] In some embodiments, the frame is a reconstructed frame of chroma samples. In some embodiments, to extract features from the frame, the processor is further configured to extract features from the reconstructed frame of chroma samples using a feature extraction portion based on the auxiliary information and the reconstructed frame of luma samples.
[0130] In some embodiments, the at least a portion of the auxiliary information includes a partition map and a QP map, both associated with the frame. In some embodiments, to apply the NN-based in-loop filter, the processor is further configured to process features based on the partition map and the QP map using at least one attention block.
[0131] In some embodiments, the at least one attention block includes a spatial attention block and a channel attention block. In some embodiments, to process the features, the processor is further configured to use the spatial attention block to locate regions of the frame with blockiness and distortion based on the partition map and the QP map, and to use the channel attention block to combine intensity attention and contrast attention of the frame based on the QP map.
[0132] In some embodiments, the RAB further comprises at least one residual block. In some embodiments, to apply the NN-based in-loop filter, the processor is further configured to process the features using the RAB to obtain local correlations between the features, and process the features using the transformer block to obtain long-range correlations between the features.
[0133] In some embodiments, the frame is a reconstructed frame of luma samples, and the backbone portion includes three attention blocks and six transformer blocks.
[0134] In some embodiments, the frame is a reconstructed frame of chroma samples, and the backbone portion includes one attention block and three transformer blocks.
[0135] According to another aspect of the present disclosure, a tangible computer-readable device is provided. The tangible computer-readable device has instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform a plurality of operations. The plurality of operations include receiving a frame and auxiliary information associated with the frame, and applying a NN-based in-loop filter to the frame based on the auxiliary information to enhance the quality of the frame. The NN-based in-loop filter includes a backbone portion, the backbone portion including at least one transformer block and at least one RAB. The at least one RAB includes at least one attention block that receives at least a portion of the auxiliary information.
[0136] The foregoing description of the various embodiments will reveal the general nature of the present disclosure so that others can easily modify and / or adapt these embodiments for various applications without departing from the general concepts of the present disclosure by applying knowledge within the art without undue experimentation. Therefore, based on the teachings and guidance given herein, such adaptations and modifications are intended to be within the meaning and range of equivalents of the disclosed embodiments. It should be understood that the wording or terminology herein is for descriptive purposes only and not for limiting purposes, and therefore the terminology or terminology of this specification will be interpreted by the skilled person based on the teachings and guidance.
[0137] The embodiments of the present disclosure have been described above with the help of functional building blocks that illustrate the implementation of specific functions and their relationships. For ease of description, the boundaries of these functional building blocks are arbitrarily defined herein. Alternative boundaries may be defined so long as the specified functions and their relationships are appropriately performed.
[0138] The Summary and Abstract sections may set forth one or more but not all exemplary embodiments of the present disclosure as contemplated by one or more inventors, and thus, are not intended to limit the present disclosure and the appended claims in any way.
[0139] Various functional blocks, modules, and steps are disclosed above. The provided arrangement is illustrative and not restrictive. Therefore, the functional blocks, modules, and steps may be reordered or combined in a manner different from the examples provided above. Similarly, some embodiments include only a subset of the functional blocks, modules, and steps, and any such subset is permitted.
[0140] The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.
Claims
1. A method for enhancing the quality of a frame, comprising: receiving, by a processor, the frame and auxiliary information associated with the frame; as well as The processor applies a neural network (NN)-based in-loop filter to the frame based on the auxiliary information to enhance the quality of the frame, wherein the NN-based in-loop filter includes a backbone portion, the backbone portion includes at least one transformer block and at least one residual attention block (RAB), and the at least one RAB includes at least one attention block, and the at least one attention block receives at least a portion of the auxiliary information.
2. The method according to claim 1, wherein The auxiliary information includes a prediction map, a partition map, and a quantization parameter (QP) map, all associated with the frame.
3. The method according to claim 1, wherein The NN-based in-loop filter also includes a feature extraction part; and Applying the NN-based in-loop filter includes extracting features from the frame based on the side information using the feature extraction part.
4. The method according to claim 3, wherein: The frame is a reconstructed frame of chroma samples; and Extracting the feature from the frame includes extracting the feature from the reconstructed frame of chrominance samples based on the side information and the reconstructed frame of luma samples using the feature extraction portion.
5. The method according to claim 3, wherein: The at least a portion of the auxiliary information comprises a partition map and a QP map each associated with the frame; and Applying the NN-based in-loop filter includes processing the features based on the partition map and the QP map using the at least one attention block.
6. The method according to claim 5, wherein: The at least one attention block includes a spatial attention block and a channel attention block; and Processing the features includes: Using the spatial attention block, locate areas of the frame having blockiness and distortion based on the partition map and the QP map; as well as Using the channel attention block, the intensity attention and contrast attention of the frame are combined based on the QP map.
7. The method according to claim 3, wherein: The RAB further includes at least one residual block; and Applying the NN-based in-loop filter includes: Processing the features using the RAB to obtain local correlations between the features; and The features are processed using the transformer block to obtain long-range correlations between the features.
8. The method according to claim 1, wherein The frame is a reconstructed frame of luma samples; and The backbone consists of three attention blocks and six transformer blocks.
9. The method according to claim 1, wherein The frame is a reconstructed frame of chroma samples; and The backbone consists of an attention block and three transformer blocks.
10. The method according to claim 3, wherein: The NN-based in-loop filter further includes a reconstruction portion; and Applying the NN-based in-loop filter includes: processing the features based on the at least a portion of the side information using the backbone portion to generate a global feature for the frame; and The frame is reconstructed based on the global features using the reconstructed portion to generate an enhanced frame.
11. A system for enhancing the quality of a frame, comprising: a memory configured to store instructions; as well as a processor coupled to the memory and configured to, when executing the instructions: receiving the frame and auxiliary information associated with the frame; as well as Applying a neural network (NN) based in-loop filter to the frame based on the auxiliary information to enhance the quality of the frame, wherein the NN based in-loop filter includes a backbone portion, the backbone portion includes at least one transformer block, at least one residual attention block (RAB), and the at least one RAB includes at least one attention block, the at least one attention block receiving at least a portion of the auxiliary information.
12. The system according to claim 11, wherein The auxiliary information includes a prediction map, a partition map, and a quantization parameter (QP) map, all associated with the frame.
13. The system according to claim 11, wherein: The NN-based in-loop filter also includes a feature extraction part; and To apply the NN-based in-loop filter, the processor is further configured to extract features from the frame based on the auxiliary information using the feature extraction part.
14. The system according to claim 13, wherein: The frame is a reconstructed frame of chroma samples; and To extract the feature from the frame, the processor is further configured to extract the feature from the reconstructed frame of chrominance samples based on the side information and the reconstructed frame of luma samples using the feature extraction portion.
15. The system according to claim 13, wherein: The at least a portion of the auxiliary information comprises a partition map and a QP map each associated with the frame; and To apply the NN-based in-loop filter, the processor is further configured to process the features based on the partition map and the QP map using the at least one attention block.
16. The system according to claim 15, wherein: The at least one attention block includes a spatial attention block and a channel attention block; and In order to process the features, the processor is further configured to: Using the spatial attention block, locate areas of the frame having blockiness and distortion based on the partition map and the QP map; as well as Using the channel attention block, the intensity attention and contrast attention of the frame are combined based on the QP map.
17. The system of claim 13, wherein: The RAB further includes at least one residual block; and To apply the NN-based in-loop filter, the processor is further configured to: Processing the features using the RAB to obtain local correlations between the features; as well as The features are processed using the transformer block to obtain long-range correlations between the features.
18. The system according to claim 11, wherein: The frame is a reconstructed frame of luma samples; and The backbone consists of three attention blocks and six transformer blocks.
19. The system according to claim 11, wherein: The frame is a reconstructed frame of chroma samples; and The backbone consists of an attention block and three transformer blocks.
20. A tangible computer-readable device having instructions stored thereon, the instructions, when executed by at least one computing device, causing the at least one computing device to perform a plurality of operations, the plurality of operations comprising: receiving a frame and auxiliary information associated with the frame; as well as Applying a neural network (NN) based in-loop filter to the frame based on the auxiliary information to enhance the quality of the frame, wherein the NN based in-loop filter comprises a backbone portion, the backbone portion comprises at least one transformer block, at least one residual attention block (RAB), and the at least one RAB comprises at least one attention block, the at least one attention block receiving at least a portion of the auxiliary information.