Convolutional neural network based on depth separable convolution for loop filter of video encoder

By using the WCDAB backbone and depth in the CNN using loop filters in the video encoder, the global features are generated to improve the quality of video frames, and the problems of information loss and compression artifacts in the prior art are solved, and higher encoding performance and quality are achieved.

CN120019663APending Publication Date: 2025-05-16GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280100921.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-13
Filing Date
2022-12-05
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing video encoding technology will lead to irreversible information loss and compression artifacts during the encoding process, especially under high compression rates, the performance of existing loop filters is limited, making it difficult to effectively suppress compression artifacts.

Method used

A weakly connected dense attention block (WCDAB) backbone in a convolutional neural network (CNN) using loop filters can separate convolution and attention mechanisms through multiple depths to generate global features to improve the quality of video frames.

Benefits of technology

It realizes that under low computing complexity, the objective quality of video encoding is significantly improved, block artifacts, band artifacts and noise are reduced, and encoding performance is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120019663A_ABST
    Figure CN120019663A_ABST
Patent Text Reader

Abstract

A method of generating an enhanced frame at a video encoder is provided. The method may include a weak connection dense attention block (WCDAB) trunk (403) of a convolutional neural network (CNN) of a loop filter (400) receiving a first set of feature extraction as input. The first set of feature extraction may be associated with the reconstructed frame (402a). The method may include applying, by a WCDAB trunk (403) of a CNN of a loop filter (400), a plurality of depth separable convolves on a first set of feature extraction to generate a set of global features. The method also provides an efficient multi-level training strategy based on progressive learning. The multi-stage training strategy maximizes the performance of the proposed network.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] The disclosed embodiments relate to video coding.

[0002] Digital video has become mainstream and is widely used in various applications such as digital television, video phones and teleconferencing. These digital video applications are feasible due to advances in computing and communication technology and efficient video coding technology. Various video coding techniques can be used to compress video data, so that the video data can be encoded using one or more video coding standards. Exemplary video coding standards may include, but are not limited to, versatile video coding (H.266 / VVC), high-efficiency video coding (H.265 / HEVC), advanced video coding (H.264 / AVC), moving picture expert group (MPEG) coding, etc. Summary of the invention

[0003] According to one aspect of the present disclosure, a method for generating an enhanced frame in a video encoder is provided. The method may include a weakly connected dense attention block (WCDAB) trunk of a convolutional neural network (CNN) of a loop filter receiving a first set of feature extractions as input. The first set of feature extractions may be associated with a reconstructed frame. The method may include a WCDAB trunk of the CNN of the loop filter applying multiple depth-separable convolutions to the first set of feature extractions to generate a set of global features.

[0004] According to another aspect of the present disclosure, a system for generating an enhanced frame in a video encoder is provided. The system may include a memory for storing instructions. The system may include a processor coupled to the memory and used when executing the instructions to: receive a first set of feature extractions as input through a WCDAB backbone of a CNN of a loop filter. The first set of feature extractions may be associated with a reconstructed frame. The system may include a processor coupled to the memory and used when executing the instructions to: apply multiple depth-separable convolutions to the first set of feature extractions through a WCDAB backbone of a CNN of a loop filter to generate a set of global features.

[0005] According to another aspect of the present disclosure, a method for training a loop filter model of a video encoder is provided. The method may include a processor obtaining a compressed data set including a reconstructed frame, a predicted frame, and a segmented frame. The compressed data set may be associated with a first set of quantization parameters (QP). The method may include a processor applying a deblocking filter (DBF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF) to the compressed data set. The method may include a processor obtaining a label set associated with an enhanced reconstructed frame as an output of the ALF. The label set may be associated with a second set of QPs that is less than the first set of QPs.

[0006] According to another aspect of the present disclosure, a system for training a loop filter model of a video encoder is provided. The system may include a memory for storing instructions. The system may include a processor coupled to the memory and used to obtain a compressed data set including a reconstructed frame, a predicted frame, and a segmented frame when executing the instructions. The compressed data set may be associated with a first set of QPs. The system may include a processor coupled to the memory and used to apply DBF, SAO, and ALF to the compressed data set when executing the instructions. The system may include a processor coupled to the memory and used to obtain a label set associated with an enhanced reconstructed frame as an output of the ALF when executing the instructions. The label set is associated with a second set of QPs that is smaller than the first set of QPs.

[0007] These illustrative embodiments are mentioned not to limit or define the present disclosure, but to provide examples to aid understanding of the present disclosure. Other embodiments are described in the detailed description and further description is provided. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The accompanying drawings, which are incorporated herein and constitute a part of the specification, illustrate embodiments of the present disclosure and, together with the description, further explain the principles of the present disclosure and enable one skilled in the relevant art to make and use the present disclosure.

[0009] Figure 1 A block diagram of a video encoder with a loop filter is shown.

[0010] Figure 2 A block diagram of an exemplary encoding system according to some embodiments of the present disclosure is shown.

[0011] Figure 3 A block diagram of an exemplary decoding system according to some embodiments of the present disclosure is shown.

[0012] Figure 4According to some aspects of the present disclosure Figure 2 Detailed block diagram of an exemplary loop filter for an encoder.

[0013] Figure 5A A detailed block diagram of an exemplary WCDAB according to some aspects of the present disclosure is shown.

[0014] Figure 5B A detailed block diagram of an exemplary residual attention block (RAB) according to some aspects of the present disclosure is shown.

[0015] Figure 5C Detailed block diagram showing a deep convolutional layer according to some aspects of the present disclosure.

[0016] Fig. 6A A detailed block diagram of an exemplary channel attention block (CAB) according to some aspects of the present disclosure is shown.

[0017] Figure 6B A detailed block diagram of an exemplary channel-space joint attention block (CSAB) according to some aspects of the present disclosure is shown.

[0018] Figure 6C A detailed block diagram of an exemplary spatial attention block (SAB) according to some aspects of the present disclosure is shown.

[0019] Figure 7 A detailed block diagram of an exemplary loop filter for training a loop filter model according to some aspects of the present disclosure is shown.

[0020] Fig. 8A A block diagram showing an example loop filter training strategy.

[0021] Figure 8B A block diagram illustrating an exemplary loop filter training strategy according to some aspects of the present disclosure.

[0022] Fig. 9 A first graphical representation of peak signal-to-noise ratio (PSNR) versus bit rate for video encoding using an exemplary loop filter according to some aspects of the present disclosure is shown.

[0023] Fig.10 A second graphical representation of PSNR versus bit rate for video encoding using an exemplary loop filter according to some aspects of the present disclosure is shown.

[0024] Fig.11 A flowchart illustrating a first exemplary video encoding method according to some aspects of the present disclosure is shown.

[0025] Fig.12 A flowchart of a second exemplary video encoding method according to some aspects of the present disclosure is shown.

[0026] Embodiments of the present disclosure will be described with reference to the accompanying drawings. DETAILED DESCRIPTION

[0027] Although some configurations and arrangements are discussed, it should be understood that this is for illustrative purposes only. Those skilled in the art will recognize that other configurations and arrangements can be used without departing from the spirit and scope of the present disclosure. Those skilled in the art will appreciate that the present disclosure can also be used in a variety of other applications.

[0028] Note that the phrases "one embodiment", "embodiment", "example embodiment", "some embodiments", "certain embodiments", etc. mentioned in the specification indicate that the described embodiments may include certain features, structures, or characteristics, but not every embodiment necessarily includes the certain features, structures, or characteristics. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a certain feature, structure, or characteristic is described in conjunction with an embodiment, whether or not it is explicitly described, the relevant technicians should know that such feature, structure, or characteristic can be implemented in conjunction with other embodiments.

[0029] In general, terms can be understood at least in part based on usage in context. For example, depending at least in part on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, or characteristics in a plural sense. Similarly, depending at least in part on the context, terms such as "one," "an," or "the" can also be understood to express singular or plural usage. In addition, also depending at least in part on the context, the term "based on" can be understood to not necessarily express a set of exclusive factors, but can allow for the presence of other factors that are not necessarily explicitly described.

[0030] Various aspects of the video encoding system will now be described with reference to various devices and methods. These devices and methods will be described in the following detailed description and illustrated in the accompanying drawings by various modules, components, circuits, steps, operations, processes, algorithms, etc. (collectively referred to as "elements"). These elements can be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether these elements are implemented as hardware, firmware, or software depends on the specific application and design constraints imposed on the overall system.

[0031] The techniques described herein can be used for various video coding applications. As described herein, video coding includes encoding and decoding of video. The encoding and decoding of video can be performed in units of blocks. For example, a coding block, a transform block, or a prediction block can be subjected to encoding / decoding processing, such as transformation, quantization, prediction, loop filtering, reconstruction, etc. As described herein, the block to be encoded / decoded will be referred to as a "current block". For example, the current block may represent a coding block, a transform block, or a prediction block according to the current encoding / decoding process. In addition, it should be understood that the term "unit" used in the present disclosure represents a basic unit for performing a specific encoding / decoding process, and the term "block" represents a sample array of a predetermined size. Unless otherwise specified, "block", "unit", "part", and "component" can be used interchangeably.

[0032] For example, existing video compression methods such as HEVC and VVC perform block segmentation and quantization during the encoding process, resulting in irreversible information loss and various compression artifacts, such as blocking, blurring, and banding. This phenomenon is particularly evident when the compression ratio is high. Currently, there are many deep learning-based methods used to improve the quality of compressed images and videos, mainly to reduce blocking artifacts, banding artifacts, and noise.

[0033] Versatile video coding (VVC) uses loop filters in the encoder to suppress compression artifacts and reduce distortion. These loop filters may include deblocking filters (DBF), sample adaptive offsets (SAO), and adaptive loop filters (ALF), to name a few. DBF and SAO are two types of filters used to reduce artifacts caused by the encoding process. DBF focuses on visual artifacts at block boundaries, while SAO complementarily reduces artifacts that may be caused by quantization of transform coefficients within blocks. ALF can enhance the adaptive filter of the reconstructed signal, using a Wiener-based adaptive filter to reduce the mean square error (MSE) between the original and reconstructed samples. Although these filters greatly mitigate compression artifacts, these filters are manually designed and developed based on signal processing theory assuming stationary signals. Since natural video sequences are usually non-stationary, the performance of these filters is limited. Therefore, there is still much room for improvement in the loop filters in VVC.

[0034] With the development of deep learning, various CNN-based image and video quality enhancement methods have emerged. Recently, some video encoders are designed with CNN-based loop filters, which include trained CNN filters embedded in the VVC loop. This can be achieved by inserting loop filter components or replacing some loop filter components.

[0035] Figure 1A block diagram of a video encoder 100 is shown with a loop filter 114 having a CNN loop filter (LF) 122. The video encoder 100 may include, for example, a video sequence component 102, a transform component 104, a quantization component 106, an inverse quantization component 108, an inverse transform component 110, an encoding component 112, a loop filter 114, a decoded image buffer 126, an inter prediction component 128, and an intra prediction component 130, to name a few. The loop filter 114 may include, for example, a luma mapping with chroma scaling (LMCS) component 116, a DBF 118, a SAO 120, a CNN LF 122, and an ALF 124.

[0036] Some video encoders use loop filters based on quantization parameter (QP) variable CNNs for VVC intra-frame coding. To avoid training and deployment in multiple networks, these encoders use a QP attention module (QPAM), which captures the compression noise level of different QPs and emphasizes meaningful features along the channel dimension. QPAM can be embedded in the residual block as part of the network architecture, which is designed for controllability of different QPs. To fine-tune the network, these video encoders can use a focal mean squared error (MSE) loss function. Since the loop filters in existing video encoders do not receive multiple inputs, the image enhancement performance is limited.

[0037] In other video encoders, loop filters based on dense residual convolutional neural networks (DRNs) can be used for VVC. These video encoders use residual learning components, dense shortcuts, and bottleneck layers to solve the gradient vanishing problem, encourage feature reuse, and reduce computational resources, respectively. Unfortunately, the performance of these video encoders cannot achieve an ideal balance between complexity and performance.

[0038] In other existing video encoders, a CNN-based filter can be used to enhance the quality of VVC intra-coded frames by taking auxiliary information such as segmentation and prediction information as input. For chroma, the auxiliary information also includes luma samples. Although this filter achieves adequate performance on the Y channel, the performance on other channels is relatively low and the encoding delay is not ideal.

[0039] To overcome these and other challenges, the present disclosure provides an exemplary lightweight loop CNN filter that uses a loop CNN filter model trained using a multi-stage training strategy. Compared with other loop CNN filters, the exemplary loop CNN filter described herein achieves higher performance with lower computational complexity.

[0040] The exemplary loop CNN filter performs depth-separable convolution and attention mechanism to improve the objective quality of VVC video frames. The loop CNN filter described in this paper is based on residual learning, which enhances the quality of the input image by learning the residual map. At the same time, the loop filter uses the predicted frame, segmented frame and quantization parameter (QP) map as additional auxiliary information to guide the proposed network to better enhance the quality of the enhanced reconstructed frame.

[0041] The loop CNN filter model described in this article can be trained using a multi-stage training strategy using progressive learning to train the model. For example, the parameter qp_dis can be set to represent the QP difference between the network input and the label. Since the smaller the QP, the higher the quality, the QP value of the label is less than the QP value of the input. The exemplary training strategy first uses a smaller qp_dis to train the model, and then gradually increases qp_dis during training after the network converges. Since the loss function is a multi-stage loss, the loss function can be combined with the training strategy to implement a multi-stage training strategy. The exemplary multi-stage training strategy achieves a model with higher performance than other loop CNN filter models.

[0042] In addition, the loop CNN filter described in this article may include a WCDAB backbone composed of multiple WCDABs. Each WCDAB may include four residual blocks (RABs) and a channel-spatial joint attention block (CSAB). The four RABs extract features from various inputs. The outputs of the second RAB and the fourth RAB are fused. Finally, important features are retained at the channel and spatial levels by CSAB. With the help of each subsequent WCDAB in the backbone, the proposed loop CNN filter achieves better performance. This is because deep features are more important for the quality of residual learning. Two depth-separable convolutions can be applied to the input of the RAB to extract features. The outputs of the two depth-separable convolutions can be fused, and the channel attention block (CAB) can be used to emphasize the important channels of the fused features. The following is combined with Figures 2 to 12 More details are provided on an exemplary in-loop CNN filter and an exemplary training strategy for its model.

[0043] Figure 2A block diagram of an exemplary encoding system 200 is shown according to some embodiments of the present disclosure. Figure 3 A block diagram of an exemplary decoding system 300 according to some embodiments of the present disclosure is shown. Each system 200 or 300 can be applied to or integrated into various systems and devices capable of data processing, such as computers and wireless communication devices. For example, the system 200 or 300 can be all or part of a mobile phone, a desktop computer, a laptop computer, a tablet computer, a car computer, a game console, a printer, a positioning device, a wearable electronic device, a smart sensor, a virtual reality (VR) device, an augmented reality (AR) device, or any other suitable electronic device with data processing capabilities. Figure 2 and Figure 3 As shown, system 200 or 300 may include processor 202, memory 204, and interface 206. These components are shown as being interconnected via a bus, but other connection types are also allowed. It is understood that system 200 or 300 may include any other suitable components for performing the functions described herein.

[0044] Processor 202 may include a microprocessor, such as a graphics processing unit (GPU), an image signal processor (ISP), a central processing unit (CPU), a digital signal processor (DSP), a tensor processing unit (TPU), a vision processing unit (VPU), a neural processing unit (NPU), a synergistic processing unit (SPU) or a physical processing unit (PPU), a microcontroller unit (MCU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device (PLD), a state machine, gated logic, discrete hardware circuits, and other suitable hardware for performing the various functions described in the present disclosure. Although Figure 2 and Figure 3Only one processor is shown, but it is understood that multiple processors may be included. Processor 202 may be a hardware device having one or more processing cores. Processor 202 may execute software. Whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise, software should be broadly understood as instructions, instruction sets, codes, code segments, program codes, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, processes, functions, etc. Software may include computer instructions written in an interpreted language, a compiled language, or machine code. Other techniques for directing hardware are also permitted under the broad category of software.

[0045] Memory 204 may broadly include memory (also known as main / system memory) and storage (also known as secondary memory). For example, memory 204 may include random-access memory (RAM), read-only memory (ROM), static RAM (SRAM), dynamic RAM (DRAM), ferroelectric RAM (FRAM), electrically erasable programmable ROM (EEPROM), compact disc read-only memory (CD-ROM) or other optical storage, hard disk drive (HDD) (such as magnetic disk storage or other magnetic storage device), flash drive, solid-state drive (SSD), or any other medium that can be used to carry or store the desired program code in the form of instructions that can be accessed and executed by processor 202. In a broad sense, memory 204 can be implemented by any computer-readable medium, such as a non-transitory computer-readable medium. Although Figure 2 and Figure 3 Only one memory is shown, but it will be appreciated that multiple memories may be included.

[0046] The interface 206 may broadly include a data interface and a communication interface for receiving and sending signals in the process of receiving and sending information with other external network elements. For example, the interface 206 may include an input / output (I / O) device and a wired or wireless transceiver. Figure 2 and Figure 3 Only one memory is shown, but it will be appreciated that multiple interfaces may be included.

[0047] The processor 202, memory 204, and interface 206 can be implemented in various forms in the system 200 or 300 to perform video encoding functions. In some embodiments, the processor 202, memory 204, and interface 206 of the system 200 or 300 are implemented (e.g., integrated) on one or more system-on-chips (SoCs). In one example, the processor 202, memory 204, and interface 206 can be integrated on an application processor (AP) SoC, which is responsible for application processing in an operating system (OS) environment, including running video encoding and decoding applications. In another example, the processor 202, memory 204, and interface 206 can be integrated on a dedicated processor chip such as a GPU or ISP chip dedicated to image and video processing in a real-time operating system (RTOS) for video encoding.

[0048] like Figure 2 As shown, in the encoding system 200, the processor 202 may include one or more modules, such as an encoder 201 (also referred to herein as a "pre-processing network"). Figure 2 The encoder 201 is shown to be located within a processor 202, but it is understood that the encoder 201 may include one or more submodules that may be implemented on different processors that are close to or far from each other. The encoder 201 (and any corresponding submodules or subunits) may be a hardware unit (e.g., part of an integrated circuit) of the processor 202 designed for use with other components, or a software unit implemented by the processor 202 by executing at least a portion (i.e., instructions) of a program. The instructions of the program may be stored on a computer-readable medium such as a memory 204, and when executed by the processor 202, a process having one or more functions related to video coding may be performed, such as image segmentation, inter-frame prediction, intra-frame prediction, transform, quantization, filtering, entropy coding, etc., as described in detail below.

[0049] Similarly, if Figure 3 As shown, in the decoding system 300, the processor 202 may include one or more modules, such as a decoder 301 (also referred to herein as a "post-processing network"). Figure 3The decoder 301 is shown to be located within one processor 202, but it is understood that the decoder 301 may include one or more submodules that may be implemented on different processors that are close to or far from each other. The decoder 301 (and any corresponding submodules or subunits) may be a hardware unit (e.g., part of an integrated circuit) of the processor 202 designed for use with other components, or a software unit implemented by the processor 202 by executing at least a portion (i.e., instructions) of a program. The instructions of the program may be stored on a computer-readable medium such as a memory 204, and when executed by the processor 202, a process having one or more functions related to video decoding may be performed, such as entropy decoding, inverse quantization, inverse transformation, inter-frame prediction, intra-frame prediction, filtering, as described in detail below.

[0050] review Figure 2 , the encoder 201 may include an exemplary loop CNN filter that performs depth-wise separable convolution and attention to improve the objective quality of VVC video frames, as described below in conjunction with Figures 4 to 12 described.

[0051] Figure 4 According to some aspects of the present disclosure Figure 2 Detailed block diagram of an exemplary loop CNN filter 400 (hereinafter referred to as “loop CNN filter 400”) of encoder 201. Figure 4 , the loop CNN filter 400 may include a feature extraction part 401 , a WCDAB backbone 403 , and a reconstruction part 405 .

[0052] The input of the loop CNN filter 400 may include, for example, a reconstructed frame (rec) 402a, a predicted frame (pred) 402b, a segmented frame (par) 402c, and a QP map (qp) 402d, all of which are generated by the encoder 201. The reconstructed frame 402a is a reconstruction of the current video frame by the encoder 201 for quality enhancement. The predicted frame 402b is a prediction made by the encoder 201 on the reconstructed frame 402a. The segmented frame 402c is segmentation information corresponding to the reconstructed frame 402a. The QP map 402d is used to indicate the QP value corresponding to the reconstructed frame. The QP map 402d can improve the quality of the reconstructed frames of different QPs at the reconstruction portion 405.

[0053] The feature extraction part 401 may include a plurality of parallel convolutional layers 406 (standard convolutions), each of which is used to integrate and extract shallow features of its corresponding input. Afterwards, the output features of each parallel convolutional layer 406 are concatenated 404 and fused using a convolutional layer 408 with a stride of 1 to obtain fused shallow features. Then, the fused features are downsampled using a convolutional layer 410 with a stride of 2 to reduce the computational effort of the proposed network. Finally, the downsampled features (e.g., the first set of feature extractions) are sent to the WCDAB backbone 403.

[0054] WCDAB backbone 403 may include multiple WCDABs 416. In some embodiments, WCDAB backbone 403 may include, for example, eight or more WCDABs 416. WCDAB backbone 403 may extract a global feature map (eg, a set of global features) from the first set of feature extractions received from feature extraction portion 401.

[0055] The reconstruction part 405 uses a 1x1 convolution layer 408 to reduce the channel dimension of the global feature map obtained from the WCDAB backbone 403. Then, the reconstruction part 405 can use a pixel shuffle component 412 to upsample the reduced dimension features to obtain a three-channel residual map. Finally, the summation component 414 adds the obtained residual map to the reconstructed frame 402a to generate an enhanced reconstructed frame 418.

[0056] Figure 5A According to some aspects of the present disclosure Figure 4 Detailed block diagram 500 of an exemplary WCDAB 416. Figure 5B A detailed block diagram 501 of an exemplary RAB 502 is shown in accordance with some aspects of the present disclosure. Figure 5C A detailed block diagram 503 of a depthwise separable convolutional layer 508 is shown in accordance with some aspects of the present disclosure. FIG. 5A to FIG. 5C Describe together.

[0057] refer to Figure 5A , WCDAB 416 may include, for example, a plurality of residual attention blocks (RABs) 502, a standard convolutional layer 510, and a CSAB 506. As a non-limiting example, Figure 5A Four RABs 502 are shown in FIG. Each RAB 502 can extract features from the input. The outputs of the second RAB 502 and the fourth RAB 502 can be connected 504 and fused by a standard convolution layer 510 with 1x1 convolution to generate a fused feature map. The fused feature map can be input to the CSAB 506. The CSAB 506 can retain important features at the channel and spatial levels.

[0058] refer to Figure 5BIn each RAB 502, two feature maps with different receptive fields are first obtained using two depth-wise separable convolutional layers 508, and then the two feature maps are concatenated 504 and fused using a standard convolutional layer 510. A channel attention block (CAB) 518 is used to emphasize the important channels generated by the two depth-wise separable convolutional layers 508 and fused by the standard convolutional layer 510. Figure 5C As shown, the depthwise separable convolution layer 508 may include a depthwise convolution 512 followed by a standard convolution layer 510 and a leaky rectified linear activation function (ReLU) layer 516 .

[0059] Fig. 6A According to some aspects of the present disclosure Figure 5B A detailed block diagram 600 of an exemplary CAB 518 is shown. Figure 6B According to some aspects of the present disclosure Figure 5A A detailed block diagram 601 of an exemplary CSAB 506 is shown. Figure 6C According to some aspects of the present disclosure Figure 6B A detailed block diagram 603 of an exemplary SAB 618 is shown. FIG. 6A to FIG. 6C Describe together.

[0060] refer to Fig. 6A , shows the architecture of CAB 518. CAB 518 can extract the weight of each channel 602 through global average pooling 604, channel compression 606 and expansion 608. Then, a sigmoid layer 610 multiplies the extracted weight with the input feature map to obtain a channel attention map 612.

[0061] refer to Figure 6B , since the quality of each WCDAB output directly affects the final performance of the loop CNN filter, the channel attention applied at the end of WCDAB 416 may still miss some suppressed channels, which contain important feature information. Therefore, CSAB 506 aims to refine the output features of WCDAB 416. Figure 5B As shown, CSAB 506 includes two parallel branches: channel attention block (CAB) 616 and spatial attention block (SAB) 618. Through these two branches, the input feature map 614 retains important feature information in the channel and spatial dimensions respectively, and then the two parts of information are fused to obtain the final output 620.

[0062] refer to Figure 6C, for the input feature map 614, parallel depthwise separable convolution layers 622a and 622b of different sizes convolve the input feature map 614. Then, the results of the two convolutions are added and activated with a ReLU layer 624. After that, another depthwise separable convolution 626 is applied, followed by a sigmoid layer 628 to obtain the spatial attention mask. Finally, a multiplier 630 multiplies the spatial attention map with the input feature map 614 to obtain the final spatial attention map 632.

[0063] Figure 7 A detailed block diagram 700 of an exemplary loop filter 702 for training a loop filter model is shown according to some aspects of the present disclosure.

[0064] refer to Figure 7 , the loop filter 702 may include LMCS 704, DBF 706, SAO 708, and ALF 710. A compressed data set (e.g., a reconstructed frame, a predicted frame, and a segmented frame) may be obtained after LMCS 704, and a label set may be obtained after ALF 710. The compressed data set and the label set may be used to train the exemplary loop CNN filter described herein. To this end, a weighted L1 loss and L2 loss may be used to train the CNN filter, for example, a weakly connected dense attention neural network (WCDANN) using a loss function f(x), as shown in equation (1) below. f(x)=8×Loss y +Loss u +Loss v (1) Where Loss represents the L1 loss or L2 loss in the y, u, and v channels. In some examples, L1 loss can be used in the early and middle stages of training, and L2 loss can be used in the late stages of training.

[0065] Fig. 8A A block diagram of an example loop filter training strategy 800 for training an example neural network (NN) filter 804a is shown. Figure 8B A block diagram of a network 801 implementing an example loop filter training strategy according to some aspects of the present disclosure is shown. Fig. 8A and Figure 8B Describe together.

[0066] refer to Fig. 8A, the network training usually uses the uncompressed image as the label 806a. However, since the input 802a is compressed with different QPs, the distance between the compressed image and the label is different, which makes the network difficult to learn for all training data. In addition, this distance severely limits the overall performance of the network.

[0067] refer to Figure 8B , the exemplary loop filter training strategy implemented by the network 801 uses inputs 802b (e.g., compressed images) with different QPs to maximize the learning ability of the network. For inputs 802b with current QP, the network 801 can use compressed images with lower QPs instead of uncompressed images as labels to train the NN filter 804b. In this way, the distance between the input 802b and the label 806b remains consistent, thereby improving the stability of the network performance and addressing the defects of traditional training methods.

[0068] Still refer to Figure 8B , the network 801 can set a parameter qp_dis, which represents the QP difference between the input 802a and the label 806b. Since the smaller the QP, the higher the quality, the QP value of the label 806b is less than the QP value of the input 802b. First, the network 801 can use a smaller qp_dis to train the NN filter 804b until convergence. Then, the network 801 can gradually increase qp_dis and continue to train the NN filter 804b. Since the loss function is a multi-level loss, the network 801 combines the loss function with the training strategy to implement a multi-level training strategy. First, the network 801 sets qp_dis=5 and trains the NN filter 804b with the L1 and L2 loss functions in turn. Then, the network 801 increases qp_dis by 10 and trains the NN filter 804b again with the L1 and L2 loss functions. After that, the network 801 increases qp_dis again and trains the NN filter 804b with the L1 and L2 loss functions. In this way, the exemplary loop filter training strategy achieves network convergence and maximizes the learning ability of network 801. Fig. 9 A first graphical representation 900 of PSNR versus bitrate for video encoding using an exemplary in-loop CNN filter is shown in accordance with some aspects of the present disclosure. Fig.10 A second graphical representation 1000 of PSNR versus bit rate for video encoding using an exemplary loop filter according to some aspects of the present disclosure is shown.

[0069] Reference again Figure 8B, since the NN filter 804b processes YUV images, the image characteristics of the UV channel and the Y channel are different. For example, when qp_dis=5, the image quality difference corresponding to the Y channel is much greater than that of the UV channel. Therefore, in actual training, qp_dis can be used only for the Y channel. For the UV channel, its label uses an uncompressed image. However, when different filters are used to process the Y channel and the UV channel independently, the image difference problem between the Y channel and the UV channel no longer exists, and the proposed training strategy can be applied to different filters.

[0070] Fig.11 1 is a flowchart of an exemplary video encoding method 1100 according to some embodiments of the present disclosure. The method 1100 may be performed by an apparatus (e.g., encoder 201, loop CNN filter 400, or any other suitable video encoding and / or compression system). The method 1100 may include operations 1102 to 1112 as described below. It is understood that some of the operations may be optional, and some of the operations may be performed simultaneously or in different order. Fig.11 Execute in the order shown.

[0071] refer to Fig.11 At 1102, the device receives a plurality of inputs generated by encoding in a feature extraction component. Figure 4 , the input of the loop CNN filter 400 may include, for example, a reconstructed frame (rec) 402a, a predicted frame (pred) 402b, a segmented frame (par) 402c, and a QP map (qp) 402d, all of which are generated by the encoder 201. The reconstructed frame 402a is a reconstruction of the current video frame by the encoder 201 for quality enhancement. The predicted frame 402b is a prediction made by the encoder 201 on the reconstructed frame 402a. The segmented frame 402c is segmentation information corresponding to the reconstructed frame 402a. The QP map 402d is used to indicate the QP value corresponding to the reconstructed frame. The QP map 402d can improve the quality of the reconstructed frames of different QPs at the reconstructed portion 405.

[0072] At 1104, the device may apply a standard convolution to the plurality of inputs via a feature extraction component to generate a first set of feature extractions. Figure 4 , the feature extraction part 401 may include multiple parallel convolutional layers 406 (standard convolutions), each of which is used to integrate and extract shallow features of its corresponding input. Afterwards, the output features of each parallel convolutional layer 406 are connected 404 and fused using a convolutional layer 408 with a stride of 1 to obtain fused shallow features. Then, a convolutional layer 410 with a stride of 2 is used to reduce the computation of the network. Finally, the downsampled features (e.g., the first set of feature extractions) are sent to the WCDAB backbone 403.

[0073] At 1106, the apparatus may receive as input a first set of feature extractions through a WCDAB backbone of a CNN of a loop filter, the first set of feature extractions being associated with a reconstructed frame. Figure 4 , the WCDAB backbone 403 may receive a first set of feature extractions from the feature extraction portion 401 .

[0074] At 1108, the device may apply a plurality of depthwise separable convolutions to the first set of feature extractions through a WCDAB backbone of a CNN with a loop filter to generate a set of global features. Figure 4 , the WCDAB backbone 403 can extract a global feature map (eg, a set of global features) from the first set of feature extractions received from the feature extraction portion 401 .

[0075] At 1110, the apparatus may generate a residual map based on a set of global features by a reconstruction component of a CNN of a loop filter. Figure 4 , the reconstruction part 405 uses a 1x1 convolution layer 408 to reduce the channel dimension of the global feature map obtained from the WCDAB backbone 403. Then, the reconstruction part 405 can use a pixel shuffling component 412 to upsample the reduced dimension features to obtain a three-channel residual map.

[0076] At 1112, the device may apply the residual map to the reconstructed frame through the reconstruction component of the CNN of the loop filter to generate an enhanced reconstructed frame. Figure 4 , the obtained residual image is added to the reconstructed frame 402a to generate an enhanced reconstructed frame 418.

[0077] Fig.12 Flowchart showing an exemplary method 1200 for training loop CNN filter patterns according to some embodiments of the present disclosure. Method 1200 may be performed by an apparatus (e.g., encoder 201, network 801, or any other suitable video encoding and / or compression system). Method 1200 may include operations 1202 to 1212 as described below. It is understood that some of these operations may be optional, and some of these operations may be performed simultaneously or in different order. Fig.12 Execute in the order shown.

[0078] refer to Fig.12 At 1202, the device may obtain, through a processor, a training data set including a reconstructed frame, a predicted frame, and a segmented frame at each QP. Figure 8B , the exemplary loop filter training strategy implemented by the network 801 uses inputs 802b with different QPs to maximize the learning ability of the network.

[0079] At 1204, the device may apply DBF, SAO, and ALF to the compressed data set via a processor. Figure 7 , the loop filter 702 may include a LMCS 704, a DBF 706, a SAO 708, and an ALF 710. A compressed data set (e.g., a reconstructed frame, a predicted frame, and a segmented frame) may be obtained after the LMCS 704, and a label set may be obtained after the ALF 710. The DBF 706, the SAO 708, and the ALF 710 may be applied to the compressed data set to obtain the label set.

[0080] At 1206, the apparatus may obtain, via a processor, a label set associated with an enhanced reconstructed frame as an output of the ALF, the label set being associated with a second set of QPs that is smaller than the first set of QPs. Figure 7 , a compressed data set (e.g., reconstructed frame, predicted frame, and segmented frame) may be obtained after LMCS 704, and a label set may be obtained after ALF 710. DBF 706, SAO 708, and ALF 710 may be applied to the compressed data set to obtain the label set.

[0081] At 1204, the device may define, via a processor, a parameter qp_dis as the difference between the input (e.g., reconstructed frames, predicted frames, and segmented frames compressed at an input QP) and the label (e.g., reconstructed frames output by the ALF compressed at an output QP). Figure 8B , the network 801 may set a parameter qp_dis, which represents the QP difference between the input 802a and the label 806b. Since a smaller QP represents a higher quality, the QP value of the label 806b is smaller than the QP value of the input 802b.

[0082] At 1206, the device may train the CNN loop filter under the current qp_dis. Figure 8B , the network 801 can use a smaller qp_dis to train the NN filter 804b.

[0083] At 1208, the device may increase the current qp_dis after the network converges. Figure 8B, the network 801 can use a smaller qp_dis to train the NN filter 804b until convergence. Then, the network 801 can gradually increase qp_dis and continue to train the NN filter 804b. Since the loss function is a multi-level loss, the network 801 combines the loss function with the training strategy to implement a multi-level training strategy. First, the network 801 sets qp_dis=5 and trains the NN filter 804b with the L1 and L2 loss functions in turn. Then, the network 801 increases qp_dis by 10 and trains the NN filter 804b again with the L1 and L2 loss functions. After that, the network 801 increases qp_dis again and trains the NN filter 804b with the L1 and L2 loss functions.

[0084] At 1210, the device may determine whether network performance is stagnant. Figure 8B , the network 801 can determine whether its performance improves with subsequent training using an increased qp_dis. If not, the operation can return to 1206; if so, the operation can move to 1212.

[0085] At 1212, the device may fix the network parameters and end the training. Figure 8B Once the network parameters are fixed, the loop CNN filter model used by the NN filter 804b is generated. In this way, the exemplary loop filter training strategy achieves network convergence and maximizes the learning ability of the network 801.

[0086] In various aspects of the present disclosure, the functions described herein may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as instructions on a non-transitory computer-readable medium. Computer-readable media include computer storage media. Storage media may be any computer program that can be processed by a processor (e.g., Figure 2 and Figure 3 As an example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, HDD (e.g., disk storage or other magnetic storage device), flash drive, SSD, or any other medium that can be used to carry or store the required program code in the form of instructions or data structures and can be accessed by a processing system (e.g., a mobile device or a computer). Disks and optical disks used herein include CDs, laser optical disks, optical disks, digital video discs (DVDs), and floppy disks, where disks typically reproduce data magnetically, while optical disks use lasers to reproduce data optically. Combinations of the above should also be included within the scope of computer-readable media.

[0087] According to one aspect of the present disclosure, a method for generating an enhanced frame in a video encoder is provided. The method may include a WCDAB backbone of a CNN of a loop filter receiving a first set of feature extractions as input. The first set of feature extractions may be associated with a reconstructed frame. The method may include a WCDAB backbone of a CNN of a loop filter applying a plurality of depthwise separable convolutions to the first set of feature extractions to generate a set of global features.

[0088] In some embodiments, the method may include a reconstruction component of the CNN of the loop filter generating a residual map based on a set of global features. In some embodiments, the method may include a reconstruction component of the CNN of the loop filter applying the residual map to the reconstructed frame to generate an enhanced reconstructed frame.

[0089] In some embodiments, the method may include a feature extraction component of a CNN of a loop filter receiving a plurality of inputs generated by encoding. In some embodiments, the method may include a feature extraction component of a CNN of a loop filter applying a standard convolution to the plurality of inputs to generate a first set of feature extractions.

[0090] In some embodiments, the plurality of inputs include reconstructed frames, predicted frames, segmented frames, and QP maps.

[0091] In some embodiments, the WCDAB backbone of the CNN of the loop filter applies multiple depthwise separable convolutions to the first set of feature extractions to generate a set of global features, which may include: applying a first depthwise convolution of the RAB to the first set of feature extractions to generate a first feature map having a first field. In some embodiments, the WCDAB backbone of the CNN of the loop filter applies multiple depthwise separable convolutions to the first set of feature extractions to generate a set of global features, which may include: applying a second depthwise convolution of the RAB to the first set of feature extractions to generate a second feature map having a second field different from the first field.

[0092] In some embodiments, the WCDAB backbone of the CNN of the loop filter applies multiple depth-separable convolutions to the first set of feature extractions to generate a set of global features, which may include: connecting the first feature map and the second feature map. In some embodiments, the WCDAB backbone of the CNN of the loop filter applies multiple depth-separable convolutions to the first set of feature extractions to generate a set of global features, which may include: generating a fused feature map by applying a standard convolution to the first feature map and the second feature map after the connection. In some embodiments, the WCDAB backbone of the CNN of the loop filter applies multiple depth-separable convolutions to the first set of feature extractions to generate a set of global features, which may include: inputting the fused feature map into the CAB to identify a set of channels from the fused feature map. In some embodiments, the fused feature map may include a set of global features.

[0093] In some embodiments, the reconstruction component of the CNN of the loop filter generates a residual map based on a set of global features, which may include: inputting the CAB and SAB of the CSAB into the group of channels identified from the fused feature map. In some embodiments, the reconstruction component of the CNN of the loop filter generates a residual map based on a set of global features, which may include: generating a set of channel dimension features from the group of channels using CAB. In some embodiments, the reconstruction component of the CNN of the loop filter generates a residual map based on a set of global features, which may include: generating a set of spatial dimension features from the group of channels using SAB. In some embodiments, the reconstruction component of the CNN of the loop filter generates a residual map based on a set of global features, which may include: fusing the group of channel dimension features and the group of spatial dimension features to generate a residual map.

[0094] In some embodiments, the set of spatial dimensional features is generated by applying a third depthwise convolution of a first size and a fourth depthwise convolution of a second size different from the first size to the set of channels.

[0095] According to another aspect of the present disclosure, a system for generating an enhanced frame in a video encoder is provided. The system may include a memory for storing instructions. The system may include a processor coupled to the memory and used when executing the instructions to: receive a first set of feature extractions as input through a WCDAB backbone of a CNN of a loop filter. The first set of feature extractions may be associated with a reconstructed frame. The system may include a processor coupled to the memory and used when executing the instructions to: apply multiple depth-separable convolutions to the first set of feature extractions through a WCDAB backbone of a CNN of a loop filter to generate a set of global features.

[0096] In some embodiments, the processor coupled to the memory can also be used when executing the instructions to: generate a residual map based on a set of global features by a reconstruction component of the CNN through a loop filter. In some embodiments, the processor coupled to the memory can also be used when executing the instructions to: apply the residual map to the reconstructed frame by the reconstruction component of the CNN through the loop filter to generate an enhanced reconstructed frame.

[0097] In some embodiments, the processor coupled to the memory can also be used when executing the instructions to: receive multiple inputs generated by encoding through the feature extraction component of the CNN of the loop filter. In some embodiments, the processor coupled to the memory can also be used when executing the instructions to: apply standard convolution to the multiple inputs through the feature extraction component of the CNN of the loop filter to generate a first set of feature extractions.

[0098] In some embodiments, the multiple inputs may include reconstructed frames, predicted frames, segmented frames, and QP maps.

[0099] In some embodiments, the processor coupled to the memory, when executing the instructions, can be used to apply multiple depth-wise separable convolutions to a first set of feature extractions through a WCDAB backbone of a CNN of a loop filter to generate a set of global features as follows: apply a first depth-wise convolution of the RAB to the first set of feature extractions to generate a first feature map having a first field. In some embodiments, the processor coupled to the memory, when executing the instructions, can be used to apply multiple depth-wise separable convolutions to a first set of feature extractions through a WCDAB backbone of a CNN of a loop filter to generate a set of global features as follows: apply a second depth-wise convolution of the RAB to the first set of feature extractions to generate a second feature map having a second field different from the first field.

[0100] In some embodiments, the processor coupled to the memory may be used, when executing the instructions, to apply multiple depth-separable convolutions to the first set of feature extractions through the WCDAB backbone of the CNN of the loop filter to generate a set of global features as follows: connect the first feature map and the second feature map. In some embodiments, the processor coupled to the memory may be used, when executing the instructions, to apply multiple depth-separable convolutions to the first set of feature extractions through the WCDAB backbone of the CNN of the loop filter to generate a set of global features as follows: generate a fused feature map by applying a standard convolution to the first feature map and the second feature map after the connection. In some embodiments, the processor coupled to the memory may be used, when executing the instructions, to apply multiple depth-separable convolutions to the first set of feature extractions through the WCDAB backbone of the CNN of the loop filter to generate a set of global features as follows: input the fused feature map into the CAB to identify a set of channels from the fused feature map. In some embodiments, the fused feature map may include a set of global features.

[0101] In some embodiments, the processor coupled to the memory can be used to generate a residual map based on a set of global features by the reconstruction component of the CNN through the loop filter when executing the instruction: the CAB and SAB of the CSAB are input to the group of channels identified from the fusion feature map. In some embodiments, the processor coupled to the memory can be used to generate a residual map based on a set of global features by the reconstruction component of the CNN through the loop filter when executing the instruction: a set of channel dimension features are generated from the group of channels using the CAB. In some embodiments, the processor coupled to the memory can be used to generate a residual map based on a set of global features by the reconstruction component of the CNN through the loop filter when executing the instruction: a set of spatial dimension features are generated from the group of channels using the SAB. In some embodiments, the processor coupled to the memory can be used to generate a residual map based on a set of global features by the reconstruction component of the CNN through the loop filter when executing the instruction: the group of channel dimension features and the group of spatial dimension features are fused to generate a residual map.

[0102] In some embodiments, the set of spatial dimensional features may be generated by applying a third depthwise convolution of a first size and a fourth depthwise convolution of a second size different from the first size to the set of channels.

[0103] According to another aspect of the present disclosure, a method for training a loop filter model of a video encoder is provided. The method may include a processor obtaining a compressed data set including a reconstructed frame, a predicted frame, and a segmented frame. The compressed data set may be associated with a first set of QPs. The method may include a processor applying DBF, SAO, and ALF to the compressed data set. The method may include a processor obtaining a label set associated with an enhanced reconstructed frame as an output of the ALF. The label set may be associated with a second set of QPs that is less than the first set of QPs.

[0104] In some embodiments, the method may include the processor generating a loop filter model based on the multi-stage loss function and a label set including a second set of QPs that is smaller than the first set of QPs.

[0105] According to another aspect of the present disclosure, a system for training a loop filter model of a video encoder is provided. The system may include a memory for storing instructions. The system may include a processor coupled to the memory and used to obtain a compressed data set including a reconstructed frame, a predicted frame, and a segmented frame when executing the instructions. The compressed data set may be associated with a first set of QPs. The system may include a processor coupled to the memory and used to apply DBF, SAO, and ALF to the compressed data set when executing the instructions. The system may include a processor coupled to the memory and used to obtain a label set associated with an enhanced reconstructed frame as an output of the ALF when executing the instructions. The label set is associated with a second set of QPs that is smaller than the first set of QPs.

[0106] In some embodiments, the processor coupled to the memory, when executing the instructions, may also be used to generate a loop filter model based on the multi-level loss function and a label set including a second set of QPs that is smaller than the first set of QPs.

[0107] The above description of the embodiments will reveal the general nature of the present disclosure, and others can easily modify and / or adjust these embodiments for various applications without departing from the general concept of the present disclosure by applying knowledge within the technical scope of the art, without excessive experimentation. Therefore, based on the teachings and guidance provided herein, such adjustments and modifications are intended to belong to the meaning and scope of the equivalents of the disclosed embodiments. It should be understood that the wording or terminology herein is for description rather than for limitation, and therefore, those skilled in the art should interpret the terms or wording of this specification in accordance with the teachings and guidance.

[0108] The disclosed embodiments have been described above with the aid of functional building blocks, which illustrate the implementation of their specific functions and their relationships. For ease of description, the boundaries of these functional building blocks are arbitrarily defined herein. As long as their specific functions and their relationships are properly performed, other boundaries may also be defined.

[0109] The Summary and Abstract sections may set forth one or more but not all exemplary embodiments of the present disclosure as contemplated by the inventor(s), and thus, are not intended to limit the present disclosure and the appended claims in any way.

[0110] Various functional blocks, modules, and steps are disclosed above. The arrangements provided are illustrative and not restrictive. Therefore, the functional blocks, modules, and steps may be reordered or combined in a manner different from the examples provided above. Likewise, some embodiments include only a subset of the functional blocks, modules, and steps, and any such subset is permitted.

[0111] The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.

Claims

1. A method for generating an enhanced frame in a video encoder, comprising: A Weakly Connected Dense Attention Block (WCDAB) backbone of a convolutional neural network (CNN) of a loop filter receives as input a first set of feature extractions associated with a reconstructed frame; and The WCDAB backbone of the CNN of the loop filter applies a plurality of depthwise separable convolutions to the first set of feature extractions to generate a set of global features.

2. The method according to claim 1, further comprising: A reconstruction component of the CNN of the loop filter generates a residual map based on the set of global features; as well as The reconstruction component of the CNN of the loop filter applies the residual map to the reconstructed frame to generate an enhanced reconstructed frame.

3. The method according to claim 2, further comprising: The feature extraction component of the CNN of the loop filter receives a plurality of inputs generated by encoding; as well as The feature extraction component of the CNN of the loop filter applies a standard convolution to the plurality of inputs to generate the first set of feature extractions.

4. The method according to claim 3, wherein: The plurality of inputs include the reconstructed frame, the predicted frame, the segmented frame, and a quantization parameter (QP) map.

5. The method according to claim 2, wherein: The WCDAB backbone of the CNN of the loop filter applying the plurality of depthwise separable convolutions to the first set of features to generate the set of global features comprises: applying a first depthwise convolution of a residual attention block (RAB) to the first set of feature extractions to generate a first feature map having a first field; and A second depthwise convolution of the RAB is applied to the first set of features to generate a second feature map having a second field different from the first field.

6. The method according to claim 5, wherein: The WCDAB backbone of the CNN of the loop filter applying the plurality of depthwise separable convolutions to the first set of features to generate the set of global features comprises: connecting the first feature map and the second feature map; generating a fused feature map by applying a standard convolution to the first feature map and the second feature map after the concatenation; and Inputting the fused feature map into a channel attention block (CAB) to identify a set of channels from the fused feature map, Wherein, the fused feature map includes the set of global features.

7. The method according to claim 6, wherein: The reconstruction component of the CNN of the loop filter generates the residual map based on the set of global features, comprising: Inputting the set of channels identified from the fused feature map into a channel attention branch (CAB) and a spatial attention branch (SAB) of a channel-spatial joint attention block (CSAB); generating a set of channel-dimensional features from the set of channels using the CAB; generating a set of spatial dimension features from the set of channels using the SAB; and The set of channel dimension features and the set of spatial dimension features are fused to generate the residual map.

8. The method according to claim 7, wherein: The set of spatial dimensional features is generated by applying a third depthwise convolution of a first size and a fourth depthwise convolution of a second size different from the first size to the set of channels.

9. A system for generating an enhanced frame in a video encoder, comprising: A memory for storing instructions; as well as A processor, coupled to the memory and configured, when executing the instructions, to: A weakly connected dense attention block (WCDAB) backbone of a convolutional neural network (CNN) through a loop filter receives as input a first set of feature extractions associated with a reconstructed frame; as well as The WCDAB backbone of the CNN through the loop filter applies a plurality of depthwise separable convolutions to the first set of feature extraction to generate a set of global features.

10. The system according to claim 9, wherein: The processor coupled to the memory, when executing the instructions, is further configured to: A reconstruction component of the CNN through the loop filter generates a residual map based on the set of global features; and The reconstruction component of the CNN through the loop filter applies the residual map to the reconstructed frame to generate an enhanced reconstructed frame.

11. The system according to claim 10, wherein: The processor coupled to the memory, when executing the instructions, is further configured to: The feature extraction component of the CNN through the loop filter receives a plurality of inputs generated by encoding; and The feature extraction component of the CNN through the loop filter applies a standard convolution to the plurality of inputs to generate the first set of feature extractions.

12. The system according to claim 11, wherein: The plurality of inputs include the reconstructed frame, the predicted frame, the segmented frame, and a quantization parameter (QP) map.

13. The system according to claim 10, wherein: The processor coupled to the memory, when executing the instructions, is to apply the plurality of depthwise separable convolutions to the first set of feature extractions through the WCDAB backbone of the CNN of the loop filter to generate the set of global features as follows: extracting and applying a first depthwise convolution of a residual attention block (RAB) to the first set of features to generate a first feature map having a first field; as well as A second depthwise convolution of the RAB is applied to the first set of features to generate a second feature map having a second field different from the first field.

14. The system according to claim 13, wherein: The processor coupled to the memory, when executing the instructions, is to apply the plurality of depthwise separable convolutions to the first set of feature extractions through the WCDAB backbone of the CNN of the loop filter to generate the set of global features as follows: connecting the first feature map and the second feature map; generating a fused feature map by applying a standard convolution to the first feature map and the second feature map after the concatenation; as well as Inputting the fused feature map into a channel attention block (CAB) to identify a set of channels from the fused feature map, Wherein, the fused feature map includes the set of global features.

15. The system of claim 14, wherein: The processor coupled to the memory is used, when executing the instructions, to generate the residual map based on the set of global features by the reconstruction component of the CNN through the loop filter as follows: Inputting the set of channels identified from the fused feature map into a channel attention branch (CAB) and a spatial attention branch (SAB) of a channel-spatial joint attention block (CSAB); generating a set of channel-dimensional features from the set of channels using the CAB; generating a set of spatial dimension features from the set of channels using the SAB; and The set of channel dimension features and the set of spatial dimension features are fused to generate the residual map.

16. The system of claim 15, wherein: The set of spatial dimensional features is generated by applying a third depthwise convolution of a first size and a fourth depthwise convolution of a second size different from the first size to the set of channels.

17. A method for training a loop filter model of a video encoder, comprising: The processor obtains a compressed data set including a reconstructed frame, a predicted frame, and a segmented frame, the compressed data set being associated with a first set of quantization parameters (QP); The processor applies a deblocking filter (DBF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF) to the compressed data set; as well as The processor obtains a set of labels associated with an enhanced reconstructed frame as an output of the ALF, the set of labels being associated with a second set of QPs that is smaller than the first set of QPs.

18. The method according to claim 17, further comprising: The processor generates the loop filter model based on a multi-level loss function and the label set including the second set of QPs that are smaller than the first set of QPs.

19. A system for training a loop filter model of a video encoder, comprising: A memory for storing instructions; as well as A processor, coupled to the memory and configured, when executing the instructions, to: Obtaining a compressed data set including a reconstructed frame, a predicted frame, and a segmented frame, wherein the compressed data set is associated with a first set of quantization parameters (QP); applying a deblocking filter (DBF), a sample adaptive offset (SAO), and an adaptive loop filter (ALF) to the compressed data set; as well as A set of labels associated with an enhanced reconstructed frame is obtained as an output of the ALF, the set of labels being associated with a second set of QPs that is smaller than the first set of QPs.

20. The system of claim 19, wherein: The processor coupled to the memory, when executing the instructions, is further configured to: The loop filter model is generated based on a multi-level loss function and the label set including the second set of QPs that are smaller than the first set of QPs.