Super-resolution convolutional neural network filter with reference image resampling function in multifunctional video coding

By using a lightweight multi-level hybrid scale and depth information and attention mechanism (LMSDA) network, the problem of difficult to maintain high-definition video transmission image quality in the prior art is solved, and efficient video encoding performance is achieved.

CN120035847APending Publication Date: 2025-05-23GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280100972.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-13
Filing Date
2022-12-05
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing video encoding technology is difficult to maintain when processing high-definition video, especially when bandwidth is limited, and traditional interpolation methods have limitations when processing video complex characteristics.

Method used

A lightweight multi-level hybrid scale and depth information and attention mechanism (LMSDA) network is used to receive input images through the head of the LMSDA network, extract features, and generate multi-scale and depth information through multiple LMSDA blocks (LMSDABs), and feature extraction and reconstruction are carried out in combination with attention mechanism.

Benefits of technology

This improves the image quality of video encoding, reduces network complexity and computing complexity, and enhances the performance of video encoding, especially when bandwidth is limited.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120035847A_ABST
    Figure CN120035847A_ABST
Patent Text Reader

Abstract

According to an aspect of the present disclosure, a video encoding method is provided. The method may include a lightweight multi-level hybrid scale and depth information and attention mechanism (LMSDA) network receiving an input image. The method may include the LMSDA network extracting a first set of features from an input image. The method may include the LMSDA network inputting a first set of features through a plurality of LMSDA blocks (LMSDABs). The method may include the LMSDA network generating a second set of features based on an output of the LMSDAB. The method may include the LMSDA network using a convolutional layer to reduce a number of channels associated with a second set of features of the LMSDAB output. The method may include the LMSDA network upsampling the second set of features to generate an enhanced output image.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] The disclosed embodiments relate to video coding.

[0002] Digital video has become mainstream and is widely used in various applications such as digital television, video phones and teleconferencing. These digital video applications are feasible due to advances in computing and communication technology and efficient video coding technology. Various video coding techniques can be used to compress video data, so that the video data can be encoded using one or more video coding standards. Exemplary video coding standards may include, but are not limited to, versatile video coding (H.266 / VVC), high-efficiency video coding (H.265 / HEVC), advanced video coding (H.264 / AVC), moving picture expert group (MPEG) coding, etc. Summary of the invention

[0003] According to one aspect of the present disclosure, a video encoding method is provided. The method may include a head of a lightweight multi-level mixed scale and depth information with attention mechanism (LMSDA) network receiving an input image. The method may include the head of the LMSDA network extracting a first set of features from the input image. The method may include a trunk portion of the LMSDA network inputting a first set of features through a plurality of LMSDA blocks (LMSDAB). The method may include a trunk portion of the LMSDA network generating a second set of features based on the output of the LMSDAB. The method may include a reconstruction portion of the LMSDA network upsampling the second set of features to generate an enhanced output image.

[0004] According to another aspect of the present disclosure, a video encoding system is provided. The system may include a memory for storing instructions. The system may include a processor coupled to the memory, which, when executing the instructions, is used to receive an input image through the head of an LMSDA network. The system may include a processor coupled to the memory, which, when executing the instructions, is used to extract a first set of features from the input image through the head of the LMSDA network. The system may include a processor coupled to the memory, which, when executing the instructions, is used to input a first set of features through a plurality of LMSDABs through the trunk portion of the LMSDA network. The system may include a processor coupled to the memory, which, when executing the instructions, is used to generate a second set of features based on the output of the LMSDAB through the trunk portion of the LMSDA network. The system may include a processor coupled to the memory, which, when executing the instructions, is used to upsample the second set of features through the reconstruction portion of the LMSDA network to generate an enhanced output image.

[0005] According to another aspect of the present disclosure, a video encoding method is provided. The method may include a feature extraction part of LMSDAB applying a first convolution layer of a first kernel size and a second convolution layer of a second kernel size to a first set of features to generate a second set of features. The method may include the feature extraction part of LMSDAB combining the second set of features in a channel dimension. The method may include the feature extraction part of LMSDAB fusing the second set of features combined in the channel dimension using a third convolution layer of the first kernel size to generate a fused feature map.

[0006] According to another aspect of the present disclosure, a video encoding system is provided. The system may include a memory for storing instructions. The system may include a processor coupled to the memory, which, when executing the instructions, is used to apply a first convolution layer of a first kernel size and a second convolution layer of a second kernel size to a first set of features through a feature extraction part of LMSDAB to generate a second set of features. The system may include a processor coupled to the memory, which, when executing the instructions, is used to combine the second set of features in a channel dimension through a feature extraction part of LMSDAB. The system may include a processor coupled to the memory, which, when executing the instructions, is used to fuse the second set of features combined in the channel dimension through a third convolution layer of a first kernel size through a feature extraction part of LMSDAB to generate a fused feature map.

[0007] These illustrative embodiments are mentioned not to limit or define the present disclosure, but to provide examples to aid understanding of the present disclosure. Other embodiments are described in the detailed description and further description is provided. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] The accompanying drawings, which are incorporated herein and constitute a part of the specification, illustrate embodiments of the present disclosure and, together with the description, further explain the principles of the present disclosure and enable one skilled in the relevant art to make and use the present disclosure.

[0009] Figure 1 A block diagram of an exemplary encoding system according to some embodiments of the present disclosure is shown.

[0010] Figure 2 A block diagram of an exemplary decoding system according to some embodiments of the present disclosure is shown.

[0011] Figure 3 A detailed block diagram showing an exemplary lightweight multi-level mixed scale and depth information and attention mechanism (LMSDA) network for the luminance channel according to some embodiments of the present disclosure.

[0012] Figure 4 A detailed block diagram of an exemplary LMSDA network for a chroma channel according to some embodiments of the present disclosure is shown.

[0013] Figure 5 A detailed block diagram of an exemplary LMSDA block (LMSDAB) is shown according to some embodiments of the present disclosure.

[0014] Fig. 6A An example multi-scale feature extraction component is shown.

[0015] Figure 6B An exemplary multi-scale feature extraction component according to some embodiments of the present disclosure is shown.

[0016] Fig. 7A A first exemplary convolutional model according to some aspects of the present disclosure is shown.

[0017] Figure 7B A second exemplary convolutional model according to some aspects of the present disclosure is shown.

[0018] Figure 7C A third exemplary convolutional model according to some aspects of the present disclosure is shown.

[0019] Figure 8 A detailed block diagram of an exemplary channel attention block (CAB) according to some aspects of the present disclosure is shown.

[0020] Fig. 9 A detailed block diagram of an exemplary multi-scale spatial attention block (MSSAB) according to some aspects of the present disclosure is shown.

[0021] Fig.10 A detailed block diagram of an exemplary spatial attention block (SAB) according to some aspects of the present disclosure is shown.

[0022] Fig.11A flowchart illustrating a first exemplary video encoding method according to some aspects of the present disclosure is shown.

[0023] Fig.12 A flowchart of a second exemplary video encoding method according to some aspects of the present disclosure is shown.

[0024] Embodiments of the present disclosure will be described with reference to the accompanying drawings. DETAILED DESCRIPTION

[0025] Although some configurations and arrangements are discussed, it should be understood that this is for illustrative purposes only. Relevant technicians will recognize that other configurations and arrangements can be used without departing from the spirit and scope of the present disclosure. Relevant technicians will understand that the present disclosure can also be used in various other applications.

[0026] Note that the phrases "one embodiment", "embodiment", "example embodiment", "some embodiments", "certain embodiments", etc. mentioned in the specification indicate that the described embodiments may include certain features, structures, or characteristics, but not every embodiment necessarily includes the certain features, structures, or characteristics. Furthermore, these phrases do not necessarily refer to the same embodiment. Furthermore, when a certain feature, structure, or characteristic is described in conjunction with an embodiment, whether or not it is explicitly described, the relevant technicians should know that such feature, structure, or characteristic can be implemented in conjunction with other embodiments.

[0027] In general, terms can be understood at least in part based on usage in context. For example, depending at least in part on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, or characteristics in a plural sense. Similarly, depending at least in part on the context, terms such as "one," "an," or "the" can also be understood to express singular or plural usage. In addition, also depending at least in part on the context, the term "based on" can be understood to not necessarily express a set of exclusive factors, but can allow for the presence of other factors that are not necessarily explicitly described.

[0028] Various aspects of the video encoding system will now be described with reference to various devices and methods. These devices and methods will be described in the following detailed description and illustrated in the accompanying drawings by various modules, components, circuits, steps, operations, processes, algorithms, etc. (collectively referred to as "elements"). These elements can be implemented using electronic hardware, firmware, computer software, or any combination thereof. Whether these elements are implemented as hardware, firmware, or software depends on the specific application and design constraints imposed on the overall system.

[0029] The techniques described herein can be used in various video coding applications. As described herein, video coding includes encoding and decoding a video. The encoding and decoding of a video can be performed in units of blocks. For example, a coding block, a transform block, or a prediction block can be subjected to encoding / decoding processing, such as transformation, quantization, prediction, loop filtering, reconstruction, etc. As described herein, the block to be encoded / decoded will be referred to as a "current block". For example, a current block may represent a coding block, a transform block, or a prediction block according to a current encoding / decoding process. In addition, it should be understood that the term "unit" used in the present disclosure represents a basic unit for performing a specific encoding / decoding process, and the term "block" represents a sample array of a predetermined size. Unless otherwise specified, "block", "unit", and "component" can be used interchangeably.

[0030] In recent years, the development of imaging and display technology has led to an explosive growth in high-definition video. Although video coding technology has improved significantly, it is still challenging to transmit high-definition video, especially in the case of limited bandwidth. To cope with this problem, one existing strategy is resampling-based video coding. In resampling-based video coding, the video is first downsampled before encoding, and then the decoded video is upsampled to the same resolution as the original video. The AOMediaVideo 1 (AV1) format includes a mode in which frames are encoded at a low resolution and then upsampled to the original resolution at the decoder through bilinear or bicubic interpolation. Versatile Video Coding (VVC) also supports a resampling-based coding scheme called reference picture resampling (RPR), which performs temporal prediction between different resolutions. The advantages of RPR include, for example, 1) reducing the video encoding bitstream and the amount of network bandwidth used to transmit the encoded bitstream, and 2) reducing video encoding and decoding delays. For example, after downsampling, the image resolution becomes smaller, thereby increasing the speed of the video encoding / decoding (codec) process.

[0031] Although RPR has certain advantages, the image quality after upsampling still needs to be maintained. Unfortunately, traditional interpolation methods have limitations when dealing with the complex characteristics of videos.

[0032] For neural network-based video coding (NNVC), the main ideas are as follows. The first is to use the difference in receptive fields brought by different convolution kernel sizes to extract different scale information from the input feature map and use it as the basic block of the operation. The network stacked by these basic blocks can indeed improve the performance of the output. But this may require a large number of network parameters as a basis. In addition, while paying attention to scale information, the other direction often ignores the layer depth information. The second is to reduce network parameters through separable convolution, which includes depth-separable convolution and spatial separable convolution. Separable convolution can reduce the number of network parameters, but the problem still exists. For example, if the dimension of the channel is not high, using the rectified linear activation function (ReLU) as the activation function will cause information loss in the depth-separable convolution. An existing solution is to perform a standard convolution before the depth-separable convolution to increase its dimension.

[0033] For some video coding technologies based on convolutional neural network (CNN), residual learning is used. If the learning ability of the network is function W, and the input of the network is f in , then the output is f out , denoted as f out =f in +W(f in ). Compared to the network directly learning the entire image, residual learning makes the network simpler by learning the residual between the input and output. This simplification is because the network learns a more accurate mapping through residual connections. Even in the worst case, residual learning can ensure that the output quality does not deteriorate, which makes the network learn faster and easier. Therefore, residual learning greatly reduces the complexity of network learning and has been widely used. However, due to the different sizes of input and output in super-resolution (SR) tasks, residual learning cannot be directly applied. Therefore, residual learning can only be used in feature space because the dimensions of feature space are consistent.

[0034] At present, CNN and generative adversarial network (GAN) are commonly used for network learning. CNN uses L1 or L2 loss to make the output gradually approach the true value during the network convergence process. For SR tasks, the high-resolution map output by the network is required to be consistent with the true value. L1 or L2 loss is a loss function that compares at the pixel level. L1 loss calculates the sum of the absolute values ​​of the difference between the output and the true value, while L2 loss calculates the sum of the squares of the difference between the output and the true value. Although CNN using L1 or L2 loss can remove block artifacts and noise in the input image, it cannot restore the lost texture in the input image. GAN can improve the perceptual quality to generate credible results. GAN-based methods such as the method implemented by deep convolutional generative adversarial network (DCGAN) can achieve ideal texture and detail information restoration. Through adversarial learning of the generator and the discriminator, GAN can generate texture information lost in the input. However, due to the randomness of the texture information generated by GAN, it may be inconsistent with the true value. Although rich textures can be generated from the input image, these textures are far from the true value. In other words, although the GAN-based method improves the perceptual quality and visual effect, it increases the difference between the output and the true value, thus reducing the peak-signal-to-noise ratio (PSNR) performance.

[0035] To overcome these and other challenges, the present disclosure provides an exemplary LMSDA network comprising a CNN for RPR-based SR in VVC. The exemplary LMSDA network is designed for residual learning to reduce network complexity and improve learning ability. The basic block of the LMSDA network combined with the attention mechanism is called "LMSDAB". Using LMSDAB, the LMSDA network can extract multi-scale and depth information of image features. For example, multi-scale information can be extracted by convolution kernels of different sizes, and depth information can be extracted from different depths of the network. For LMSDAB, convolutional layers are shared to greatly reduce the number of network parameters.

[0036] For example, the exemplary LMSDA network effectively extracts low-level features in the U-Net structure by stacking LMSDAB, and transmits the low-level features in the U-Net structure to the high-level feature extraction module through the U-Net connection. High-level features may include global semantic information, while low-level features contain local detail information. Therefore, the U-Net connection further reuses low-level features while restoring local details. After extracting multi-scale and layer information, LMSDAB can implement an attention mechanism to enhance important information while weakening unimportant information. In addition, the present disclosure provides an exemplary multi-scale attention mechanism that combines the multi-scale spatial attention maps obtained by convolution after performing spatial attention on each scale information. Then, channel attention can be combined to enhance the feature maps extracted by LMSDAB in the spatial and channel domains. The following is combined with Figures 1 to 12 Provide more details of an exemplary LMSDA network and its multi-scale attention mechanism.

[0037] Figure 1 A block diagram of an exemplary encoding system 100 is shown according to some embodiments of the present disclosure. Figure 2 A block diagram of an exemplary decoding system 200 according to some embodiments of the present disclosure is shown. Each system 100 or 200 can be applied to or integrated into various systems and devices capable of data processing, such as computers and wireless communication devices. For example, the system 100 or 200 can be all or part of a mobile phone, a desktop computer, a laptop computer, a tablet computer, a car computer, a game console, a printer, a positioning device, a wearable electronic device, a smart sensor, a virtual reality (VR) device, an augmented reality (AR) device, or any other suitable electronic device with data processing capabilities. Figure 1 and Figure 2 As shown, system 100 or 200 may include processor 102, memory 104, and interface 106. These components are shown as being interconnected via a bus, but other connection types are also allowed. It is understood that system 100 or 200 may include any other suitable components for performing the functions described herein.

[0038] Processor 102 may include a microprocessor, such as a graphics processing unit (GPU), an image signal processor (ISP), a central processing unit (CPU), a digital signal processor (DSP), a tensor processing unit (TPU), a vision processing unit (VPU), a neural processing unit (NPU), a synergistic processing unit (SPU) or a physical processing unit (PPU), a microcontroller unit (MCU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device (PLD), a state machine, gated logic, discrete hardware circuits, and other suitable hardware for performing the various functions described in the present disclosure. Although Figure 1 and Figure 2 Only one processor is shown, but it is understood that multiple processors may be included. Processor 102 may be a hardware device having one or more processing cores. Processor 102 may execute software. Whether referred to as software, firmware, middleware, microcode, hardware description language, or otherwise, software should be broadly understood as instructions, instruction sets, codes, code segments, program codes, programs, subroutines, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, processes, functions, etc. Software may include computer instructions written in an interpreted language, a compiled language, or machine code. Other techniques for directing hardware are also permitted under the broad category of software.

[0039] Memory 104 may broadly include memory (also referred to as main / system memory) and storage (also referred to as secondary memory). For example, memory 104 may include random-access memory (RAM), read-only memory (ROM), static RAM (SRAM), dynamic RAM (DRAM), ferroelectric RAM (FRAM), electrically erasable programmable ROM (EEPROM), compact disc read-only memory (CD-ROM) or other optical storage, hard disk drive (HDD) (such as magnetic disk storage or other magnetic storage device), flash drive, solid-state drive (SSD), or any other medium that can be used to carry or store the desired program code in the form of instructions that can be accessed and executed by processor 102. In a broad sense, memory 104 can be implemented by any computer-readable medium, such as a non-transitory computer-readable medium. Although Figure 1 and Figure 2 Only one memory is shown, but it will be appreciated that multiple memories may be included.

[0040] The interface 106 may broadly include a data interface and a communication interface for receiving and sending signals in the process of receiving and sending information with other external network elements. For example, the interface 106 may include an input / output (I / O) device and a wired or wireless transceiver. Figure 1 and Figure 2 Only one memory is shown, but it will be appreciated that multiple interfaces may be included.

[0041] The processor 102, the memory 104, and the interface 106 can be implemented in various forms in the system 100 or 200 to perform video encoding functions. In some embodiments, the processor 102, the memory 104, and the interface 106 of the system 100 or 200 are implemented (e.g., integrated) on one or more system-on-chip (SoC). In one example, the processor 102, the memory 104, and the interface 106 can be integrated on an application processor (AP) SoC, which is responsible for application processing in an operating system (OS) environment, including running video encoding and decoding applications. In another example, the processor 102, the memory 104, and the interface 106 can be integrated on a dedicated processor chip such as a GPU or an ISP chip for image and video processing in, for example, a real-time operating system (RTOS) dedicated to video encoding.

[0042] As Figure 1 shown, in the encoding system 100, the processor 102 can include one or more modules, such as the encoder 101 (also referred to herein as the "preprocessing network"). Although Figure 1 the encoder 101 is shown to be located within one processor 102, it can be understood that the encoder 101 can include one or more sub-modules, which can be implemented on different processors that are close to or far from each other. The encoder 101 (and any corresponding sub-modules or sub-units) can be a hardware unit (e.g., a part of an integrated circuit) of the processor 102 designed to be used with other components, or a software unit implemented by the processor 102 by executing at least a part of a program (i.e., instructions). The instructions of the program can be stored on a computer-readable medium such as the memory 104, and when executed by the processor 102, can perform a process having one or more functions related to video encoding, such as image partitioning, inter-frame prediction, intra-frame prediction, transformation, quantization, filtering, entropy encoding, etc., as described in detail below.

[0043] Similarly, as Figure 2 shown, in the decoding system 200, the processor 102 can include one or more modules, such as the decoder 201 (also referred to herein as the "post-processing network"). Although Figure 2The decoder 201 is shown to be located within one processor 102, but it is understood that the decoder 201 may include one or more submodules that may be implemented on different processors that are close to or far from each other. The decoder 201 (and any corresponding submodules or subunits) may be a hardware unit (e.g., part of an integrated circuit) of the processor 102 designed for use with other components, or a software unit implemented by the processor 102 by executing at least a portion (i.e., instructions) of a program. The instructions of the program may be stored on a computer-readable medium such as a memory 104, and when executed by the processor 102, a process having one or more functions related to video decoding may be performed, such as entropy decoding, inverse quantization, inverse transformation, inter-frame prediction, intra-frame prediction, filtering, as described in detail below.

[0044] Back to Figure 1 , for the RPR function, the encoder 101 first downsamples the current video frame to reduce the transmission bit stream in the limited bandwidth. When the decoder 201 restores the current frame, the current frame is upsampled to its original resolution. In order to process the complex characteristics of the video with high precision, the encoder 101 may include an exemplary LMSDA network (e.g., an SR neural network) that replaces the upsampling algorithm in the RPR configuration. The exemplary LMSDA network of the encoder 101 adopts residual learning to reduce the network learning complexity and improve performance. Residual learning restores image details with high precision because the image details are contained in the residual.

[0045] The basic block of the LMSDA network is LMSDAB, which applies different sizes of convolution kernels and convolution layer depths while using fewer parameters. LMSDAB extracts multi-scale information and depth information, which is combined with the attention mechanism to complete feature extraction. Since residual learning cannot be directly applied to SR, the LMSDA network first upsamples the input image to the same resolution as the output by interpolation. Then, the LMSDA network enhances the image quality through residual learning. In some embodiments, LMSDAB uses 1x1 and 3x3 convolution operators while using shared convolution layers to reduce the number of parameters. This can enable LMSDAB to extract larger scale features to obtain multi-scale information.

[0046] At the same time, layer depth information is also captured by sharing convolutional layers. The LMSDA network enhances image features through multi-scale spatial attention block (MSSAB) and channel attention block (CAB). MSSAB and CAB learn attention maps in the spatial and channel domains, respectively, and apply attention operations to these dimensions of the acquired feature maps to enhance important spatial and channel information. Figures 3 to 12Provide more details about the LMSDA network and LMSDAB.

[0047] Figure 3 FIG. 4 shows a detailed block diagram of an exemplary LMSDA network 300 (hereinafter referred to as "LMSDA network 300") for a luminance channel (e.g., Y channel) according to some embodiments of the present disclosure. As Figure 3 shown, the LMSDA network 300 includes, for example, a head 301, a backbone portion 303, and a reconstruction portion 305.

[0048] The head 301 includes a convolutional layer 304 for extracting shallow features of the input image. The convolutional layer 304 in the head 301 is followed by a rectified linear activation function (ReLU) activation function (not shown). Given an input Y LR , the shallow feature f is obtained through the head network ψ according to Equation (1) 0 . f 0 = ψ(Y LR ) (1).

[0049] The backbone portion 303 may include M LMSDABs 306. The backbone portion 303 uses f 0 as an input. The connector 310 connects the LMSDAB outputs, and finally the number of channels is reduced through a 1x1 convolutional layer 312 to obtain f according to Equation (3) ft . f ft can be used as an input to the reconstruction portion 305. To utilize low-level features, the backbone portion 303 uses the connection method in U-Net to add the outputs f i and f M-i of the i-th and (M - i)-th LMSDABs as inputs to ω M-i+1 . f M-i+1 = ω M-i+1 (f i + f M-i ) 0 < i < M / 2 (2); and f ft = Conv(C[ω M , ω M-1 , … ω 1 (f 0 )]) + f 0 (3), where ω i represents the M-th LMSDAB 306, C[.] represents channel concatenation, and f irepresents the output of the Mth LMSDAB 306. Channel connection may refer to stacking features on the channel dimension. For example, suppose the dimensions of two feature maps are B x C1 x H x W and B x C2 x H x W. After connection, the dimensions become B x (C1 + C2) x H x W.

[0050] The reconstruction part 305 (eg, upsampling network) includes a convolution layer 304 and a pixel shuffling layer 316. The upsampling network can be expressed according to equation (4). Y HR =PS(Conv(f ft ))+Y LR (4) where Y HR It is an upsampled image, PS is a pixel shuffle layer, Conv represents a convolutional layer, and the upsampling part does not use the ReLU activation function.

[0051] In addition to the above three parts, the input can also be upsampled by the upsampling bicubic component 308, and the input image can be directly added to the output. In this way, the LMSDA network 300 only needs to learn the global residual information to enhance the quality of the output image 318, thereby reducing the training and computational complexity.

[0052] Figure 4 Detailed block diagram of an exemplary LMSDA network 400 (hereinafter referred to as “LMSDA network 400”) for a chroma channel according to some embodiments of the present disclosure is shown. Figure 4 As shown, the LMSDA network 400 includes, for example, a header portion 401 , a trunk portion 403 , and a reconstruction portion 405 .

[0053] The input of the LMSDA network 400 includes Y, U, and V channels. Since the chrominance component contains less information and is prone to losing key information after compression, it is difficult for CNN to learn the information lost in the input. Therefore, relying solely on a single U or V channel for SR may not work well. Therefore, the LMSDA network 400 uses all three Y, U, and V channels to solve the problem of insufficient information of a single chrominance component. The luminance channel (e.g., the Y channel) may carry more information than the chrominance channels (e.g., the U and V channels), so the luminance channel guides the SR of the chrominance channels.

[0054] like Figure 4As shown, the head 401 may include two 3x3 convolutional layers 404, one for downsampling and the other for extracting shallow features after mixing the chroma component 402a and the luma component 402b. The U channel and the V channel may be concatenated together to generate the chroma component 402a. The luma channel (e.g., the Y channel) is twice the size of the chroma channel (e.g., the U / V channel). Therefore, the Y channel needs to be downsampled first. To this end, a 3x3 convolutional layer 406 with a stride of 2 may be used for downsampling. The output f of the head 401 is 0 It can be expressed by formula (5). f 0 =Conv(Conv(C[U LR ,V LR ])+dConv(Y LR )) (5), where f 0 represents the output of the head, dConv() represents the downsampling convolution layer 406, and Conv() represents the convolution layer 404 with a stride of 1.

[0055] The backbone portion 403 may include M LMSDABs 408. The backbone portion 403 uses f 0 The connector 412 connects the LMSDAB output and finally passes through a 1x1 convolutional layer 414 to reduce the number of channels to obtain f according to equation (3) shown above. ft .f ft Can be used as input to the reconstruction part 405.

[0056] The reconstruction part 405 (e.g., upsampling network) includes a convolution layer 404 and a pixel shuffling layer 416. In addition to the above three parts, the input can be upsampled by the upsampling bicubic component 410, and the input image can be directly added to the output. In this way, the LMSDA network 400 only needs to learn the global residual information to enhance the quality of the output image 418, thereby reducing the training and computational complexity.

[0057] Figure 5 A detailed block diagram of an exemplary LMSDAB 500 (hereinafter referred to as “LMSDAB 500 ”) is shown according to some embodiments of the present disclosure.

[0058] refer to Figure 5, LMSDAB 500 can extract multi-scale and deep features from a large receptive field using stacked convolutional layers 502, 504. Important spatial and channel information can be extracted from the features extracted from the stacked convolutional layers 502 and 504 using MSSAB 506 and CAB 508. When extracting features with various receptive fields, it may be beneficial to use parallel convolutions with different receptive fields. In order to increase the receptive field and capture multi-scale and deep information while reducing network parameters, LMSDAB 500 can be designed with three parts, for example, a feature extraction part, a feature fusion part, and an attention enhancement part.

[0059] The feature extraction part includes a 1x1 convolution layer 504 and three 3x3 convolution layers 502. The feature fusion part can use the connector 514 to connect the features in the channel dimension and use the 1x1 convolution layer 504 for fusion and dimensionality reduction. The attention enhancement part uses MSSAB 506 and CAB 508 to enhance the fused features in the spatial and channel dimensions. Note that each convolution layer 502, 504 is followed by a ReLU activation function to improve the performance of the network. For example, the ReLU activation function performs nonlinear mapping with high accuracy, solves the gradient vanishing problem in the neural network, and reduces the network convergence delay. The overall operation performed by LMSDAB 500 is described below.

[0060] For example, the feature extraction part can be used to extract scale and depth features. The feature extractor is based on a 1x1 convolution layer 504 and a 3x3 convolution layer 502. Larger scale features are obtained by another 3x3 convolution layer 502, and the output of the 3x3 convolution layer 502 of the previous stage is used as the input of the next stage. The features extracted from the feature extraction part are spliced ​​together in the channel dimension. Then, the extracted features are fused by the 1x1 convolution layer 504 to generate a fused feature map, which reduces the number of dimensions and reduces the computational complexity. The attention enhancement part uses MSSAB 506 as input to obtain three branches from the feature extraction part to obtain a multi-scale spatial attention map. The multi-scale spatial attention map can be generated by the fused feature map through pixel-by-pixel multiplication 510. Then, CAB508 obtains the output feature map of channel attention enhancement. Finally, the input of LMSDAB 500 and the output of CAB 508 can be combined using pixel-by-pixel addition 512.

[0061] Fig. 6A An example multi-scale feature extraction component 600 is shown. Figure 6B An exemplary multi-scale feature extraction component 601 according to some embodiments of the present disclosure is shown. Fig. 6A and Figure 6B Describe together.

[0062] refer to Fig. 6A, shows the architecture of an example multi-scale feature extraction component 600 for extracting multi-scale feature information. The example multi-scale feature extraction component 600 may include, for example, a 1x1 convolution / ReLU layer 602, two 3x3 convolution / ReLU layers 604, and two 5x5 convolution / ReLU layers 606. The architecture includes four branches, each of which independently extracts information of different scales without interfering with each other. As the layers go deeper from top to bottom, the size and number of convolution kernels increase. Therefore, with Fig. 6A The number of parameters associated with the illustrated architecture may be excessive, resulting in undesirable computational complexity and network delays. The number of parameters associated with such existing structures may be excessive and redundant, resulting in undesirable computational complexity and network delays.

[0063] on the other hand, Figure 6B The architecture of an exemplary multi-scale feature extraction component 601 with shared convolution for multi-scale feature extraction is shown, which includes one 1x1 convolution / ReLU layer 602 and three 3x3 convolution / ReLU layers 604. Figure 6B One advantage of the architecture shown is that it takes into account the depth information of the convolutional layers while obtaining multi-scale information. Figure 6B The architecture shown in FIG. 1 can obtain larger scale features by using the 3x3 convolution / ReLU layer 604 with the 3x3 output of the previous stage as input, thereby reducing the number of parameters without reducing performance by sharing the convolution / ReLU layer 604. This is because, in the convolution operation, the receptive field of the large convolution kernel is obtained through two or more convolution connections, such as 7A to 7C shown.

[0064] In addition, the exemplary multi-scale feature extraction component 601 generates deep feature information. In a connected CNN, different network depths can produce different feature information. That is, shallower network layers produce low-level information, such as rich textures and edges, while deeper network layers can extract high-level semantic information, such as contours. The exemplary LMSDAB not only extracts scale information, but also obtains depth information in different depth convolutions. Therefore, LMSDAB extracts scale information and deep feature information, thereby generating rich feature extraction for SR.

[0065] Fig. 7A A first exemplary convolutional model 700 is shown in accordance with some aspects of the present disclosure. Figure 7B A second exemplary convolutional model 701 is shown in accordance with some aspects of the present disclosure. Figure 7C A third exemplary convolutional model 703 is shown in accordance with some aspects of the present disclosure. 7A to 7C Describe together.

[0066] refer to 7A to 7C, the receptive field of the kernel of the 7x7 convolution / ReLU layer 702 is 7x7, which is equivalent to connecting a 5x5 convolution / ReLU layer 704 and a 3x3 convolution / ReLU layer 706 (e.g. Figure 7B ) or three 3x3 convolution / ReLU layers 706 (as shown Figure 7C The receptive field obtained by

[0067] Figure 8 A detailed block diagram of an exemplary CAB 800 (hereinafter referred to as “CAB 800 ”) according to some aspects of the present disclosure is shown.

[0068] In conventional convolution calculations, each output channel corresponds to a separate convolution kernel, which are independent of each other. That is, the output channel does not fully consider the correlation between the input channels. CAB 800 can be included in LMSDAB to solve / improve this problem. The operations performed by CAB 800 can be divided into three steps, such as squeezing, excitation, and scaling.

[0069] Regarding the squeezing operation of CAB 800, global average pooling is performed on the input feature map F 802 to obtain f sq . For example, CAB 800 first squeezes the global spatial information into channel descriptors. This is achieved through global average pooling to generate channel statistics. CAB 800 can be excited to better obtain the dependencies of each channel. Two conditions need to be met during excitation. One is that the nonlinear relationship between each channel can be learned, and the other is to ensure that each channel has a non-zero output. Therefore, the activation function here is S-type (sigmoid) instead of the commonly used ReLU. The excitation process is f sq After two fully connected layers, the channels are compressed and restored. In image processing, in order to avoid the conversion between matrices and vectors, a 1x1 convolution layer 804 is used instead of a fully connected layer. Finally, CAB 800 performs scaling using a dot product (e.g., C / rx1x1 convolution layer 806) to generate an enhanced input feature map F'808.

[0070] Fig. 9 A detailed block diagram of an exemplary MSSAB 900 is shown in accordance with some aspects of the present disclosure.

[0071] refer to Fig. 9 ,MSSAB consists of three spatial attention blocks (SABs) 902, a connection layer 906, and a convolution layer 908. The inputs of the three SABs 902 are the outputs of the three branches of the feature extraction part of LMSDAB, and a spatial attention operation is performed on each proposed feature. Then, the three spatial attention maps are connected together in the connection layer 906 and fused through a 3x3 convolution layer 908 to obtain the final multi-scale spatial attention map.

[0072] Fig.10 FIG. 1 is a detailed block diagram of an exemplary SAB 1000 according to some aspects of the present disclosure. Fig.10 , the operations performed by SAB1000 include, for example, pooling, convolution, and normalization. First, SAB 1000 may perform average pooling and maximum pooling on the input feature map F 1002 to obtain two 1xHxW feature maps 1004 to reflect spatial information, such as average information and maximum information. Then, SAB 1000 may connect the two feature maps 1004 and use a convolutional layer to fuse the average information and the maximum information to generate a spatial information map with a channel number of 1. The feature map obtained in the second step is then fed into a sigmoid activation function 1006 to normalize the value of the feature map to 0 to 1 as the weight of the corresponding position of the input feature map 1002 to generate a spatial attention map 1008.

[0073] To train the exemplary LMSDA network, L2 loss is used to facilitate gradient descent. When the error is large, it descents faster, and when the error is small, it descents slower, which is conducive to convergence. The loss function f(x) can be expressed as Equation (6). f(x)=L2 (6).

[0074] Fig.11 1 is a flowchart of an exemplary video encoding method 1100 according to some embodiments of the present disclosure. The method 1100 may be performed by an apparatus (e.g., encoder 101, LMSDA network 300, 400, or any other suitable video encoding and / or compression system). The method 1100 may include operations 1102 to 1110 as described below. It is understood that some of the operations may be optional, and some of the operations may be performed simultaneously or in different order. Fig.11 Execute in the order shown.

[0075] refer to Fig.11 At 1102, the device may receive an input image through a head of the LMSDA network. For example, referring to Figure 3 , the head of the LMSDA network 300 receives an input image 302 .

[0076] At 1104, the device may extract a first set of features from the input image through the head of the LMSDA network. Figure 3 , the head 301 includes a convolutional layer 304 for extracting shallow features of the input image 302. The convolutional layer 304 in the head 301 is followed by a rectified linear activation function (ReLU) activation function (not shown). Given an input Y LR , the shallow feature f is obtained through the head network ψ according to the formula (1) shown above 0 .

[0077] At 1106, the device may input a first set of features through a plurality of LMSDABs via a backbone portion of the LMSDA network. Figure 3 , shallow feature f 0 (eg, a first set of features) may be received as input to the backbone portion 303 .

[0078] At 1108, the device can generate a second set of features based on the output of LMSDAB through the backbone of the LMSDA network. Figure 3 , the backbone portion 303 may include M LMSDABs 306. The backbone portion 303 uses f 0 The connector 310 connects the LMSDAB output and finally reduces the number of channels through the 1x1 convolutional layer 312 to obtain f according to equation (3) shown above. ft (e.g., the second set of features). ft It can be used as the input of the reconstruction part 305. In order to utilize the low-level features, the backbone part 303 uses the connection method in U-Net according to formula (2) to convert f i and f M-i Add as ω M-i+1 Input.

[0079] At 1110, the apparatus may upsample the second set of features through the reconstruction portion of the LMSDA network to generate an enhanced output image. Figure 3 , the reconstruction part 305 (e.g., upsampling network) includes a convolution layer 304 and a pixel shuffling layer 316. The upsampling network can be expressed according to equation (4) shown above. In addition to the above three parts, the input can also be upsampled by the upsampling bicubic component 308, and the input image can be directly added to the output. In this way, the LMSDA network 300 only needs to learn the global residual information to enhance the quality of the output image 318, thereby reducing the training and computational complexity.

[0080] Fig.12 1 is a flowchart of an exemplary video encoding method 1200 according to some embodiments of the present disclosure. The method 1200 may be performed by an apparatus (e.g., LMSDAB 500 or any other suitable video encoding and / or compression system). The method 1200 may include operations 1202 to 1214 as described below. It is understood that some of the operations may be optional, and some of the operations may be performed simultaneously or in different modes. Fig.12 Execute in the order shown.

[0081] refer to Fig.12At 1202, the device may apply a first convolution layer of a first kernel size and a second convolution layer of a second kernel size to the first set of features through a feature extraction portion of LMSDAB to generate a second set of features. Figure 5 , LMSDAB 500 can use stacked convolutional layers 502, 504 to extract multi-scale and deep features from a large receptive field.

[0082] At 1204, the device may combine the second set of features in the channel dimension through the feature extraction part of LMSDAB. Figure 4 , the feature fusion part can use connector 514 to connect features in the channel dimension.

[0083] At 1206, the device may generate a fused feature map by fusing the second set of features combined in the channel dimension using a third convolutional layer of the first kernel size through the feature extraction portion of LMSDAB. Figure 5 , the feature fusion part can use 1x1 convolution layer 504 for fusion and dimension reduction.

[0084] At 1208, the device may use MSSAB to obtain multiple outputs of multiple stacked convolutional layers through the feature fusion portion of LMSDAB. Figure 5 ,The attention enhancement part uses MSSAB 506 as input to obtain three branches from the feature extraction part.

[0085] At 1210, the apparatus may generate an MSSAB output by applying a plurality of stacked spatial attention layers to a plurality of outputs of a plurality of stacked convolutional layers through a feature fusion portion of LMSDAB. Figure 5 ,The attention enhancement part uses MSSAB506 as input to obtain three branches from the feature extraction part to obtain the multi-scale spatial attention map.

[0086] At 1212, the device may perform pixel-by-pixel multiplication on the fused feature map and the MSSAB output of the MSSAB through the feature fusion portion of the LMSDAB to generate a multi-scale spatial attention map. Figure 5 , a multi-scale spatial attention map can be generated by fusion feature map through pixel-by-pixel multiplication 510.

[0087] At 1214, the device may use CAB to obtain a channel attention map based on the multi-scale spatial attention map through the attention enhancement part. Figure 5 , CAB 508 obtains a channel-attention-enhanced output feature map (e.g., a channel attention map).

[0088] In various aspects of the present disclosure, the functions described herein may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions may be stored as instructions on a non-transitory computer-readable medium. Computer-readable media include computer storage media. Storage media may be any computer program that can be processed by a processor (e.g., Figure 1 and Figure 2 As an example and not limitation, such computer-readable media may include RAM, ROM, EEPROM, CD-ROM, or other optical disk storage, HDD (e.g., disk storage or other magnetic storage device), flash drive, SSD, or any other medium that can be used to carry or store the required program code in the form of instructions or data structures and can be accessed by a processing system (e.g., a mobile device or a computer). Disks and optical disks used in this article include CDs, laser optical disks, optical disks, digital video discs (DVDs), and floppy disks, where disks typically reproduce data magnetically, while optical disks use lasers to reproduce data optically. Combinations of the above should also be included within the scope of computer-readable media.

[0089] According to one aspect of the present disclosure, a video encoding method is provided. The method may include a head of an LMSDA network receiving an input image. The method may include a head of the LMSDA network extracting a first set of features from the input image. The method may include a trunk portion of the LMSDA network inputting the first set of features through a plurality of LMSDABs. The method may include a trunk portion of the LMSDA network generating a second set of features based on outputs of the LMSDABs. The method may include a reconstruction portion of the LMSDA network upsampling the second set of features to generate an enhanced output image.

[0090] In some embodiments, the LMSDA network may be associated with a luma channel or a chroma channel.

[0091] In some embodiments, the backbone portion of the LMSDA network generating the second set of features based on the output of LMSDAB may include applying a first convolutional layer of a first kernel size and a second convolutional layer of a second kernel size to the first set of features to generate a third set of features. In some embodiments, the backbone portion of the LMSDA network generating the second set of features based on the output of LMSDAB may include combining the third set of features in the channel dimension. In some embodiments, the backbone portion of the LMSDA network generating the second set of features based on the output of LMSDAB may include fusing the third set of features combined in the channel dimension using a third convolutional layer of the first kernel size to generate a fused feature map.

[0092] In some embodiments, the backbone portion of the LMSDA network generating the second set of features based on the output of LMSDAB may include using MSSAB to obtain multiple outputs of multiple stacked convolutional layers. In some embodiments, the backbone portion of the LMSDA network generating the second set of features based on the output of LMSDAB may include performing pixel-by-pixel multiplication of the fused feature map and the MSSAB output of MSSAB to generate a multi-scale spatial attention feature map

[0093] In some embodiments, generating a second set of features based on the output of LMSDAB by the backbone portion of the LMSDA network may include generating an MSSAB output by applying a plurality of stacked spatial attention layers to a plurality of outputs of a plurality of stacked convolutional layers.

[0094] In some embodiments, the backbone portion of the LMSDA network generating a second set of features based on the output of LMSDAB may include using CAB to obtain a channel attention map based on a multi-scale spatial attention map.

[0095] In some embodiments, an enhanced output image may be generated based at least in part on a multi-scale spatial attention map and a channel attention map.

[0096] According to another aspect of the present disclosure, a video encoding system is provided. The system may include a memory for storing instructions. The system may include a processor coupled to the memory, which, when executing the instructions, is used to receive an input image through the head of an LMSDA network. The system may include a processor coupled to the memory, which, when executing the instructions, is used to extract a first set of features from the input image through the head of the LMSDA network. The system may include a processor coupled to the memory, which, when executing the instructions, is used to input a first set of features through a plurality of LMSDABs through the trunk portion of the LMSDA network. The system may include a processor coupled to the memory, which, when executing the instructions, is used to generate a second set of features based on the output of the LMSDAB through the trunk portion of the LMSDA network. The system may include a processor coupled to the memory, which, when executing the instructions, is used to upsample the second set of features through the reconstruction portion of the LMSDA network to generate an enhanced output image.

[0097] In some embodiments, the LMSDA network may be associated with a luma channel or a chroma channel.

[0098] In some embodiments, the processor coupled to the memory can be used to generate a second set of features based on the output of LMSDAB through the backbone portion of the LMSDA network as follows when executing the instructions: apply a first convolution layer of a first kernel size and a second convolution layer of a second kernel size to the first set of features to generate a third set of features. In some embodiments, the processor coupled to the memory can be used to generate a second set of features based on the output of LMSDAB through the backbone portion of the LMSDA network as follows when executing the instructions: combine the third set of features in the channel dimension. In some embodiments, the processor coupled to the memory can be used to generate a second set of features based on the output of LMSDAB through the backbone portion of the LMSDA network as follows when executing the instructions: fuse the third set of features combined in the channel dimension using a third convolution layer of the first kernel size to generate a fused feature map.

[0099] In some embodiments, the processor coupled to the memory can be used to generate a second set of features based on the output of LMSDAB through the trunk portion of the LMSDA network as follows when executing the instructions: multiple outputs of multiple stacked convolutional layers are obtained using MSSAB. In some embodiments, the processor coupled to the memory can be used to generate a second set of features based on the output of LMSDAB through the trunk portion of the LMSDA network as follows when executing the instructions: perform pixel-by-pixel multiplication of the MSSAB output of the fused feature map and MSSAB to generate a multi-scale spatial attention feature map.

[0100] In some embodiments, a processor coupled to the memory, when executing instructions, can be used to generate a second set of features based on the output of LMSDAB through the backbone portion of the LMSDA network as follows: generate MSSAB outputs by applying multiple stacked spatial attention layers to multiple outputs of multiple stacked convolutional layers.

[0101] In some embodiments, a processor coupled to the memory, when executing the instructions, can be used to generate a second set of features based on the output of LMSDAB through the backbone part of the LMSDA network as follows: obtain a channel attention map based on the multi-scale spatial attention map using CAB.

[0102] In some embodiments, an enhanced output image may be generated based at least in part on a multi-scale spatial attention map and a channel attention map.

[0103] According to another aspect of the present disclosure, a video encoding method is provided. The method may include a feature extraction part of LMSDAB applying a first convolution layer of a first kernel size and a second convolution layer of a second kernel size to a first set of features to generate a second set of features. The method may include the feature extraction part of LMSDAB combining the second set of features in a channel dimension. The method may include the feature extraction part of LMSDAB fusing the second set of features combined in the channel dimension using a third convolution layer of the first kernel size to generate a fused feature map.

[0104] In some embodiments, the method may include the feature fusion portion of LMSDAB using MSSAB to obtain multiple outputs of multiple stacked convolutional layers. In some embodiments, the method may include the feature fusion portion of LMSDAB generating MSSAB outputs by applying multiple stacked spatial attention layers to the multiple outputs of the multiple stacked convolutional layers. In some embodiments, the method may include the feature fusion portion of LMSDAB performing pixel-by-pixel multiplication of the fused feature map and the MSSAB output of MSSAB to generate a multi-scale spatial attention feature map.

[0105] In some embodiments, the method may include an attention enhancement part using CAB to obtain a channel attention map based on a multi-scale spatial attention map.

[0106] According to another aspect of the present disclosure, a video encoding system is provided. The system may include a memory for storing instructions. The system may include a processor coupled to the memory, which, when executing the instructions, is used to apply a first convolution layer of a first kernel size and a second convolution layer of a second kernel size to a first set of features through a feature extraction part of LMSDAB to generate a second set of features. The system may include a processor coupled to the memory, which, when executing the instructions, is used to combine the second set of features in a channel dimension through a feature extraction part of LMSDAB. The system may include a processor coupled to the memory, which, when executing the instructions, is used to fuse the second set of features combined in the channel dimension through a third convolution layer of a first kernel size through a feature extraction part of LMSDAB to generate a fused feature map.

[0107] In some embodiments, the processor coupled to the memory may also be used to obtain multiple outputs of multiple stacked convolutional layers using MSSAB through the feature fusion portion of LMSDAB when executing the instructions. In some embodiments, the processor coupled to the memory may also be used to generate MSSAB outputs by applying multiple stacked spatial attention layers to multiple outputs of multiple stacked convolutional layers through the feature fusion portion of LMSDAB when executing the instructions. In some embodiments, the processor coupled to the memory may also be used to perform pixel-by-pixel multiplication of the fused feature map and the MSSAB output of MSSAB through the feature fusion portion of LMSDAB when executing the instructions to generate a multi-scale spatial attention feature map.

[0108] In some embodiments, the processor coupled to the memory may also be used, when executing the instructions, to obtain a channel attention map based on the multi-scale spatial attention map using CAB through the attention enhancement part.

[0109] The above description of the embodiments will reveal the general nature of the present disclosure, and others can easily modify and / or adjust these embodiments for various applications without departing from the general concept of the present disclosure by applying knowledge within the technical scope of the art, without excessive experimentation. Therefore, based on the teachings and guidance provided herein, such adjustments and modifications are intended to belong to the meaning and scope of the equivalents of the disclosed embodiments. It should be understood that the wording or terminology herein is for description rather than for limitation, and therefore, those skilled in the art should interpret the terms or wording of this specification in accordance with the teachings and guidance.

[0110] The disclosed embodiments have been described above with the aid of functional building blocks, which illustrate the implementation of their specific functions and their relationships. For ease of description, the boundaries of these functional building blocks are arbitrarily defined herein. As long as their specific functions and their relationships are properly performed, other boundaries may also be defined.

[0111] The Summary and Abstract sections may set forth one or more but not all exemplary embodiments of the present disclosure as contemplated by the inventor(s), and thus, are not intended to limit the present disclosure and the appended claims in any way.

[0112] Various functional blocks, modules, and steps are disclosed above. The arrangements provided are illustrative and not restrictive. Therefore, the functional blocks, modules, and steps may be reordered or combined in a manner different from the examples provided above. Likewise, some embodiments include only a subset of the functional blocks, modules, and steps, and any such subset is permitted.

[0113] The breadth and scope of the present disclosure should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.

Claims

1. A video encoding method, include: The head of the lightweight multi-level mixed scale and depth information with attention mechanism (LMSDA) network receives the input image; The head of the LMSDA network extracts a first set of features from the input image; The backbone of the LMSDA network inputs the first set of features through a plurality of LMSDA blocks (LMSDAB); The backbone portion of the LMSDA network generates a second set of features based on the output of the LMSDAB; and The reconstruction portion of the LMSDA network upsamples the second set of features to generate an enhanced output image.

2. The method according to claim 1, in, The LMSDA network is associated with a luma channel or a chroma channel.

3. The method according to claim 1, in, The backbone of the LMSDA network generates the second set of features based on the output of the LMSDAB including: applying a first convolutional layer of a first kernel size and a second convolutional layer of a second kernel size to the first set of features to generate a third set of features; combining the third set of features in a channel dimension; and The third set of features combined in the channel dimension is fused by using a third convolutional layer with the first kernel size to generate a fused feature map.

4. The method according to claim 3, in, The backbone of the LMSDA network generates the second set of features based on the output of the LMSDAB including: Using Multi-Scale Spatial Attention Blocks (MSSAB) to obtain multiple outputs of multiple stacked convolutional layers; and A pixel-by-pixel multiplication is performed on the fused feature map and the MSSAB output of the MSSAB to generate a multi-scale spatial attention feature map.

5. The method according to claim 4, in, The backbone of the LMSDA network generates the second set of features based on the output of the LMSDAB including: The MSSAB output is generated by applying multiple stacked spatial attention layers to the multiple outputs of the multiple stacked convolutional layers.

6. The method according to claim 4, in, The backbone of the LMSDA network generates the second set of features based on the output of the LMSDAB including: A channel attention block (CAB) is used to obtain a channel attention map based on the multi-scale spatial attention map.

7. The method according to claim 6, in, The enhanced output image is generated based at least in part on the multi-scale spatial attention map and the channel attention map.

8. A video encoding system, include: A memory for storing instructions; as well as A processor, coupled to the memory and configured, when executing the instructions, to: The input image is received through the head of a lightweight multi-level mixed scale and depth information and attention mechanism (LMSDA) network; extracting a first set of features from the input image via the head of the LMSDA network; Inputting the first set of features through a backbone portion of the LMSDA network through a plurality of LMSDA blocks (LMSDAB); generating a second set of features based on the output of the LMSDAB by the backbone portion of the LMSDA network; as well as The second set of features is upsampled by a reconstruction portion of the LMSDA network to generate an enhanced output image.

9. The system according to claim 8, in, The LMSDA network is associated with a luma channel or a chroma channel.

10. The system according to claim 8, in, The processor is coupled to the memory and, when executing the instructions, is configured to generate the second set of features based on the output of the LMSDAB by the backbone portion of the LMSDA network as follows: applying a first convolutional layer of a first kernel size and a second convolutional layer of a second kernel size to the first set of features to generate a third set of features; Combining the third set of features in a channel dimension; as well as The third set of features combined in the channel dimension is fused by using a third convolutional layer with the first kernel size to generate a fused feature map.

11. The system according to claim 10, in, The processor is coupled to the memory and, when executing the instructions, is configured to generate the second set of features based on the output of the LMSDAB by the backbone portion of the LMSDA network as follows: Use Multi-Scale Spatial Attention Block (MSSAB) to obtain multiple outputs of multiple stacked convolutional layers; as well as A pixel-by-pixel multiplication is performed on the fused feature map and the MSSAB output of the MSSAB to generate a multi-scale spatial attention feature map.

12. The system according to claim 11, in, The processor is coupled to the memory and, when executing the instructions, is configured to generate the second set of features based on the output of the LMSDAB by the backbone portion of the LMSDA network as follows: The MSSAB output is generated by applying multiple stacked spatial attention layers to the multiple outputs of the multiple stacked convolutional layers.

13. The system according to claim 11, in, The processor is coupled to the memory and, when executing the instructions, is configured to generate the second set of features based on the output of the LMSDAB by the backbone portion of the LMSDA network as follows: A channel attention block (CAB) is used to obtain a channel attention map based on the multi-scale spatial attention map.

14. The system according to claim 13, in, The enhanced output image is generated based at least in part on the multi-scale spatial attention map and the channel attention map.

15. A video encoding method, include: A feature extraction part of a lightweight multi-level mixed scale and depth information and attention mechanism block (LMSDAB) applies a first convolution layer of a first kernel size and a second convolution layer of a second kernel size to the first set of features to generate a second set of features; The feature extraction portion of the LMSDAB combines the second set of features in a channel dimension; as well as The feature extraction part of the LMSDAB generates a fused feature map by fusing the second set of features combined in the channel dimension using a third convolutional layer of the first kernel size.

16. The method according to claim 15, further comprising: include: The feature fusion part of the LMSDAB uses a multi-scale spatial attention block (MSSAB) to obtain multiple outputs of multiple stacked convolutional layers; The feature fusion portion of the LMSDAB generates an MSSAB output by applying a plurality of stacked spatial attention layers to the plurality of outputs of the plurality of stacked convolutional layers; as well as The feature fusion part of the LMSDAB performs pixel-by-pixel multiplication on the fused feature map and the MSSAB output of the MSSAB to generate a multi-scale spatial attention feature map.

17. The method according to claim 16, further comprising: include: The attention enhancement part uses a channel attention block (CAB) to obtain a channel attention map based on the multi-scale spatial attention map.

18. A video encoding system, include: A memory for storing instructions; as well as A processor, coupled to the memory and configured, when executing the instructions, to: Applying a first convolutional layer of a first kernel size and a second convolutional layer of a second kernel size to the first set of features to generate a second set of features through a feature extraction part of a lightweight multi-level mixed scale and depth information and attention mechanism block (LMSDAB); combining the second set of features in a channel dimension by the feature extraction portion of the LMSDAB; as well as The feature extraction part of the LMSDAB generates a fused feature map by fusing the second set of features combined in the channel dimension using a third convolutional layer of the first kernel size.

19. The system according to claim 18, in, The processor coupled to the memory, when executing the instructions, is further configured to: Obtain multiple outputs of multiple stacked convolutional layers using a multi-scale spatial attention block (MSSAB) through the feature fusion part of the LMSDAB; generating, by the feature fusion portion of the LMSDAB, a MSSAB output by applying a plurality of stacked spatial attention layers to the plurality of outputs of the plurality of stacked convolutional layers; as well as The feature fusion part of the LMSDAB performs pixel-by-pixel multiplication on the fused feature map and the MSSAB output of the MSSAB to generate a multi-scale spatial attention feature map.

20. The system according to claim 19, in, The processor coupled to the memory, when executing the instructions, is further configured to: A channel attention map is obtained based on the multi-scale spatial attention map using a channel attention block (CAB) through the attention enhancement part.