Super-resolution convolutional neural network (CNN) filter with reference image resampling (RPR) function
By using multi-mixed scale and deep information attention neural network (MMSDANet) in video encoding and decoding, the problem of traditional methods dealing with complex characteristic videos during upsampling is solved, achieving more efficient video reconstruction and image quality improvement.
Patent Information
- Application Number
- CN202280097777.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-05
- Publication Date
- 2025-05-16
AI Technical Summary
Traditional video encoding and decoding methods are difficult to effectively process videos with complex characteristics during upsampling, resulting in a decline in image quality.
Multi-mixed scale and deep information attention neural network (MMSDANet) are used for video compression, multi-scale information and depth-related layer information are extracted during the upsampling process through convolutional neural network filter, and feature extraction is enhanced by using attention mechanism.
It improves video reconstruction performance and efficiency, significantly improves image quality after upsampling, and can better restore the details and boundary information of the video.
Smart Images

Figure CN120019404A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to video compression schemes capable of improving video reconstruction performance and efficiency. More specifically, the present disclosure relates to systems and methods for providing convolutional neural network filters for upsampling processes. Background Art
[0002] Video codecs for high-definition video have been a focus of attention for nearly a decade. Although codec technology has improved, it remains challenging to transmit high-definition video using limited bandwidth. Approaches to address this problem include resampling-based video codecs, in which (i) the original video is first "downsampled" to form an encoded video before encoding, (ii) the encoded video is transmitted as a bitstream and then decoded to form a decoded video; and (iii) the decoded video is then "upsampled" to the same resolution as the original video. For example, Versatile Video Coding (VVC) supports a resampling-based codec scheme (reference picture resampling, RPR), which enables temporal prediction between different resolutions. However, conventional methods cannot effectively handle the upsampling process, especially for videos with complex characteristics. Therefore, it is advantageous to have an improved system and method to address the above needs. Summary of the invention
[0003] The present disclosure relates to systems and methods for video compression using neural networks to improve the image quality of videos. More specifically, the present disclosure provides a multi-mixed scale and depth information attention neural network (MMSDANet) to perform an upsampling process (which may be referred to as a super-resolution (SR) process). Although the following systems and methods are described with respect to video processing, in some embodiments, the systems and methods may be used for other image processing systems and methods. The convolutional neural network (CNN) framework can be trained by deep learning and / or artificial intelligence solutions.
[0004] MMSDANet is a CNN filter for RPR-based SR in VVC. MMSDANet can be embedded in the VVC codec. MMSDANet includes multi-mixed scale and depth information attention blocks (MMSDAB). MMSDANet is based on residual learning to accelerate network convergence and reduce training complexity. MMSDANet effectively extracts low-level features in a "U-Net" structure by stacking MMSDABs, and transmits the extracted low-level features to the high-level feature extraction module through U-Net connections. High-level features contain global semantic information, while low-level features contain local detail information. U-Net connections can also reuse low-level features when restoring local details.
[0005] More specifically, MMSDANet adopts residual learning to reduce network complexity and improve learning ability. MMSDAB is designed as a basic block combined with an attention mechanism to extract multi-scale information and depth-related layer information of image features. Multi-scale information can be extracted by convolution kernels of different sizes, while depth-related layer information can be extracted according to different depths of the network. For MMSDAB, sharing the parameters of the convolution layer can reduce the number of overall network parameters, thereby significantly improving the overall system efficiency.
[0006] In some embodiments, the method may be implemented by a tangible, non-transitory computer-readable medium having processor instructions stored thereon, which, when executed by one or more processors, cause one or more processors to perform one or more aspects / features of the method described herein. In other embodiments, the method may be implemented by a system comprising a computer processor and a non-transitory computer-readable storage medium having instructions stored thereon, which, when executed by a computer processor, cause the computer processor to perform one or more actions of the method described herein. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] In order to more clearly describe the technical solutions in the embodiments of the present disclosure, the following is a brief description of the drawings. The drawings only show some aspects or embodiments of the present disclosure, and ordinary technicians in this field can still deduce other drawings based on these drawings without creative work.
[0008] Figure 1 is a schematic diagram showing an MMSDANet framework according to one or more embodiments of the present disclosure.
[0009] Figure 2 is a schematic diagram showing another MMSDANet framework according to one or more embodiments of the present disclosure.
[0010] Figure 3is a schematic diagram illustrating an MMSDAB according to one or more embodiments of the present disclosure.
[0011] Figure 4 is a schematic diagram showing a convolutional model with an equivalent receptive field according to one or more embodiments of the present disclosure.
[0012] Figure 5 is a schematic diagram of a Squeeze and Excitation (SE) attention mechanism according to one or more embodiments of the present disclosure.
[0013] Figure 6 a to Figure 6 e and Figure 7 a to Figure 7 e is an image showing test results according to one or more embodiments of the present disclosure.
[0014] Figure 8 is a schematic diagram of a wireless communication system according to one or more embodiments of the present disclosure.
[0015] Fig. 9 is a schematic block diagram of a terminal device according to one or more embodiments of the present disclosure.
[0016] Fig.10 is a flow chart of a method according to one or more embodiments of the present disclosure. DETAILED DESCRIPTION
[0017] In order to more clearly describe the technical solutions in the embodiments of the present disclosure, the following is a brief description of the drawings. The drawings only show some aspects or embodiments of the present disclosure, and ordinary technicians in this field can still deduce other drawings based on these drawings without creative work.
[0018] Figure 1 1 is a schematic diagram showing an MMSDANet 100 according to one or more embodiments of the present disclosure. To implement the RPR function, the current frame for encoding is first downsampled to reduce bitstream transmission and then restored at the decoding end. The current frame will be upsampled to its original resolution. MMSDANet 100 includes an SR neural network to replace the traditional upsampling algorithm in the traditional RPR configuration. MMSDANet 100 is a multi-level mixed scale and depth related layer information and an attention mechanism (e.g., see Figure 3 ) CNN filters. The MMSDANet framework 100 uses residual learning to reduce the complexity of network learning in order to improve performance and efficiency.
[0019] Residual learning can recover image details well at least because the residual contains image details. Figure 1 As shown, the MMSDANet 100 includes a plurality of MMSDABs 101 for using different convolutional kernel sizes and convolutional layer depths. The MMSDAB 101 extracts multi-scale information and depth information, and then combines an attention mechanism (for example, see Figure 3 ) to complete feature extraction. The MMSDANet 100 upsamples the input image 10 to the same resolution as the output image 20 by interpolation, and then enhances the image quality through residual learning.
[0020] The MMSDAB 101 is a basic block for extracting multi-scale information and convolutional depth information of the input feature map. Then, an attention mechanism is applied to enhance important information and suppress noise. The MMSDAB 101 shares convolutional layers to effectively reduce the number of parameters generated by using convolutional kernels of different sizes. When sharing convolutional layers, layer depth information is also introduced.
[0021] As Figure 1 shown, the MMSDANet 100 includes three parts: a head part 102, a backbone part 104, and an upsampling part 106. The head part 102 includes a convolutional layer 105 for extracting shallow features of the input image 10. Behind the convolutional layer 105 is a ReLU (Rectified Linear Unit) activation function. Using "Y LR " to indicate the input image 10, and using "ψ" to represent the head part 102, the shallow feature f0 can be expressed as follows:
[0022] f0 = ψ(Y LR ) Equation (1).
[0023] The backbone part 104 includes "M" MMSDABs 103. In some embodiments, "M" can be an integer greater than 2. In some embodiments, "M" can be "8". The backbone part 104 uses f0 as the input, cascades the outputs of the MMSDAB 103 (in the cascade block 107), and then reduces the number of channels through a "1×1" convolution 109 to obtain f that can subsequently be put into the upsampling part 106 (or the reconstruction part) ft . To make full use of low-level features, the connection method in "U-Net" is used to add f i and f M-i as the input of ω M-i+1 , as shown in the following equation:
[0024] f M-i+1 = ω M-i+1 (f i + f M-i ) 0 < i < M / 2 Equation (2).
[0025] fft =Conv(C[ω M ,ω M-1 ,…ω1(f0)])+f0 Equation (3).
[0026] Among them, ω i Indicates the “M”th MMSDAB. “C[.]” indicates the channel concatenation process. The channel concatenation process refers to stacking features in the channel dimension. For example, the dimensions of two feature maps can be “B×C1×H×W” and “B×C2×H×W”. After the concatenation process, the dimensions become “B×(C1+C2)×H×W”. The parameter “f i " represents the output of the "M"th MMSDAB.
[0027] The upsampling part 106 includes a convolution layer 111 and a pixel shuffle process 113. The upsampling part 106 can be represented as follows:
[0028] Y HR =PS(Conv(f ft ))+Y LR Equation (4).
[0029] Among them, Y HR is the upsampled image, PS is a pixel shuffle layer, Conv represents a convolutional layer, and the ReLU activation function is not used in the upsampling part 106.
[0030] In some embodiments, in addition to the three parts, the input image 10 can also be added to the output of the upsampling part 106. With this arrangement, MMSDANet 100 only needs to learn global residual information to enhance the quality of the input image 10. This significantly reduces the training difficulty and burden of MMSDANet 100.
[0031] In some embodiments, when MMSDANet 100 is used for chroma and luma channels, the backbone portion 104 and the upsampling portion 106 may be the same. In such embodiments, the input of the network and the head portion 102 may be different.
[0032] Figure 2is a schematic diagram showing an MMSDANet framework 200 for a chrominance channel. The input of the MMSDANet framework 20 includes three channels, including a luminance (luminance or luma) channel Y and a chrominance (chrominance or chroma) channel U and a chrominance channel V. In some embodiments, the chrominance channels (U, V) contain less information and are prone to losing key information after compression. Therefore, when designing the chrominance component network MMSDANet 200, all three channels Y, U, and V are used to provide sufficient information. The luminance channel Y includes more information than the chrominance channels U and V, so it is beneficial to use the luminance channel Y to guide the upsampling process (i.e., SR process) of the chrominance channels U and V.
[0033] like Figure 2 As shown, the head part 202 of MMSDANet 200 includes two 3×3 convolutional layers 205a, 205b. The 3×3 convolutional layer 205a is used to extract shallow features, while the 3×3 convolutional layer 205b is used to extract shallow features after mixing the chrominance channel and the luminance channel. First, the two channels U and V are cascaded together and passed through the 3×3 convolutional layer 205a. Then, the shallow features are extracted by the convolutional layer 205b.
[0034] The size of the guidance component Y may be twice that of the UV channel, so the Y channel needs to be downsampled first. Therefore, the head part 202 includes a 3×3 convolutional layer 201 with a stride of 2 for downsampling. The head part 202 can be represented as follows:
[0035] f0=Conv(Conv(C[U LR ,V LR ])+dConv(Y LR )) Equation (5).
[0036] Wherein, f0 represents the output of the head part 202, dConv() represents the downsampling convolution, and Conv() represents the standard convolution with a stride of 1.
[0037] like Figure 1 and Figure 2 As shown, MMSDAB 103 is a basic unit of network 100 and network 200 . Figure 3is a schematic diagram showing an MMSDAB 300 according to one or more embodiments of the present disclosure. The MMSDAB 300 is designed to extract features from a large receptive field and emphasize important channels from the extracted features through SE (squeeze and excite) attention. When extracting features using various receptive fields, parallel convolutions with different receptive fields are considered effective. In order to increase the receptive field and capture multi-scale information and depth information, the MMSDAB 300 includes a structure having three layers 302, 304 and 306.
[0038] The first layer 302 includes three convolutional layers 301 (1×1), 303 (3×3), and 305 (5×5). The second layer 304 includes a concatenation block, a 1×1 convolutional layer 307, and two 3×3 convolutional layers 309.
[0039] The third layer 306 includes four parts: a cascade block 311, a channel shuffle block 313, a 1×1 convolution layer 315, and a SE attention block 317. In the illustrated embodiment, each of the convolution layers is followed by a ReLU activation function to improve the performance of the MMSDAB 300. The ReLU activation function has good nonlinear mapping capabilities, so it can solve the gradient vanishing problem in the neural network and accelerate the network convergence.
[0040] The overall process of MMSDAB 300 can be described as follows:
[0041] first step : Three convolutions (e.g., 301, 303, and 305) with kernel sizes of 1×1, 3×3, and 5×5, respectively, are used to extract features of different scales of the input image 30.
[0042] Step 2 : Two 3×3 convolutions (e.g., 309) are used to further extract the depth information and scale information of the input image 30 by combining the multi-scale information in the first step. Before this step, the multi-scale information in the first step is cascaded and a 1×1 convolution layer (e.g., 307) is used for dimensionality reduction to reduce computational cost. Since the input of the second step is the output of the first step, no additional convolution operation is required, thereby further reducing the required computational resources.
[0043] Step 3 :The outputs of the first two steps are first fused through a cascade operation (e.g., 311) and a channel shuffle operation (e.g., 313). Then, the dimension of the layer is reduced through a 1×1 convolution layer (e.g., 315). Finally, a squeeze and excitation (SE) attention block 317 is used to enhance important channel information and suppress weak channel information. Then, the output image 33 can be generated.
[0044] Another aspect of MMSDAB 300 is that it provides an architecture with shared convolution parameters, so that computational efficiency can be significantly improved. By considering the depth information of the convolution layer when obtaining multi-scale information (e.g., the second layer 304 of MMSDAB 300), coding performance and efficiency can be significantly improved. In addition, the number of convolution parameters used in MMSDAB 300 is significantly less than the number of convolution parameters used in other conventional methods.
[0045] In conventional methods, a typical convolutional layer module may include four branches, and each branch independently extracts different scale information without interfering with each other. As the layer goes deeper from top to bottom, the size and number of convolution kernels required increase significantly. Such multi-scale modules require a large number of parameters to support their computation of scale information. Compared with conventional methods, MMSDAB 300 is advantageous at least because: (1) the branches of MMSDAB 300 are not independent of each other; and (2) the convolutional layer can utilize the small-scale information obtained from the previous layer to obtain large-scale information. Figure 4 As explained, in the convolution operation, the receptive field of a large convolution kernel can be obtained by cascading two or more convolutions.
[0046] Figure 4 is a schematic diagram showing a convolution model with an equivalent receptive field according to one or more embodiments of the present disclosure. For example, the receptive field of a 7×7 convolution kernel is 7×7, which is equivalent to the receptive field obtained by cascading a 5×5 convolution layer and a 3×3 convolution layer or cascading three 3×3 convolution layers. Therefore, by sharing the small-scale convolution output as an intermediate result of the large-scale convolution, the required convolution parameters are greatly reduced.
[0047] For example, the dimension of the input feature map may be 64×64×64. For a 7×7 convolution, the number of required parameters is “7×7×64×64”. For a 3×3 convolution, the number of required parameters is “3×3×64×64”. For a 5×5 convolution, the number of required parameters is “5×5×64×64”. For a 1×1 convolution, the number of required parameters is “1×1×64×64”. As can be seen from the above examples, using “3×3” convolution and / or “5×5” convolution instead of “7×7” convolution can significantly reduce the number of required parameters.
[0048] In some embodiments, MMSDAB 300 can generate deep feature information. In a cascaded CNN, different network depths can produce different feature information. In other words, the "shallower" network layers produce low-level information, including rich textures and edges, while the "deeper" network layers can extract high-level semantic information, such as contours.
[0049] After MMSDAB 300 uses a "1×1" convolution (e.g., 307) to reduce the dimension, MMSDAB 300 connects two 3×3 convolutions in parallel (e.g., 309), so that both larger scale information and deep feature information can be obtained. Therefore, the entire MMSDAB 300 can extract scale information as well as deep feature information. Therefore, the entire MMSDAB 300 achieves rich feature extraction capabilities.
[0050] Figure 5 is a schematic diagram of a squeeze and excitation (SE) attention mechanism according to one or more embodiments of the present disclosure. In order to better capture channel information, MMSDAB 300 uses the SE attention mechanism, such as Figure 5 As shown. In traditional convolution calculations, each output channel corresponds to a separate convolution kernel, and these convolution kernels are independent of each other, so the output channel does not fully consider the correlation between input channels. To solve this problem, the proposed SE attention mechanism has three steps, namely, the "squeeze" step, the "excitation" step, and the "scale" step.
[0051] extrusion: First, global average pooling is performed on the input feature map to obtain f sq Each of the multiple learned filters operates with a local receptive field, so each unit of the transform output cannot exploit contextual information outside of that region. To alleviate this problem, the SE attention mechanism first “squeezes” the global spatial information into the channel descriptor. This is achieved by generating channel-dependent statistics through global average pooling.
[0052] excitation: The motivation of this step is to better obtain the dependency of each channel. Two conditions need to be met: the first condition is that the nonlinear relationship between each channel can be learned, and the second condition is that each channel has an output (for example, the value cannot be 0). The activation function in the embodiment shown can be "sigmoid (S function)" instead of the commonly used ReLU. The excitation process is f sq The channels are compressed and restored through two fully connected layers. In image processing, in order to avoid the conversion between matrices and vectors, 1×1 convolutional layers are used instead of fully connected layers.
[0053] Zoom: Finally, a dot product is performed between the output after excitation and the SE attention.
[0054] In some embodiments, CNN uses L1 loss or L2 loss to make the output gradually approach the ground truth as the network converges. For upsampling (or SR) tasks, the high-resolution map output by MMSDANet needs to be consistent with the ground truth. L1 loss or L2 loss is a loss function that compares at the pixel level. L1 loss calculates the sum of the absolute values of the difference between the output and the ground truth, while L2 loss calculates the sum of the squares of the difference between the output and the ground truth. Although CNN uses L1 loss or L2 loss to remove block effects and noise in the input image, it cannot restore the lost texture in the input image. In some embodiments, MMSDANet is trained using L2 loss, and the loss function f(x) can be expressed as follows:
[0055] f(x)=L2 Equation (6)
[0056] L2 loss facilitates gradient descent. When the error is large, it decreases faster, and when the error is small, it decreases slower, which is conducive to convergence.
[0057] Figure 6 a to Figure 6 e (i.e. "basketball") and Figure 7 a to Figure 7 e (i.e., "RHorses") is an image showing test results according to one or more embodiments of the present disclosure. The images are described as follows: (a) a low-resolution image compressed with a QP (quantization parameter) of 32 after downsampling the original image; (b) an uncompressed high-resolution image; (c) a high-resolution image compressed with a QP of 32; (d) a high-resolution image of (a) after upsampling using the RPR process; (e) a high-resolution image of (a) after upsampling using MMSDANet.
[0058] like Figure 6 (e) and Figure 7 As shown in (e), the upsampling performance using MMSDANet is better than that using RPR (e.g., Figure 6 (d) and Figure 7 (d)). Obviously, MMSDANet recovers more details and boundary information than RPR upsampling.
[0059] The following Tables 1 to 4 show the quantization measurements using MMSDANet. Tables 1 to 4 show the test results under the "allintra (AI)" configuration and the "random access (RA)" configuration. Among them, the "shaded area" represents positive gain, and the "bold / underlined" numbers represent negative gain. These tests were all conducted under "CTC (common test condition)". "VTM (VVC test model)-11.0" using the new "MCTF (Motion Compensated Temporal Filtering)" was used as the benchmark for the test.
[0060] Tables 1 and 2 show the results compared with the VTM 11.0RPR anchor. MMSDANet achieves BD (Bjontegaard-delta) rate reductions ({Y, Cb, Cr}) of {-8.16%, -25.32%, -26.30%} and {-6.72%, -26.89%, -28.19%} under AI and RA configurations, respectively.
[0061] Tables 3 and 4 show the results compared with the VTM 11.0NNVC (neural network-based video coding)-1.0 anchor. MMSDANet achieves BD rate reduction ({Y, Cb, Cr}) of {-8.5%, 18.78%, -12.61%} and {-4.21%, 4.53%, -9.55%} in RA and AI configurations, respectively.
[0062] Table 1 Results of the proposed AI-configured approach compared with the RPR anchor.
[0063]
[0064]
[0065] Table 2 Results of the proposed RA configuration method compared with RPR anchor.
[0066]
[0067] Table 3 Results of the proposed AI-configured method compared with NNVC anchors.
[0068]
[0069]
[0070] Table 4 Results of the proposed RA configuration method compared with NNVC anchors.
[0071]
[0072] Figure 8 8 is a schematic diagram of a wireless communication system 800 according to one or more embodiments of the present disclosure. The wireless communication system 800 may implement the MMSDANet framework discussed herein. Figure 8 As shown, the wireless communication system 800 may include a network device (or base station) 801. Examples of the network device 801 include a base transceiver station (Base Transceiver Station, BTS), a node B (NodeB, NB), an evolved node B (eNB or eNodeB), a next generation node B (gNB or gNode B), a wireless fidelity (Wireless Fidelity, Wi-Fi) access point (access point, AP), etc. In some embodiments, the network device 801 may include a relay station, an access point, a vehicle-mounted device, and a wearable device, etc. The network device 801 may include a wireless connection device for a communication network, such as: a Global System for Mobile Communication (GSM) network, a Code Division Multiple Access (CDMA) network, a Wideband CDMA (WCDMA) network, an LTE (Long Term Evolution) network, a Cloud Radio Access Network (CRAN), a network based on the Institute of Electrical and Electronics Engineers (IEEE) 802.11 (e.g., a Wi-Fi network), an Internet of Things (IoT) network, a device-to-device (D2D) network, a next generation network (e.g., a 5G network), or a future evolved Public Land Mobile Network (PLMN), etc. A 5G system or network may be referred to as a New Radio (NR) system or network.
[0073] exist Figure 8In the wireless communication system 800, a terminal device 803 is also included. The terminal device 803 may be a terminal user device configured to facilitate wireless communication. The terminal device 803 may be configured to be wirelessly connected to the network device 801 according to one or more corresponding communication protocols / standards (e.g., via a wireless channel 805). The terminal device 803 may be mobile or fixed. The terminal device 803 may be a user equipment (UE), an access terminal, a user unit, a user station, a mobile station, a mobile station, a remote station, a remote terminal, a mobile device, a user terminal, a terminal, a wireless communication device, a user agent or a user device. Examples of the terminal device 803 include a modem, a cellular phone, a smart phone, a cordless phone, a Session Initiation Protocol (SIP) phone, a wireless local loop (WLL) station, a personal digital assistant (PDA), a handheld device with wireless communication capabilities, a computing device or another processing device connected to a wireless modem, a vehicle-mounted device, a wearable device, an Internet of Things (IoT) device, a device used in a 5G network, or a device used in a public land mobile network, etc. For illustrative purposes, Figure 8 Only one network device 801 and one terminal device 803 are shown in the wireless communication system 800. However, in some examples, the wireless communication system 800 may include additional network devices 801 and / or terminal devices 803.
[0074] Fig. 9It is a schematic block diagram of a terminal device 903 (for example, which can implement the method discussed herein) according to one or more embodiments of the present disclosure. As shown in the figure, the terminal device 903 includes a processing unit 910 (for example, a DSP, a CPU (Central Processing Unit), a GPU (graphics processing unit, a graphics processor), etc.) and a memory 920. The processing unit 910 can be configured to implement instructions corresponding to the methods discussed herein and / or other aspects of the embodiments described above. It should be understood that the processor 910 in the embodiments of the present technology can be an integrated circuit chip and has signal processing capabilities. During implementation, each step in the aforementioned method can be implemented by using an integrated logic circuit of hardware in the processor 910 or instructions in the form of software. The processor 910 can be a general-purpose processor, a digital signal processor (digital signal processor, DSP), an application specific integrated circuit (application specific integrated circuit, ASIC), a field programmable gate array (field programmable gate array, FPGA) or another programmable logic device, a discrete gate or transistor logic device, a discrete hardware component. The methods, steps, and logic block diagrams disclosed in the embodiments of the present technology can be implemented or executed. The general processor 910 may be a microprocessor, or alternatively, the processor 910 may be any conventional processor, etc. The steps in the method disclosed with reference to the embodiments of the present technology may be directly executed or completed by a decoding processor implemented as hardware, or by using a combination of hardware modules and software modules in the decoding processor. The software module may be located in a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, or another mature storage medium in the art. The storage medium is located in the memory 920, and the processor 910 reads the information in the memory 920, and completes the steps in the aforementioned method in combination with its hardware.
[0075] It is understood that the memory 920 in the embodiment of the present technology may be a volatile memory or a non-volatile memory, or may include both a volatile memory and a non-volatile memory. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random-access memory (RAM) and is used as an external cache. For exemplary and non-limiting descriptions, various forms of RAM may be used, and these RAMs are, for example, static random-access memory (SRAM), dynamic random-access memory (DRAM), synchronous dynamic random-access memory (SDRAM), double data rate synchronous dynamic random-access memory (DDR SDRAM), enhanced synchronous dynamic random-access memory (ESDRAM), synchronous link dynamic random-access memory (SLDRAM), and direct Rambus random-access memory (DR RAM). It should be noted that the memory in the systems and methods described herein is intended to include, but is not limited to, these memories and any other suitable types of memory. In some embodiments, the memory may be a non-transitory computer-readable storage medium that stores instructions that can be executed by a processor.
[0076] Fig.101000 is a flow chart of a method according to one or more embodiments of the present disclosure. Method 1000 may be implemented by a system (such as a system having the MMSDANet discussed herein). Method 1000 is for enhancing image quality (particularly for upsampling processes). Method 1000 includes: At block 1001, receiving an input image.
[0077] At block 1003, method 1000 continues by processing the input image through a first convolutional layer. In some embodiments, the first convolutional layer is a "3×3" convolutional layer and is included in a first part of a multi-mixed scale and depth information attention neural network (MMSDANet). Figure 1 and Figure 2 The implementation example of MMSDANet is discussed in detail.
[0078] At box 1005, method 1000 continues to process the input image through a plurality of multi-mixed scale and depth information attention blocks (MMSDABs). Each MMSDAB in the plurality of MMSDABs includes more than two convolution branches that share convolution parameters. In some embodiments, the plurality of MMSDABs includes 8 MMSDABs. In some embodiments, each MMSDAB in the plurality of MMSDABs includes a first layer, a second layer, and a third layer. In some embodiments, the first layer includes three convolution layers with different dimensions. In some embodiments, the second layer includes one "1×1" convolution layer and two "3×3" convolution layers. In some embodiments, the third layer includes a concatenation block, a channel shuffle block, a "1×1" convolution layer, and a squeeze and excitation (SE) attention block. Reference Figure 3 , discusses embodiments of MMSDAB in detail. In some embodiments, a plurality of MMSDABs are included in a second portion of the MMSDANet, and wherein the second portion of the MMSDANet includes a cascade module.
[0079] At box 1007, method 1000 continues to concatenate the output of MMSDAB to form a concatenated image. At box 1009, method 1000 continues to process the concatenated image through a second convolutional layer to form an intermediate image. The second convolution kernel size of the second convolutional layer is smaller than the first convolution kernel size of the first convolutional layer. In some embodiments, the second convolutional layer is a "1×1" convolutional layer. At box 1011, method 1000 continues to process the intermediate image through a third convolutional layer and a pixel shuffle layer to generate an output image. In some embodiments, the third convolutional layer is a "3×3" convolutional layer, and wherein the third convolutional layer is included in the third part of MMSDANet.
[0080] Additional considerations
[0081] The above detailed description of the examples of the disclosed technology is not intended to be exhaustive or to limit the disclosed technology to the precise form disclosed above. Although the specific examples of the disclosed technology are described above for illustrative purposes, it will be recognized by those skilled in the relevant art that various equivalent modifications can be made within the scope of the described technology. For example, although the process or frame is presented in a given order, alternative embodiments can perform routines with steps or adopt systems with frames in different orders, and some processes or frames can be deleted, moved, added, subdivided, combined and / or modified to provide alternative embodiments or sub-combinations. Each of these processes or frames can be implemented in various different ways. In addition, although the process or frame is sometimes shown as being performed in sequence, these processes or frames can also be performed or implemented in parallel, or can be performed at different times. In addition, any specific numbers mentioned herein are examples only; alternative embodiments can adopt different values or ranges.
[0082] In the specific embodiments, many specific details are set forth to provide a thorough understanding of the technology currently discussed. In other embodiments, the technology introduced here can be practiced without these specific details. In other examples, well-known features (such as specific functions or routines) are not described in detail to avoid unnecessarily obscuring the present disclosure. References to "embodiment / embodiment" or "one embodiment / embodiment" etc. in this description mean that the specific features, structures, materials or characteristics described are included in at least one embodiment of the described technology. Therefore, such phrases appearing in this specification do not necessarily all refer to the same embodiment / embodiment. On the other hand, these references do not necessarily exclude each other. In addition, specific features, structures, materials or characteristics can be combined in any suitable manner in one or more embodiments / embodiments. It should be understood that the various embodiments shown in the figures are merely illustrative representations and are not necessarily drawn to scale.
[0083] For the sake of clarity, some details describing structures or processes are not elaborated herein, which are well known and commonly associated with communication systems and subsystems, but which may unnecessarily obscure some important aspects of the disclosed technology. In addition, although the following disclosure describes several embodiments of different aspects of the present disclosure, many other embodiments may have different configurations or different components than those described in this section. Therefore, the disclosed technology may have other embodiments with additional elements or without many of the elements described below.
[0084] Many embodiments or aspects of the technology described herein can take the form of computer executable instructions or processor executable instructions, including routines performed by programmable computers or processors. It will be appreciated by those skilled in the relevant art that the described technology can be practiced on a computer or processor system other than the computer or processor system shown and described below. The technology described herein can be implemented in a special-purpose computer or data processor, which is specially programmed, configured or constructed to perform one or more computer executable instructions in each computer executable instruction described below. Therefore, the terms "computer" and "processor" commonly used herein refer to any data processor. The information processed by these computers and processors can be presented on any suitable display medium. Instructions for performing computer executable tasks or processor executable tasks can be stored in or on any suitable computer-readable medium (including hardware, firmware, or a combination of hardware and firmware). Instructions can be included in any suitable memory device (including, for example, flash drive and / or other suitable media).
[0085] The term "and / or" in this specification is only used to describe the association relationship of associated objects, and indicates that there may be three relationships. For example, A and / or B may indicate the following three situations: A exists alone, both A and B exist, and B exists alone.
[0086] These and other changes can be made to the disclosed technology in accordance with the above-described detailed description. Although the detailed description describes specific examples of the disclosed technology and the expected best mode, no matter how detailed the above description appears in the text, the disclosed technology can be practiced in a variety of ways. The details of the system may vary greatly in its specific implementation, but are still included in the technology disclosed herein. As described above, specific terms used to describe specific features or aspects of the disclosed technology should not be taken as implying that the term is redefined herein to be limited to any specific characteristics, features or aspects of the disclosed technology associated with the term. Therefore, the present invention is not limited except by the limitations of the appended claims. In general, the terms used in the claims should not be interpreted as limiting the disclosed technology to the specific examples disclosed in the specification unless the above-mentioned detailed description section explicitly defines these terms.
[0087] Those skilled in the art will appreciate that, in conjunction with the examples described in the embodiments disclosed in this specification, the units and algorithm steps can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but the implementation should not be considered to be beyond the scope of this application.
[0088] Although certain aspects of the invention are presented in certain claim forms, the applicants contemplate the various aspects of the invention in any number of claim forms. Accordingly, the applicants reserve the right to issue additional claims after filing this application, to issue these additional claim forms in this or a continuation application.
Claims
1. A method for video processing, the method comprising: receiving an input image; Processing the input image through a first convolutional layer; Processing the input image by a plurality of multi-mixed scale and depth information attention blocks MMSDAB, wherein each of the plurality of MMSDABs comprises more than two convolution branches sharing convolution parameters; cascading outputs of the plurality of MMSDABs to form a cascaded image; Processing the concatenated image through a second convolutional layer to form an intermediate image, wherein a second convolutional kernel size of the second convolutional layer is smaller than a first convolutional kernel size of the first convolutional layer; and The intermediate image is processed through a third convolutional layer and a pixel shuffling layer to generate an output image.
2. The method according to claim 1, wherein: The input image is received through the first part of the multi-mixed scale and depth information attention neural network MMSDANet.
3. The method according to claim 2, wherein: The first convolutional layer is a 3×3 convolutional layer, and wherein the first convolutional layer is included in the first part of the MMSDANet.
4. The method according to claim 3, wherein: The plurality of MMSDABs are included in a second portion of the MMSDANet, and wherein the second portion of the MMSDANet includes a cascading module.
5. The method according to claim 1, wherein: The plurality of MMSDABs includes 8 MMSDABs.
6. The method according to claim 4, wherein: The second convolutional layer is a 1×1 convolutional layer, and wherein the second convolutional layer is included in the second part of the MMSDANet.
7. The method according to claim 6, wherein: The third convolutional layer is a 3×3 convolutional layer, and wherein the third convolutional layer is included in a third part of the MMSDANet.
8. The method according to claim 1, wherein: Each MMSDAB of the plurality of MMSDABs includes a first layer, a second layer, and a third layer.
9. The method according to claim 8, wherein: The first layer includes three convolutional layers with different dimensions.
10. The method according to claim 8, wherein: The second layer includes one 1×1 convolutional layer and two 3×3 convolutional layers.
11. The method according to claim 8, wherein: The third layer includes a concatenation block, a channel shuffle block, a 1×1 convolution layer, and a squeeze and excitation SE attention block.
12. A system for video processing, the system comprising: processor; as well as a memory configured to store instructions that, when executed by the processor, are used to: receiving an input image; Processing the input image through a first convolutional layer; Processing the input image by a plurality of multi-mixed scale and depth information attention blocks MMSDAB, wherein each of the plurality of MMSDABs comprises more than two convolution branches sharing convolution parameters; cascading outputs of the plurality of MMSDABs to form a cascaded image; Processing the cascaded image through a second convolutional layer to form an intermediate image, wherein a second convolutional kernel size of the second convolutional layer is smaller than a first convolutional kernel size of the first convolutional layer; Processing the intermediate image through a third convolutional layer and a pixel shuffling layer; and Generates the output image.
13. The system according to claim 12, wherein: The input image is received through the first part of the multi-mixed scale and depth information attention neural network MMSDANet.
14. The system according to claim 13, wherein: The first convolutional layer is a 3×3 convolutional layer, and wherein the first convolutional layer is included in the first part of the MMSDANet.
15. The system according to claim 14, wherein: The plurality of MMSDABs are included in a second portion of the MMSDANet, and wherein the second portion of the MMSDANet includes a cascading module.
16. The system according to claim 12, wherein: The plurality of MMSDABs includes 8 MMSDABs.
17. The system of claim 15, wherein: The second convolutional layer is a 1×1 convolutional layer, wherein the second convolutional layer is included in the second part of the MMSDANet, wherein the third convolutional layer is a 3×3 convolutional layer, and wherein the third convolutional layer is included in the third part of the MMSDANet.
18. The system according to claim 12, wherein: Each MMSDAB of the plurality of MMSDABs includes a first layer, a second layer, and a third layer.
19. The system of claim 18, wherein: The first layer includes three convolutional layers with different dimensions, wherein the second layer includes one 1×1 convolutional layer and two 3×3 convolutional layers, and wherein the third layer includes a cascade block, a channel shuffle block, a 1×1 convolutional layer, and a squeeze and excitation SE attention block.
20. A method for video processing, the method comprising: receiving an input image; Processing the input image through a 3×3 convolutional layer; Processing the input image by a plurality of multi-mixed scale and depth information attention blocks MMSDAB, wherein each of the plurality of MMSDABs comprises more than two convolution branches sharing convolution parameters; cascading outputs of the plurality of MMSDABs to form a cascaded image; Processing the concatenated images through a 1×1 convolutional layer to form an intermediate image; Processing the intermediate image through a third convolutional layer and a pixel shuffling layer; and Generates the output image.