An image high-definition restoration method, device and storage medium
By introducing a multi-level multi-structure attention mechanism in deep learning, combining windows, mobile windows and global attention operations, the problem of difficulty in capturing global and local attention dependencies in the existing technology is solved, and more efficient image processing performance and computing efficiency are achieved.
Patent Information
- Application Number
- CN202211310311.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-25
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2042-10-25
AI Technical Summary
The prior art is difficult to capture both global and local attention dependencies, resulting in incomplete information acquisition and low computing efficiency in image processing in the field of deep learning.
A multi-level multi-structure attention mechanism is adopted, including window attention, moving window attention and global attention operation. Through multiple multi-scale operations, combined with GELU activation function and residual connection, effective extraction of global and local features of the image is achieved.
It realizes the capture of global and local attention dependencies simultaneously in deep learning, improves the performance and computing efficiency of image processing, and solves the problem that attention mechanisms in the existing technology cannot fully obtain information.
Smart Images

Figure CN115660984B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and in particular to an image high-definition restoration method, device, and computer storage medium. Background Art
[0002] In the existing technology, for a given image, the attention mechanism focuses on obtaining dependency relationships for each pixel. It has been proven that using the attention mechanism in neural networks can bring better effects in various visual image processing tasks. However, the performance brought by attention is often highly correlated with the computational complexity. Currently, most works tend to use local attention to reduce the computational complexity of attention. Correspondingly, adopting this method will greatly weaken the ability of the attention mechanism to obtain information from the entire image.
[0003] With the development of scientific theories and technologies, the effectiveness of deep learning and the attention mechanism has been fully verified in many visual tasks. However, considering the above-mentioned computational complexity problem, there are mainly two solutions in the current computer vision field: one is the block pixel fusion mechanism represented by ViT, which takes a pixel block with a side length of 16 pixels as a token, thereby performing the fusion of the entire image and extracting long-range dependency relationships; the other is the method represented by Swin that performs local attention operations and approximates global dependency relationships by superimposing and moving non-overlapping windows. However, both of these methods have their own problems. Although ViT can capture global information, it also loses a lot of information. While Swin captures accurately, it only captures local relationships and seriously loses long-range relationships. Therefore, currently in the field of deep learning, there is no all-rounder that can make up for the shortcomings of various popular methods, and this problem has seriously hindered the development of this field. Summary of the Invention
[0004] To this end, the technical problem to be solved by the present invention is to overcome the problem in the prior art that it is difficult to capture both global and local attention dependency relationships simultaneously.
[0005] To solve the above technical problem, the present invention provides an image high-definition restoration method, including:
[0006] Performing preliminary feature extraction on the low-resolution image to be restored through convolution to obtain a first feature map;
[0007] Performing multiple multi-scale multi-structure attention operations on the first feature map to obtain a target feature map, where the i-th multi-structure attention operation is:
[0008] Perform a shift-conv operation on the feature map output by the (i-1)-th multi-structure attention operation. After passing through the GELU activation function, perform the shift-conv operation again, and then perform a residual connection with the feature map output by the (i-1)-th multi-structure attention operation. Divide the finally output feature map into three parts in the channel dimension, and perform window attention operation, shifted window attention operation, and global attention operation respectively. Finally, add the three outputs obtained in the channel dimension to get the output of the i-th multi-structure attention operation, where the global attention operation is as follows:
[0009] Dot product the result of horizontal information extraction of the third-channel feature, the result of horizontal information extraction and then vertical information extraction of the third-channel feature, and the result of vertical information extraction of the third-channel feature to obtain the global attention feature;
[0010] Perform a residual connection between the target feature map and the first feature map, then perform upsampling, and finally perform information extraction through convolution and a resolution magnification operation to obtain the restored high-resolution image.
[0011] Preferably, perform preliminary feature extraction on the low-resolution image X to be restored through a 3×3 convolution to obtain the first feature map F0 = Conv 3×3 (X).
[0012] Preferably, the multiple multi-scale multi-structure attention operations are cyclically executed in sequence with three relatively prime window sizes.
[0013] Preferably, the specific formula for the global attention operation is:
[0014]
[0015] Where is the third-channel feature, θ() and g() represent two convolution operations, R h () and R v () represent horizontal and vertical structural changes respectively, f() represents the softmax operation, and T represents the transpose operation.
[0016] Preferably, the window attention operation divides the image into multiple small windows, and then performs traditional attention calculation on each window. The specific calculation formula is:
[0017]
[0018] Where is the first-channel feature, R w() represents the window partitioning operation, θ() and g() represent two convolution operations, f() represents the softmax operation, and T represents the transpose operation.
[0019] Preferably, the moving window attention operation first performs a window movement on the image, then divides the image into multiple small windows, and then performs traditional attention calculation on each window. The specific calculation formula is:
[0020]
[0021] where, is the second-channel feature, R w () represents the window partitioning operation, θ() and g() represent two convolution operations, f() represents the softmax operation, S() and US() represent the window movement and inverse window movement operations, and T represents the transpose operation.
[0022] Preferably, the target feature map F K is upsampled after residual connection with the first feature map F0, and then the final information extraction is performed through a 3×3 convolution, and the resolution is amplified through pixel shuffle to obtain the restored high-resolution image Y = PS(Conv 3×3 (U(F0 + F K )))
[0023] where, U() is the upsampling operation and PS() is the pixel shuffle operation.
[0024] The present invention also provides an image high-definition restoration device, including:
[0025] A preliminary feature extraction module for performing preliminary feature extraction on the low-resolution image to be restored through convolution to obtain a first feature map;
[0026] A multi-scale multi-structure attention operation module for performing multiple multi-scale multi-structure attention operations on the first feature map to obtain a target feature map. Among them, the i-th multi-structure attention operation is:
[0027] Perform a shift-conv operation on the feature map output by the (i - 1)-th multi-structure attention operation, and after passing through the GELU activation function, perform a shift-conv operation again, then perform a residual connection with the feature map output by the (i - 1)-th multi-structure attention operation. The finally output feature map is divided into three parts in the channel dimension, and window attention operation, moving window attention operation and global attention operation are respectively performed. Finally, the three outputs obtained are added in the channel dimension to obtain the output of the i-th multi-structure attention operation. Among them, the global attention operation is:
[0028] Dot product is performed on the result of horizontal information extraction of the third-channel feature, the result of horizontal information extraction of the third-channel feature followed by vertical information extraction, and the result of vertical information extraction of the third-channel feature to obtain the global attention feature;
[0029] An image restoration module is used to perform residual connection between the target feature map and the first feature map, then perform upsampling, and finally perform information extraction through convolution and perform a resolution magnification operation to obtain the restored high-resolution image.
[0030] Preferably, the image high-definition restoration device is applied to image magnification, high-definition of old photos, and video enhancement services.
[0031] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned image high-definition restoration method are implemented.
[0032] The above technical solution of the present invention has the following advantages compared with the prior art:
[0033] The image high-definition restoration method of the present invention proposes and designs multi-level and multi-structure attention. The multi-structure attention includes the existing window attention, moving window attention, and the newly introduced global attention operation. The newly introduced global attention operation decouples the image in the horizontal and vertical directions, and then calculates the global attention dependence relationship at a very low cost. The self-calculation and combined calculation of the three kinds of attention enable the neural network to simultaneously make up for the deficiencies in local and global attention, better compensate the performance of the existing attention mechanism, and its most prominent global attention module has very good performance and very low complexity, perfectly solving the high-complexity problem encountered by the current attention structure and greatly improving the calculation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to make the content of the present invention easier to be clearly understood, the following further details the present invention according to the specific embodiments of the present invention and in conjunction with the accompanying drawings, where:
[0035] Figure 1 is the implementation flowchart of an image high-definition restoration method of the present invention;
[0036] Figure 2 is the structural diagram of the image high-definition restoration network provided by the present invention;
[0037] Figure 3 is the implementation flowchart of the multi-level and multi-scale multi-structure attention provided by an embodiment of the present invention;
[0038] Figure 4This is the structural diagram of the multi-structure attention of the present invention;
[0039] Figure 5 This is the structural block diagram of an image high-definition restoration device provided by an embodiment of the present invention. Detailed implementation manners
[0040] The core of the present invention is to provide an image high-definition restoration method, device and computer storage medium, which can capture global and local attention dependencies simultaneously and improve performance.
[0041] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0042] Please refer to Figure 1 and Figure 2 , Figure 1 This is the implementation flowchart of an image high-definition restoration method provided by the present invention, Figure 2 This is the structural diagram of the image high-definition restoration network provided by the present invention; the specific operation steps are as follows:
[0043] S101: Perform preliminary feature extraction on the low-resolution image to be restored through convolution to obtain a first feature map;
[0044] The low-resolution image X to be restored is subjected to preliminary feature extraction through a 3×3 convolution, and the input 3-channel image is expanded to 64 channels, so that the original RGB three-channel information becomes the feature channel information required by the neural network, and a first feature map F0 = Conv 3×3 (X).
[0045] S102: Perform multiple multi-scale multi-structure attention operations on the first feature map to obtain a target feature map,
[0046] The present invention uses multi-scale window sizes to detect different object sizes, and does not use the conventional power of 2, because the receptive field of the power of 2 is not as good as the relatively prime window size under the same computational complexity; the multiple multi-scale multi-structure attention operations are executed in sequence with three relatively prime window sizes. Numbers without a common divisor are called relatively prime. Selecting relatively prime numbers for the window is called a relatively prime window size;
[0047] As Figure 3 shown:
[0048] In one embodiment, the 30×30 first feature map is subjected to 24 multi-scale multi-structural attention operations. For each multi-level attention module, it is divided into one of the sizes 5×5, 10×10, and 15×15 according to the execution order. If the window is 4, 8, or 16, then the receptive field is 16. If the window size is 5, 7, or 9, then the receptive field is 315. However, the size of the input image of the neural network is 30×30, so the window sizes adopted in this embodiment are 5, 10, and 15. In this way, the receptive field is 30, exactly the same size as the input image, and the receptive field information can be utilized maximally;
[0049] Among them, as Figure 4 shown, the i-th multi-structural attention operation is:
[0050] Perform a shift-conv operation on the feature map output by the (i - 1)-th multi-structural attention operation to pre-extract and fuse features. The function is to further extract feature information without adding too many parameters. After passing through the GELU activation function, perform the shift-conv operation again, and then perform a residual connection with the feature map output by the (i - 1)-th multi-structural attention operation. The finally output feature map is divided into three parts in the channel dimension, and window attention operation, moving window attention operation, and global attention operation are performed respectively. Local fine-grained dependence relationships, semi-global progressive dependence relationships, and global positioning dependence relationships are extracted respectively in the three parts. Finally, the three outputs obtained are added in the channel dimension to obtain the output of the i-th multi-structural attention operation. Among them, the global attention operation is:
[0051] Dot product the result of horizontal information extraction of the third-channel feature, the result of horizontal information extraction of the third-channel feature followed by vertical information extraction, and the result of vertical information extraction of the third-channel feature to obtain the global attention feature;
[0052] S103: After performing a residual connection between the target feature map and the first feature map, perform upsampling, and then perform final information extraction through convolution and a resolution magnification operation to obtain the restored high-resolution image.
[0053] Perform a residual connection between the target feature map F K and the first feature map F0, then perform upsampling, and then perform final information extraction through a 3×3 convolution, and perform a resolution magnification function through pixel shuffle to obtain the restored high-resolution image Y = PS(Conv 3×3 (U(F0 + F K )))
[0054] Among them, U() is the upsampling operation, and PS() is the pixel shuffle operation.
[0055] The image high-definition restoration method described in the present invention proposes and designs multi-level and multi-structure attention. The multi-structure attention includes the existing window attention, moving window attention, and the newly introduced global attention operation. The newly introduced global attention operation decouples the image in the horizontal and vertical directions, and then calculates the global attention dependence relationship at a very low cost. The self-calculation and combined calculation of the three attentions enable the neural network to simultaneously make up for the deficiencies of local and global attentions, perform better performance compensation for the existing attention mechanism, and its most prominent global attention module has very good performance and very low complexity, perfectly solving the problem of high complexity encountered by the current attention structure and greatly improving the calculation efficiency.
[0056] Based on the above embodiments, the above step S102 is further described as follows:
[0057] The specific formula of the global attention operation is:
[0058]
[0059] Among them, is the third-channel feature, θ() and g() represent two convolution operations, R h () and R v () represent the horizontal and vertical structural changes respectively, f() represents the softmax operation, and T represents the transpose operation.
[0060] The window attention operation divides the image into multiple small windows, and then performs traditional attention calculation on each window. The specific calculation formula is:
[0061]
[0062] Among them, is the first-channel feature, R w () represents the window division operation, θ() and g() represent two convolution operations, f() represents the softmax operation, and T represents the transpose operation.
[0063] The moving window attention operation first performs a window movement on the image, then divides the image into multiple small windows, and then performs traditional attention calculation on each window to ensure that the result of the subsequent window division is different from that of the window attention, facilitating the transmission of information to the entire image. The specific calculation formula is:
[0064]
[0065] Among them, is the second-channel feature, R w() represents the window partitioning operation, θ() and g() represent two convolutional operations, f() represents the softmax operation, S() and US() represent the window shifting and inverse window shifting operations, and T represents the transpose operation.
[0066] The present invention is a neural network-based attention mechanism model applied to the low-level single-image super-resolution task under computer vision image reconstruction. It can better compensate for the performance of the existing attention mechanism. At the same time, the computational efficiency is also slightly improved. It solves the problem that the attention mechanism cannot capture long-distance and short-distance dependencies simultaneously, resulting in a significant improvement in performance. It is believed that in the future, this patent can achieve better results in more fields.
[0067] Please refer to Figure 5 , Figure 5 which is the structural block diagram of an image high-definition restoration device provided by an embodiment of the present invention; the specific device may include:
[0068] The preliminary feature extraction module 100 is used to perform preliminary feature extraction on the low-resolution image to be restored through convolution to obtain the first feature map;
[0069] The multi-scale multi-structure attention operation module 200 is used to perform multiple multi-scale multi-structure attention operations on the first feature map to obtain the target feature map, where the i-th multi-structure attention operation is:
[0070] Perform a shift-conv operation on the feature map output by the (i - 1)-th multi-structure attention operation, and after passing through the GELU activation function, perform a shift-conv operation again, then perform a residual connection with the feature map output by the (i - 1)-th multi-structure attention operation. Divide the finally output feature map into three parts in the channel dimension, and perform window attention operation, moving window attention operation, and global attention operation respectively. Finally, add the three obtained outputs in the channel dimension to obtain the output of the i-th multi-structure attention operation, where the global attention operation is:
[0071] Dot product the result of horizontal information extraction of the third-channel feature, the result of horizontal information extraction of the third-channel feature followed by vertical information extraction, and the result of vertical information extraction of the third-channel feature to obtain the global attention feature;
[0072] The image restoration module 300 is used to perform upsampling after performing a residual connection between the target feature map and the first feature map, then perform final information extraction through convolution, and perform a resolution magnification operation to obtain the restored high-resolution image.
[0073] The image high-definition restoration device of this embodiment is used to implement the aforementioned image high-definition restoration method. Therefore, the specific implementation manners in the image high-definition restoration device can be seen in the embodiment part of the previous image high-definition restoration method. For example, the preliminary feature extraction module 100, the multi-scale multi-structure attention operation module 200, and the image restoration module 300 are respectively used to implement steps S101, S102, and S103 in the above image high-definition restoration method. Therefore, the specific implementation manners can refer to the descriptions of the corresponding individual part embodiments and will not be elaborated here.
[0074] The specific embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned image high-definition restoration method are implemented.
[0075] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0076] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0077] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the specified functions in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0078] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the steps specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 in one block or a plurality of blocks.
[0079] Obviously, the above embodiments are only examples for clear illustration and are not limitations on the implementation manners. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all implementation manners here. The obvious changes or modifications derived therefrom still fall within the protection scope of the present invention.
Claims
1. An image high-definition restoration method, characterized in that, Including: Performing preliminary feature extraction on the low-resolution image to be restored through convolution to obtain a first feature map; Performing multiple multi-scale multi-structural attention operations on the first feature map to obtain a target feature map, where the i-th multi-structural attention operation is: Performing a shift-conv operation on the feature map output by the (i - 1)-th multi-structural attention operation, and after passing through the GELU activation function, performing a shift-conv operation again, then performing a residual connection with the feature map output by the (i - 1)-th multi-structural attention operation. Dividing the finally output feature map into three parts in the channel dimension, respectively performing window attention operation, shifted window attention operation and global attention operation, and finally adding the three obtained outputs in the channel dimension to obtain the output of the i-th multi-structural attention operation, where the global attention operation is: Taking the dot product of the result of horizontal information extraction of the third-channel feature, the result of horizontal information extraction of the third-channel feature followed by vertical information extraction, and the result of vertical information extraction of the third-channel feature to obtain a global attention feature; Performing a residual connection between the target feature map and the first feature map, then performing upsampling, and finally performing information extraction through convolution and a resolution magnification operation to obtain the restored high-resolution image.
2. The image high-definition restoration method according to claim 1, wherein Perform preliminary feature extraction on the low-resolution image X to be restored through 3×3 convolution to obtain the first feature map F0 = Conv 3×3 (X).
3. The image high-definition restoration method according to claim 1, characterized in that The multiple multi-scale multi-structural attention operations are cyclically executed in sequence with three relatively prime window sizes.
4. The image high-definition restoration method according to claim 1, characterized in that The specific formula of the global attention operation is: Among them, is the third-channel feature, where θ() and g() represent two convolutional operations, and R h () and R v () represent horizontal and vertical structural changes respectively, f() represents the softmax operation, and T represents the transpose operation.
5. The image high-definition restoration method according to claim 1, wherein The window attention operation divides the image into multiple small windows, and then performs traditional attention calculation on each window. The specific calculation formula is: Among them, is the first channel feature, R w () represents the window partitioning operation, θ() and g() represent two convolution operations, f() represents the softmax operation, and T represents the transpose operation.
6. The image high-definition restoration method according to claim 1, wherein The shifted window attention operation first performs a window shift on the image, then divides the image into multiple small windows, and then performs traditional attention calculation on each window. The specific calculation formula is: Among them, is the second channel feature, R w () represents the window partitioning operation, θ() and g() represent two convolution operations, f() represents the softmax operation, S() and US() represent the window shift and inverse window shift operations, and T represents the transpose operation.
7. The image high-definition restoration method according to claim 1, characterized in that Upsample the target feature map F K after performing residual connection with the first feature map F0, then perform final information extraction through 3×3 convolution, and perform resolution magnification through pixelshuffle to obtain the restored high-resolution image Y = PS(Conv 3×3 (U(F0 + F K ))) Where, U() is the upsampling operation, and PS() is the pixel shuffle operation.
8. An image high-definition restoration device, characterized in that, Including: A preliminary feature extraction module for performing preliminary feature extraction on the low-resolution image to be restored through convolution to obtain a first feature map; A multi-scale multi-structural attention operation module for performing multiple multi-scale multi-structural attention operations on the first feature map to obtain a target feature map, where the i-th multi-structural attention operation is: Performing a shift-conv operation on the feature map output by the (i - 1)-th multi-structural attention operation, and after passing through the GELU activation function, performing a shift-conv operation again, then performing a residual connection with the feature map output by the (i - 1)-th multi-structural attention operation. Dividing the finally output feature map into three parts in the channel dimension, respectively performing window attention operation, shifted window attention operation and global attention operation, and finally adding the three obtained outputs in the channel dimension to obtain the output of the i-th multi-structural attention operation, where the global attention operation is: Dot product the result of horizontal information extraction of the third-channel feature, the result of horizontal information extraction of the third-channel feature followed by vertical information extraction, and the result of vertical information extraction of the third-channel feature to obtain the global attention feature; An image restoration module, configured to perform residual connection between the target feature map and the first feature map and then perform upsampling, and then perform final information extraction through convolution and perform a resolution magnification operation to obtain a restored high-resolution image.
9. The image high-definition restoration device according to claim 8, characterized in that, Applied to image magnification, old photo high definition, and video enhancement services.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of an image high-definition restoration method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Multi-scale residual attention network image super-resolution reconstruction method based on attention
CN110992270A
Image super-resolution reconstruction method based on residual module and attention mechanism
CN111179171A