A method and device for realizing image super-resolution

By using a combination method of encoder network, self-attention network and implicit neural network in image super-resolution processing, the problem that image super-resolution processing in the prior art is difficult to retain global information, and efficient image super-resolution processing is realized, ensuring image quality and computing efficiency.

CN119251052BActive Publication Date: 2025-05-23INSPUR SOFTWARE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411764620.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-05-23
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

The prior art is difficult to effectively retain global information of the image in image super-resolution processing, resulting in weak performance of high-frequency feature domains and the image quality after image super-resolution cannot be further improved.

Method used

Image super-resolution neural network including encoder network, self-attention network and implicit neural network are used to extract features through two-dimensional convolutional layer and Fourier residual modules, and focus on the global features of the image in combination with the self-attention mechanism, and use the implicit neural network for pixel prediction.

Benefits of technology

The high-frequency information in the image is effectively retained, which significantly saves computing resources, improves computing speed, ensures the structural consistency of the image and the retention of high-frequency characteristics. The image still maintains high definition and contrast.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119251052B_ABST
    Figure CN119251052B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for implementing image super-resolution, which relates to the field of image processing; it includes: Step 1: Establish an image super-resolution neural network, Step 2: Crop the images in the image dataset into image patches of a preset size, Step 3: Input the image patches into the encoder network to extract features of the image patches and obtain a feature matrix F r , Step 5: Input the feature matrix F r into the implicit neural network to perform the mapping operation from coordinate points to pixels, predict the pixels for the coordinate points input by the image patches, and achieve image super-resolution; compared with the original image, as the magnification factor increases, the structural consistency of the image is effectively guaranteed, and less high-frequency features are lost, and the image still maintains high clarity and contrast.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention discloses a method and a device for realizing image super-resolution, and relates to the field of image processing. Background Art

[0002] Currently, the way machines process images is usually to save the image as a sequence of pixels and perform a series of operations based on it. This method is strictly limited by resolution. Especially in the field of deep learning, when processing data sets, it is often necessary to resize all images to a uniform size, and this resizing is a lossy process that may come at the expense of reduced image fidelity.

[0003] With the continuous advancement of deep learning technology, the image super-resolution task based on deep learning has achieved remarkable results. Image super-resolution is the process of restoring a high-resolution image from a low-resolution image or image sequence. Initially, the deconvolution operation in the convolutional neural network was used to achieve image upsampling, and the image was enlarged to a specified multiple and then pixel prediction was performed. This method has advantages over the traditional interpolation-based super-resolution reconstruction method and can achieve faster image super-resolution processing. However, its limitation is that it can only achieve fixed magnification, and each adjustment of the magnification requires retraining the neural network.

[0004] The emergence of implicit neural networks provides a new idea for image representation, using local implicit neural image functions to represent images as potential encoding sets and predict corresponding pixel values ​​given coordinates. However, its limitations are that it ignores the global information of the image and has weak performance in the high-frequency feature domain, and cannot restore the feature information lost in the encoder, thus failing to further improve the image quality after image super-resolution. Summary of the invention

[0005] In view of the problems of the prior art, the present invention provides a method and device for realizing image super-resolution, which maps an image from a discrete domain to a continuous domain, pays more attention to the global information of the image, realizes the image super-resolution task with lower computing resources, and ensures the structural consistency of the image.

[0006] The specific scheme proposed by the present invention is:

[0007] A method for realizing image super-resolution, comprising:

[0008] Step 1: Establish an image super-resolution neural network. The image super-resolution neural network includes an encoder network, a self-attention network, and an implicit neural network. The encoder network includes a two-dimensional convolutional layer and a Fourier residual module in sequence.

[0009] Step 2: Crop the images in the image dataset into image blocks of preset size.

[0010] Step 3: Input the image block into the encoder network and extract features from the image block:

[0011] Step 31: First input the image block into the two-dimensional convolution layer for two-dimensional convolution operation to obtain the feature matrix F 1 ,

[0012] Step 32: Separately perform feature matrix F 1 The zero padding operation is performed on the two-dimensional convolution layer, and the two-dimensional convolution layer after the zero padding operation is Fourier transformed. At the same time, the Fourier residual module is used to transform the feature matrix F 1 Perform Fourier convolution operation to obtain the feature matrix F 傅 ,

[0013] Step 33: Use the Fourier residual module to transform the feature matrix F 傅 Perform inverse Fourier transform to obtain the characteristic matrix F 逆傅 ,

[0014] Step 34: Use the Fourier residual module to transform the feature matrix F 逆傅 With the feature matrix F 1 Cascade to get the potential feature matrix F L ,

[0015] Step 4: Convert the latent feature matrix F L Extract key features and obtain key feature matrix F k , and then the key feature matrix F k Input the self-attention network and reconstruct the latent feature matrix F L Feature matrices of the same scale F r ,

[0016] Step 5: Convert the feature matrix F r Input to the implicit neural network to perform the mapping operation from coordinate points to pixels. For the coordinate points input by the image block, use the following formula:

[0017] Predict pixels to achieve image super-resolution, where I o For image output, S t is the area of ​​the image block at the input coordinate point, S is the area of ​​the rectangular region formed by the four neighboring points of the input coordinate point. t The value range is {00,01,10,11}, f is the mapping function, z * t is the potential code of the area where the four neighboring points are located, v * t are the coordinates of the neighboring points, x q are the input point coordinates.

[0018] Furthermore, in step 31 of the method for implementing image super-resolution, a two-dimensional convolution operation is performed using the following formula:

[0019] Get the feature matrix F 1 , where the input image is I(x,y) , the size of the two-dimensional convolutional layer is k(m,n) , k i and k j Represents the coordinates of the convolution kernel.

[0020] Furthermore, step 4 of the method for implementing image super-resolution specifically includes:

[0021] Step 41: Use the maximum pooling operation to transform the potential feature matrix F L Extract key features and obtain key feature matrix F k ,

[0022] Step 42: Matrix the key features F k Input into the self-attention network for global feature extraction, and then reconstruct the feature matrix obtained after global feature extraction to obtain the latent feature matrix F L Feature matrices of the same scale F r .

[0023] Furthermore, in step 1 of the method for implementing image super-resolution, while establishing an image super-resolution neural network, the L1 norm is used as the loss function of the image super-resolution neural network, using the formula:

[0024] Perform L1 norm on the output image I o and the input image I The loss calculation is used to improve the prediction accuracy of image super-resolution neural networks at high frequencies.

[0025] The present invention also provides an image super-resolution implementation device, comprising a network management module, a cropping module, a feature extraction module, a self-attention calculation module and a pixel prediction module.

[0026] The network management module establishes an image super-resolution neural network, which includes an encoder network, a self-attention network, and an implicit neural network. The encoder network includes a two-dimensional convolutional layer and a Fourier residual module in sequence.

[0027] The cropping module crops the images in the image dataset into image blocks of preset sizes.

[0028] The feature extraction module inputs the image block into the encoder network and extracts features from the image block:

[0029] Step 31: First input the image block into the two-dimensional convolution layer for two-dimensional convolution operation to obtain the feature matrix F 1 ,

[0030] Step 32: Separately perform feature matrix F 1 The zero padding operation is performed on the two-dimensional convolution layer, and the two-dimensional convolution layer after the zero padding operation is Fourier transformed. At the same time, the Fourier residual module is used to transform the feature matrix F 1 Perform Fourier convolution operation to obtain the feature matrix F 傅 ,

[0031] Step 33: Use the Fourier residual module to transform the feature matrix F 傅 Perform inverse Fourier transform to obtain the characteristic matrix F 逆傅 ,

[0032] Step 34: Use the Fourier residual module to transform the feature matrix F 逆傅 With the feature matrix F 1 Cascade to get the potential feature matrix F L ,

[0033] The self-attention calculation module converts the latent feature matrix F L Extract key features and obtain key feature matrix F k , and then the key feature matrix F k Input the self-attention network and reconstruct the latent feature matrix F L Feature matrices of the same scale F r ,

[0034] The pixel prediction module converts the feature matrix F r Input to the implicit neural network to perform the mapping operation from coordinate points to pixels. For the coordinate points input by the image block, use the following formula:

[0035] Predict pixels to achieve image super-resolution, where I o For image output, S t is the area of ​​the image block at the input coordinate point, S is the area of ​​the rectangular region formed by the four neighboring points of the input coordinate point. t The value range is {00,01,10,11}, f is the mapping function, z * t is the potential code of the area where the four neighboring points are located, v * t are the coordinates of the neighboring points, x q are the input point coordinates.

[0036] Furthermore, when the feature extraction module of the image super-resolution implementation device performs the two-dimensional convolution operation in step 31, the following formula is used:

[0037] Get the feature matrix F 1 , where the input image is I(x,y) , the size of the two-dimensional convolutional layer is k(m,n) , k i and k j Represents the coordinates of the convolution kernel.

[0038] Furthermore, the self-attention calculation module of the image super-resolution implementation device uses max pooling operation to process the latent feature matrix F L for key feature extraction to obtain a key feature matrix F k ,

[0039] and inputs the key feature matrix F k into the self-attention network for global feature extraction, and then reconstructs a feature matrix of the same scale as the latent feature matrix F L based on the feature matrix obtained after global feature extraction. F r .

[0040] Furthermore, when the network management module of the image super-resolution implementation device constructs the image super-resolution neural network, it uses the L1 norm as the loss function of the image super-resolution neural network, and uses the formula:

[0041] to calculate the loss of the L1 norm for the output image I o and the input image I , thereby improving the prediction accuracy of the image super-resolution neural network in the high frequency range.

[0042] The advantages of the method of the present invention are as follows:

[0043] It mainly uses Fourier convolution for feature extraction, effectively retains the high-frequency information in the image, and uses the residual connection method for learning. It introduces self-attention for key features of the encoded features, enabling the network to connect the context information of the image, focus on its global features, significantly save computing resources, and improve the computing speed. Finally, the feature matrix containing local and global feature information and the query coordinate points are input into the implicit neural network for pixel prediction, realizing image super-resolution. Compared with the original image, as the magnification increases, the structural consistency of the image is effectively guaranteed, and less high-frequency features are lost, and the image still maintains high clarity and contrast. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1 is a schematic diagram of the method flow of the present invention.

[0045] Figure 2 is a schematic diagram of the coordinate area in pixel prediction. DETAILED DESCRIPTION OF THE INVENTION

[0046] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.

[0047] Embodiment 1: The present invention provides a method for realizing image super-resolution, comprising:

[0048] Step 1: Establish an image super-resolution neural network. The image super-resolution neural network includes an encoder network, a self-attention network, and an implicit neural network. The encoder network includes a two-dimensional convolutional layer and a Fourier residual module in sequence.

[0049] Step 2: Crop the images in the image dataset into blocks of a preset size, such as 48×48 blocks. The purpose of preprocessing the data is to allow the network to pay more attention to the local area of ​​the image and extract more detailed features, which is conducive to the network's pixel prediction of details. Cropping the image can divide the image into blocks of the same size, which is equivalent to allowing the network to pay attention to the local details of the image. Adjusting the image size can also increase the processing speed of the network and reduce the burden on computing resources.

[0050] Step 3: Input the image block into the encoder network and extract features from the image block:

[0051] Step 31: First input the image block into the two-dimensional convolution layer for two-dimensional convolution operation to obtain the feature matrix F 1 . A two-dimensional convolution operation is performed using the following formula:

[0052] Get the feature matrix F 1 , where the input image is I(x,y) , the size of the two-dimensional convolutional layer is k(m,n) , k i and k j Represents the coordinates of the convolution kernel.

[0053] Step 32: Separately perform feature matrix F 1 The zero padding operation is performed on the two-dimensional convolution layer, and the two-dimensional convolution layer after the zero padding operation is Fourier transformed. At the same time, the Fourier residual module is used to transform the feature matrix F 1 Perform Fourier convolution operation to obtain the feature matrix F 傅 。 The zero padding operation makes the characteristic matrix F 1The two-dimensional convolution layer can satisfy the size of an integer power of 2. It is conducive to the rapid Fourier transform. Then the feature matrix and convolution kernel are Fourier transformed respectively to obtain their feature performance in the frequency domain. After performing the Fourier convolution operation, the feature matrix is ​​obtained. F 傅 .

[0054] Step 33: Use the Fourier residual module to transform the feature matrix F 傅 Perform inverse Fourier transform to obtain the characteristic matrix F 逆傅 ,

[0055] Step 34: Use the Fourier residual module to transform the feature matrix F 逆傅 With the feature matrix F 1 Cascade to get the potential feature matrix F L The encoder network can be set with a two-dimensional convolution layer and six Fourier residual modules. The potential feature matrix obtained by the first Fourier residual module is input to the second Fourier residual module, and Fourier transform is continued. The Fourier convolution operation is performed, and the obtained feature matrix is ​​cascaded with the first potential feature matrix to obtain the second potential feature matrix. And so on, after six Fourier residual modules, the final potential feature matrix is ​​obtained. F L .

[0056] Step 4: Convert the latent feature matrix F L Extract key features and obtain key feature matrix F k , and then the key feature matrix F k Input the self-attention network and reconstruct the latent feature matrix F L Feature matrices of the same scale F r .

[0057] The specific steps may include:

[0058] Step 41: First, use the maximum pooling operation to transform the latent feature matrix F L Extract key features. Divide the input potential features into multiple rectangular areas evenly, take the maximum value of the pixel value in the area, and perform the maximum pooling operation on all rectangular areas at once to obtain the key feature matrix. F k .exist F kAttention is calculated on the network, which will greatly reduce resource consumption.

[0059] Step 42: F k Input into the self-attention network for global feature extraction. Get three features q, k, v, query (Query, Q), key (Key, K) and value (Value, V), then perform matrix operations on the three features q, k, v in pairs. First, perform matrix operations on k and q to get matrix A:

[0060]

[0061] Here It represents the scaling factor. Then the obtained matrix A is operated by the softmax function to obtain the matrix Right now:

[0062]

[0063] Then the resulting matrix Perform matrix operations with matrix v to obtain the output matrix B:

[0064]

[0065] The obtained matrix B is upsampled to obtain the same matrix as the potential feature matrix F L Reconstructed feature matrix of the same scale F r .

[0066] Step 5: Convert the feature matrix F r Input to the implicit neural network to perform the mapping operation from coordinate points to pixels. For the coordinate points input by the image block, use the following formula:

[0067] Predict pixels to achieve image super-resolution, where I o For image output, S t is the area of ​​the image block at the input coordinate point, S is the area of ​​the rectangular region formed by the four neighboring points of the input coordinate point. t The value range is {00,01,10,11}, f is the mapping function, z * t is the potential code of the area where the four neighboring points are located, v * tare the coordinates of the neighboring points, x q is the input point coordinate. The implicit neural network can be implemented using a multilayer perceptron consisting of 5 fully connected layers. When the predicted coordinate value moves in the image, the latent code will suddenly jump. The method of the present invention expands the representation range of each latent code, so that the image block of the latent code overlaps with the adjacent image block, avoiding the disadvantage of pixel value jump and solving the problem of image discontinuity. For example Figure 2 The image is divided into 9 image blocks. x q are the input point coordinates, S t Is the input point x q The area of ​​the image block where it is located, S yes S 00 + S 01 + S 10 + S 11 The area obtained by adding, z * 00 、z * 01 、z * 10 、z * 11 is the potential code of the region where the four neighboring points are located.

[0068] Embodiment 2: Based on Embodiment 1, while establishing the image super-resolution neural network in step 1, the L1 norm is used as the loss function of the image super-resolution neural network, using the formula:

[0069] Perform L1 norm on the output image I o and the input image I The loss calculation is used to improve the prediction accuracy of image super-resolution neural networks at high frequencies.

[0070] The method of the present invention was experimentally verified on the RTX 3090 GPU. After arbitrary magnification on different natural images, good magnification effects were achieved. Compared with the original image, as the magnification increases, the structural consistency of the image is effectively guaranteed, and the loss of high-frequency features is also less. Quantitative experiments on peak signal-to-noise ratio, structural similarity, and perceptual loss were also carried out. The experimental results are shown in Table 1.

[0071] Table 1:

[0072]

[0073] Quantitative experimental results show that the proposed method performs well in all three metrics.

[0074] Embodiment 3: The present invention also provides an apparatus for realizing image super-resolution, comprising a network management module, a cropping module, a feature extraction module, a self-attention calculation module and a pixel prediction module.

[0075] The network management module establishes an image super-resolution neural network, which includes an encoder network, a self-attention network, and an implicit neural network. The encoder network includes a two-dimensional convolutional layer and a Fourier residual module in sequence.

[0076] The cropping module crops the images in the image dataset into image blocks of preset sizes.

[0077] The feature extraction module inputs the image block into the encoder network and extracts features from the image block:

[0078] Step 31: First input the image block into the two-dimensional convolution layer for two-dimensional convolution operation to obtain the feature matrix F 1 ,

[0079] Step 32: Separately perform feature matrix F 1 The zero padding operation is performed on the two-dimensional convolution layer, and the two-dimensional convolution layer after the zero padding operation is Fourier transformed. At the same time, the Fourier residual module is used to transform the feature matrix F 1 Perform Fourier convolution operation to obtain the feature matrix F 傅 ,

[0080] Step 33: Use the Fourier residual module to transform the feature matrix F 傅 Perform inverse Fourier transform to obtain the characteristic matrix F 逆傅 ,

[0081] Step 34: Use the Fourier residual module to transform the feature matrix F 逆傅 With the feature matrix F 1 Cascade to get the potential feature matrix F L ,

[0082] The self-attention calculation module converts the latent feature matrix F L Extract key features and obtain key feature matrix F k , and then the key feature matrix F kInput the self-attention network and reconstruct the latent feature matrix F L Feature matrices of the same scale F r ,

[0083] The pixel prediction module converts the feature matrix F r Input to the implicit neural network to perform the mapping operation from coordinate points to pixels. For the coordinate points input by the image block, use the following formula:

[0084] Predict pixels to achieve image super-resolution, where I o For image output, S t is the area of ​​the image block at the input coordinate point, S is the area of ​​the rectangular region formed by the four neighboring points of the input coordinate point. t The value range is {00,01,10,11}, f is the mapping function, z * t is the potential code of the area where the four neighboring points are located, v * t are the coordinates of the neighboring points, x q are the input point coordinates.

[0085] As the information interaction and execution process between the modules of the above-mentioned device are based on the same concept as the embodiment of the method of the present invention, the specific contents can be found in the description of the embodiment of the method of the present invention and will not be repeated here.

[0086] Likewise, the benefits of the device of the present invention are:

[0087] Fourier convolution is mainly used for feature extraction, which effectively retains the high-frequency information in the image, and residual connection is used for learning. The key features of the encoded features are introduced into the self-attention learning, so that the network can connect the contextual information of the image and pay attention to its global features, which significantly saves computing resources and improves the computing speed. Finally, the feature matrix containing local and global feature information and the query coordinate points are input into the implicit neural network for pixel prediction, realizing image super-resolution. Compared with the original image, with the increase of magnification, the structural consistency of the image is effectively guaranteed, and the loss of high-frequency features is also less, and the image still maintains a high clarity and contrast.

[0088] It should be noted that not all steps and modules in the above-mentioned processes and device structures are necessary, and some steps or modules can be ignored according to actual needs. The execution order of each step is not fixed and can be adjusted as needed. The system structure described in the above-mentioned embodiments can be a physical structure or a logical structure, that is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities, or some components in multiple independent devices may be implemented together.

[0089] The above-described embodiments are only preferred embodiments for fully illustrating the present invention, and the protection scope of the present invention is not limited thereto. Equivalent substitutions or changes made by those skilled in the art based on the present invention are within the protection scope of the present invention. The protection scope of the present invention shall be subject to the claims.

Claims

1. A method for realizing image super-resolution, characterized in that include: Step 1: Establish an image super-resolution neural network. The image super-resolution neural network includes an encoder network, a self-attention network, and an implicit neural network. The encoder network includes a two-dimensional convolutional layer and a Fourier residual module in sequence. Step 2: Crop the images in the image dataset into image blocks of preset size. Step 3: Input the image block into the encoder network and extract features from the image block: Step 31: First input the image block into the two-dimensional convolution layer for two-dimensional convolution operation to obtain the feature matrix F 1 , Step 32: Separately perform feature matrix F 1 The zero padding operation is performed on the two-dimensional convolution layer, and the two-dimensional convolution layer after the zero padding operation is Fourier transformed. At the same time, the Fourier residual module is used to transform the feature matrix F 1 Perform Fourier convolution operation to obtain the feature matrix F 傅 , Step 33: Use the Fourier residual module to transform the feature matrix F 傅 Perform inverse Fourier transform to obtain the characteristic matrix F 逆傅 , Step 34: Use the Fourier residual module to transform the feature matrix F 逆傅 With the feature matrix F 1 Cascade to get the potential feature matrix F L , Step 4: Convert the latent feature matrix F L Extract key features and obtain key feature matrix F k , and then the key feature matrix F k Input the self-attention network and reconstruct the latent feature matrix F L Feature matrices of the same scale F r , specifically including: Step 41: Use the maximum pooling operation to transform the potential feature matrix F L Extract key features and obtain key feature matrix F k , Step 42: Key Features Matrix F k Input into the self-attention network for global feature extraction, and then reconstruct the feature matrix obtained after global feature extraction to obtain the latent feature matrix F L Feature matrices of the same scale F r ; Step 5: Convert the feature matrix F r Input to the implicit neural network to perform the mapping operation from coordinate points to pixels. For the coordinate points input by the image block, use the following formula: , Predict pixels to achieve image super-resolution, where I o For image output, S t is the area of ​​the image block at the input coordinate point, S is the area of ​​the rectangular region formed by the four neighboring points of the input coordinate point. t The value range is {00,01,10,11}, f is the mapping function, is the potential code of the area where the four neighboring points are located, are the coordinates of the neighboring points, x q are the input point coordinates.

2. The method for realizing image super-resolution according to claim 1, wherein a two-dimensional convolution operation is performed in step 31, using the following formula: , Get the feature matrix F 1 , where the input image is I(x,y) , the size of the two-dimensional convolutional layer is k(m,n) , k i and k j Represents the coordinates of the convolution kernel.

3. The method for realizing image super-resolution according to claim 1, characterized in that While establishing the image super-resolution neural network in step 1, the L1 norm is used as the loss function of the image super-resolution neural network, using the formula: , Perform L1 norm on the output image I oi and the input image I i The loss calculation of n represents the total number of times the loss calculation is performed, which improves the prediction accuracy of the image super-resolution neural network at high frequencies.

4. A device for realizing image super-resolution, characterized in that It includes network management module, cropping module, feature extraction module, self-attention calculation module and pixel prediction module. The network management module establishes an image super-resolution neural network, which includes an encoder network, a self-attention network, and an implicit neural network. The encoder network includes a two-dimensional convolutional layer and a Fourier residual module in sequence. The cropping module crops the images in the image dataset into image blocks of preset sizes. The feature extraction module inputs the image block into the encoder network and extracts features from the image block: Step 31: First input the image block into the two-dimensional convolution layer for two-dimensional convolution operation to obtain the feature matrix F 1 , Step 32: Separately perform feature matrix F 1 The zero padding operation is performed on the two-dimensional convolution layer, and the two-dimensional convolution layer after the zero padding operation is Fourier transformed. At the same time, the Fourier residual module is used to transform the feature matrix F 1 Perform Fourier convolution operation to obtain the feature matrix F 傅 , Step 33: Use the Fourier residual module to transform the feature matrix F 傅 Perform inverse Fourier transform to obtain the characteristic matrix F 逆傅 , Step 34: Use the Fourier residual module to transform the feature matrix F 逆傅 With the feature matrix F 1 Cascade to get the potential feature matrix F L , The self-attention calculation module converts the latent feature matrix F L Extract key features and obtain key feature matrix F k , and then the key feature matrix F k Input the self-attention network and reconstruct the latent feature matrix F L Feature matrices of the same scale F r , where the self-attention calculation module uses the maximum pooling operation to transform the potential feature matrix F L Extract key features and obtain key feature matrix F k , The key feature matrix F k Input into the self-attention network for global feature extraction, and then reconstruct the feature matrix obtained after global feature extraction to obtain the latent feature matrix F L Feature matrices of the same scale F r ; The pixel prediction module converts the feature matrix F r Input to the implicit neural network to perform the mapping operation from coordinate points to pixels. For the coordinate points input by the image block, use the following formula: , Predict pixels to achieve image super-resolution, where I o For image output, S t is the area of ​​the image block at the input coordinate point, S is the area of ​​the rectangular region formed by the four neighboring points of the input coordinate point. t The value range is {00,01,10,11}, f is the mapping function, is the potential code of the area where the four neighboring points are located, are the coordinates of the neighboring points, x q are the input point coordinates.

5. The device for realizing image super-resolution according to claim 4, characterized in that feature extraction When the module performs the two-dimensional convolution operation in step 31, the following formula is used: , Get the feature matrix F 1 , where the input image is I(x,y) , the size of the two-dimensional convolutional layer is k(m,n) , k i and k j Represents the coordinates of the convolution kernel.

6. The device for realizing image super-resolution according to claim 4, characterized in that When the network management module establishes the image super-resolution neural network, it uses the L1 norm as the loss function of the image super-resolution neural network, using the formula: , Perform L1 norm on the output image I oi and the input image I i The loss calculation of n represents the total number of times the loss calculation is performed, which improves the prediction accuracy of the image super-resolution neural network at high frequencies.

Citation Information

Patent Citations

  • Image super-resolution reconstruction method and device based on implicit neural network

    CN117057985A

  • Lightweight multi-scale image super-resolution reconstruction method

    CN118333857A