A method for improving image resolution

The image super-resolution model is trained through deep learning technology, which solves the problem of low image resolution on the Internet, and achieves efficient and automated image resolution improvement, improving the image quality and viewing experience.

CN114549314BActive Publication Date: 2025-06-13NANJING UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210154779.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-21
Publication Date
2025-06-13
Estimated Expiration
2042-02-21

AI Technical Summary

Technical Problem

The low resolution of images on the Internet leads to poor visual experience, and traditional methods require professional skills and a lot of time, which is not efficient.

Method used

Deep learning technology is used to train the image super-resolution model using the DIV2K dataset, feature extraction and image reconstruction are performed through convolutional neural networks and Transformer blocks, and training is combined with the minimum absolute value deviation loss function and the Adam optimizer to obtain the final image super-resolution model.

Benefits of technology

It realizes automation to improve picture resolution, improve image quality and viewing experience, greatly improves processing efficiency, and has a small amount of model parameters and good image reconstruction effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114549314B_ABST
    Figure CN114549314B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for improving image resolution, including: training an image super-resolution model using the DIV2K dataset; after the training is completed to obtain the final image super-resolution model, cropping the image to be tested into image patches of 96*96 size, and then sequentially inputting each image patch into the image super-resolution model; the input image patch undergoes shallow feature extraction by a convolutional neural network, deep feature extraction by a Transformer block, and pixel recombination for image reconstruction to obtain a high-resolution output image patch; the obtained series of high-resolution output image patches are stitched together in the order in which they appear in the input picture to form the final high-resolution output image. Compared with the prior art, the image super-resolution model proposed in the present invention contains fewer parameters and can achieve better image reconstruction effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image optimization method, in particular to a method for improving image resolution. Background Art

[0002] In recent years, with the rapid development of deep learning, computer vision tasks have achieved very good results. Currently, a large number of pictures are generated on the Internet, and there are a huge number of pictures in places such as social media, news web pages, and e-commerce shopping websites. Pictures have become an indispensable part of everyone's life. However, due to various reasons, including relatively low original shooting resolution and reduced quality after multiple transmissions, the resolution of pictures on the Internet is often very low. This not only makes our visual experience of viewing poor, but may even make it impossible to accurately judge some key information. Therefore, improving the resolution of pictures is of great significance.

[0003] Traditional means of improving picture resolution include using software such as Photoshop to edit pictures to improve picture resolution. However, these software require users to have certain professional skills to use proficiently, and manual picture processing takes a lot of time and energy, and the efficiency is often not high. In this case, using a machine to automatically improve the resolution of images can effectively improve work efficiency. Moreover, through appropriate deep learning algorithms, low-resolution images processed by machines can often achieve very good visual effects. Using deep learning technology, only need to train an image super-resolution model in advance. This model is very small in size and has moderate resource occupancy during actual operation. Inputting the low-resolution image into the trained model can obtain an image with better resolution. High-resolution images have a better visual experience and can effectively improve the viewing experience. Summary of the Invention

[0004] Object of the Invention: The technical problem to be solved by the present invention is to provide a method for improving image resolution in view of the deficiencies of the prior art.

[0005] To solve the above technical problem, the present invention discloses a method for improving image resolution, including the following steps:

[0006] Step 1, training an image super-resolution model using the DIV2K dataset. Image super-resolution is to improve image resolution. Randomly crop image patches of 48*48 size from the pictures in the DIV2K dataset as training data;

[0007] Step 2, using a convolutional neural network to perform shallow feature extraction on the image patches to obtain a feature map;

[0008] Step 3, using a Transformer block to perform deep feature extraction on the feature map to obtain a new feature map;

[0009] Step 4, use the pixel recombination method to reconstruct the feature map obtained in Step 3 to obtain the output image O;

[0010] Step 5, use the least absolute deviation loss function to calculate the difference between the output image O and the high-resolution image HR, and use the Adam optimizer for gradient descent; repeat Steps 1 to 5 for a total of 500,000 rounds to obtain the final image super-resolution model;

[0011] Step 6, during testing, divide the low-resolution image into several image patches of size 96*96;

[0012] Step 7, sequentially input the image patches in Step 6 into the image super-resolution model described in Step 5 to obtain high-resolution image patches; sequentially splice the obtained high-resolution image patches in the order in which they appear in the low-resolution image to obtain the final high-resolution image.

[0013] In the present invention, Step 1 includes:

[0014] The training dataset uses the DIV2K dataset, which includes high-resolution images HR and low-resolution images LR; among them, the number of high-resolution images is 800, and there are 800 x2-times low-resolution images, 800 x3-times low-resolution images, and 800 x4-times low-resolution images respectively;

[0015] Form image pairs from the high-resolution images and low-resolution images according to different tasks;

[0016] Among them, in the 2-fold image super-resolution task, the high-resolution image HR and the x2-fold low-resolution image LR form an image pair;

[0017] In the 3-fold image super-resolution task, the high-resolution image HR and the x3-fold low-resolution image LR form an image pair;

[0018] In the 4-fold image super-resolution task, the high-resolution image HR and the x4-fold low-resolution image LR form an image pair;

[0019] The training batch size is set to 32, that is, 32 image patches are processed at one time during the training of the image super-resolution model.

[0020] In the present invention, Step 2 includes:

[0021] The length and width of the image patches during training are set to 48, then the dimension of the input image data X is [32, 3, 48, 48], where 3 represents that the number of channels of the image is 3, that is, 3 RGB channels;

[0022] For the input image data X, a convolutional neural network is used for shallow feature extraction to obtain the feature map F 1 , F 1 has a dimension of [32, 180, 48, 48]; the process is as follows:

[0023] F 1 = CNN(X)

[0024] Among them, the convolutional neural network CNN of this process includes a 3x3 convolution, its input dimension is 3, the output dimension is 180, the convolution kernel size is 3, the number of edge padding pixels is 1, and the stride is 1.

[0025] In the present invention, step 3 includes:

[0026] Input the feature map F obtained in step 2 1 into the Transformer block for deep feature extraction, and the output is a new feature map F 2 , whose dimension is [32, 180, 48, 48]; this process includes the following steps:

[0027] Step 3-1, the input feature map F 1 has a dimension of [32, 180, 48, 48], and its dimension is converted to [32, 180, 2304] using the flatten operation; its dimension is converted to [32, 2304, 180] using the transpose operation. After the operation, the matrix X 0 is obtained. The process is as follows:

[0028] X 0 = F 1 .flatten(-2).transpose(-1, -2)

[0029] Among them, the flatten operation is to flatten the matrix, and the transpose operation is to transpose the matrix;

[0030] Step 3-2, input X 0 into the position encoding convolutional neural network PosCNN to calculate the position encoding of X 0 . The position encoding convolutional neural network is implemented using a 3x3 convolution. The input dimension of this convolution is 180, the output dimension is 180, the convolution kernel size is 3, the number of edge padding pixels is 1, the stride is 1, and the number of groups is 180; the position encoding pos of X 0 is obtained using the position encoding convolutional neural network. The dimension of the position encoding pos is also [32, 2304, 180]. Then, the position encoding pos and X 0 are added to obtain the matrix X 1 . This process is expressed as follows:

[0031] pos = PosCNN(X 0 )

[0032] X 1 = X 0 + pos;

[0033] Step 3-3: Input X 1 into the Transformer block. The Transformer block consists of a total of 36 Transformer layer structures. Each Transformer layer structure is composed of two parts. The first part is the multi-head attention method MSA or the effective wide-area multi-head attention method EWMSA, and the second part is the multi-layer perceptron method MLP. According to the sequence number (1, 2,..., 36) of each Transformer Layer, if it is odd, the first part is MSA, and if it is even, the first part is EWMSA. The calculation process of each Transformer layer structure is as follows:

[0034] X 2 = MSA(LN(X 1 )) + X 1 ...(1) or X 2 = EWMSA(LN(X 1 )) + X 1 ...(2)

[0035] F 2 = MLP(LN(X 2 )) + X 2 ...(3)

[0036] where MSA represents the Multi-head Self-Attention method, i.e., the multi-head attention method; EWMSA represents the Effective Wide-area Multi-head Self-Attention method, i.e., the effective wide-area multi-head attention method; LN represents the LayerNorm operation, i.e., the layer normalization operation; MLP represents the Multi-Layer Perceptron method, i.e., the multi-layer perceptron method. The new feature map F 2 obtained after being processed by the Transformer block method has a dimension of [32, 180, 48, 48];

[0037] Step 3-4: Add the feature maps F1 and F2 together through residual connection to fuse the features of F1 and F2 and obtain the feature map F 3 , and this process is expressed as follows:

[0038] F 3 = F 1+conv(F 2 )

[0039] Among them, conv in this process represents the convolution operation. The input dimension of this convolution is 180, the output dimension is 180, the convolution kernel size is 3, the number of pixels for edge padding is 1, and the stride is 1.

[0040] In the present invention, step 4 includes:

[0041] Take the feature map F obtained in step 3 3 Perform upsampling using the pixel recombination method; this process specifically includes three operations, namely conv_before_upsample processing, upsampling processing, and conv_last processing; where conv_before_upsample is a convolution operation, upsampling is an upsampling operation, and conv_last is also a convolution operation. The dimension of the feature map F 3 is [32, 180, 48, 48]. After being processed by the pixel recombination method, the dimension of the obtained output image O is [32, 180, 96, 96], and the dimension of its corresponding high-resolution image is [32, 180, 96, 96]; this process is expressed as follows:

[0042] F 4 = conv_before_upsample(F 3 )

[0043] F 5 = upsampling(F 4 )

[0044] O = conv_last(F 5 )

[0045] Among them, both F 4 and F 5 represent the feature maps obtained in the intermediate calculation steps.

[0046] In the present invention, step 5 includes the following steps:

[0047] Step 5-1, use the least absolute deviation loss function L1 to calculate the difference loss between the output image O and its corresponding high-resolution image HR; this process is expressed as follows:

[0048] loss = L1(O, HR).

[0049] In the present invention, step 5 includes the following steps:

[0050] Step 5-2, use the Adam optimizer to update the network parameters, where the parameter settings of the Adam optimizer are as follows:

[0051] learning rate = 0.0002; weight decay = 0; milestones = [250000, 400000, 450000, 475000, 500000]; gamma = 0.5.

[0052] In the present invention, step 5 includes the following steps:

[0053] Step 5-3, repeat steps 1 to 5 for a total of 500,000 rounds. After the training is completed, the final image super-resolution model M is obtained.

[0054] In the present invention, step 6 includes:

[0055] Prepare a low-resolution image LR. During testing, the resolution of the low-resolution image LR can be of any size; divide the low-resolution image LR into several image blocks of size 96*96. Suppose there are finally n image blocks of size 96*96, and these image blocks are denoted as X i , where i = 1, 2...n, and n is a natural number.

[0056] In the present invention, step 7 includes the following steps:

[0057] Step 7-1, sequentially input the image blocks obtained in step 6 into the image super-resolution model M obtained in step 5; obtain n image blocks of size 192*192; denote these image blocks as Y i ; The method is as follows:

[0058]

[0059]

[0060]

[0061]

[0062]

[0063]

[0064]

[0065]

[0066] Among them, represents the feature map obtained after the i-th image block is processed by the convolutional CNN, represents the matrix obtained by flattening and transposing the matrix of the feature map after the operation. Denote the matrix and the matrix obtained after adding the stop encoding Denote the feature map obtained after being processed by the MLP method Denote the feature map and the feature map the feature map obtained after residual connection Denote the feature map obtained after being processed by the cony_before_upsample operation Denote the feature map obtained after the upsampling operation, O denotes the final output image patch obtained after the conv_last operation;

[0067] The formula for the odd layer is: where Denote the matrix obtained after being processed by the multi-head attention mechanism operation;

[0068] The formula for the even layer is: where Denote the matrix obtained after being processed by the efficient global multi-head attention mechanism operation;

[0069] Among them, in the Transformer block method, MSA and EWMSA are alternately used; that is, the odd layer formula (1) is used for calculation in the odd layer, and the even layer formula (2) is used in the even layer;

[0070] Step 7-2, step 7-1 obtains n high-resolution image patches Y i , and these n high-resolution image patches are stitched together in the order of the image patch X i in the low-resolution image LR to obtain the final high-resolution image Y.

[0071] Beneficial effects:

[0072] By using deep learning technology to train an image super-resolution model, once the final model is obtained after training, the machine can automatically process low-resolution pictures to obtain high-resolution pictures, making the picture quality higher and the viewing experience better, which can greatly improve the processing efficiency. The image super-resolution model proposed in the present invention contains fewer parameters and can achieve better image reconstruction effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0074] Figure 1 It is a schematic diagram of the training and testing work processes in a method for improving image resolution provided in the embodiment part of the present invention.

[0075] Figure 2 It is a schematic diagram of the DIV2K dataset structure adopted in a method for improving image resolution provided in the embodiment part of the present invention.

[0076] Figure 3 It is a schematic diagram of the low-resolution LR images in the DIV2K dataset in a method for improving image resolution provided in the embodiment part of the present invention.

[0077] Figure 4 It is a schematic diagram of the high-resolution HR images in the DIV2K dataset in a method for improving image resolution provided in the embodiment part of the present invention.

[0078] Figure 5 It is a schematic diagram of the low-resolution pictures input during testing in a method for improving image resolution provided in the embodiment part of the present invention.

[0079] Figure 6 It is a schematic diagram of the high-resolution pictures output by the model during testing in a method for improving image resolution provided in the embodiment part of the present invention.

[0080] Figure 7 It is a schematic diagram of partial details of the high-resolution pictures output by the model during testing in a method for improving image resolution provided in the embodiment part of the present invention. Detailed implementation manners

[0081] The present invention discloses a method for improving image resolution, which is applied to scenarios where it is necessary to use automated means to improve the resolution of pictures to make their visual effects better.

[0082] The present invention provides a method for improving image resolution, including the following steps:

[0083] Step 1: Use the DIV2K dataset (the DIV2K dataset is an image super-resolution dataset) to train an image super-resolution model (image super-resolution refers to improving image resolution), and randomly crop image patches of 48*48 in size from the pictures in the DIV2K dataset as training data;

[0084] Step 2: Use a convolutional neural network to perform shallow feature extraction on the image patches to obtain feature maps;

[0085] Step 3: Use the Transformer block to perform deep feature extraction on the feature map to obtain a new feature map, where the Transformer is a deep learning network structure;

[0086] Step 4: Use the pixel reorganization method to perform image reconstruction on the feature map obtained in the previous step to obtain the output image O;

[0087] Step 5: Use the least absolute deviation loss function L1 to calculate the difference loss between the output image O and the high-resolution image HR, and use the Adam optimizer (the Adam optimizer is a commonly used optimizer in deep learning) for gradient descent. Repeat steps 1 to 5 for a total of 500,000 rounds to obtain the final image super-resolution model M;

[0088] Step 6: Divide the low-resolution image into several image patches of size 96*96, and sequentially input these image patches into the image super-resolution model obtained above;

[0089] Step 7: After the above image patches are sequentially subjected to shallow feature extraction by the convolutional neural network, deep feature extraction by the Transformer block, and pixel reorganization image reconstruction, high-resolution image patches are obtained; the obtained high-resolution image patches are sequentially stitched together according to their order in the low-resolution image to obtain the final high-resolution image.

[0090] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0091] In an embodiment of the present invention, as Figure 1 shown, the workflow for improving the image resolution constructed by the method of the present invention is roughly divided into two major stages:

[0092] The first stage: Train the image super-resolution model, including preparing the DIV2K dataset. As Figure 2 shown, each time 32 low-resolution LR pictures are selected from the DIV2K dataset, and each LR image is as Figure 3 shown. Then, image patches of size 48*48 are randomly cropped from each picture, and these 32 image patches are used as a batch for training. After preparing the training data, the batch is input into the image super-resolution model, and after passing through shallow feature extraction by the convolutional neural network, deep feature extraction by the Transformer block, and pixel reorganization image reconstruction in sequence, the output O is obtained. Then, the least absolute deviation loss function L1 is used to calculate the difference between the output image O and the high-resolution image HR, where the high-resolution image is as Figure 4 shown. Finally, the Adam optimizer is used to optimize the parameters of the entire network. The above process needs to be repeated 500,000 times to end. After training is completed, the final image super-resolution model M is obtained.

[0093] In the second stage, the image super-resolution model obtained in the first stage is used to process the input image. This includes first cropping the input low-resolution image into several image patches of size 96*96, as shown in the input low-resolution image Figure 5 . Then, these image patches are sequentially input into the image super-resolution model M. After passing through the shallow feature extraction of the convolutional neural network, the deep feature extraction of the Transformer block, and the pixel recombination for image reconstruction, several image patches of size 192*192 can be obtained. Then, these image patches of size 192*192 are stitched together in the order in which they appear in the input image, and a higher-resolution output image can be obtained, as shown in Figure 6 . By magnifying a partial area of the output image, as shown in Figure 7 , we can see that it is clearer than the same area of the input low-resolution image.

[0094] In the method for improving image resolution described in this embodiment, step 1 includes:

[0095] The training dataset uses the DIV2K dataset. The DIV2K dataset includes high-resolution images HR and low-resolution images LR. Among them, the number of high-resolution images is 800, and there are 800 x2-fold, x3-fold, and x4-fold low-resolution images each. Taking the 2-fold image super-resolution task as an example, the high-resolution image HR and the x2-fold LR image form an image pair. The training batch size is set to 32, that is, 32 image patches are processed at one time during the training of the image super-resolution model.

[0096] In the method for improving image resolution described in this embodiment, step 2 includes:

[0097] The batch size during training is set to 32, and the patch size is set to 48. Here, batch size is the number of image patches processed by the model at one time, and the image patch size patch size is the length and width of the image patch. Then, the dimension of the input image data X is [32, 3, 48, 48], where 3 represents that the number of channels of the image is 3, that is, the three RGB channels. For the input data X, first use the convolutional neural network to perform shallow feature extraction on X to obtain the feature map F 1 , and the dimension of F 1 is [32, 180, 48, 48]. This process can be expressed by the following formula:

[0098] F 1 = CNN(X)

[0099] The Convolutional Neural Network method CNN in this process includes a 3x3 convolution with an input dimension of 3, an output dimension of 180, a convolution kernel size of 3, 1 pixel of padding on the edges, and a stride of 1.

[0100] In the method for improving image resolution described in this embodiment, step 3 includes the following steps:

[0101] Input the feature map F obtained in step 2 1 into the Transformer block for deep feature extraction. After being processed by this method, the output is a new feature map F 2 , whose dimension is [32, 180, 48, 48]. This process specifically includes the following steps:

[0102] Step 3-1, the input feature map F 1 has a dimension of [32, 180, 48, 48]. First, use the flatten operation to convert its dimension to [32, 180, 2304], and then use the transpose operation to convert its dimension to [32, 2304, 180]. This process can be expressed by the following formula:

[0103] X 0 = F 1 .flatten(-2).transpose(-1, -2)

[0104] Among them, the flatten operation flattens the matrix, and the transpose operation transposes the matrix.

[0105] Step 3-2, input X 0 into the position encoding convolutional neural network PosCNN to calculate the position encoding of X 0 . The PosCNN method is implemented using a 3x3 convolution. The input dimension of this convolution is 180, the output dimension is 180, the convolution kernel size is 3, the number of padding pixels on the edges is 1, the stride is 1, and the number of groups is 180. Use PosCNN to obtain the position encoding pos of X 0 . The dimension of pos is also [32, 2304, 180]. Then add pos and X 0 to get X 1 , and this process can be expressed by the following formula:

[0106] pos = PosCNN(X 0 )

[0107] X 1 = X 0 + pos

[0108] Step 3-3, take X 1Input into the Transformer block method, the Transformer block method contains a total of 36 Transformer Layer structures. Each Transformer Layer consists of two parts. The first part is MSA or EWMSA, and the second part is MLP. According to the sequence number of each Transformer Layer (1, 2,..., 36), if it is odd, the first part is MSA, and if it is even, the first part is EWMSA. The calculation process of each Transformer Layer can be expressed by the following formula:

[0109] X 2 = MSA(LN(X 1 )) + X 1 …(1) or X 2 = EWMSA(LN(X 1 )) + X 1 …(2)

[0110] F 2 = MLP(LN(X 2 )) + X 2 …(3)

[0111] Among them, MSA in the formula represents Multi-head Self-Attention, that is, the multi-head attention mechanism method, EWMSA represents Effective Wide-area Multi-head Self-Attention, that is, the efficient global multi-head attention mechanism method, LN represents the LayerNorm operation, and MLP represents Multi-Layer Perceptron, that is, the multi-layer perceptron. The dimension of F 2 obtained after being processed by the Transformer block is [32, 180, 48, 48].

[0112] Step 3-4, add the feature maps F1 and F2 together through residual connection to fuse the features of F1 and F2. This process can be expressed by the following formula:

[0113] F 3 = F 1 + conv(F 2 )

[0114] Among them, the input dimension of conv in this process is 180, the output dimension is 180, the convolution kernel size is 3, the number of edge padding pixels is 1, and the stride is 1.

[0115] In the method for improving image resolution described in this embodiment, the step 4 includes:

[0116] Input the F obtained in step 3 3 into the pixel recombination method for upsampling. This process specifically includes three operations, namely conv_before_upsample processing, upsampling processing, and conv_last processing. F 3 has a dimension of [32, 180, 48, 48]. After being processed by the pixel recombination method, the obtained output O has a dimension of [32, 180, 96, 96], and the dimension of its corresponding high-resolution image is [32, 180, 96, 96]. This process can be expressed by the following formula:

[0117] F 4 = conv_before_upsample(F 3 )

[0118] F 5 = upsampling(F 4 )

[0119] O = conv_last(F 5 )

[0120] In the method for improving image resolution described in this embodiment, step 5 includes the following steps:

[0121] Step 5-1, use the least absolute deviation loss function to calculate the difference loss between the output O and its corresponding high-resolution image HR. This process can be expressed by the following formula:

[0122] loss = L1(O, HR)

[0123] Step 5-2, use the Adam optimizer to update the network parameters, and some of the main parameter settings of the Adam optimizer are as follows:

[0124] learn rate = 0.0002; weight decay = 0; milestones = [250000, 400000, 450000, 475000, 500000]; gamma = 0.5;

[0125] Step 5-3, repeat steps 1 to 5 for a total of 500,000 rounds. After the training is completed, the final image super-resolution model M is obtained.

[0126] In the method for improving image resolution described in this embodiment, step 6 includes the following steps:

[0127] Prepare a low-resolution image LR. During testing, the resolution of LR can be of any size. Divide LR into several image patches of size 96*96. Suppose n image patches of size 96*96 are finally obtained, and let these image patches be X i (i = 1, 2...n).

[0128] In the method for improving image resolution described in this embodiment, step 7 includes the following steps:

[0129] Step 7-1, sequentially input the image patches obtained in step 6 into the image super-resolution model M obtained in step 5. This step will obtain n image patches of size 192*192. Let these image patches be Y i (i = 1, 2...n). Specifically, this step can be represented by the following formula:

[0130]

[0131]

[0132]

[0133] Or

[0134]

[0135]

[0136]

[0137]

[0138]

[0139] In the Transformer block method, MSA and EWMSA are used alternately. That is, the odd layers use formulas (1) and (3) for calculation, and the even layers use formulas (2) and (3). The meanings of the symbols in this step are the same as those in step 3.

[0140] Step 7-2, the above step obtains n high-resolution image patches, which are respectively Y i (i = 1, 2...n). Concatenate these n high-resolution image patches together in the order of X i (i = 1, 2...n) in LR to obtain the final high-resolution image Y. Compared with the original low-resolution image LR, Y has a higher resolution and a clearer picture.

[0141] There are a large number of images on the Internet. However, due to various reasons, the resolution of many images is relatively low, which results in a poor viewing experience for us. In some cases, even the key information in the images cannot be clearly seen. Therefore, improving the resolution of images is of great significance. Traditional methods such as using software like Photoshop to process images often require the operator to master relatively professional skills and are time-consuming and laborious. Moreover, the image quality of the obtained high-resolution images is not always satisfactory. By using deep learning technology to train an image super-resolution model, once the final model is obtained after training, the machine can automatically process low-resolution images to obtain high-resolution images, resulting in higher image quality and a better viewing experience, which can greatly improve the processing efficiency.

[0142] In a specific implementation, the present invention further provides a computer storage medium. The computer storage medium can store a program, and when the program is executed, it may include some or all of the steps in the embodiments of a method for improving the resolution of images provided by the present invention. The storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0143] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.

[0144] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. The above-described embodiments of the present invention do not constitute a limitation to the protection scope of the present invention.

[0145] The present invention provides an idea and method for a method of improving image resolution. There are many ways to specifically implement this technical solution. The above is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by existing technologies.

Claims

1. A method for improving image resolution, characterized in that, it includes the following steps: Step 1, use the DIV2K dataset to train an image super-resolution model, and randomly crop 48*48-sized image patches from the pictures in the DIV2K dataset as training data; Step 2, use a convolutional neural network to perform shallow feature extraction on the image patches to obtain a feature map; Step 3, use a Transformer block to perform deep feature extraction on the feature map to obtain a new feature map; Step 4, use a pixel reorganization method to perform image reconstruction on the feature map obtained in Step 3 to obtain an output image O; Step 5, use the least absolute deviation loss function to calculate the difference between the output image O and the high-resolution image HR, and use the Adam optimizer for gradient descent; repeat Steps 1 to 5 for a total of 500,000 rounds to obtain the final image super-resolution model; Step 6, during testing, divide the low-resolution image into several 96*96-sized image patches; Step 7, sequentially input the image patches in Step 6 into the image super-resolution model in Step 5 to obtain high-resolution image patches; splice the obtained high-resolution image patches in sequence according to their order in the low-resolution image to obtain the final high-resolution image; Among them, Step 3 includes: Input the feature map F obtained in step 2 1 into the Transformer block for deep feature extraction, and the output is a new feature map F 2 , whose dimension is [32, 180, 48, 48]; this process includes the following steps: Step 3-1, input the feature map F 1 with a dimension of [32, 180, 48, 48]. Use the flatten operation to convert its dimension to [32, 180, 2304]; use the transpose operation to convert its dimension to [32, 2304, 180]. After the operation, obtain the matrix X 0 , and the process is as follows: X 0 = F 1 .flatten(-2).transpose(-1,-2) Among them, the flatten operation is to flatten the matrix, and the transpose operation is to transpose the matrix; Step 3-2, input X 0 into the position encoding convolutional neural network PosCNN to calculate the position encoding of X 0 . The position encoding convolutional neural network is implemented using a 3x3 convolution. The input dimension of this convolution is 180, the output dimension is 180, the convolution kernel size is 3, the number of padding pixels on the edges is 1, the stride is 1, and the number of groups is 180. Use the position encoding convolutional neural network to obtain the position encoding pos of X 0 . The dimension of the position encoding pos is also [32, 2304, 180]. Then add the position encoding pos and X 0 to obtain the matrix X 1 . This process is shown as follows: pos = PosCNN(X 0 ) X 1 = X 0 + pos; Step 3-3, input X 1 into the Transformer block. The Transformer block contains a total of 36 Transformer layer structures. Each Transformer layer structure consists of two parts. The first part is the multi-head attention method MSA or the efficient global multi-head attention method EWMSA, and the second part is the multi-layer perceptron method MLP. According to the sequence number (1, 2…, 36) of each Transformer Layer, if it is odd, the first part is MSA, and if it is even, the first part is EWMSA. The calculation process of each Transformer layer structure is as follows: X 2 = MSA(LN(X 1 )) + X 1 …(1) or X 2 = EWMSA(LN(X 1 )) + X 1 …(2) F 2 = MLP(LN(X 2 )) + X 2 …(3) Among them, MSA represents the Multi-head Self-Attention method, that is, the multi-head attention method; EWMSA represents the Effective Wide-area Multi-head Self-Attention method, that is, the efficient global multi-head attention method; LN represents the LayerNorm operation, that is, the layer normalization operation; MLP represents the Multi-Layer Perceptron method, that is, the multi-layer perceptron method; the new feature map F obtained by processing through the Transformer block method 2 has a dimension of [32, 180, 48, 48]; Step 3-4, add the feature maps F1 and F2 together through residual connection, fuse the features of the feature maps F1 and F2, and obtain the feature map F 3 , and this process is expressed as follows: F 3 = F 1 + conv(F 2 ) Among them, conv in this process represents a convolution operation. The input dimension of this convolution is 180, the output dimension is 180, the convolution kernel size is 3, the number of edge padding pixels is 1, and the stride is 1.

2. A method for improving image resolution according to claim 1, characterized in that, Step 1 includes: The training dataset uses the DIV2K dataset, which includes high-resolution images HR and low-resolution images LR; among them, the number of high-resolution images is 800, and there are 800 x2-times low-resolution images, 800 x3-times low-resolution images, and 800 x4-times low-resolution images respectively; Form image pairs from high-resolution images and low-resolution images according to different tasks; Among them, in the 2-fold image super-resolution task, the high-resolution image HR and the x2-fold low-resolution image LR form an image pair; In the 3-fold image super-resolution task, the high-resolution image HR and the x3-fold low-resolution image LR form an image pair; In the 4-fold image super-resolution task, the high-resolution image HR and the x4-fold low-resolution image LR form an image pair; The training batch size is set to 32, that is, 32 image patches are processed at one time during the training process of the image super-resolution model.

3. A method for improving image resolution according to claim 2, characterized in that, Step 2 includes: The length and width of the image patches during training are set to 48, then the dimension of the input image data X is [32, 3, 48, 48], where 3 represents that the number of channels of the image is 3, that is, 3 RGB channels; For the input image data X, a convolutional neural network is used for shallow feature extraction to obtain the feature map F 1 , F 1 has a dimension of [32, 180, 48, 48]; the process is as follows: F 1 = CNN(X) Among them, the convolutional neural network (CNN) of this process includes a 3x3 convolution, with an input dimension of 3, an output dimension of 180, a convolution kernel size of 3, 1 pixel of padding on the edges, and a stride of 1.

4. A method for improving image resolution according to claim 3, wherein, step 4 includes: Take the feature map F obtained in step 3 3 Perform upsampling using the pixel recombination method; this process specifically includes three operations, namely conv_before_upsample processing, upsampling processing, and conv_last processing; where conv_before_upsample is a convolution operation, upsampling is an upsampling operation, and conv_last is also a convolution operation; the feature map F 3 has a dimension of [32, 180, 48, 48]. After being processed by the pixel recombination method, the dimension of the output image O obtained is [32, 180, 96, 96], and the dimension of its corresponding high-resolution image is [32, 180, 96, 96]; this process is represented as follows: F 4 = conv_before_upsample(F 3 ) F 5 = upsample(F 4 ) O = conv_last(F 5 ) Among them, F 4 and F 5 both represent the feature maps obtained in the intermediate steps of the calculation.

5. A method for improving image resolution according to claim 4, wherein, step 5 includes the following steps: Step 5-1, use the least absolute deviation loss function L1 to calculate the difference loss between the output image O and its corresponding high-resolution image HR; this process is expressed as follows: loss = L1(O, HR).

6. A method for improving image resolution according to claim 5, wherein, step 5 includes the following steps: Step 5-2, use the Adam optimizer to update the network parameters, and the parameter settings of the Adam optimizer are as follows: learn rate = 0.0002; weight decay = 0; milestones = [250000, 400000, 450000, 475000, 500000]; gamma = 0.

5.

7. A method for improving image resolution according to claim 6, wherein, step 5 includes the following steps: Step 5-3, repeat steps 1 to 5 for a total of 500,000 rounds. After the training is completed, the final image super-resolution model M is obtained.

8. A method for improving image resolution according to claim 7, wherein, step 6 includes: Prepare a low-resolution image LR. During testing, the resolution of the low-resolution image LR can be of any size. Divide the low-resolution image LR into several image patches of size 96*96. Suppose finally n image patches of size 96*96 are obtained, and let these image patches be X i , where i = 1, 2... n, and n is a natural number.

9. A method for improving image resolution according to claim 8, wherein, step 7 includes the following steps: Step 7-1: Input the image patches obtained in Step 6 into the image super-resolution model M obtained in Step 5 in sequence; n image patches of size 192*192 are obtained; let these image patches be Y i, i = 1, 2... n, where n is a natural number; the specific process is as follows: Among them, represents the feature map obtained after the \(i\)-th image patch is processed by the convolutional CNN, represents the matrix obtained by flattening and transposing the matrix of the feature map after the matrix flattening and matrix transpose operations, represents the matrix after adding the position encoding, represents the feature map obtained after being processed by the MLP method, represents the feature map and the feature map after residual connection, represents the feature map obtained after being processed by the conv_before_upsample operation, represents the feature map obtained after the upsampling operation, and \(O\) represents the final output image patch obtained after being processed by the conv_last operation; The formula for the odd layer is: where represents the matrix obtained after the multi-head attention mechanism operation; The formula for the even layer is as follows: where denotes the matrix obtained after the operation of the efficient global multi-head attention mechanism; Among them, in the Transformer block method, MSA and EWMSA are alternately used; that is, the odd layers use the odd-layer formula (1) for calculation, and the even layers use the even-layer formula (2); Step 7-2, the n high-resolution image patches Y are obtained in Step 7-1 i , and these n high-resolution image patches are stitched together in the order of the image patch X i in the low-resolution image LR to obtain the final high-resolution image Y.

Citation Information

Patent Citations

  • An image super-resolution reconstruction method based on a convolutional neural network

    CN109903228A

  • Image reconstruction method and system

    CN113256497A