Image Reconstruction Method Based on Multi-Granularity Vit Autoencoder

Through the image reconstruction method based on multi-grained Vit automatic encoder, the problem of information loss and model size and resolution trade-offs during image reconstruction is solved, effective noise reduction processing and image super resolution are achieved, and image reconstruction quality is improved.

CN115619681BActive Publication Date: 2025-06-13FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211374681.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-04
Publication Date
2025-06-13
Estimated Expiration
2042-11-04

AI Technical Summary

Technical Problem

The prior art has a trade-off between information loss and model size and resolution in the image reconstruction process, and the convolution reduces the feature map size to cause the encoding process to lose information.

Method used

The image reconstruction method based on a multi-grained Vit automatic encoder is adopted to realize effective noise reduction processing and image super-resolution under image reconstruction tasks through a multi-grained Vit image refiner network. The method includes constructing an image reconstruction training set, training a Vit-based image reconstruction refiner, using an encoder and a decoder to extract and restore image features, and fusing global information and local information through a jump connection module.

Benefits of technology

It effectively reduces the information loss in the image encoding process, improves the high-resolution effect after image reconstruction, maintains scale invariance through attention mechanism, replaces the characteristic of smaller feature maps during convolution, and reduces information loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115619681B_ABST
    Figure CN115619681B_ABST
Patent Text Reader

Abstract

The present invention relates to an image reconstruction method based on a multi-granularity Vit autoencoder, comprising the following steps: Step S1: Construct an image reconstruction training set and train a Vit-based image reconstruction refiner, the Vit-based image reconstruction refiner including an encoder, a decoder and a skip connection module; Step S2: Input the original image into the encoder to obtain intermediate features, and sample the encoder local information in each encoding layer; Step S3: Input the obtained intermediate features into the decoder to restore the image information, and fuse the decoded information with the global information in each decoding process. The present invention realizes effective noise reduction processing and image super-resolution under the task of image reconstruction through this multi-granularity Vit image refiner network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of pattern recognition and computer vision, and particularly to an image reconstruction method based on a multi-granularity Vit autoencoder. Background Art

[0002] In recent years, deep convolutional neural networks have achieved great success in the field of computer vision, and deep learning-based image generation methods have also utilized the advantages of deep convolutional neural networks in feature extraction. Various GAN models have continuously obtained more accurate and better generation results in the fields of image generation and image denoising. In addition, more and more researchers have published papers related to image reconstruction in various computer vision conferences. Because image reconstruction has a wide range of application fields and huge commercial value, both academia and industry have been continuously exploring new image reconstruction technologies. In recent years, with the major breakthroughs achieved by deep learning and attention mechanisms in the field of computer vision, more breakthroughs can be obtained in image reconstruction tasks by combining vision Transformer methods.

[0003] An autoencoder is a classic generative model. For example, a convolutional network with a Unet structure is commonly used as a refiner in the field of image generation. Although convolutional neural networks have achieved good results in the GAN field, there are still some problems, such as the model needs to balance image resolution and model size, and the characteristic of convolutional shrinking the feature map size causes some information loss in the encoding process. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide an image reconstruction method based on a multi-granularity Vit autoencoder, which realizes effective noise reduction processing and image super-resolution under the image reconstruction task through a multi-granularity Vit image refiner network.

[0005] To achieve the above purpose, the present invention adopts the following technical solutions:

[0006] An image reconstruction method based on a multi-granularity Vit autoencoder, comprising the following steps:

[0007] Step S1: Construct an image reconstruction training set and train a Vit-based image reconstruction refiner, where the Vit-based image reconstruction refiner includes an encoder, a decoder, and a skip connection module;

[0008] Step S2: Input the original image into the encoder to obtain intermediate features, and sample the encoder local information in each encoding layer;

[0009] Step S3: Input the obtained intermediate features into the decoder to restore the image information, and fuse the decoding information with the global information in each decoding process.

[0010] Further, the construction of the image reconstruction training set is specifically as follows: Obtain a publicly available image set from the network as the image reconstruction training set, apply similarity transformation to the images in the image reconstruction training set, subtract the mean value from the pixel values of all input images, and perform normalization.

[0011] Further, the encoder structure adopts 4 Vit encoding layers different from the traditional ones, and uses shard sizes of (64, 32, 16, 8) as the input of different granularity features to the encoding layer respectively. Similarly, the corresponding encoding layers adopt the number of encoding heads of (2, 4, 8, 12);

[0012] During the encoding process of each layer, convolve on the feature map with the convolutional kernel of the next layer's granularity size to replace the downsampling process in the traditional convolutional network, which is used to change the feature granularity. At the same time, sample the local information of the encoder, and send the features of the intermediate encoding layer as the encoded global information into the skip connection module.

[0013] Further, the decoder structure adopts 4 Vit decoding layers, and uses shard sizes of (8, 16, 32, 64) as the input of different granularity features to the decoding layer respectively. Similarly, the corresponding decoding layers adopt the number of decoding heads of (12, 8, 4, 2);

[0014] Send the features of the intermediate decoding layer as the decoded local information into the skip connection module. During the decoding process of each layer, fuse the global information and the decoded local information through the skip connection module, and send the output fused features into the lower decoding module.

[0015] Further, the skip connection module fuses the decoded local information as the key and value vectors, and the global information as the query vector through the cross-attention mechanism to obtain the fused features and input them into the lower decoding layer for further analysis. The calculation formula is as follows:

[0016] F d,k+1 = F d,k + coAttn(Norm(F e,k ), Norm(F d,k ))

[0017] Where F d,k represents the k-th layer decoding feature map, F e,k represents the k-th layer encoding feature map, F d,k+1 represents the decoding feature map sent to the lower layer after being fused through the skip connection. Norm(·) represents the normalization layer, and coAttn(·) represents the cross-attention.

[0018] A refiner for image reconstruction based on a multi-granularity Vit autoencoder, comprising an encoder, a decoder, and a skip connection module; the encoder adopts a multi-granularity Vit structure, which is a network model for extracting deep features of an image and outputs intermediate features; the decoder adopts a multi-granularity Vit structure, which is used to decode and restore the intermediate features obtained by the encoder and outputs an image with the same scale as the original image; the skip connection module fuses the decoded local information as key and value vectors and the global information as a query vector through a cross-attention mechanism to obtain a fused feature and input it into the lower decoding layer for further analysis.

[0019] The present invention has the following beneficial effects compared with the prior art:

[0020] 1. The present invention can effectively denoise an image to complete an image reconstruction task, reduce the loss in the image encoding process, and improve the high-resolution effect of the image after reconstruction.

[0021] 2. The present invention uses the Transfomer method to form the encoder and decoder structures, maintains the scale invariance in the encoding process through the attention mechanism method, replaces the characteristic that the feature map becomes smaller in the convolution process, and reduces the information loss in the encoding process.

[0022] 3. The present invention uses multi-granularity slicing for feature extraction, so that each encoding layer extracts different granularity features. Without changing the scale of the feature map, the features of the image are summarized through the multi-granularity method, making the extracted features more complete and plump.

[0023] 4. The present invention connects each layer of the encoding and decoding parts through a skip connection module, so that each granularity feature realizes the fusion of features before and after encoding through this module. While combining the advantages of the feature pyramid, a cross-attention mechanism is added to make the features before encoding guide the features after encoding, optimize the fusion process, and improve the feature matching degree. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a schematic diagram of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] The present invention will be further described below with reference to the drawings and embodiments.

[0026] Please refer to Figure 1 , the present invention provides an image reconstruction method based on a multi-granularity Vit autoencoder, comprising the following steps:

[0027] Step S1: Construct an image reconstruction training set and train an image reconstruction refiner based on Vit. The image reconstruction refiner based on Vit includes an encoder, a decoder, and a skip connection module of Vit;

[0028] Step S2: Input the original image into the Vit encoder to obtain intermediate features, and sample the encoder local information in each encoding layer;

[0029] Step S3: Input the obtained intermediate features into the decoder to restore the image information, and fuse the decoding information with the global information in each decoding process.

[0030] In this embodiment, the construction of the image reconstruction training set is specifically as follows: Obtain a publicly available image set from the network as the image reconstruction training set and use it for the self-supervised training of the network model; Apply similarity transformation to the images in the image reconstruction training set, subtract the mean value from the pixel values of all input images, and perform normalization. Apply similarity transformation to the images in the image reconstruction training set, subtract the mean value from the pixel values of all input images, and perform normalization.

[0031] In this embodiment, the network structure of the refiner is composed of a Vit-based encoder, decoder, and skip module, replacing the traditional Unet network structure composed of CNN convolutions. The information loss in the image encoding process using Vit is lower than that of the traditional method, and finally, better image reconstruction quality is obtained.

[0032] Let D = {d 1 , d 2 ,..., d N} be the images in the training set, and d i be the i-th image. Calculate using MSELoss, and its calculation formula is as follows:

[0033]

[0034] where is the network prediction output image;

[0035] Then, use the gradient descent and backpropagation algorithms to update the network parameters, and use the trained model for image denoising. Preprocess the sample images in the image set in the same way as S12 and input them into the trained network model. Finally, output new image data to obtain the denoised image, realizing image reconstruction. And this structure can be used in various image generation fields and inserted as a refinement module to continuously correct the image features in the generation process.

[0036] Preferably, in this embodiment, a multi-granularity Vit structure is used as the encoder, a network model for extracting image depth features, and intermediate features are output; The Vit encoder structure uses 4 Vit encoding layers different from the traditional ones, and uses shard sizes of (64, 32, 16, 8) as the input encoding layers for different granularity features respectively. Similarly, the corresponding encoding layers use the number of encoding heads of (2, 4, 8, 12);

[0037] During each layer of the encoding process, a convolutional kernel with the granularity size of the next layer is used to perform convolution on the feature map to replace the downsampling process in the traditional convolutional network, which is used to change the feature granularity. Meanwhile, the local information of the encoder is sampled, and the features of the intermediate encoding layer are used as the encoded global information and sent to the skip connection module.

[0038] Preferably, in this embodiment, a multi-granularity ViT structure is used as the decoder to decode and restore the intermediate features and output an image with the same scale as the original image. The decoder structure adopts 4 layers of ViT decoding layers, and the shard sizes of (8, 16, 32, 64) are respectively used as the input of different granularity features to the decoding layer. Correspondingly, the number of encoding heads of the decoding layers is (12, 8, 4, 2) respectively.

[0039] The features of the intermediate decoding layer are used as the decoded local information and sent to the skip connection module. During each layer of the decoding process, the global information and the decoded local information are fused through the skip connection module, and the output fused features are sent to the lower-level decoding module.

[0040] In this embodiment, the local information of the encoding layer is used as the global information and input into the skip connection module; during the middle of the decoding layer, the decoded output feature information is used as the decoded local information and sent to the skip connection module.

[0041] Different from the skip connection path of the traditional feature pyramid, in the skip connection module, the decoded local information is used as the key and value vectors, and the global information is used as the query vector through the cross-attention mechanism for fusion, and the fused features are input into the lower-level decoding layer for further analysis. The calculation formula is as follows:

[0042] F d,k+1 =F d,k +coAttn(Norm(F e,k ),Norm(F d,k ))

[0043] Where F d,k represents the k-th layer decoding feature map, F e,k represents the k-th layer encoding feature map, F d,k+1 represents the decoding feature map sent to the lower layer after being fused through the skip connection. Norm(·) represents the normalization layer, and coAttn(·) represents the cross-attention.

[0044] The above are only the preferred embodiments of the present invention. All equivalent changes and modifications made according to the scope of the patent application of the present invention shall fall within the scope covered by the present invention.

Claims

1. An image reconstruction method based on a multi-granularity Vit autoencoder, characterized in that, it includes the following steps: Step S1: Construct an image reconstruction training set and train a Vit-based image reconstruction refiner, where the Vit-based image reconstruction refiner includes an encoder, a decoder, and a skip connection module; Step S2: Input the original image into the encoder to obtain intermediate features, and sample the encoder local information in each encoding layer; Step S3: Input the obtained intermediate features into the decoder to restore the image information, and fuse the decoded information with the global information in each decoding process; The structure of the encoder adopts 4 Vit encoding layers different from the traditional ones, and uses shard sizes of (64, 32, 16, 8) as different granularity features to input into the encoding layer. Similarly, the corresponding encoding layers respectively adopt the number of encoding heads of (2, 4, 8, 12); In each encoding process, a convolutional kernel with the granularity size of the next layer is used to convolve on the feature map to replace the downsampling process in the traditional convolutional network, which is used to change the feature granularity. At the same time, sample the encoder local information, and send the features of the intermediate encoding layer as the encoded global information into the skip connection module; The structure of the decoder adopts 4 Vit decoding layers, and uses shard sizes of (8, 16, 32, 64) as different granularity features to input into the decoding layer. Similarly, the corresponding decoding layers respectively adopt the number of decoding heads of (12, 8, 4, 2); Send the features of the intermediate decoding layer as the decoded local information into the skip connection module. In each decoding process, fuse the global information and the decoded local information through the skip connection module, and send the output fused features into the lower decoding module; The skip connection module fuses the decoded local information as the key and value vectors, and the global information as the query vector through the cross-attention mechanism to obtain the fused features and input them into the lower decoding layer for further analysis. The calculation formula is as follows: F d,k+1 = F d,k + coAttn(Norm(F e,k ), Norm(F d,k )) Among them, F d,k represents the decoding feature map of the k-th layer, and F e,k represents the encoding feature map of the k-th layer, and F d,k+1 represents the decoding feature map sent to the lower layer after being fused through skip connections. Norm(·) represents the normalization layer, and coAttn(·) represents cross-attention.

2. The image reconstruction method based on a multi-granularity Vit autoencoder according to claim 1, characterized in that, the construction of the image reconstruction training set is specifically: obtain a publicly available image set from the network as the image reconstruction training set, apply similarity transformation to the images in the image reconstruction training set, subtract the mean value from the pixel values of all input images, and perform normalization.

3. A refiner for image reconstruction based on a multi-granularity Vit autoencoder, characterized in that, it includes an encoder, a decoder, and a skip connection module; the encoder adopts a multi-granularity Vit structure, which is a network model for extracting deep features of images and outputs intermediate features; the decoder adopts a multi-granularity Vit structure, which is used to decode and restore the intermediate features obtained by the encoder, and outputs an image with the same scale as the original image; the skip connection module fuses the decoded local information as the key and value vectors, and the global information as the query vector through the cross-attention mechanism to obtain the fused features and input them into the lower decoding layer for further analysis.

Citation Information

Patent Citations

  • Motion-compensated compression of dynamic voxelized point clouds

    CN109196559A

  • Pulmonary nodule image detection method and system based on CT image

    CN113888466A