Hybrid semantic based dual-stream image reconstruction system and method
By combining low-level features and high-level semantic features in a dual-stream image reconstruction system, the problem of semantic information loss in traditional image compression methods is solved, high-quality image reconstruction is achieved, and the structural and semantic information of the image is preserved.
Patent Information
- Application Number
- CN202210520908.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-13
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2042-05-13
AI Technical Summary
Traditional image compression methods lose global semantic information, resulting in block artifacts and blurred edges in decoded images at low bit rates. Existing image reconstruction methods based on convolutional neural networks have difficulty in obtaining semantic segmentation maps, and the reconstructed image quality is unclear.
A dual-stream image reconstruction system based on hybrid semantics is adopted, combining the low-level feature extraction module and the high-level semantic feature extraction module. Through the hybrid feature stream module and reconstruction module, the PiDiNet network and the ResNet-152 network are used to extract edge features and high-level semantic features, and the image reconstruction is performed in combination with deformable convolution and feature fusion operations.
The structural and semantic information of the image is preserved, obvious distortion is avoided, and high-quality image reconstruction is achieved.
Smart Images

Figure CN114972942B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of image reconstruction, in particular to a dual-flow image reconstruction system and method based on mixed semantics. BACKGROUND
[0002] In today's rapid development of the Internet, there will be a large amount of data every day, of which image data accounts for a large part of Internet data. Every day, a large amount of image data is generated and transmitted, and the network transmission pressure is increasing. An image reconstruction method that can save storage space and reduce network transmission pressure is of great significance to people's life and work. According to different tasks or generation results, the generalized image reconstruction includes denoising, deblurring, image restoration, image inpainting, image super-resolution, image compression and image translation, etc.
[0003] The encoding process of traditional block-based image compression methods such as JPEG includes transformation, quantization and encoding, and the decoding process includes corresponding inverse transformation, including decoding, inverse quantization and inverse transformation process. The decoding process is the reconstruction process of the image, which mainly reduces the pixel-level redundancy to improve the coding efficiency, and is a pixel-based image compression technology. In recent years, with the improvement of computing power and the powerful nonlinear expression ability of deep learning, convolutional neural networks have made major breakthroughs in computer vision tasks (image classification, image segmentation and target detection, etc.). At present, the image reconstruction method based on convolutional neural network CNN mainly includes three types: 1) image reconstruction using hyper-prior network and auto-encoder structure (see literature: Ballé, Minnen D, Singh S, et al. Variational image compression with a scale hyperprior [J]. 2018); 2) end-to-end image reconstruction using feature extraction and convolutional network (see literature: Agustsson E, Tschannen M, Mentzer F, et al. Generative Adversarial Networks for Extreme Learned Image Compression [C] / / 2019IEEE / CVF International Conference on Computer Vision (ICCV). IEEE, 2019.); 3) image reconstruction using feature and generative adversarial network GAN network (see literature: Isola P, Zhu J Y, Zhou T, et al. Image-to-Image Translation with Conditional Adversarial Networks [J]. IEEE, 2016.). In method 1), the auto-encoder is used to extract the feature points to obtain the latent feature points, and then the hyper-prior network is used to capture the structural information of the latent feature points. The latent feature points are also adaptively modeled, and the entropy decoding result is input to the main decoding end to obtain the final reconstructed picture. The parameter optimization of the whole network still uses the overall rate-distortion function L = λ * D + R. In method 2), an encoder (feature extraction), a generator and a multi-scale discriminator are used to jointly train a generative model to reconstruct the image. The specific process is that the source image and the semantic segmentation map are input into the encoder to extract the latent features, and the quantized features are obtained by quantizing the latent features. The quantized features are input into the generator to obtain the reconstructed image. The generator can input the semantic segmentation map to preserve the details of the user's specific area, and the training uses an adversarial loss to obtain global semantic information.In the method 3), a conditional generative adversarial network (CGAN) is generated based on a condition, the image is guided to be reconstructed by adding the condition information, and edge features, semantic label images and the like are used as the condition information.
[0004] However, the traditional image compression method takes a pixel block as a unit, and global semantic information is lost during compression, and especially at a low bit rate, the decoded image often has serious distortion, such as block artifacts and edge blur. In the method 1), a hyper-prior architecture is used to model each feature point without using visual features. In the method 2), a high compression rate and a satisfactory reconstructed image can be obtained on a natural image, but a semantic segmentation image needs to be used during reconstruction and discrimination, and it is difficult to obtain an accurate semantic segmentation image of the image, especially for an image with a complex background. In the method 3), it is also difficult to obtain a semantic segmentation label, and the reconstructed image has unclear quality and is blurred. SUMMARY
[0005] Based on the above problems, the application provides a dual-flow image reconstruction system based on mixed semantics, which comprises a low-level feature extraction module, a high-level semantic feature extraction module, a mixed feature flow module, a low-level feature flow module and a reconstruction module.
[0006] The low-level feature extraction module is used to extract edge features of an image as low-level features according to an edge detection algorithm.
[0007] The high-level semantic feature extraction module is used to extract high-level semantic features of the image according to a variational autoencoder and a reparameterization method.
[0008] The low-level feature flow module is used to reconstruct image features x' out according to the low-level features.
[0009] The mixed feature flow module is used to reconstruct image features x nout according to the low-level features and the high-level semantic features.
[0010] The reconstruction module is used to reconstruct an image according to the features x' out , x nout .
[0011] The edge detection algorithm in the low-level feature extraction module uses a PiDiNet network.
[0012] The variational autoencoder in the high-level semantic feature extraction module is constructed by using a 152-layer residual network (abbreviated as ResNet-152) network.
[0013] The low-level feature flow module is constructed by using a deformable convolution block.
[0014] The mixed feature flow module comprises a deformable convolution, a fully connected layer and a progressive block Gi The convolution kernel size of the deformable convolution is 3*3; the progressive block G i An attention mechanism residual network block (referred to as SE-ResNet block) is adopted for construction.
[0015] The reconstruction module comprises a feature fusion concat() operation, a 3*3 convolution, and a To RGB module, the To RGB module comprising a 3*3 two-dimensional convolution and a spectral norm regularization operation.
[0016] The progressive block G i The backbone network is a bottleneck residual block, wherein the bottleneck residual block is composed of three convolution layers of 1*1 convolution, 3*3 convolution and 1*1 convolution, and is activated by using a LeakyRelu() activation function.
[0017] The application provides a dual-flow image reconstruction method based on mixed semantics, and the method is realized based on the dual-flow image reconstruction system based on mixed semantics.
[0018] Step 1: obtaining an image X, and pre-processing the image X to obtain a to-be-reconstructed image x with N*N pixels;
[0019] Step 2: sending the to-be-reconstructed image x into a low-level feature extraction module to extract low-level features l of the image;
[0020] Step 3: sending the to-be-reconstructed image x into a high-level semantic feature extraction module to extract high-level semantic features h of the image;
[0021] Step 4: sending the low-level features l extracted in step 2 and the high-level semantic features h extracted in step 3 into a mixed feature flow module to extract features x nout for reconstructing the image;
[0022] Step 5: sending the low-level features l extracted in step 2 into a low-level feature flow module to extract features x′ out for reconstructing the image;
[0023] Step 6: sending the features x nout and x′ out into a reconstruction module to output a reconstructed image
[0024] The step 3 comprises:
[0025] Step 3.1: inputting the to-be-reconstructed image x into a ResNet-152 backbone network to extract features;
[0026] Step 3.2: performing a 256-dimensional full connection layer on the features obtained in step 3.1, and performing an n zThe mean μ and log variance log e σ 2 ;
[0027] Step 3.3: The mean μ and log variance log e σ 2 The advanced semantic feature h is obtained using the reparameterization method.
[0028] The step 4 includes:
[0029] Step 4.1: The low-level feature l is input into the deformable convolution, and the average pooling is performed to obtain the partial input x0 of the progressive block G1;
[0030] Step 4.2: The advanced semantic feature h is input into the first fully connected layer to obtain a group of affine transformation parameters α0 and β0;
[0031] Step 4.3: The low-level feature l is down-sampled to obtain l0 which has the same dimension as x0, and then the feature fusion concat() operation is performed on x0 and l0 to obtain the fusion feature map [l0, x0];
[0032] Step 4.4: [l0, x0] and the affine transformation parameters α0 and β0 are input into the first progressive block G1 to obtain the output feature map x1;
[0033] Step 4.5: The feature fusion add() operation is performed on x0 and the output feature map x1 to obtain the fusion feature map x 1out ;
[0034] Step 4.6: The feature map x i-1 output by the previous progressive block G i-1 is obtained, the low-level feature l is down-sampled to obtain l i-1 which has the same dimension as x i-1 , the advanced semantic feature h is input into the i-th fully connected layer to obtain a group of affine transformation parameters α i-1 and β i-1 (i≥2);
[0035] Step 4.7: Along the depth direction of the reconstruction network, the feature map x i-1 and the low-level feature l i-1 are subjected to the feature fusion concat() operation to obtain the fusion feature map [l i-1 , x i-1 ], and then the fusion feature map [l i-1 , x i-1 ] and the affine transformation parameters α i-1 and β i-1 are input into the i-th progressive block G i to obtain the output feature map xi The output feature map x i and the fusion feature map x i-1 of the previous progressive block G (i-1)out perform a feature fusion add() operation to obtain the fusion feature map x iout .
[0036] Step 4.8: let i add 1, repeat steps 4.6-4.7 to obtain the fusion feature map x nout after n-1 times, wherein n=log2N-1.
[0037] The step 6 includes:
[0038] Step 6.1: perform a feature fusion concat() operation on the feature map x nout and x' out to obtain the feature concat(x nout , x' out );
[0039] Step 6.2: send the feature concat(x nout , x' out ) into a 3*3 convolution and use LeakyRelu() for activation;
[0040] Step 6.3: send the feature output by step 6.2 into a To RGB module to obtain the reconstructed image
[0041] The present application has the following beneficial effects:
[0042] The present application proposes a dual-flow image reconstruction system and method based on mixed semantics, designs a mixed feature flow branch and a low-level feature flow branch to build a reconstruction network, and uses a dual-feature flow network design to ensure that semantic information and structural information are better used for image reconstruction, wherein the mixed feature flow branch outputs a mixed semantic information feature map x nout , and the low-level feature flow branch outputs a structural information feature map x' out ; the low-level feature flow branch uses a deformable convolution module to extract features, and uses deformable convolution to adaptively change the receptive field to capture some important information; in order to strengthen the propagation and reuse of features, the feature fusion technology concat() operation and add() operation are used; the image reconstructed by the method of the present application can retain complete structural information and semantic information, and will not appear obvious distortion. BRIEF DESCRIPTION OF DRAWINGS
[0043] Figure 1 The figure is a block diagram of the dual-flow image reconstruction system based on mixed semantics in the present application;
[0044] Figure 2A principle diagram of two feature extraction principles in the application;
[0045] Figure 3 A principle diagram of high-level semantic feature extraction in the application;
[0046] Figure 4 A structure diagram of reconstruction based on semantic features in the application;
[0047] Figure 5 A principle diagram of two feature fusion modes in the application;
[0048] Figure 6 A reconstruction effect diagram of the method in the application on different data sets, wherein (a) is a reconstruction effect diagram for the CelebA-HQ and FFHQ data sets, and (b) is a reconstruction effect diagram for the alps seasons data set. DETAILED DESCRIPTION
[0049] The application will be further described below in combination with the drawings and specific implementation examples. In view of the problems existing in the prior image technology, there is an urgent need for a new image reconstruction method and system that can not only save storage space and reduce network transmission pressure, but also fully utilize semantic features to obtain good reconstruction effect. Selecting appropriate semantic features is the key to saving storage space and reducing network transmission pressure, and designing a better image reconstruction network is the key to reconstructing high-quality images. The application provides a thought based on which semantic features can be used for image reconstruction, and also provides a better image reconstruction network architecture, including a low-level feature extractor, a high-level semantic feature extractor, a deformable convolution layer, a plurality of fully connected layers, a plurality of progressive blocks G i concat(), an add() operation, a deformable convolution block, a normal convolution layer, and a To RGB module. At the same time, the application can be trained according to different applications to optimize the reconstruction process, so that it is suitable for a variety of application scenarios.
[0050] As shown in Figure 1 , a dual-flow image reconstruction system based on mixed semantics includes a low-level feature extraction module, a high-level semantic feature extraction module, a mixed feature flow module, a low-level feature flow module, and a reconstruction module.
[0051] The low-level feature extraction module is configured to extract edge features of an image as low-level features according to an edge detection algorithm.
[0052] The high-level semantic feature extraction module is configured to extract high-level semantic features of the image according to a variational autoencoder and a reparameterization method.
[0053] The low-level feature flow module is configured to reconstruct image features x′ out according to the low-level features.
[0054] The mixed feature flow module is configured to reconstruct the image feature x according to the low-level feature and the high-level semantic feature nout ;
[0055] The reconstruction module is configured to reconstruct the image according to the feature x′ out , x nout .
[0056] The low-level feature extraction module extracts the edge feature of the image as the low-level feature l using an edge detection algorithm, and the edge detection algorithm uses a PiDiNet network, which can be expressed by the following formula:
[0057] l = f edge (x) = PiDiNet(x), where f edge () represents the low-level feature extraction algorithm.
[0058] The high-level semantic feature extraction module extracts the high-level semantic feature h of the image using a variational autoencoder (VAE) and a reparameterization method, wherein the VAE is based on ResNet-152; the principle diagram is as shown in Figure 3 .
[0059] The high-level semantic feature h extraction operation can be expressed by the following mathematical model:
[0060] h = f h (x) where f h (x) represents the high-level semantic feature h extraction module
[0061] f h (x) uses a variational autoencoder VAE to first obtain the mean μ and the log e σ 2 variance vector log e σ 2 of the input image x, and then uses a reparameterization method to obtain the final high-level semantic feature h. Wherein the variational autoencoder VAE uses four layers of 2D convolution and one fully connected layer to constitute an Encoder, and uses the Encoder to obtain the mean μ and the log e σ 2 variance log e σ 2 of the input image x.
[0062] The mixed feature flow module is implemented using a deformable convolution, multiple fully connected layers, and multiple progressive blocks G i to obtain the feature x nout of the reconstructed image. Wherein the convolution kernel size of the deformable convolution is 3*3; the progressive block G i is implemented based on the SE-ResNet block, and the convolution kernel size is 3*3, using the LeakyRelu() activation function. The principle diagrams of the two feature extraction methods are as shown in Figure 2 .
[0063] The progressive block G i The main network is a bottleneck residual block, wherein the bottleneck residual block is composed of three convolutional layers of 1*1 convolution, 3*3 convolution and 1*1 convolution, and is activated by using a LeakyRelu() activation function.
[0064] The low-level feature flow module is implemented based on a deformable convolution block, the deformable convolution block is composed of three layers of deformable convolution, a convolution kernel size of which is 3*3, and a convolution kernel size of which obtains an offset is also 3*3, and input and output channel numbers of the three layers of deformable convolution are (1, 16, 3*3) -> (16, 8, 3*3) -> (8, 4, 3*3) in turn.
[0065] The reconstruction module is composed of a feature fusion concat() operation, a 3*3 convolution and a To RGB module, and the reconstructed image The specific operation is as follows: first, the feature maps x nout and x′ out are subjected to a feature fusion concat() operation to obtain fused features concat(x nout , x′ out ); the fused features are sent into a 3*3 convolution, and then subjected to LeakyRelu() activation; finally, the features obtained above are sent into a To RGB module to obtain the reconstructed image The To RGB module includes a 3*3 2D convolution and a spectral norm regularization operation (Spectral Norm Regularization). The feature fusion concat() operation is used to splice the low-level features l i and x i in the channel direction, so that the channel number is increased; the add() operation is used to superimpose the feature maps x i-1 and x i in an element-wise manner, so that the channel number remains unchanged. Through the two fusion modes, the features extracted by each layer of the network are better utilized, and the propagation and reuse of the features are strengthened. The reconstruction structure diagram is shown in Figure 4 , and the principle diagrams of the two feature fusion modes are shown in Figure 5 .
[0066] A dual-flow image reconstruction method based on mixed semantics, implemented based on the dual-flow image reconstruction system based on mixed semantics, the method comprising:
[0067] Step 1: Obtain an image X, and pre-process the image X to obtain a to-be-reconstructed image x with N*N pixels;
[0068] Step 2: send the image to be reconstructed x into the low-level feature extraction module to extract the low-level features l of the image;
[0069] Step 3: send the image to be reconstructed x into the high-level semantic feature extraction module to extract the high-level semantic features h of the image; including:
[0070] Step 3.1: input the image to be reconstructed x into the ResNet-152 backbone network to extract features;
[0071] Step 3.2: pass the features obtained in step 3.1 through a 256-dimensional fully connected layer, through an n z -dimensional fully connected layer, to obtain the mean μ and log e σ 2 of the high-level semantic features h; considering the computational complexity and the compact representation of the image, here take nz=64.
[0072] Step 3.3: use the reparameterization method to obtain the high-level semantic features h using the mean μ and log e σ 2 .
[0073] Step 4: send the low-level features l extracted in step 2 and the high-level semantic features h extracted in step 3 into the mixed feature flow module to extract the features x nout for reconstructing the image; including:
[0074] Step 4.1: send the low-level features l into deformable convolution, and perform average pooling to obtain the partial input x0 of the progressive block G1;
[0075] Step 4.2: input the high-level semantic features h into the first fully connected layer to obtain a group of affine transformation parameters α0 and β0;
[0076] Step 4.3: downsample the low-level features l to obtain l0 which has the same dimension as x0, and then perform feature fusion concat() operation on x0 and l0 to obtain the fusion feature map [l0, x0];
[0077] Step 4.4: send [l0, x0] and the affine transformation parameters α0 and β0 into the first progressive block G1 to obtain the output feature map x1;
[0078] Step 4.5: perform feature fusion sum() operation on x0 and the output feature map x1 to obtain the fusion feature map x 1out ;
[0079] Step 4.6: obtain the feature map x i-1 output by the previous progressive block G i-1 ; downsample the low-level features l to obtain l i-1 which has the same dimension as xi-1 ; input the high-level semantic feature h into the i-th fully connected layer to obtain a set of affine transformation parameters a i-1 and b i-1 (i > 2);
[0080] Step 4.7: along the depth direction of the reconstruction network, perform feature fusion concat() operation on the feature map x i-1 and the low-level feature l i-1 to obtain the fused feature map [l i-1 , x i-1 ], and then input the fused feature map [l i-1 , x i-1 ] and the affine transformation parameters a i-1 and b i-1 into the i-th progressive block G i to obtain the output feature map x i , and perform feature fusion add() operation on the output feature map x i and the fused feature map x i-1 of the previous progressive block G (i-1)out to obtain the fused feature map x iout ;
[0081] Step 4.8: let i be incremented by 1, and repeat steps 4.6-4.7 to obtain the fused feature map x nout after n-1 times of fusion, where n = log2N-1.
[0082] Step 5: input the low-level feature l extracted in step 2 into the low-level feature flow module to extract the feature x' for reconstructing the image out ;
[0083] Step 6: input the features x nout and x' out into the reconstruction module to output the reconstructed image comprising:
[0084] Step 6.1: perform feature fusion concat() operation on the feature maps x nout and x' out to obtain the feature concat(x nout , x' out );
[0085] Step 6.2: input the feature concat(x nout , x' out ) into a 3*3 convolution and use LeakyRelu() for activation;
[0086] Step 6.3: input the feature output by step 6.2 into the To RGB module to obtain the reconstructed image
[0087] In training the network, data sets of different resolutions (256*256, 512*512, 1024*1024) are used for training to enhance the generalization ability of the reconstruction network; in addition, the following techniques are used during training: ① celebaHQ dataset is trained alone, ② celebaHQ dataset and FFHQ dataset are mixed and then trained;
[0088] The image reconstruction network training is completed on an Ubuntu 18.04.5 LTS system, using an NVIDIA RTX 3090 GPU and a pytorch deep learning framework, and the training data uses paired image-low-level feature l and real image x, the batchsize is set to 2 each time, the Adam optimizer is used during training, the initial learning rate is 0.001, and after 40 epochs, the learning rate is adjusted to alpha*0.001 (where alpha=0.1, and a total of 200 epochs are trained). The image reconstruction algorithm proposed in the application can be tested on CPU or GPU, and the training results include edge features A, real images B and reconstructed images The test result is a reconstructed image The application can perform efficient end-to-end image reconstruction (i.e., input low-level feature l and high-level semantic feature h, and output is a reconstructed image), and the reconstructed image has good visual effect. The reconstruction effect on different data sets is as shown in Figure 6 .
[0089] Three data sets, the alps seasons dataset, celebaHQ dataset and FFHQ dataset, are used. According to the characteristics of the data sets, different edge extraction methods are used to extract low-level features (edge features), among which the alps seasons dataset uses a canny operator to extract, and the Celeba-HQ dataset and FFHQ dataset use a PiDiNet network to extract;
[0090] The peak signal-to-noise ratio PSNR and structural similarity SSIM values and subjective evaluation indexes of the method of the application and other reconstruction methods on the Celeba-HQ and FFHQ data sets are shown in Tables 1 and 2, respectively.
[0091] Table 1 PSNR and SSIM values on the Celeba-HQ dataset
[0092]
[0093] Table 2 PSNR and SSIM values on the FFHQ dataset
[0094]
[0095] In the table, pix2pix, cycleGAN, LCIC respectively represent the methods given in the following references:
[0096] [1] Isola P, Zhu J Y, Zhou T, et al. Image-to-image translation with conditional adversarial networks [C]. Proceedings of the IEEE conference on computer vision and pattern recognition. 2017: 1125-1134.
[0097] [2] Zhu J Y, Park T, Isola P, et al. Unpaired image-to-image translation using cycle-consistent adversarial networks [C]. Proceedings of the IEEE International Conference on Computer Vision. 2017: 2223-2232.
[0098] [3] Chang J, Mao Q, Zhao Z, et al. Layered conceptual image compression via deep semantic synthesis [C]. 2019 IEEE International Conference on Image Processing (ICIP). IEEE, 2019: 694-698.
[0099] The traditional image compression method takes a pixel block as a unit, and when compression is performed, the global semantic information is lost, especially at low bit rate, which often leads to serious distortion of the decoded image, such as block artifacts and edge blur. However, the dual-flow image reconstruction method based on hybrid semantics proposed by the present application uses two complementary semantic features, which can save storage space (edge features are single-channel, and high-level semantic features are 64-dimensional vectors), and due to the use of two complementary semantic features, the reconstructed image retains complete structural information and semantic information, and does not have obvious distortion.
[0100] Compared with the prior image reconstruction method, the image reconstruction scheme provided by the application fully uses two complementary semantic features of the image, i.e., low-level features (edge features) and high-level semantic features, the reconstructed image can not only maintain good detail information but also not lose the semantics of the whole image. Under this image reconstruction mechanism, the semantic features can be stored in a small space, and the original image can be reconstructed by sending the semantic features into the trained reconstruction network.
Claims
1. A dual-stream image reconstruction system based on hybrid semantics, characterized in that: include: Low-level feature extraction module, high-level semantic feature extraction module, hybrid feature flow module, low-level feature flow module, reconstruction module; The low-level feature extraction module is used to extract edge features of the image as low-level features according to the edge detection algorithm; The high-level semantic feature extraction module is used to extract high-level semantic features of the image based on the variational autoencoder and the reparameterization method; The low-level feature stream module is used to reconstruct image features x′ based on low-level features out ; The hybrid feature stream module is used to reconstruct image features x based on low-level features and high-level semantic features. nout ; The reconstruction module is used to out 、x nout Reconstruct the image; The hybrid semantics-based dual-stream image reconstruction system uses the following methods to achieve image reconstruction, including: Step 1: Obtain image X and preprocess image X to obtain the image x to be reconstructed with N*N pixels; Step 2: Send the image to be reconstructed x into the low-level feature extraction module to extract the low-level features l of the image; Step 3: Send the image to be reconstructed x to the high-level semantic feature extraction module to extract the high-level semantic features h of the image; Step 4: The low-level features l extracted in step 2 and the high-level semantic features h extracted in step 3 are fed into the hybrid feature stream module to extract the features x used to reconstruct the image. nout ; Step 4.1: Feed the low-level feature l into the deformable convolution and perform average pooling to obtain the partial input x0 of the progressive block G1; Step 4.2: Input the high-level semantic feature h into the first fully connected layer to obtain a set of affine transformation parameters α0 and β0; Step 4.3: Downsample the low-level feature l to obtain l0 of the same dimension as x0, and then perform feature fusion (concat()) operation on x0 and l0 to obtain the fused feature map [l0, x0]; Step 4.4: Feed [l0,x0] and affine transformation parameters α0 and β0 into the first progressive block G1 to obtain the output feature map x1; Step 4.5: Fuse x0 with the output feature map x1 to obtain the fused feature map x 1out ; Step 4.6: Get the previous progressive block G i-1 Output feature map x i-1 ; Downsample the low-level feature l to obtain i-1 l of the same dimension i-1 ; Input the high-level semantic feature h into the i-th fully connected layer to obtain a set of affine transformation parameters α i-1 and β i-1 ; Step 4.7: Along the depth direction of the reconstruction network, the feature map x i-1 and low-level features l i-1 Perform feature fusion concat() operation to obtain fusion feature map [l i-1 ,x i-1 ], and then fuse the feature map [l i-1 ,x i-1 ] and the affine transformation parameter α i-1 and β i-1 Feed the i-th progressive block G i In the output feature map x i , the output feature map x i and the previous progressive block G i-1 The fusion feature map x (i-1)out Perform feature fusion add() operation to obtain the fusion feature map x iout ; Step 4.8: Let i increase by 1, repeat steps 4.6 to 4.7, and obtain the feature map x after fusion n-1 times nout , where n = log2N-1; Step 5: Send the low-level features l extracted in step 2 to the low-level feature flow module to extract the features x′ used to reconstruct the image out ; Step 6: Transform feature x nout and x′ out Send to the reconstruction module and output the reconstructed image 2. The dual-stream image reconstruction system based on hybrid semantics according to claim 1, characterized in that: The edge detection algorithm in the low-level feature extraction module uses the PiDiNet network; The variational autoencoder in the high-level semantic feature extraction module is constructed using the ResNet-152 network; The low-level feature flow module is constructed using deformable convolution blocks; The hybrid feature flow module includes deformable convolution, fully connected layer and progressive block G i , the convolution kernel size of the deformable convolution is 3*3; the progressive block G i Built using SE-ResNet blocks; The reconstruction module includes a feature fusion concat() operation, a 3*3 convolution and a To RGB module. The To RGB module includes a 3*3 two-dimensional convolution and a spectral norm regularization operation.
3. The dual-stream image reconstruction system based on hybrid semantics according to claim 2, characterized in that: The progressive block G i The backbone network is the Bottleneck residual block, which is composed of three convolutional layers: 1*1 convolution, 3*3 convolution, and 1*1 convolution, and is activated using the LeakyRelu() activation function.
4. The dual-stream image reconstruction system based on hybrid semantics according to claim 1, characterized in that: The step 3 comprises: Step 3.1: Input the image to be reconstructed x into the ResNet-152 backbone network to extract features; Step 3.2: Pass the features obtained in step 3.1 through a 256-dimensional fully connected layer and then through an n z The fully connected layer of dimension 1 is used to obtain the mean μ and logarithmic variance log of the high-level semantic feature h. e σ 2 ; Step 3.3: Substitute the mean μ and logarithmic variance log e σ 2 The high-level semantic feature h is obtained using the reparameterization method.
5. The dual-stream image reconstruction system based on hybrid semantics according to claim 1, characterized in that: The step 6 comprises: Step 6.1: Convert the feature map x nout and x′ out Perform feature fusion concat() operation to obtain feature concat(x nout ,x′ out ); Step 6.2: Concat(x nout ,x′ out ) is fed into a 3*3 convolution and activated using LeakyRelu(); Step 6.3: Send the features output from step 6.2 to the To RGB module to obtain the reconstructed image
Citation Information
Patent Citations
Super-resolution multi-scale residual fusion model of single image and restoration method thereof
CN111861961A