Real world image super-resolution reconstruction method and system based on hybrid experts
By adopting a combination method of hybrid expert network and adversarial generation network in image super-resolution reconstruction, the problem of poor super-resolution reconstruction of real-world images is solved, achieving a more efficient and stable image reconstruction effect.
Patent Information
- Application Number
- CN202510172291.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-17
- Publication Date
- 2025-06-10
AI Technical Summary
The prior art has poor super-resolution reconstruction when processing real-world images, especially when facing diverse images and different scene contents, the reconstruction effect is unstable.
Using a real-world image super-resolution reconstruction method based on hybrid experts, a hybrid expert generator network based on convolutional neural network is constructed and an adversarial generation discriminator network is used to optimize model parameters using joint loss function to generate high-quality super-resolution images.
It improves the reconstruction efficiency and quality of super-score technology on real-world images, adapts to different scene contents, enhances the generalization ability of the network, and the generated images have better visual perception.
Smart Images

Figure CN120125439A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer bottom-layer vision image super-resolution reconstruction. Specifically, it relates to a method for real-world image super-resolution reconstruction based on a mixture of experts. Background Art
[0002] In real life, due to the hardware limitations of optical devices and the diffraction effect of light, observers usually obtain low-resolution (LR) images. Low-resolution images limit the accuracy of information transmission to a certain extent. This problem can be solved by improving the image quality. The ways to improve the image quality mainly include two aspects: hardware and software. On the one hand, one can try to break the hardware limitations of the device, but this method is costly and has very limited ability to improve image quality. On the other hand, from a software perspective, image super-resolution (SR) reconstruction technology can be used to improve image quality. This method has a lower cost and better effects. SR aims to overcome or effectively compensate for problems such as digital image blurring and low quality caused by the limitations of digital image acquisition and processing systems or the acquisition and processing environment itself. SR has extensive technical applications in actual scenarios such as medical imaging, remote sensing imaging, video imaging monitoring and security, and image compression.
[0003] Currently, the deep learning-based method is the mainstream method for image SR reconstruction technology. Among them, the methods based on convolutional neural networks, residual connection networks, and generative adversarial networks have achieved good results. The methods based on convolutional neural networks and residual connection networks can build large-scale deep neural networks to efficiently extract high-frequency detail features of images, and can perform well in pixel-level image reconstruction. However, due to the single network training mode and generally using low-resolution pictures with fixed degradation kernels for training, which is unreasonable in the real world, it will lead to poor reconstruction effects when the network faces images with various feature differentiations such as edges, textures, colors, and structures, as well as images with multi-scene contents. The method based on generative adversarial networks can compare with real images, making the SR-reconstructed images have better visual perception effects, but it also faces the same problems when dealing with multi-type multi-scene images and real-world images with unknown degradation kernels. Summary of the Invention
[0004] The present invention aims to provide a method for real-world image super-resolution reconstruction based on a mixture of experts, to solve the problem of poor super-resolution reconstruction effects of existing technologies for diverse real-world images, and to improve the reconstruction efficiency of the SR reconstruction model.
[0005] Another object of the present invention is to provide a real-world image super-resolution reconstruction system based on a mixture of experts.
[0006] To achieve the above object, on the one hand, the present invention provides a real-world image super-resolution reconstruction method based on mixture of experts, including:
[0007] S100. Obtain a real-world image data set;
[0008] S200. Estimate the degradation kernel and noise blocks of the images from the data set, add them to the degradation pool of the data constructor, construct low-resolution images based on the degradation pool and the obtained data set, and construct data sample pairs with the original images or the images obtained by cleaning the noise from the original images;
[0009] S300. Construct a mixture of experts generator network Generator and an adversarial generation discriminator network Discriminator based on a convolutional neural network to obtain a super-resolution reconstruction network model;
[0010] S400. Input the constructed sample data pairs into the generator network and the discriminator network respectively for training, optimize the model parameters with a joint loss function, and obtain a real-world image super-resolution reconstruction model based on mixture of experts;
[0011] S500. Input real-world images into the real-world image super-resolution reconstruction model based on mixture of experts, and output corresponding super-resolution images.
[0012] A further preferred technical solution of the present invention is that in step S200, the method for estimating the degradation kernel and noise blocks of the images from the data set and adding them to the degradation pool of the data constructor is as follows:
[0013] S210. Estimate the degradation kernel from the original images of the data set based on the kernel estimation method. The conditions that the estimated degradation kernel needs to satisfy are:
[0014]
[0015] Wherein, (I src *k)↓ s is the low-resolution image obtained by downsampling the original image I src using the degradation kernel k. ↓ s represents the downsampling process, and I src ↓ s is the low-resolution image obtained by bicubic interpolation downsampling using the ideal degradation kernel function; the second term of this formula is to limit the sum of k to 1, and k i,j represents the pixel value of the i-th row and the j-th column of the blur kernel k; the third term is the penalty boundary of k, and m i,j is the penalty constraint optimization method for k i,jThe penalty factor, which belongs to the hyperparameter value, is used to limit the value of the blur kernel at the boundary and avoid excessive values or discontinuous changes of the blur kernel in the boundary region. In the fourth term, L(·) is the contrast loss between the low-resolution image and the original image.
[0016] S220. Collect the noise block n from the original images in the dataset based on the method of noise injection.
[0017] S230. Add the estimated degradation kernels {k 1 , k 2 , …, k l} and the collected noise blocks {n 1 , n 2 , …, n +} into the degradation pool to expand the degradation pool.
[0018] Preferably, in step S200, the low-resolution image is constructed based on the degradation pool and the obtained dataset, and a data sample pair is constructed with the original image or the image obtained by cleaning the noise from the original image. The specific method is as follows:
[0019] S240. Based on the estimated degradation kernels {k 1 , k 2 , …, k l} and the collected noise blocks {n 1 , n 2 , …, n +}, the low-resolution image is obtained by the following formula:
[0020] I LR = (I HR * k i )↓ s + n j
[0021] where I LR represents the constructed low-resolution image, I HR represents the original image or the image obtained by cleaning the noise from the original image, is the high-resolution image, the degradation kernel k i and the noise block n j are randomly selected from the degradation pool {k 1 , k 2 , …, k l}, {n 1 , n 2 , …, n +};
[0022] S250. Combine the constructed low-resolution image and the high-resolution image into a corresponding sample data pair.
[0023] Preferably, the image obtained by cleaning the noise from the original image is obtained by performing source domain noise cleaning on the original image, which is expressed as:
[0024] I HR =(I src *k)↓ c
[0025] where k is the ideal bicubic sampling kernel and c is the sampling multiple.
[0026] Preferably, in step S300, the generator network Generator is composed of at least one convolutional neural network layer plus an activation function, a gating network layer, multiple mixture-of-experts networks, a weight fusion network layer, two sub-pixel convolutional layers, and a reconstruction convolutional layer;
[0027] The gating network layer uses a learnable fully connected layer to calculate the weights assigned to each mixture-of-experts network. Each mixture-of-experts network is a deep convolutional neural network using residual connections. Multiple mixture-of-experts networks are trained separately to extract different types of features of images in different scenarios. The weight fusion network layer fuses the features extracted by multiple mixture-of-experts networks according to the weights assigned by the gating network layer. Sub-pixel convolution uses feature recombination to improve the resolution of the output image. The reconstruction convolutional layer outputs the features as RGB image data with three channels;
[0028] In step S300, the adversarial generation discriminator network Discriminator is composed of an image feature extraction network and a discriminant block. The image feature extraction network is composed of multiple hierarchical convolutional network layers. Each hierarchical convolutional network layer consists of a convolutional layer, a BN layer, and a relu activation function; the discriminant block consists of a dense fully connected layer with 1024 dimensions, a relu activation function, a fully connected layer with dimension 1, and a sigmoid activation function.
[0029] Preferably, when the constructed sample data pairs are respectively input into the generator network and the discriminator network for training in step S400, the processing method for the sample data pairs is specifically as follows:
[0030] The low-resolution image is input into the generator network Generator. Shallow image features are extracted through a convolutional neural network layer and then input into the gating network layer for weight distribution calculation based on the extracted shallow features. The calculation method of the gating network layer adopts the top-k strategy, routing the shallow image features to the top k expert networks with the highest scores and performing normalized weight distribution according to the scores. Each mixture expert network performs deep and different types of feature extraction, outputs the obtained deep feature values to the weight fusion network layer. The weight fusion network layer fuses the features extracted by each mixture expert network according to the weights calculated by the gating network layer and outputs them to sub-pixel convolution for pixel magnification, and finally outputs to the reconstruction convolution layer for image reconstruction;
[0031] The low-resolution image and the high-resolution image are input into the image feature extraction network. Multiple hierarchical convolutional network layers extract the details of the image in a way that the number of convolutional kernels doubles, respectively obtaining feature maps with 512 channels. The feature maps are simultaneously input into the discriminant block. The dense fully connected layer of the discriminant block receives the feature maps of the low-resolution image and the high-resolution image. The fully connected layer adds a sigmoid function to generate a discrimination probability and obtains a discrimination score. According to the feedback of the discrimination score result, the parameters of the generator network are updated.
[0032] Preferably, when optimizing the model parameters with the joint loss function in step S400, the joint loss function includes a pixel loss function, a perceptual loss function, and an adversarial loss function, expressed as:
[0033]
[0034] Among them, is the pixel loss function, L 6 is the perceptual loss function, λ 7 is the adversarial loss function, λ 1 , λ 6 , λ 7 are balance factor hyperparameters.
[0035] On the other hand, the present invention provides a real-world image super-resolution reconstruction system based on mixture experts, including a data acquisition module for obtaining a real-world image dataset;
[0036] A data processing module for estimating the degradation kernel and noise block of the image from the dataset, adding them to the degradation pool of the data constructor, constructing a low-resolution image based on the degradation pool and the obtained dataset, and constructing a data sample pair with the original image or the image obtained by cleaning the noise from the original image;
[0037] A model construction module, configured to construct a hybrid expert generator network Generator and an adversarial generation discriminator network Discriminator based on a convolutional neural network, so as to obtain a super-resolution reconstruction network model;
[0038] A model training module, configured to separately input the constructed sample data pairs into the generator network and the discriminator network for training, optimize the model parameters with a joint loss function, and obtain a real-world image super-resolution reconstruction model based on hybrid experts;
[0039] A super-resolution image acquisition module, configured to input a real-world image into the real-world image super-resolution reconstruction model based on hybrid experts, and output a corresponding super-resolution image.
[0040] On the other hand, the present invention provides a non-transitory computer-readable storage medium, on which computer instructions are stored, and the computer instructions cause a computer to execute the above-mentioned real-world image super-resolution reconstruction method based on hybrid experts.
[0041] On another aspect, the present invention provides an electronic device, including: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The processor calls the logic instructions in the memory to execute the above-mentioned real-world image super-resolution reconstruction method based on hybrid experts.
[0042] On still another aspect, the present invention provides a computer program product, the computer program product includes a computer program, the computer program is stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer executes the above-mentioned real-world image super-resolution reconstruction method based on hybrid experts.
[0043] Beneficial effects: The present invention provides a real-world image super-resolution reconstruction method and system based on hybrid experts, which solves the problems that the current super-resolution technology has poor super-resolution effects for real-world images, and the super-resolution effects are unstable for images with different scene contents.
[0044] The constructed model can use multiple hybrid expert networks to perform separate training on images with different scene contents, disassemble sub-tasks for different feature extractions such as edges, textures, colors, and structures, and perform feature fusion in combination with the weight allocation mechanism of a learnable gating network, so that the reconstruction network can adapt to the scene content and increase the network generalization ability.
[0045] The network model structure based on generative adversarial can perform adversarial training based on high-resolution images with noise cleaned, enhancing the ability of the reconstruction network to generate images with better visual sensory effects. The construction of low-resolution image data based on degradation kernel estimation and noise injection enables the reconstruction network to train with real-world image datasets instead of based on a fixed pair of degradation kernels, and also enhances the reconstruction ability of the reconstruction model for detailed information of images belonging to different domains. Brief Description of the Drawings
[0046] Figure 1 It is a flowchart of the method for real-world image super-resolution reconstruction based on mixture of experts of the present invention;
[0047] Figure 2 It is a structural diagram of the Generator of the mixture of experts generator network based on convolutional neural network of the present invention;
[0048] Figure 3 It is a structural diagram of the mixture of experts network of the present invention;
[0049] Figure 4 It is a structural diagram of the adversarial generation discriminator network Discriminator of the present invention. Detailed Description of the Invention
[0050] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention, and they should not be construed as limitations to the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present invention without creative efforts shall fall within the scope of protection of the present invention. In the description of the present invention, it should be understood that the terms used are only for the purpose of description and cannot be construed as indicating or implying relative importance.
[0051] The following combines Figures 1 - 4 to describe the method and system for real-world image super-resolution reconstruction based on mixture of experts provided by the present invention.
[0052] Embodiment 1: This embodiment provides a method for real-world image super-resolution reconstruction based on mixture of experts, as Figure 1 shown, including:
[0053] S100. Obtain a real-world image dataset;
[0054] S200. Estimate the degradation kernel and noise blocks of the image from the dataset, add them to the degradation pool of the data constructor, construct a low-resolution image based on this degradation pool and the obtained dataset, and construct a data sample pair with the original image or the image obtained by cleaning the noise from the original image;
[0055] S300. Construct a hybrid expert generator network Generator and an adversarial generation discriminator network Discriminator based on a convolutional neural network to obtain a super-resolution reconstruction network model;
[0056] S400. Input the constructed sample data pairs into the generator network and the discriminator network respectively for training, optimize the model parameters with the joint loss function, and obtain a real-world image super-resolution reconstruction model based on hybrid experts;
[0057] S500. Input a real-world image into the real-world image super-resolution reconstruction model based on hybrid experts, and output the corresponding super-resolution image.
[0058] The above steps are described in detail.
[0059] Step S100 obtains a real-world image dataset.
[0060] In this step, the data in the dataset can be sourced from any real image data. Of course, it is preferably high-resolution images obtained based on hardware acquisition devices, and the image data can be directly used without downsampling.
[0061] In step S200, first use a method based on kernel estimation and noise injection to expand the degradation pool from the image data. The estimated degradation kernel needs to satisfy equation (1):
[0062]
[0063] Among them, the first term in this equation, (I src *k)↓ s is the low-resolution image obtained by downsampling the original image I src using the degradation kernel k. ↓ s represents the downsampling process. I src ↓ s is the low-resolution image obtained by bicubic interpolation downsampling using an ideal degradation kernel function. This term motivates the downsampled image to retain as much low-frequency information of the source image as possible by minimizing the error between the two. The second term in this equation restricts the sum of k to 1, and k i,j represents the pixel value at the i-th row and j-th column of the blur kernel k. The third term is the penalty boundary of k, and m i,j is for k in the penalty constraint optimization method i,jThe penalty factor, which belongs to the hyperparameter value, is used to limit the value of the blur kernel at the boundary and avoid excessive values or discontinuous changes of the blur kernel in the boundary region. In the fourth term, L(·) is the contrast loss between the low-resolution image and the original image, and this term is to ensure the source domain consistency of the degradation kernel. The obtained degradation kernel k from solving equation (1) for the input image data is added to the degradation pool.
[0064] Noise injection adopts the method of explicit injection. Since the high-frequency information of the image is lost during the downsampling process, the noise distribution of the degraded image will also change accordingly. To make the downsampled image have a similar noise distribution to the original image, the noise block n is directly collected from the original image data. The estimated degradation kernels {k 1 , k 2 , …, k l} and the collected noises {n 1 , n 2 , …, n +} are added to the degradation pool.
[0065] To obtain more high-resolution images, source domain noise cleaning can be performed on the original image. This process is optional, and the specific implementation is as shown in equation (2):
[0066] I HR = (I src * k)↓ c #(2)
[0067] where k is the ideal bicubic sampling kernel and c is the sampling multiple.
[0068] Based on the degradation kernel estimation and noise injection, a construction method for the LR image is obtained, as shown in equation (3):
[0069] I LR = (I HR * k i )↓ s + n j #(3)
[0070] where the degradation kernel k i and the noise block n j are randomly selected from the degradation pools {k 1 , k 2 , …, k l}, {n 1 , n 2 , …, n +}. The constructed LR images and HR images form sample data pairs.
[0071] Step S300 inputs the constructed sample data pairs into the hybrid expert generator network Generator and the adversarial generation discriminator network Discriminator based on the convolutional neural network respectively, and trains them in combination with the joint loss function, and finally obtains the hybrid expert real-world image super-resolution reconstruction network model. The execution process of step S300 and the structure of the generator network are as Figure 2 shown.
[0072] In the generator network Generator, first is a 9*9*64 convolutional neural network plus a relu activation function for extracting the shallow features of the image. It should be noted that more convolutional layers or residual connection layers can be added to this shallow feature extraction network module to improve the learning effect of the weight distribution of the next gating network. The gating network (GatingNetwork) consists of a 1*1 convolutional layer, a dense fully connected layer plus a softmax activation function. The convolutional layer is used to reduce the dimension of the extracted shallow features, and the fully connected layer is used to learn the weight calculation for allocating expert networks to different images. The softmax function is designed as shown in Equation (4):
[0073]
[0074] where n is the number of expert networks participating in the weight calculation. This function will calculate n probability values between (0,1) and the sum is 1. In this embodiment, the top-k strategy is adopted to select experts, that is, the top k maximum values are selected from the n probability values. The expert networks corresponding to the numbers of these k values are the selected expert networks. Then, these k values are normalized to obtain k weight values between (0,1) and the sum is 1. For the other n-k expert networks that are not assigned weights, their weight values are set to 0 and will not participate in the deep feature calculation. Then, the weight values of all expert networks are {w 1 ,w 2 ,…,w k ,0,…,0}. After the shallow features are subjected to deep and category-specific feature extraction by the selected expert networks, they will be input into the weights fusion network (Weights Network). This network will fuse the deep features based on the weight values calculated by the gating network, as shown in Equation (5):
[0075]
[0076] where m is the fused feature, m j is the deep feature extracted by the expert network, and the deep feature is passed to the subsequent resolution magnification module.
[0077] The resolution magnification module consists of a convolutional layer, a sub-pixel convolutional layer, and a relu activation function. The deep features of the LR image extracted first pass through a convolutional layer of 3*3*256, and then are input into two consecutive sub-pixel convolutional layers (PixelShuffler). The sub-pixel convolution calculation is as shown in Equation (6):
[0078] I MR =F(W L *f L (m)+b L )#(6)
[0079] where W L , b L are a learnable network weight and bias respectively, m is the image feature output by the previous convolution operation, and f L is to magnify the number of channels of the image feature. Specifically, r is the image magnification factor, and f L will expand the number of channels of m to C×r 2 , F is a periodic pixel reorganization operation that reshapes a vector of H×W×C×r 2 into a vector of rH×rW×C. Finally, after passing through a relu activation function and a convolutional layer of 9*9*3, a three-channel RGB super-resolution image with a specified magnification factor is output.
[0080] Regarding different resolution magnification methods, generally, the magnification factor r is set to 2 in each sub-pixel convolutional layer, and the image resolution can be magnified by 2 times. If a larger magnification factor such as ×4 or ×8 is required, it can be achieved by stacking sub-pixel convolutional layers. As Figure 2 shown, two reconstruction layers are stacked in the generator network structure, indicating that this structure is used for ×4 super-resolution reconstruction.
[0081] The construction of the mixture of experts network is as Figure 3, the expert network is composed of a deep convolutional neural network layer based on residual connections and a feature fusion layer. The VGG network is used in the network. The characteristic of the VGG network is to replace large-scale convolutional kernels by stacking multiple 3*3 convolutional kernels, which can reduce the parameters of the convolutional layer without reducing the receptive field of the convolution. All convolutional layers use the same number of channels, 64. After the convolutional layer, there are a Batch Normalization (BN) layer and a relu activation function, which are used to reduce the training difficulty of the network and accelerate the training speed of the network. A residual connection is established between the output of each convolutional layer and the output of the previous layer using a feature fusion layer, which can solve the training difficulty and overfitting problem of the deep convolutional network. It should be noted that the core of the method based on mixture of experts is the task division strategy, that is, the task of image feature extraction is decomposed into more targeted subtasks, such as edge detection, texture recognition, color feature extraction, and image structure feature extraction. The expert network structure described in this embodiment is a network structure with better feature extraction effect in current computer vision, but it is not limited to this one structure. The most ideal situation is to use different pre-trained network models, including but not limited to convolutional neural networks, networks based on attention mechanisms, and GAN networks, etc.
[0082] In step S300, the structure of the discriminator network Discriminator is as Figure 4 , the image feature extraction network of the discriminator network is composed of a VGG network. There is a BN layer and a relu activation function after each convolutional layer. The number of channels of each convolutional layer is multiplied by 2 successively from 64. After each convolution, a convolutional layer with a kernel size and number of channels unchanged and a stride of 2 is used to reduce the resolution of the image and accelerate the network training. The input LR image and HR image pass through the VGG network to obtain feature maps with 512 channels respectively. The feature maps are input into the discriminant block at the same time. The discriminant block consists of a dense fully connected layer with 1024 dimensions, a relu activation function, a fully connected layer with 1 dimension plus a sigmoid activation function. The 1024-dimensional fully connected layer is used to receive the feature maps of the LR image and the HR image, and the 1-dimensional fully connected layer plus the sigmoid function is used to generate the discrimination probability.
[0083] In step S400, the constructed sample data pairs are input into the generator network and the discriminator network for training respectively, and the model parameters are optimized in combination with the joint loss function. The joint loss function used is as shown in Equation (7):
[0084]
[0085] Among them, is the pixel loss, L 6is the perceptual loss, λ 7 is the adversarial loss, λ 1 , λ 6 , λ 7 is the balance factor hyperparameter. Based on the experience of existing research, in this embodiment, the balance factor is initialized as: λ 1 = 0.01, λ 6 = 1, λ 7 = 0.005.
[0086] Furthermore, L s+2234L5 is the pixel loss. Generally, pixel loss is divided into L1 loss and L2 loss. L1 and L2 are levels of mathematical norms. L1 loss is also called mean absolute error (MAE) loss, and its calculation method is as shown in Equation (8):
[0087]
[0088] where G(·) is the pixel value of the SR image output by the generator network, rW and rH are the width and height of the generated SR image respectively, and I \,] is the pixel point coordinate of the image.
[0089] L2 loss is also called mean squared error (MSE), and its calculation method is as shown in Equation (9):
[0090]
[0091] L1 loss and L2 loss have complementary advantages and disadvantages. The L1 loss function has a stable gradient, and the training process will not cause gradient explosion, with strong robustness. However, it is not differentiable at 0, and for points with small losses, relatively large loss values may be obtained; the L2 function is continuous and smooth, differentiable everywhere, convenient for derivative calculation, and the error value decreases as the gradient decreases accordingly, which is beneficial to function convergence. However, when the error between the predicted value and the actual value is greater than 1, the error will be amplified, and gradient explosion may occur when solving the gradient by gradient descent. Therefore, this embodiment adopts the smoothL1 loss that combines the advantages and disadvantages of the two loss functions, and its calculation method is as shown in Equation (10):
[0092]
[0093] Furthermore, L 6 is the perceptual loss. The perceptual loss calculates the difference between two images through a pre-trained neural network. Usually, a pre-trained convolutional neural network is used to obtain their eigenvalue in the network, and then the Euclidean distance between them is calculated. In the present invention, the pre-trained VGG-19 network is used to obtain the perceptual loss between the HR image and the generated SR image, and its calculation method is as shown in Equation (11):
[0094]
[0095] Among them, φ i,j (·) represents the image eigenvalue calculated by the j-th convolutional layer before the i-th pooling layer in the given VGG network, and W i,j and H i,j represent the image feature dimensions generated by the VGG network.
[0096] Furthermore, λ 7 is the adversarial loss, which is a loss function defined based on adversarial training of GAN. By deceiving the discriminator network, it motivates the generator network to generate more natural images and improves the visual effect of the images generated by the generator network. The calculation method is as shown in Equation (12):
[0097]
[0098] Among them, P n (G(I LR ), D(I HR )) is the probability that the SR image generated by the generator network output by the discriminator network is a natural HR image. By training to reduce the adversarial loss, the generator network can generate more realistic SR images in terms of visual effect.
[0099] In step S500, the real-world image is input into the trained generator Generator, and the corresponding super-resolution image is output. The image reconstruction process is essentially as shown in Equation (13):
[0100] I MR = G r (I LR ) #(13)
[0101] Among them, I MR represents the generated super-resolution image, I LR represents the LR image constructed by estimating the degradation kernel and noise block of the original image and adding them to the degradation pool of the data constructor. G r (·) represents the generator network after training and optimization.
[0102] Example 2: This example provides a real-world image super-resolution reconstruction system based on a mixture of experts, including
[0103] A data acquisition module for obtaining a real-world image dataset;
[0104] A data processing module for estimating the image degradation kernel and noise block from the dataset, adding them to the degradation pool of the data constructor, constructing a low-resolution image based on the degradation pool and the obtained dataset, and constructing a data sample pair with the original image or the image obtained by cleaning the noise from the original image;
[0105] A model construction module, configured to construct a hybrid expert generator network Generator and an adversarial generation discriminator network Discriminator based on a convolutional neural network, so as to obtain a super-resolution reconstruction network model;
[0106] A model training module, configured to respectively input the constructed sample data pairs into the generator network and the discriminator network for training, optimize the model parameters with a joint loss function, and obtain a real-world image super-resolution reconstruction model based on hybrid experts;
[0107] A super-resolution image acquisition module, configured to input a real-world image into the real-world image super-resolution reconstruction model based on hybrid experts, and output a corresponding super-resolution image.
[0108] Embodiment 3: This embodiment provides a non-transitory computer-readable storage medium, on which computer instructions are stored, and the computer instructions cause the computer to execute a real-world image super-resolution reconstruction method based on hybrid experts. The method includes the following steps:
[0109] S100. Obtain a real-world image dataset;
[0110] S200. Estimate the degradation kernel and noise block of the image from the dataset, add them to the degradation pool of the data constructor, construct a low-resolution image based on the degradation pool and the obtained dataset, and construct a data sample pair with the original image or the image obtained by cleaning the noise from the original image;
[0111] S300. Construct a hybrid expert generator network Generator and an adversarial generation discriminator network Discriminator based on a convolutional neural network, so as to obtain a super-resolution reconstruction network model;
[0112] S400. Respectively input the constructed sample data pairs into the generator network and the discriminator network for training, optimize the model parameters with a joint loss function, and obtain a real-world image super-resolution reconstruction model based on hybrid experts;
[0113] S500. Input a real-world image into the real-world image super-resolution reconstruction model based on hybrid experts, and output a corresponding super-resolution image.
[0114] Embodiment 4: This embodiment provides an electronic device, which may include: a processor, a communications interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus. The processor can call the logical instructions in the memory to execute a real-world image super-resolution reconstruction method based on a mixture of experts. The method includes the following steps:
[0115] S100. Obtain a real-world image dataset;
[0116] S200. Estimate the degradation kernel and noise blocks of the image from the dataset, add them to the degradation pool of the data constructor, construct a low-resolution image based on the degradation pool and the obtained dataset, and construct a data sample pair with the original image or the image obtained by cleaning the noise from the original image;
[0117] S300. Construct a mixture of experts generator network Generator and an adversarial generation discriminator network Discriminator based on a convolutional neural network to obtain a super-resolution reconstruction network model;
[0118] S400. Input the constructed sample data pairs into the generator network and the discriminator network for training respectively, optimize the model parameters with a joint loss function, and obtain a real-world image super-resolution reconstruction model based on a mixture of experts;
[0119] S500. Input the real-world image into the real-world image super-resolution reconstruction model based on a mixture of experts, and output the corresponding super-resolution image.
[0120] In addition, when the logical instructions in the above-mentioned memory are implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0121] Embodiment 5: The present embodiment provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a real-world image super-resolution reconstruction method based on mixture of experts. The method includes the following steps:
[0122] S100. Obtain a real-world image data set;
[0123] S200. Estimate the degradation kernel and noise blocks of the image from the data set, add them to the degradation pool of the data constructor, construct a low-resolution image based on the degradation pool and the obtained data set, and construct a data sample pair with the original image or the image obtained by cleaning the noise from the original image;
[0124] S300. Construct a mixture of experts generator network Generator and an adversarial generation discriminator network Discriminator based on a convolutional neural network to obtain a super-resolution reconstruction network model;
[0125] S400. Input the constructed sample data pairs into the generator network and the discriminator network respectively for training, optimize the model parameters with a joint loss function, and obtain a real-world image super-resolution reconstruction model based on mixture of experts;
[0126] S500. Input the real-world image into the real-world image super-resolution reconstruction model based on mixture of experts, and output the corresponding super-resolution image.
[0127] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0128] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0129] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A real-world image super-resolution reconstruction method based on mixed experts, characterized in that: include: S100, obtaining a real-world image dataset; S200, estimating the degradation kernel and noise block of the image from the data set, adding them to the degradation pool of the data constructor, constructing a low-resolution image based on the degradation pool and the acquired data set, and constructing a data sample pair with the original image or the image after the noise is cleaned from the original image; S300, constructing a hybrid expert generator network Generator based on a convolutional neural network and an adversarial generation discriminator network Discriminator to obtain a super-resolution reconstruction network model; S400, inputting the constructed sample data pairs into the generator network and the discriminator network for training respectively, optimizing the model parameters with a joint loss function, and obtaining a real-world image super-resolution reconstruction model based on mixed experts; S500 , inputting a real-world image into the real-world image super-resolution reconstruction model based on mixed experts, and outputting a corresponding super-resolution image.
2. The real-world image super-resolution reconstruction method based on mixed experts according to claim 1, characterized in that The degradation kernel and noise block of the image are estimated from the data set in step S200 and added to the degradation pool of the data constructor. The specific method is: S210, based on the kernel estimation method, the degradation kernel is estimated from the original image of the data set, and the conditions that the estimated degradation kernel needs to meet are: Among them, the first term of this formula (I src *k)↓ s is the original image I src The low-resolution image obtained by downsampling using the degenerate kernel k,↓ s represents the downsampling process, I src ↓ s is a low-resolution image obtained by bicubic interpolation downsampling using an ideal degradation kernel function; the second term of this formula is to limit the sum of k to 1, k i,j represents the pixel value of the i-th row and j-th column of blur kernel k; the third term is the penalty boundary of k, m i,j is the pair k in the penalty constrained optimization method i,j The penalty factor is a hyperparameter value. In the fourth term, L(·) is the contrast loss between the low-resolution image and the original image. S220, collecting noise blocks n from the original image of the data set based on a noise injection method; S230, the estimated degradation kernel {k1, k2, …, k l } and the collected noise blocks {n1,n2,…,n m }Add it to the degradation pool to expand the degradation pool.
3. The real-world image super-resolution reconstruction method based on mixed experts according to claim 2, characterized in that: The step S200 constructs a low-resolution image based on the degradation pool and the acquired data set, and constructs a data sample pair with the original image or the image after the noise is cleaned from the original image. The specific method is: S240, based on the estimated degradation kernel {k1, k2, …, k l } and the collected noise blocks {n1,n2,…,n m }, and use the following formula to get the low-resolution image: I LR =(I HR *k i )↓ s +n j Among them, I LR Represents the constructed low-resolution image, I HR represents the original image or the image after cleaning the noise from the original image, which is a high-resolution image, and the degradation kernel k i and noise block n j From the degradation pool {k1, k2, …, k l },{n1,n2,…,n m } randomly selected; S250: Combining the constructed low-resolution image and high-resolution image into corresponding sample data pairs.
4. The real-world image super-resolution reconstruction method based on mixed experts according to claim 3, characterized in that: The image after cleaning the noise of the original image is obtained by cleaning the source domain noise of the original image, which is expressed as: I HR =(I src *k)↓ c Among them, k is the ideal bicubic sampling kernel and c is the sampling multiple.
5. The real-world image super-resolution reconstruction method based on mixed experts according to claim 1, characterized in that: The generator network Generator in step S300 is composed of at least one convolutional neural network layer plus an activation function, a gated network layer, multiple hybrid expert networks, a weight fusion network layer, two sub-pixel convolutional layers and a reconstruction convolutional layer; The gated network layer is composed of a learnable fully connected layer to calculate the weights assigned to each hybrid expert network. Each hybrid expert network is a deep convolutional neural network using residual connections. Multiple hybrid expert networks are trained separately to extract different types of features from different scene images. The weight fusion network layer fuses the features extracted by multiple hybrid expert networks according to the weights assigned by the gated network layer. Sub-pixel convolution uses feature recombination to improve the resolution of the output image. The reconstruction convolution layer outputs the features as RGB image data of three channels. The adversarial generative discriminator network Discriminator described in step S300 is composed of an image feature extraction network and a discriminant block. The image feature extraction network is composed of multiple hierarchical convolutional network layers, each of which is composed of a convolutional layer, a BN layer and a relu activation function; the discriminant block is composed of a 1024-dimensional dense fully connected layer, a relu activation function, a dimension 1 fully connected layer and a sigmoid activation function.
6. The real-world image super-resolution reconstruction method based on mixed experts according to claim 5, characterized in that When the constructed sample data pairs are input into the generator network and the discriminator network for training in step S400, the method for processing the sample data pairs is specifically as follows: The low-resolution image is input into the generator network Generator, and the shallow image features are extracted through a layer of convolutional neural network. Then, the features are input into the gated network layer to calculate the weight distribution according to the extracted shallow features. The gated network layer calculation method adopts the top-k strategy to route the shallow image features to the top k expert networks with the highest scores, and normalize the weight distribution according to the scores. Each hybrid expert network extracts different types of deep features, and outputs the obtained deep feature values to the weight fusion network layer. The weight fusion network layer fuses the features extracted by each hybrid expert network according to the weights calculated by the gated network layer, outputs them to the sub-pixel convolution for pixel amplification, and finally outputs them to the reconstruction convolution layer for image reconstruction. The low-resolution image and the high-resolution image are input into the image feature extraction network. Multiple levels of convolutional network layers extract the details of the image by doubling the number of convolution kernels, and feature maps with 512 channels are obtained respectively. The feature maps are input into the discriminant block at the same time. The dense fully connected layer of the discriminant block receives the feature maps of the low-resolution image and the high-resolution image. The fully connected layer adds the sigmoid function to generate the discriminant probability and obtain the discriminant score. According to the feedback of the discriminant score result, the generator network parameters are updated.
7. The real-world image super-resolution reconstruction method based on mixed experts according to claim 1, characterized in that: When the model parameters are optimized by the joint loss function in step S400, the joint loss function includes a pixel loss function, a perceptual loss function and an adversarial loss function, which is expressed as: in, is the pixel loss function, L p is the perceptual loss function, λ a To combat the loss function, λ1,λ p ,λ a is the balance factor hyperparameter.
8. A real-world image super-resolution reconstruction system based on mixed experts, characterized in that include A data acquisition module, used to obtain real-world image datasets; A data processing module, used for estimating the degradation kernel and noise block of the image from the data set, adding the estimated image to the degradation pool of the data constructor, constructing a low-resolution image based on the degradation pool and the acquired data set, and constructing a data sample pair with the original image or the image after the noise is cleaned from the original image; The model building module is used to build a hybrid expert generator network Generator based on a convolutional neural network and an adversarial generation discriminator network Discriminator to obtain a super-resolution reconstruction network model; A model training module is used to input the constructed sample data pairs into the generator network and the discriminator network for training respectively, and optimize the model parameters by a joint loss function to obtain a real-world image super-resolution reconstruction model based on mixed experts; The super-resolution image acquisition module is used to input the real-world image into the real-world image super-resolution reconstruction model based on mixed experts and output the corresponding super-resolution image.
9. A non-transitory computer-readable storage medium, characterized in that: Computer instructions are stored thereon, and the computer instructions enable the computer to execute the real-world image super-resolution reconstruction method based on mixed experts as described in any one of claims 1-7.
10. An electronic device, characterized in that: include: A processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus, and the processor calls the logic instructions in the memory to execute the real-world image super-resolution reconstruction method based on mixed experts as described in any one of claims 1-7.
Citation Information
Cited By
Data-enhanced text image super-resolution reconstruction method and device
CN120807292A