Image super-resolution reconstruction method based on semantic joint perception

The semantic joint perception method in image super-resolution uses a conditional GAN with text input to improve image quality by aligning semantic features, addressing over-sharpening and text-driven deviations, achieving better visual and objective quality.

CN120318078AActive Publication Date: 2025-07-15SOUTHWEAT UNIV OF SCI & TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510809484.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-15
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

There are problems with oversharpening and artifacts in existing image super-resolution reconstruction techniques, and the problems that the results of text-driven super-resolution reconstruction models may be completely deviated from the original high-resolution images.

Method used

The image super-resolution reconstruction method is adopted with semantic joint perception, and the adversarial network is generated by building conditions, combining text information, and semantic alignment discrimination in the discriminator is introduced. Adversarial training is used by generators and discriminators, and the loss function is optimized to balance visual quality and objective quality.

Benefits of technology

The balance of visual quality and objective quality is achieved, the potential of text information is realized, the effect of image super-resolution reconstruction is improved, oversharpening and artifacts are reduced, and the natural sense of image and detail recovery is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318078A_ABST
    Figure CN120318078A_ABST
Patent Text Reader

Abstract

The invention discloses an image super-resolution reconstruction method based on semantic joint perception, and belongs to the technical field of image super-resolution reconstruction. The method comprises the following steps: constructing an original high-resolution image data set and a low-resolution image data set; constructing an image super-resolution model; calculating a loss function of the image super-resolution model; training an image super-resolution model by using the original high-resolution image data set and the low-resolution image data set; and inputting the low-resolution image and the text information corresponding to the low-resolution image into the trained image super-resolution model, and outputting an image super-resolution reconstruction result. According to the method, text information is introduced and a conditional generative adversarial network is used for carrying out a super-resolution reconstruction task, specifically, the text information is introduced into a discriminator to assist an image super-resolution reconstruction task, and an input layer is newly added into the discriminator to support condition input; the visual quality and the objective quality are balanced, and the potential of the text information is exerted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image super-resolution reconstruction, and in particular relates to a semantic joint perception image super-resolution reconstruction method. Background Art

[0002] Image super-resolution (ISR) is a technical means that aims to reconstruct low-resolution images that are degraded due to performance limitations of imaging devices or during transmission, processing, storage, etc. into high-resolution images with rich details and higher clarity, making their quality as close as possible to real high-resolution images.

[0003] The image super-resolution reconstruction task is essentially an ill-posed (ill-posed) problem. Specifically, given a low-resolution image block, due to the loss of some information during the degradation process, the low-resolution image may correspond to multiple different high-resolution image blocks. From a mathematical perspective, the low-resolution image can be and high-resolution images The relationship is simplified to: ,in is the downsampling operator. The pathological nature of this relationship is that, given a low-resolution image Solving high-resolution images , this solution is not unique, and the tiny low-resolution image Changes in (such as noise interference) may result in the restored high-resolution image There are great changes. In order to solve this pathological problem, it is usually necessary to introduce prior knowledge. The image super-resolution method based on deep learning enables the network to learn the intrinsic prior information of the image through a large amount of training data. In this way, when a low-resolution image is given, the network can predict a reasonable high-resolution image that conforms to the statistical characteristics of natural images as much as possible based on these prior knowledge, rather than randomly selecting from many possible solutions. For example, the super-resolution generative adversarial network (SRGAN) uses the adversarial training mechanism of the generator and the discriminator to allow the generator to learn how to generate more realistic and more consistent with the distribution of natural images. High-resolution images. This adversarial training process not only improves the visual quality of the image, but also effectively alleviates the pathological nature of the super-resolution task, making the reconstructed high-resolution image closer to the real image in terms of details and overall quality.

[0004] The main goal of image super-resolution reconstruction is to improve the visual and objective quality of images. On the one hand, visual effects focus on the subjective perception of the reconstructed image by the human eye, that is, whether the image looks natural, clear, and has realistic details; on the other hand, objective evaluation relies on quantitative metrics such as Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), and Structural Similarity (SSIM) to measure the quality difference between the reconstructed image and the true high-resolution image. However, in practical applications, there may be a contradiction between these two evaluation methods. For example, some reconstructed images perform excellently in objective metrics but are visually inferior to those with slightly lower objective metrics. This phenomenon is called the "higher but not necessarily better" resolution problem, indicating that relying solely on objective metrics cannot fully reflect the visual quality of images. Therefore, the image super-resolution reconstruction task needs to strike a balance between these two aspects. By designing a reasonable loss function to guide model training, it can optimize objective metrics while also taking into account visual effects. For example, SRGAN introduced perceptual loss to optimize image quality by mimicking the perceptual characteristics of the human visual system. ESRGAN further improved on this basis and tried to generate super-resolution results that are visually closer to real images. However, this improvement also brought some new problems, such as oversharpening and artifacts, which may affect the naturalness and visual quality of the image. To solve these problems, Real-ESRGAN extended ESRGAN and tried to reduce negative effects such as artifacts and oversharpening while improving visual effects. However, most existing models are based on traditional single-modal image super-resolution (SR) frameworks and are optimized by combining Generative Adversarial Networks (GANs), but these methods still cannot fully solve the above contradictions, indicating that there is still room for further exploration and improvement in the field of image super-resolution reconstruction.

[0005] The text-to-image (T2I) diffusion model has developed rapidly. It generates high-quality images based on text information and involves natural language processing and computer vision. seesr uses the T2I model for image super-resolution tasks and introduces image information as a condition into the T2I model, namely the Controlled T2I diffusion model. Using a pre-trained text-to-image (T2I) model to perform image super-resolution tasks is currently attracting increasing attention. However, fundamentally speaking, the output results of this model largely depend on the accuracy of the text input. It can be said that the correctness of the text information directly determines the correctness of the output results. Once the input text information is incorrect, it will have a catastrophic impact on the super-resolution reconstruction results of the target low-resolution image. For example, in Figure 1In this case, the model wrongly super-resolved "a person driving a boat" into "a bird". This fully demonstrates that when a text-driven image generation model is applied to the image super-resolution task, the result may deviate completely from the original high-resolution image. Summary of the Invention

[0006] An object of the present invention is to provide a semantic joint perception-based image super-resolution reconstruction method for the above-mentioned deficiencies in the prior art, so as to solve the problems of oversharpening and artifacts in existing image super-resolution reconstructions, and the problem that in the task model of text-driven super-resolution reconstruction image generation, the result may deviate completely from the original high-resolution image.

[0007] To achieve the above object, the technical solution adopted by the present invention is:

[0008] A semantic joint perception-based image super-resolution reconstruction method, which includes the following steps:

[0009] S1. Construct an original high-resolution image dataset;

[0010] S2. Downsample the original high-resolution image to construct a low-resolution image dataset;

[0011] S3. Construct an image super-resolution model based on a conditional generative adversarial network and text information;

[0012] S4. Calculate the loss function of the image super-resolution model;

[0013] S5. Use the original high-resolution image dataset and the low-resolution image dataset to train the image super-resolution model;

[0014] S6. Input the low-resolution image and its corresponding text information into the trained image super-resolution model, and output the image super-resolution reconstruction result.

[0015] Further, in S2, the image super-resolution model includes:

[0016] A generator, which inputs the low-resolution image into the generator and outputs a reconstructed high-resolution image;

[0017] A discriminator, which inputs the reconstructed high-resolution image and its corresponding text information, the original high-resolution image and its corresponding text information into the discriminator for semantic alignment discrimination, and transmits the semantic alignment discrimination result to the generator for adversarial training.

[0018] Further, inputting the low-resolution image into the generator and outputting a reconstructed high-resolution image specifically includes:

[0019] The input low-resolution image is subjected to multiple convolutional operations and dynamic residual processing to generate the core features of the image;

[0020] Meanwhile, the input low-resolution image is upsampled to the target resolution through bilinear interpolation to obtain a reference high-resolution image;

[0021]

[0022] The core features of the image and the reference high-resolution image are merged to generate a reconstructed high-resolution image:

[0023]

[0024] In the formula, is the reconstructed high-resolution image; is the core feature of the image; is the reference high-resolution image; represents bilinear interpolation; is the input low-resolution image.

[0025] Furthermore, the input low-resolution image is subjected to multiple convolutional operations and dynamic residual processing to generate the core features of the image, specifically including:

[0026] The low-resolution image is input into the first layer of convolution to project the low-resolution image dynamically into the feature space by dynamic convolution:

[0027]

[0028] In the formula, is the first image feature; is the weight; is the activation function; is the dynamic convolution operation;

[0029] The first image feature is input into the residual block for dynamic residual processing:

[0030]

[0031] In the formula, is the second image feature; represents dynamic residual processing;

[0032] The second image feature is input into the second layer of convolution for pixel rearrangement and dynamic convolution processing:

[0033]

[0034] In the formula, is the third image feature; represents pixel rearrangement;

[0035] Input the third image feature into the third layer of convolution for dynamic convolution processing and adjust the channel dimension:

[0036]

[0037] In the formula, is the fourth image feature;

[0038] Input the fourth image feature into the fourth layer of convolution to map the fourth image feature to the output channels:

[0039]

[0040] In the formula, is the core feature of the image.

[0041] Furthermore, input the reconstructed high-resolution image and its corresponding text information into the discriminator for semantic alignment discrimination, specifically including:

[0042] Project the semantic features of the text information corresponding to the reconstructed high-resolution image onto the target shape:

[0043]

[0044] In the formula, s is the semantic feature of the text information; , , are the semantic features of the text information introduced in different layers of convolution respectively; is to convert the input semantic features of the text information from its original representation into a multi-dimensional tensor suitable for network processing;

[0045] Input the reconstructed high-resolution image into the first-stage convolutional layer, perform feature transformation and combine it with the attention mechanism:

[0046]

[0047]

[0048] In the formula, is the fifth image feature; is the first concatenated feature; is the activation function; is the convolution operation; is the attention output;

[0049] Input the first concatenated feature into the second-stage convolutional layer, extract the image features and combine them with the attention mechanism:

[0050]

[0051]

[0052] In the formula, is the sixth image feature; is the second stitching feature;

[0053] Input the second stitching feature into the third-stage convolutional layer, extract the image feature and combine it with the attention mechanism:

[0054]

[0055]

[0056] In the formula, is the seventh image feature; is the third stitching feature;

[0057] Project the third stitching feature through the last convolutional layer to a single-channel output:

[0058]

[0059] In the formula, is the final output of the discriminator, representing the prediction result of the reconstructed high-resolution image input.

[0060] Furthermore, in the discriminator, the attention mechanism is introduced in the first-stage convolutional layer, the second-stage convolutional layer and the third-stage convolutional layer, and then the semantic features of the text information are embedded into the image features;

[0061] Among them, the attention mechanism in the first-stage convolutional layer is expressed as:

[0062] Normalize the input fifth image feature and the semantic features of the text information:

[0063]

[0064]

[0065] In the formula, is the normalized fifth image feature; is the semantic feature of the normalized text information in the first stage; represents the normalization process;

[0066] Use a 1×1 convolution to map the normalized fifth image feature and the semantic feature of the normalized text information in the first stage to the embedding dimension:

[0067]

[0068] In the formula, is the query matrix, is the key matrix, is the value matrix;

[0069] Calculate the attention weight matrix:

[0070]

[0071] where, is the attention weight matrix; is the dimension of the key; is the transpose of the matrix; is the normalization operation;

[0072] Weight the attention weight matrix to the value matrix to generate the attention output .

[0073] Furthermore, in S4, the loss function of the image super-resolution model includes:

[0074] The loss function of the generator is:

[0075]

[0076] where, is the pixel loss, is the perceptual loss, is the style loss, is the GAN loss;

[0077] The loss function of the discriminator is:

[0078]

[0079] where, is the loss of the original high-resolution image, is the loss of the reconstructed high-resolution image.

[0080] Furthermore, the perceptual loss is:

[0081]

[0082]

[0083] where, is the perceptual loss of the layer under the loss criterion, is the mean absolute error; is the number of elements of the feature map of the Value of and are the feature representations of the generated image and the real image at the layer respectively; is the original high-resolution image; is the reconstructed high-resolution image; is the weight corresponding to the layer, is the overall weight of the perceptual loss;

[0084] Style loss is:

[0085]

[0086]

[0087]

[0088] In the formula, is the Gram matrix calculated for the feature map of the layer; is the feature map of the layer; is the number of channels, and and are the height and width of the feature map respectively; represents the activation value of each channel in the deep feature map at each spatial position; is the style loss of the layer under the loss criterion; is the number of elements of the Gram matrix of the layer; and are the Gram matrices of the feature representations of the reconstructed high-resolution image and the original high-resolution image at the layer respectively;

[0089] GAN loss is:

[0090]

[0091] In the formula, is the final output of the discriminator, which is the discrimination result of the discriminator on the reconstructed high-resolution image.

[0092] Furthermore, the original high-resolution image loss is:

[0093]

[0094] Loss of the reconstructed high - resolution image is:

[0095]

[0096] In the formula, is the final output of the discriminator, is the discrimination result of the discriminator on the original high - resolution image, is the discrimination result of the discriminator on the reconstructed high - resolution image.

[0097] The semantic joint perception - based image super - resolution reconstruction method provided by the present invention has the following beneficial effects:

[0098] The present invention introduces text information and uses a conditional generative adversarial network for the super - resolution reconstruction task. Specifically, text information is introduced into the discriminator to assist the image super - resolution reconstruction task, and a new input layer is added to the discriminator to support conditional input; a balance is achieved between visual quality and objective quality, and the potential of text information is exerted. BRIEF DESCRIPTION OF THE DRAWINGS

[0099] Figure 1 is the reconstruction result of the SISR model using the T2I - based model in the background art, where Figure 1 in (a) is the result image of image super - resolution reconstruction using the SISR model based on the T2I model; Figure 1 in (b) is the original high - resolution image.

[0100] Figure 2 is the network result diagram of the generator in the embodiment of the present invention.

[0101] Figure 3 is the network result diagram of the discriminator in the embodiment of the present invention.

[0102] Figure 4 is the flow chart of the combination of convolution and attention mechanism in the embodiment of the present invention.

[0103] Figure 5 is the flow chart of the semantic joint perception - based image super - resolution reconstruction method in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0104] The specific embodiments of the present invention will be described below to facilitate those skilled in the art to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions made using the concept of the present invention are within the scope of protection.

[0105] Example 1

[0106] The semantic joint perception image super-resolution reconstruction method of this embodiment effectively utilizes text information and can well balance the image visual quality and objective quality. Refer to Figure 5 , and it specifically includes the following contents:

[0107] S1. Construct an original high-resolution image dataset;

[0108] In this embodiment, a large number of high-resolution (HR) images are collected. These images come from public datasets. The training datasets include CelebA-HQ, coco, DIV2K, Flickr2k; the test datasets include Set5, Set14, BSD100, Urban100. Among them, real low-resolution images are also collected to test the performance of the model.

[0109] S2. Downsample the original high-resolution images to construct a low-resolution image dataset;

[0110] In this embodiment, the original high-resolution images are specifically downsampled by bilinear interpolation to generate corresponding low-resolution images;

[0111] And the original high-resolution images and low-resolution images are cropped. The large images are cropped into small pieces for easy training, and the images are normalized to accelerate the training process.

[0112] S3. Construct an image super-resolution model based on a conditional generative adversarial network and text information;

[0113] The image super-resolution model of this embodiment includes:

[0114] A generator. The low-resolution image is input into this generator. After multiple convolutional operations and dynamic residual processing, a reconstructed high-resolution image is output;

[0115] Specifically, refer to Figure 2 , the generator is a convolutional neural network architecture. The input is a low-resolution image. After dynamic convolution, residual processing, upsampling, and interpolation upsampling and other operations, the output is a reconstructed high-resolution image. Its specific process is:

[0116] Input the low-resolution image into the first convolutional layer, and project the dynamic convolution of the low-resolution image into the feature space:

[0117]

[0118] Wherein, is the first image feature, which contains the preliminarily extracted low-level image features. It is the first feature representation of the network and is used for subsequent residual block processing; is the weight; is the activation function; is the dynamic convolution operation; is the input low-resolution image;

[0119] Input the first image feature into the residual block for dynamic residual processing, and this process is repeated multiple times:

[0120]

[0121] Wherein, is the second image feature, which contains richer high-level feature information. It is the result of stacking multiple residual blocks and is used for the main part of the network; represents the dynamic residual processing;

[0122] Next, perform upsampling, input the second image feature into the second convolutional layer, and perform pixel rearrangement and dynamic convolution processing:

[0123]

[0124] Wherein, is the third image feature, which is the high-resolution image feature after pixel rearrangement and dynamic convolution processing. It represents a higher-resolution feature map, which has been upsampled from a lower resolution to a higher resolution, but still needs further convolutional processing to enhance its feature information; is the pixel rearrangement;

[0125] Input the third image feature into the third convolutional layer, perform dynamic convolution processing and adjust the channel dimension. The LeakyReLU activation function is used for non-linear processing of the output of the dynamic convolution to enhance the feature representation ability:

[0126]

[0127] Wherein, is the fourth image feature, which contains a feature map with a higher resolution and rich features, providing input for subsequent final convolutional layers and output image generation;

[0128] Input the fourth image feature into the fourth convolutional layer to map the fourth image feature to the output channel:

[0129]

[0130] In the formula, is the core feature of the image.

[0131] Meanwhile, the input low-resolution image is upsampled to the target resolution through bilinear interpolation to obtain a reference high-resolution image;

[0132]

[0133] The core feature of the image and the reference high-resolution image are merged to generate a reconstructed high-resolution image:

[0134]

[0135] In the formula, is the reconstructed high-resolution image, which is the final output result of the generator network. This image contains both the high-frequency features generated by the network and the structural information of the low-resolution input; is the core feature of the image; is the reference high-resolution image; is the bilinear interpolation operation.

[0136] The discriminator takes the reconstructed high-resolution image and its corresponding text information, and the original high-resolution image and its corresponding text information as inputs, performs semantic alignment discrimination, and transmits the semantic alignment discrimination result to the generator for adversarial training;

[0137] Reference Figure 3 , the discriminator is a hybrid network architecture combining a convolutional neural network and a Transformer. Among them, the semantic feature extractor can effectively extract the semantic features of text information. After the image passes through the first layer of convolution, it passes through the LeakyReLU activation function, then passes through three layers of convolution in sequence, and the text information is fused with the image features through a modified spatial transformer (MST) after each layer. After fusion, it is further processed through a convolutional layer, and finally, the feature map is converted into the output of the original channels through the last layer of convolution as the final discrimination result of the model.

[0138] In the discriminator structure designed in this embodiment, the text information and the image information are fused by introducing an attention mechanism in three consecutive convolutional layers. The core of the design lies in feature extraction and semantic enhancement. By extracting low-level to high-level features layer by layer, its complexity is increased, and semantic context is dynamically introduced through the attention mechanism to improve the understanding of image details until the discriminator combines local details and global information to generate the final prediction result. The main input of the hybrid network architecture is the image information, with its shape being the number of input samples, width, height, and number of channels, and the text information is used as an auxiliary input, containing semantic information closely related to the image;

[0139] Since the processes experienced by the reconstructed high-resolution image and its corresponding text information, and the original high-resolution image and its corresponding text information are the same when used as inputs, this embodiment will use the reconstructed high-resolution image and its corresponding text information as inputs for illustration here:

[0140] Project the semantic features of the text information corresponding to the reconstructed high-resolution image onto the target shape:

[0141]

[0142] In the formula, s is the semantic feature of the text information; , , are respectively the semantic features of the text information introduced in different layers of convolution; The role of is to convert the semantic features of the input text information from its original representation into a multi-dimensional tensor that can adapt to network processing;

[0143] Input the reconstructed high-resolution image into the first-stage convolutional layer, and after feature transformation, combine it with the attention mechanism:

[0144]

[0145]

[0146] In the formula, is the fifth image feature; is the first concatenated feature; is the activation function; is the convolution operation; is the attention output;

[0147] Here, the LeakyReLU activation function can handle negative-valued features and avoid the problem of gradient disappearance; the semantic features of the text information and the fifth image feature Combine, add context information to the features through the attention mechanism. The attention mechanism can adaptively select useful semantic features, enhance the model's understanding of details, make the feature vectors more abundant, and then use convolution to compress the feature channels to enhance the non-linear fitting ability.

[0148] Input the first concatenated feature into the second-stage convolutional layer. After extracting the image features, combine them with the attention mechanism:

[0149]

[0150]

[0151] In the formula, is the sixth image feature; is the second concatenated feature;

[0152] The sixth image feature Introduce the second-stage semantic features to it , which can improve the modeling ability for complex semantic relationships. After embedding the semantic information, ensure the richness of the convolutional features. After convolution, maintain the channel consistency and increase the non-linear transformation at the same time.

[0153] Input the second concatenated feature into the third-stage convolutional layer. After extracting the image features, combine them with the attention mechanism:

[0154]

[0155]

[0156] In the formula, is the seventh image feature; is the third concatenated feature;

[0157] is the seventh image feature. Introduce the third-stage semantic features to it , establish the deepest context relationship, so that the model can better understand the overall image semantics.

[0158] Project the third concatenated feature through the last convolutional layer to a single-channel output:

[0159]

[0160] In the formula, is the final output of the discriminator, representing the prediction result of the reconstructed high-resolution image of the input , that is, whether the image input to the discriminator is the high-resolution image reconstructed by the generator or the original high-resolution image.

[0161] As a preference of this embodiment, in the discriminator, an attention mechanism is introduced in the first-stage convolutional layer, the second-stage convolutional layer, and the third-stage convolutional layer, so as to embed the semantic features of the text information into the image features;

[0162] Taking the first-stage convolution as an example, referring to Figure 4 , the attention mechanism in the first-stage convolutional layer is expressed as:

[0163] Normalize the semantic features of the input fifth image feature and the text information:

[0164]

[0165]

[0166] In the formula, is the normalized fifth image feature; is the semantic feature of the normalized text information in the first stage; represents the normalization process;

[0167] Use a 1×1 convolution to map the normalized fifth image feature and the semantic feature of the normalized text information in the first stage to the embedding dimension:

[0168]

[0169] In the formula, is the query matrix, is the key matrix, is the value matrix;

[0170] Calculate the attention weight matrix:

[0171]

[0172] In the formula, is the attention weight matrix; is the dimension of the key; is the transpose of the matrix; is ;

[0173] Weight the attention weight matrix to the value matrix , and generate the attention output

[0174] S4. Calculate the loss function of the image super-resolution model;

[0175] The loss function of the image super-resolution model in this embodiment includes:

[0176] The loss function of the generator is:

[0177]

[0178] In the formula, is the pixel loss, is the perceptual loss, is the style loss, is the GAN loss;

[0179] Among them, the pixel loss is used to measure the pixel-level difference between the generated image and the real image. The human eye can clearly perceive many subtle details on the image. For example, in a face image, what is her / his skin condition? Is the skin color uniform, are there freckles, age spots, etc., is the facial expression a smile, crying, or angry, and are the hair strands distinct?

[0180] The perceptual loss is to try to simulate the way the human visual system perceives images. Generally speaking, whether the colors are distinct, whether the edges are clear, whether the textures are detailed and real, etc. The perceptual loss is used to measure the perceptual difference between the generated image and the real image. Based on the pre-trained VGG network, features of different levels of the generated image and the real image are extracted, and the similarity between the generated image and the real image in perception is reflected by comparing the differences between these features.

[0181] The perceptual loss is calculated as follows:

[0182] First, calculate the perceptual loss of the layer under the loss criterion:

[0183]

[0184] The total perceptual loss is:

[0185]

[0186] In the formula, is the perceptual loss of the layer under the loss criterion, is the L1 loss, also known as the mean absolute error; is the number of elements in the feature map of the layer; i is the summation index, which is used to traverse the values from 1 to and are the feature representations of the generated image and the real image in the layer respectively; is the original high-resolution image; is the reconstructed high-resolution image; is the The weight corresponding to the layer is the overall weight of the perceptual loss; is represented as a specific layer (a certain layer) in the VGG network.

[0187] Style loss is calculated as follows:

[0188] The style loss is used to measure the style difference between the generated image and the real image, and calculates by comparing the Gram matrices of the generated image and the real image. The Gram matrix is used to capture the correlation between different channels in the feature map, and its calculation formula is as follows:

[0189]

[0190] Under the loss criterion, the style loss of the layer is:

[0191]

[0192] The total style loss is:

[0193]

[0194] In the formula, is the Gram matrix calculated for the feature map of the layer; is the feature map of the layer; is the number of channels, and and are the height and width of the feature map respectively; represents the activation value of each channel in the deep feature map at each spatial position, which is obtained by abstracting the input image through the convolutional network. It reflects the existence and intensity of local patterns (such as edges, textures, etc.) in the image, and is a mapping from the original pixels to the high-level semantic information. By calculating the Gram matrix of , the correlation statistics between various features can be obtained, which is the key index for capturing the image style. The feature map is flattened in the spatial dimension to obtain , is the batch size; is the style loss of the layer under the loss criterion; is the number of elements in the Gram matrix of the and are the reconstructed high-resolution image and the original high-resolution image At the Gram matrix of the layer feature representation; Is the overall weight of the style loss;

[0195] GAN loss Is:

[0196]

[0197] In the formula, Is the final output of the discriminator, which is the discrimination result of the reconstructed high-resolution image , and is a probability value. The goal of the generator is to maximize this probability value, that is, to make the discriminator unable to correctly identify whether the input image is a reconstructed high-resolution image or an original high-resolution image.

[0198] The loss function of the discriminator Is:

[0199]

[0200] In the formula, Is the original high-resolution image loss, Is the reconstructed high-resolution image loss.

[0201] Among them, the original high-resolution image loss Is:

[0202]

[0203] Here Is the discrimination result of the discriminator on the original high-resolution image, and is a probability value. The goal of the discriminator is to maximize this probability value, that is, to correctly identify that the input is the original high-resolution image.

[0204] The reconstructed high-resolution image loss Is:

[0205]

[0206] In the formula, Is the discrimination result of the discriminator on the reconstructed high-resolution image, and is also a probability value. The goal of the discriminator is to minimize this probability value, that is, to correctly identify that the input is the reconstructed high-resolution image.

[0207] S5. Train the image super-resolution model using the original high-resolution image dataset and the low-resolution image dataset. The training process includes:

[0208] Load the pre-trained model and the pre-trained weights of the generator and discriminator. Initialize the training settings, including the loss function, optimizer, etc.;

[0209] Data input: Input the low-resolution image, the original high-resolution image, and the corresponding text information into the model;

[0210] Train the generator: Input the low-resolution image into the generator, output the reconstructed high-resolution image, and calculate and optimize the generator loss;

[0211] Train the discriminator: Input the reconstructed high-resolution image, the original high-resolution image, and the corresponding text confidence into the discriminator, calculate the discriminator loss, and update the discriminator parameters;

[0212] Loop iteration: Repeat the above training loop until the stop condition is met.

[0213] S6: Input the low-resolution image and its corresponding text information into the trained image super-resolution model, and output the image super-resolution reconstruction result.

[0214] This embodiment conducts an ablation experiment to verify the effect of the image super-resolution model of the present invention;

[0215] Among them, the baseline is set to be similar to the traditional SISR model, which is a super-resolution model from low-resolution images to high-resolution images. The ablation experiment results are shown in Table 1. It can be seen that whether it is PSNR or SSIM, the indicators of the image super-resolution model proposed in the present invention are better than the baseline. Introducing text information into the baseline improves the quality of the target image super-resolution reconstruction result. At the same time, the influence of two different types of text inputs on the SISR task is considered to further prove the improvement of the quality of the image super-resolution reconstruction result by introducing text information. For an image selected from the CelebA-HQ dataset, it contains two text description information.

[0216] The two text information descriptions are as follows. It can be seen more superficially that the first text information description is more literally detailed and easier for human reading and understanding. The two descriptions have the same theme, both describing the same female image, and expanding on her appearance features, such as hair, eyebrows, nose, makeup, etc. The keywords all mention "hair", "eyebrows", "nose", "lipstick", "heavy makeup", and "cosmetic", etc.

[0217] The first type of text information is a natural language description, consisting of multiple short sentences, which specifically and detailedly describes the appearance characteristics of a person. The information provided is relatively scattered. Among them, some characteristics are repeated multiple times, but more details are provided, including more implicit context information, enabling the model to pay more attention to image details and overall features. For example, "busy eyebrows" and "arched eyebrows", thick arched eyebrows. The second type of text description is a list of keywords separated by a comma, which is concise and concentrated, with a high information density. Each keyword appears only once and covers all the key information in the first type. It is closer to a machine-processable format, reducing redundant information input. As a result, the model training and inference speeds will be faster, but it lacks detailed descriptions.

[0218] The ablation experiment results of the two different types of text information are shown in Table 2. It can be seen that the SSIM results are consistent, but there are differences in PSNR and LPIPS; using the first type of text information is slightly better in PSNR, indicating higher reconstruction accuracy at the pixel level, but the gap is very small and almost negligible. Using the second type of text information performs better in LPIPS, indicating better visual perception and being more suitable for scenarios with higher requirements for image details and textures. The SSIM results of using both types of text information are consistent, indicating comparable performance in terms of structure recovery. Finally, whether using the first type of text information or the second type of text information, PSNR, LPIPS, and SSIM are all better than the baseline without using text information, fully demonstrating that introducing text information indeed improves the quality of the target image super-resolution reconstruction results.

[0219] The first type of text information: "She is wearing lipstick. She is young, and smiling and has big lips, mouth slightly open, pointy nose, and high cheekbones."

[0220] The second type of text information: "beautiful, crown, hair, laugh, mouth, smile, tiara, wear, woman,"

[0221] Table 1 Ablation Experiment Results

[0222]

[0223] Table 2 Ablation Experiment Results (for two types of text descriptions, dataset: CelebA-HQ)

[0224]

[0225] Comparative experiment

[0226] The method proposed in the present invention was compared with several state-of-the-art SISR methods, traditional image-to-image SISR methods, BSRGAN, Real-ESRGAN; SISR methods that introduce other reference information for assistance (e.g., the text information proposed in the present invention), seesr.

[0227] Table 3 shows the results of the comparative experiments conducted on several large datasets, and the selected datasets are all commonly used datasets in the SR field; in order to further prove the method proposed in the present invention and improve the credibility at the same time, comparative experiments were also carried out on 4 common standard test datasets, Set5, Set14, BSD100, and Urban100. The results are shown in Table 4. The methods selected for the comparative experiments in the present invention are all advanced models in the current super-resolution field. Among them, seesr is an image super-resolution method based on a diffusion model. seesr uses a controllable text-to-image diffusion model and takes the low-resolution image as additional conditional control information to achieve the purpose of controlling the generation of specific target images from the text, that is, generating a high-resolution image from a low-resolution image. As can be seen from Table 4, in terms of the quantitative comparison results on the Set5, Set14, BSD100, and Urban100 datasets, the method proposed in the present invention performs best in terms of PSNR, LPIPS, and SSIM. That is, in terms of the reconstruction accuracy at the pixel level, the similarity in visual perception, and the structural similarity, the method of introducing text information to assist image super-resolution reconstruction proposed in the present invention is better than other methods. Combining the data in Table 3 and Table 4, it can be seen that the method proposed in the present invention is better than other methods.

[0228] Table 3 Comparative experiment results

[0229]

[0230] Table 4 Comparative experiment results on common test datasets

[0231]

[0232] Although the specific implementation manners of the invention have been described in detail with reference to the accompanying drawings, it should not be construed as a limitation on the protection scope of this patent. Within the scope described in the claims, various modifications and deformations that can be made by those skilled in the art without creative efforts still fall within the protection scope of this patent.

Claims

1. A method for semantic joint perception-based image super-resolution reconstruction, characterized in that It includes the following steps: S1. Construct an original high-resolution image dataset; S2. Downsample the original high-resolution image to construct a low-resolution image dataset; S3. Construct an image super-resolution model based on a conditional generative adversarial network and text information; S4. Calculate the loss function of the image super-resolution model; S5. Use the original high-resolution image dataset and the low-resolution image dataset to train the image super-resolution model; S6. Input the low-resolution image and its corresponding text information into the trained image super-resolution model, and output the image super-resolution reconstruction result.

2. The semantic joint perception-based image super-resolution reconstruction method according to claim 1, wherein In S2, the image super-resolution model includes: A generator that inputs the low-resolution image into the generator and outputs a reconstructed high-resolution image; A discriminator that inputs the reconstructed high-resolution image and its corresponding text information, the original high-resolution image and its corresponding text information into the discriminator for semantic alignment discrimination, and transmits the semantic alignment discrimination result to the generator for adversarial training.

3. The method for semantic joint perception-based image super-resolution reconstruction according to claim 2, wherein Inputting the low-resolution image into the generator and outputting a reconstructed high-resolution image specifically includes: Performing multi-layer convolution operations and dynamic residual processing on the input low-resolution image to generate the core features of the image; At the same time, upsampling the input low-resolution image to the target resolution through bilinear interpolation to obtain a reference high-resolution image; ; Merging the core features of the image and the reference high-resolution image to generate a reconstructed high-resolution image: ; Wherein, is the reconstructed high-resolution image; is the core feature of the image; is the reference high-resolution image; represents bilinear interpolation; is the input low-resolution image.

4. The method for semantic joint perception-based image super-resolution reconstruction according to claim 3, characterized in that, Performing multi-layer convolution operations and dynamic residual processing on the input low-resolution image to generate the core features of the image specifically includes: Inputting the low-resolution image into the first layer of convolution, and dynamically projecting the low-resolution image convolution into the feature space: ; In the formula, is the first image feature; is the weight; is the activation function; is the dynamic convolution operation; Inputting the first image feature into the residual block for dynamic residual processing: ; In the formula, is the second image feature; represents dynamic residual processing; Inputting the second image feature into the second layer of convolution for pixel rearrangement and dynamic convolution processing: ; In the formula, is the third image feature; represents pixel rearrangement; Inputting the third image feature into the third layer of convolution for dynamic convolution processing and adjusting the channel dimension: ; In the formula, is the fourth image feature; Inputting the fourth image feature into the fourth layer of convolution to map the fourth image feature to the output channel: ; In the formula, is the core feature of the image.

5. The semantic joint perception-based image super-resolution reconstruction method according to claim 2, wherein Inputting the reconstructed high-resolution image and its corresponding text information into the discriminator for semantic alignment discrimination specifically includes: Projecting the semantic features of the text information corresponding to the reconstructed high-resolution image to the target shape: ; Where s is the semantic feature of the text information; , , are respectively the semantic features of the text information introduced in different layers of convolution; is to convert the semantic feature of the input text information from its original representation into a multi-dimensional tensor adapted to network processing; The reconstructed high-resolution image is input into the first-stage convolutional layer, and after feature transformation, it is combined with the attention mechanism: ; ; In the formula, is the fifth image feature; is the first splicing feature; is the activation function; is the convolution operation; is the attention output; Inputting the first concatenated feature into the second-stage convolutional layer, extracting the image features and combining them with the attention mechanism: ; ; In the formula, is the sixth image feature; is the second splicing feature; Inputting the second concatenated feature into the third-stage convolutional layer, extracting the image features and combining them with the attention mechanism: ; ; In the formula, is the seventh image feature; is the third splicing feature; Projecting the third concatenated feature through the last convolutional layer to a single-channel output: ; In the formula, is the final output of the discriminator, representing the prediction result of the reconstructed high-resolution image inputted.

6. The semantic joint perception-based image super-resolution reconstruction method according to claim 5, characterized in that, In the discriminator, an attention mechanism is introduced in the first-stage convolutional layer, the second-stage convolutional layer, and the third-stage convolutional layer, so as to embed the semantic features of the text information into the image features; Among them, the attention mechanism in the first-stage convolutional layer is expressed as: Normalizing the input fifth image feature and the semantic features of the text information; ; ; In the formula, is the fifth image feature after standardization; is the semantic feature of the text information after standardization in the first stage; represents the standardization process; Using a 1×1 convolution to map the normalized fifth image feature and the normalized semantic features of the text information in the first stage to the embedding dimension: ; In the formula, is the query matrix, is the key matrix, is the value matrix; Calculating the attention weight matrix: ; In the formula, is the attention weight matrix; is the dimension of the key; is the transpose of the matrix; is the normalization operation; Weight the attention weight matrix to the value matrix to generate an attention output .

7. The method for semantic joint perception-based image super-resolution reconstruction according to claim 1, wherein In S4, the loss function of the image super-resolution model includes: Loss function of the generator is as follows: ; In the formula, is the pixel loss, is the perceptual loss, is the style loss, is the GAN loss; Loss function of the discriminator is as follows: ; In the formula, is the loss of the original high-resolution image, is the loss of the reconstructed high-resolution image.

8. The semantic joint perception-based image super-resolution reconstruction method according to claim 7, characterized in that Perceptual loss is as follows: ; ; In the formula, is the perceptual loss of the th layer under the loss criterion, where the mean absolute error is used; is the number of elements in the feature map of the th layer; i is the summation index, used to iterate through the values from 1 to ; and are the feature representations of the generated image and the real image in the th layer, respectively; is the original high-resolution image; is the reconstructed high-resolution image; is the weight corresponding to the th layer, is the overall weight of the perceptual loss; Style loss is as follows: ; ; ; Wherein, is the Gram matrix calculated for the layer feature map; ; is the feature map of the layer; is the number of channels, and are the height and width of the feature map respectively; represents the activation value of each channel in each spatial position of the deep layer feature map; is the style loss of the layer under the loss criterion; is the number of elements of the Gram matrix of the layer; and are the Gram matrices of the reconstructed high-resolution image and the original high-resolution image in the layer feature representation respectively; is the overall weight of the style loss; GAN loss is as follows: ; In the formula, is the final output of the discriminator, which is the discrimination result of the discriminator on the reconstructed high-resolution image.

9. The semantic joint perception-based image super-resolution reconstruction method according to claim 7, wherein Original high-resolution image loss is as follows: ; Reconstructed high-resolution image loss is as follows: ; In the formula, is the final output of the discriminator, is the discrimination result of the discriminator on the original high-resolution image, is the discrimination result of the discriminator on the reconstructed high-resolution image.

Citation Information

Patent Citations

  • Image equalization enhancement method, system, device and storage medium

    CN113421188A

  • Image super-resolution and defogging fusion method and system based on loop network

    CN114581304A

  • Multi-scale connection generative adversarial network medical image super-resolution reconstruction method

    CN116612009A

  • Thermal infrared image optimization method based on multi-channel fusion and semantic information

    CN120163710A