Real world image super-resolution reconstruction method and system based on diffusion converter

By designing an interaction control mechanism for the characteristics of the diffusion converter architecture, bidirectional information interaction between the latent features of noise, text feature representation and image potential representation is realized, and low-resolution local information is injected across the current convolution layer, the shortcomings of the DiT architecture in local information capture are solved, and the super-resolution image generation quality is significantly improved.

CN120198294APending Publication Date: 2025-06-24NANKAI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510367333.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing image super-resolution method based on diffusion converter performs poorly in local information capture and lacks special design for DiT architecture characteristics, resulting in a one-way and insufficient information interaction between the control information and the generated information.

Method used

An interactive control mechanism for the architectural characteristics of the diffusion converter is designed, and a bidirectional information interaction between the noise potential features, text feature representations and image potential representations is realized through the multimodal diffusion converter control module, and low-resolution local information is injected into the noise potential flow through a cross-current convolution layer.

Benefits of technology

It realizes more effective information interaction and image detail capture, improves the quality of super-resolution image generation, especially in the restoration tasks of text and architectural structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198294A_ABST
    Figure CN120198294A_ABST
Patent Text Reader

Abstract

The invention discloses a real world image super-resolution reconstruction method and system based on a diffusion converter, and belongs to the technical field of computer vision. Comprising the following steps: acquiring a real world image and coding to a potential space to generate an image potential representation; obtaining a noise potential feature and a text feature representation corresponding to the real world image, performing diffusion denoising on the noise potential feature, the text feature representation and the image potential representation through a diffusion converter network, and reconstructing a super-resolution image; the diffusion converter network utilizes a multi-mode diffusion converter control module to realize bidirectional information interaction among the noise potential feature, the text feature representation and the image potential representation, and injects low-resolution local information from the image potential representation into a noise potential stream through a cross-stream convolutional layer. According to the method, higher-quality image detail recovery and visual authenticity improvement can be realized, and the problems of insufficient interaction and single injection mode of super-resolution reconstruction information of an existing real image are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular, to a real-world image super-resolution reconstruction method and system based on a diffusion transformer. Background Art

[0002] The statements in this section merely mention the background art related to the present invention and do not necessarily constitute prior art.

[0003] Image Super-Resolution (ISR) aims to recover high-resolution (HR) images from low-resolution (LR) images through algorithms, and is one of the important topics in the field of computer vision.

[0004] With the popularization of digital devices and the diversification of image data acquisition channels, images captured in the real environment often have various complex degradations (such as blurring, noise, compression distortion, etc.), thus giving rise to the problem of real-world image super-resolution (Real-World Image Super-Resolution, Real-ISR). Compared with traditional super-resolution tasks, Real-ISR not only requires the model to eliminate complex degradations, but also needs to generate real and perceptually realistic image details to improve the visual effect. This high-difficulty inverse problem places higher requirements on the prior knowledge of the model.

[0005] Traditional super-resolution techniques usually adopt structures based on convolutional neural networks (CNNs). These methods rely on learning the structural information of images through large-scale training data to recover high-resolution images. However, these methods perform poorly in the face of real and complex degradation scenarios, usually unable to effectively recover details and prone to producing artificial artifacts and over-smoothing effects. To solve these problems, researchers have proposed a series of methods based on generative adversarial networks (GANs), such as Real-ESRGAN, SwinIR, etc., by constructing more complex degradation models or adversarial training mechanisms, and attempting to generate perceptually realistic detail information. However, these GAN-based methods suffer from training instability and the risk of generating false artifacts.

[0006] In recent years, with the rise of diffusion models, especially large-scale pre-trained text-to-image generation models represented by Stable Diffusion (SD), super-resolution technology has achieved new breakthroughs. These diffusion models are trained on large-scale datasets, possess rich generation priors, can generate realistic high-quality images, and demonstrate extremely strong generation potential. Most existing studies are based on the Stable Diffusion model with a UNet structure, and use low-resolution image information as conditional input into the diffusion model through mechanisms such as ControlNet or similar ones to generate corresponding high-resolution images. Although these methods improve the quality of the generated images, due to the limitations of the UNet architecture itself in capturing local image information, there are still certain bottlenecks in the reconstruction of details and structures.

[0007] The newly developed Diffusion Transformer (DiT) architecture further enhances the generation performance of diffusion models. Models built with DiT as the core, such as Stable Diffusion 3 (SD3) and Flux, etc., show better detail generation capabilities than traditional UNet architectures. In particular, the Multimodal Diffusion Transformer (MM-DiT) significantly enhances the model's ability to capture and generate cross-modal information by introducing a bidirectional attention mechanism for two streams of text and visual features. This enables the DiT architecture to achieve excellent results in image generation tasks.

[0008] However, the current image super-resolution methods based on the Diffusion Transformer (DiT) are still in the initial exploration stage and usually directly follow the ControlNet control mechanism based on the UNet architecture. Although this mechanism can incorporate low-resolution image information into the generation process, due to the lack of a dedicated design for the characteristics of the DiT architecture, the information interaction between the control information and the generation information is one-way and insufficient, thus limiting the potential performance of the DiT architecture, especially in terms of capturing local information. In addition, the existing methods often inject low-resolution image information in a relatively single way and cannot fully utilize the interaction between information flows during the diffusion process to gradually refine and optimize the generation results. Summary of the Invention

[0009] To address the deficiencies of the prior art, the present invention provides a real-world image super-resolution reconstruction method, system, electronic device, computer-readable storage medium, and computer program product based on a diffusion transformer, designing an interactive control mechanism for the characteristics of the diffusion transformer architecture to better achieve real-image super-resolution reconstruction.

[0010] In a first aspect, the present invention provides a real-world image super-resolution reconstruction method based on a diffusion transformer;

[0011] A real-world image super-resolution reconstruction method based on a diffusion transformer includes:

[0012] Obtain a real-world image and encode it into the latent space to generate an image latent representation;

[0013] Obtain noise latent features and text feature representations corresponding to the real-world image, and perform diffusion denoising on the noise latent features, the text feature representations, and the image latent representation through a diffusion transformer network to reconstruct a super-resolution image;

[0014] Among them, the diffusion transformer network uses a multimodal diffusion transformer control module to achieve two-way information interaction between the noise latent features, the text feature representations, and the image latent representation, and injects low-resolution local information from the image latent representation into the noise latent flow through a cross-flow convolutional layer.

[0015] In some embodiments, obtaining the text feature representation corresponding to the real-world image specifically includes: processing the real-world image through a multimodal large model to generate a corresponding text description; and sequentially processing the text description through multiple large language models to obtain the text feature representation.

[0016] In some embodiments, the use of the multimodal diffusion transformer control module to achieve two-way information interaction between the noise latent features, the text feature representations, and the image latent representation specifically includes:

[0017] Linearly process the noise latent features, the text feature representations, and the image latent representation processed by the previous-level multimodal diffusion transformer control module respectively, fuse the linear processing results in pairs, and perform global joint attention calculation on the fusion results through an attention mechanism;

[0018] Among them, the image latent representation processed by the previous-level multimodal diffusion transformer control module is residually connected to the corresponding output of the attention mechanism.

[0019] In some embodiments, injecting the low-resolution local information from the image latent representation into the noise latent flow through the cross-flow convolutional layer is specifically as follows: processing the result of the residual connection through a linearly connected layer and a multi-layer perceptron in sequence and outputting it to the cross-flow convolutional layer, and injecting it into the multi-layer perceptron for processing noise information through the cross-flow convolutional layer.

[0020] In some embodiments, the multi-modal diffusion transformer control module includes a first linear layer, a zero linear layer, a second linear layer, an attention mechanism, a third linear layer, a fourth linear layer and a fifth linear layer, a first multi-layer perceptron, a second multi-layer perceptron and a third multi-layer perceptron;

[0021] The first linear layer, the attention mechanism, the third linear layer and the first multi-layer perceptron are connected in sequence, the zero linear layer, the attention mechanism, the fourth linear layer and the second multi-layer perceptron are connected in sequence, and the second linear layer, the attention mechanism, the fifth linear layer and the third multi-layer perceptron are connected in sequence;

[0022] The output of the zero linear layer is connected with the corresponding output of the attention mechanism by residual connection, and a cross-flow convolutional layer is arranged between the first multi-layer perceptron and the second multi-layer perceptron.

[0023] In some embodiments, before the image latent representation is input into the diffusion transformer network, it further includes: performing image tokenization and position encoding on the image latent representation.

[0024] In a second aspect, the present invention provides a real-world image super-resolution reconstruction system based on a diffusion transformer;

[0025] A real-world image super-resolution reconstruction system based on a diffusion transformer includes:

[0026] An image encoding module, configured to: acquire a real-world image and encode it into a latent space to generate an image latent representation;

[0027] A super-resolution reconstruction module, configured to: acquire noise latent features and a text feature representation corresponding to the real-world image, and perform diffusion denoising on the noise latent features, the text feature representation and the image latent representation through a diffusion transformer network to reconstruct a super-resolution image;

[0028] Wherein, the diffusion transformer network uses a multi-modal diffusion transformer control module to realize bidirectional information interaction between the noise latent features, the text feature representation and the image latent representation, and injects low-resolution local information from the image latent representation into the noise latent flow through a cross-flow convolutional layer.

[0029] In a third aspect, the present invention provides an electronic device;

[0030] An electronic device includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the above-mentioned real-world image super-resolution reconstruction method based on a diffusion transformer.

[0031] In a fourth aspect, the present invention provides a computer-readable storage medium;

[0032] A computer-readable storage medium has a computer program / instructions stored thereon. When the computer program / instructions are executed by a processor, the steps of the above-mentioned real-world image super-resolution reconstruction method based on a diffusion transformer are implemented.

[0033] In a fifth aspect, the present invention provides a computer program product;

[0034] A computer program product includes a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the above-mentioned real-world image super-resolution reconstruction method based on a diffusion transformer are implemented. Compared with the prior art, the beneficial effects of the present invention are as follows:

[0035] 1. The technical solution provided by the present invention adds an additional LR latent feature path to the diffusion transformer, enabling the simultaneous fusion of the latent features of the noisy image (Noise Stream), the latent features of the low-resolution image (LRStream), and the text description features (Text Stream) in the attention calculation, achieving global joint attention calculation among the three information flows; enabling the low-resolution image features to continuously interact with the noise latent features during the diffusion process and co-evolve, ensuring that the low-resolution image information can dynamically adapt to the diffusion process, providing more accurate generation guidance, and significantly improving the quality of the generated super-resolution images.

[0036] 2. The technical solution provided by the present invention designs a cross-stream information injection mechanism. Between the multi-layer perceptron (MLP) of the noise information and the image information, a depth convolutional layer is introduced to inject the local features from the LR information stream into the MLP of the noise information stream, thereby making up for the deficiency of the attention mechanism in capturing local details of the image; this design of local information enhancement further strengthens the ability to capture and restore image details, especially showing outstanding performance in the restoration tasks of fine structures such as text and building structures. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings forming a part of the present invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.

[0038] Figure 1 Schematic flowchart of the real-world image super-resolution reconstruction method based on the diffusion transformer provided by an embodiment of the present invention;

[0039] Figure 2 Schematic diagram of the network architecture of the real-world image super-resolution reconstruction method based on the diffusion transformer provided by an embodiment of the present invention;

[0040] Figure 3 Example diagram of the structural comparison between the real-world image super-resolution reconstruction method based on the diffusion transformer provided by an embodiment of the present invention and the prior art SD3-ControlNet method, where Figure 3 in (a) is the network structure diagram of SD3-ControlNet, Figure 3 in (b) is the network architecture diagram of the method described in this embodiment;

[0041] Figure 4 Example diagram of the effect of the LR residual connection provided by an embodiment of the present invention; where Figure 4 in (a) is an example diagram of the attention interaction matrix between the latent features X of the noisy image and the latent features L of the low-resolution image, Figure 4 in (b) is an example diagram of the performance of LR information in deeper modules without adding the LR residual connection (upper row) and with adding the LR residual connection (lower row);

[0042] Figure 5 Example diagram of the effect comparison of the actual application provided by an embodiment of the present invention, where Figure 5 in (a) is an example diagram of the original low-resolution (LR) input image, Figure 5 in (b) is an example diagram of the model output result without using the "local information injection between MLPs" mechanism, Figure 5 in (c) is an example diagram of the model result after injecting LR information using a linear layer, Figure 5 in (d) is an example diagram of the output result of the method described in this embodiment, where the linear layer is replaced by a convolutional layer for injecting LR information. Detailed implementation manner

[0043] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0044] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0045] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0046] Embodiment 1

[0047] The existing super-resolution reconstruction technology applied to real-world images cannot fully utilize the interaction between information flows during the diffusion process to gradually refine and optimize the generated results; therefore, this embodiment provides a real-world image super-resolution reconstruction method based on a diffusion transformer, and designs an interaction control mechanism for the characteristics of the DiT architecture to better achieve the super-resolution reconstruction of real images.

[0048] Next, in combination with Figures 1 - 5 , a real-world image super-resolution reconstruction method based on a diffusion transformer disclosed in this embodiment will be described in detail. The real-world image super-resolution reconstruction method based on a diffusion transformer includes the following steps:

[0049] S1. Obtain a real-world image and encode it into the latent space to generate an image latent representation; at the same time, obtain a text feature representation corresponding to the real-world image.

[0050] Here, the real-world image is a low-resolution image input by the user. Specifically, use a pre-trained VAE encoder to encode the real-world image into the latent space to form a low-resolution image latent representation; process the real-world image through a multimodal large model to generate a corresponding text description; process the text description through multiple large language models in sequence to obtain a text feature representation.

[0051] In this embodiment, the multimodal large model is the LLAVA model, and the number of large language models is 3, which are the T5 XXL model, the CLIP-G model, and the CLIP-L model in sequence according to the processing order.

[0052] To make the latent representation of the image structurally consistent with the original noise latent features of the diffusion transformer and facilitate effective fusion in the attention mechanism in the subsequent stage; as an implementation, it further includes: performing image tokenization and positional encoding on the latent representation of the image. The specific process is as follows:

[0053] First, encode the input real-world image through a pre-trained VAE encoder into the latent space to obtain the latent representation Z of the image, expressed as:

[0054]

[0055] In the formula, E(·) represents the pre-trained VAE encoder provided by SD3, and h, w, and c respectively represent the height, width, and number of channels of the latent representation of the image.

[0056] Subsequently, perform image tokenization on the latent representation of the image. The specific process is as follows:

[0057] (1) Divide the latent representation Z of the image into multiple image patches of size p×p, and then flatten them into a one-dimensional vector, expressed as:

[0058]

[0059] In the formula, z i represents the i-th image patch, Flatten represents the flattening operation, and Z i represents the i-th latent representation of the image.

[0060] (2) Convert each image patch into a D-dimensional feature vector, that is, the image token x i , expressed as:

[0061]

[0062] In the formula, respectively represent the weight and bias terms of the linear projection layer.

[0063] (3) Combine all the image tokens to form an image token sequence X, expressed as:

[0064]

[0065] In the formula, K represents the length of the image token sequence,

[0066] Next, add positional encoding to the image token sequence to provide positional awareness information for each position in the image token sequence. The image token sequence X' after adding positional encoding is expressed as:

[0067]

[0068] In the formula, represents the position encoding.

[0069] Specifically, a common sine position encoding method is expressed as:

[0070]

[0071] where pos represents the token position index (from 0 to K - 1), and i represents the feature dimension index.

[0072] Based on the above processing flow, the method in this embodiment ensures that the spatial structure information of the image latent features is efficiently transformed into a sequence representation suitable for the diffusion transformer model for subsequent information interaction and image restoration processes.

[0073] S2. Perform diffusion denoising on the noise latent features, text feature representations, and image latent representations through a diffusion transformer network to reconstruct a super-resolution image.

[0074] In order to enable the low-resolution image features to continuously interact with the noise latent features during the diffusion process, co-evolve, ensure that the low-resolution image information can dynamically adapt to the diffusion process, and provide more accurate generation guidance, in this embodiment, an additional LR latent feature path is added to the attention mechanism, so that the noise latent features, low-resolution image latent representations, and text feature representations are simultaneously fused during attention calculation.

[0075] As an implementation manner, the diffusion transformer network includes a plurality of cascaded multi-modal diffusion transformer control modules. The multi-modal diffusion transformer control module includes a first linear layer, a zero linear layer, a second linear layer, an attention mechanism, a third linear layer, a fourth linear layer, and a fifth linear layer, a first multi-layer perceptron, a second multi-layer perceptron, and a third multi-layer perceptron; the outputs of the first linear layer, the zero linear layer, and the second linear layer are fused and then input into the attention mechanism in parallel. The attention mechanism, the third linear layer, and the first multi-layer perceptron are connected in sequence. The attention mechanism, the fourth linear layer, and the second multi-layer perceptron are connected in sequence. The attention mechanism, the fifth linear layer, and the third multi-layer perceptron are connected in sequence.

[0076] Furthermore, the output of the zero linear layer is residually connected to the corresponding output of the attention mechanism to more effectively fuse local detail information; a cross-flow convolutional layer is arranged between the first multi-layer perceptron and the second multi-layer perceptron, and the LR local information is injected into the noise latent flow through the cross-flow convolutional layer to compensate for the deficiency of the diffusion transformer in capturing local information.

[0077] In this embodiment, the cross-flow convolutional layer is a ZeroConv convolutional layer, that is, a depth convolutional layer initialized with zeros.

[0078] As an implementation, S2 specifically includes:

[0079] S201. Process the noise latent features, text feature representations, and image latent representations through a diffusion transformer network to obtain high-quality image latent features.

[0080] Furthermore, the noise latent features, text feature representations, and image latent representations are sequentially processed through a cascade of multiple multi-modal diffusion transformer control modules, and the final output of the last multi-modal diffusion transformer control module is the high-quality image latent features.

[0081] Let X n-1 represent the noise stream output by the (n - 1)-th multi-modal diffusion transformer control module, let L n-1 represent the low-resolution image stream output by the (n - 1)-th multi-modal diffusion transformer control module, and let C n-1 represent the text stream output by the (n - 1)-th multi-modal diffusion transformer control module. Exemplarily, the data processing flow of the n-th modal diffusion transformer control module is as follows:

[0082] (1) The noise stream X n-1 , the low-resolution image stream L n-1 , and the text stream C n-1 are respectively passed through the first linear layer, the zero linear layer, and the second linear layer to generate the query (Q), key (K), and value (V) of the attention mechanism, which are expressed as:

[0083]

[0084] In the formula, respectively represent the weights of each linear layer, represents the concatenation operation.

[0085] (2) The features are fused through the scaled dot-product attention mechanism, which is expressed as:

[0086]

[0087] In the formula, D represents the length of the feature.

[0088] (3) Add L n-1 to the output of the attention mechanism, and then sequentially process it through the fourth linear layer and the second multi-layer perceptron to output L n ; process the output of the attention mechanism through the third linear layer and the first multi-layer perceptron. At the same time, inject the intermediate features processed by the second multi-layer perceptron into the first multi-layer perceptron through a 3×3 depth convolution layer with an initial weight of zero, and output X n; The output of the attention mechanism is sequentially processed by the fifth linear layer and the third multi-layer perceptron, and the output is C n 。

[0089] In this embodiment, to enhance the stability of LR information in the deep module, an LR residual connection is introduced to directly add the LR input feature, i.e., L n-1 to the output of the attention mechanism.

[0090] In addition, since the attention mechanism mainly focuses on global information and has limited local information capture ability, an LR local information injection mechanism is designed between the intermediate features of the MLP of the noise stream and the LR stream. Through a 3×3 depth convolution layer with an initial weight of zero, the intermediate features of the LR stream are effectively injected into the features of the noise stream to make up for the deficiency of the Transformer structure in capturing local details. Specifically, both the first multi-layer perceptron and the second multi-layer perceptron include three linearly connected layers in sequence. The output of the second linear layer of the second multi-layer perceptron is input to the second linear layer of the first multi-layer perceptron through a 3×3 depth convolution layer. The second linear layer of the first multi-layer perceptron processes the output of the first linear layer in the first multi-layer perceptron and the output of the 3×3 depth convolution layer, and outputs X after being processed by the third linear layer of the first multi-layer perceptron n 。

[0091] S202. Linearly process the latent features of the high-quality image and perform an Unpatching operation, and decode the processed latent features of the high-quality image through a VAE decoder to reconstruct a super-resolution image.

[0092] In this embodiment, Unpatching refers to the inverse operation of the image tokenization (patching) process, which is used to recombine the image tokens in the form of a sequence output by the model into a complete latent space feature map or image representation.

[0093] Embodiment 2

[0094] This embodiment discloses a real-world image super-resolution reconstruction system based on a diffusion transformer, including:

[0095] An image encoding module, configured to: obtain a real-world image and encode it into a latent space to generate an image latent representation;

[0096] A super-resolution reconstruction module, configured to: obtain noise latent features and a text feature representation corresponding to the real-world image, and perform diffusion denoising on the noise latent features, the text feature representation, and the image latent representation through a diffusion transformer network to reconstruct a super-resolution image;

[0097] Among them, the diffusion transformer network uses a multi-modal diffusion transformer control module to achieve two-way information interaction between the noise latent feature, the text feature representation, and the image latent representation, and injects low-resolution local information from the image latent representation into the noise latent flow through a cross-flow convolutional layer.

[0098] It should be noted here that the above image encoding module and super-resolution reconstruction module correspond to the steps in Embodiment 1. The examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.

[0099] Embodiment 3

[0100] Embodiment 3 of the present invention provides an electronic device, including a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps of the above real-world image super-resolution reconstruction method based on a diffusion transformer are completed.

[0101] Embodiment 4

[0102] Embodiment 4 of the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the above real-world image super-resolution reconstruction method based on a diffusion transformer are completed.

[0103] Embodiment 5

[0104] Embodiment 5 of the present invention provides a computer program product, including a computer program / instructions. When the computer program / instructions are executed by a processor, the steps of the above real-world image super-resolution reconstruction method based on a diffusion transformer are implemented.

[0105] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0106] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to work in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction device that implements the functions specified in one or more processes and / or blocks Figure 1 in the process Figure 1 or processes and / or boxes

[0107] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to perform a series of operational steps on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one or more processes and / or blocks Figure 1 in the process Figure 1 or processes and / or boxes

[0108] In the above embodiments, the descriptions of the various embodiments have their respective focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments

[0109] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention

Claims

1. A real-world image super-resolution reconstruction method based on a diffusion transformer, characterized in that: include: Obtain real-world images and encode them into latent space to generate latent representations of images; Acquire noise latent features and text feature representations corresponding to the real-world image, perform diffusion denoising on the noise latent features, the text feature representations, and the image latent representations through a diffusion transformer network, and reconstruct a super-resolution image; Among them, the diffusion transformer network uses a multimodal diffusion transformer control module to realize bidirectional information interaction between the noise latent features, the text feature representation and the image latent representation, and injects low-resolution local information from the image latent representation into the noise latent stream through a cross-stream convolutional layer.

2. The real-world image super-resolution reconstruction method based on diffusion transformer according to claim 1, characterized in that: Acquiring the text feature representation corresponding to the real-world image is specifically as follows: processing the real-world image through a large multimodal model to generate a corresponding text description; and sequentially processing the text descriptions through multiple large language models to obtain the text feature representation.

3. The real-world image super-resolution reconstruction method based on diffusion transformer according to claim 1, characterized in that: The method of using a multimodal diffusion transformer control module to realize bidirectional information interaction between the noise potential feature, the text feature representation and the image potential representation is specifically as follows: The noise latent feature, the text feature representation and the image latent representation processed by the previous multimodal diffusion transformer control module are linearly processed respectively, and the linear processing results are fused in pairs, and the fusion results are globally joint-attended by the attention mechanism; Among them, the potential representation of the image processed by the previous level multimodal diffusion transformer control module is connected to the corresponding output residual of the attention mechanism.

4. The real-world image super-resolution reconstruction method based on diffusion transformer according to claim 3, characterized in that: The step of injecting low-resolution local information from the potential representation of the image into the noise potential stream through the cross-stream convolution layer is specifically as follows: the residual connection result is processed through a linear layer and a multilayer perceptron connected in sequence and outputted to the cross-stream convolution layer, and the multilayer perceptron for processing noise information is injected through the cross-stream convolution layer.

5. The real-world image super-resolution reconstruction method based on diffusion transformer according to claim 1, characterized in that: The multimodal diffusion transformer control module includes a first linear layer, a zero linear layer, a second linear layer, an attention mechanism, a third linear layer, a fourth linear layer, a fifth linear layer, a first multilayer perceptron, a second multilayer perceptron, and a third multilayer perceptron; The first linear layer, the attention mechanism, the third linear layer and the first multilayer perceptron are connected in sequence, the zero linear layer, the attention mechanism, the fourth linear layer and the second multilayer perceptron are connected in sequence, and the second linear layer, the attention mechanism, the fifth linear layer and the third multilayer perceptron are connected in sequence; The output of the zero linear layer is residually connected to the corresponding output of the attention mechanism, and a cross-stream convolution layer is arranged between the first multi-layer perceptron and the second multi-layer perceptron.

6. The real-world image super-resolution reconstruction method based on diffusion transformer according to claim 1, characterized in that: Before the image potential representation is input into the diffusion transformer network, it also includes: performing image tokenization and position encoding on the image potential representation.

7. A real-world image super-resolution reconstruction system based on a diffusion transformer, characterized in that: include: The image encoding module is configured to: obtain a real-world image and encode it into a latent space to generate an image latent representation; A super-resolution reconstruction module is configured to: obtain noise latent features and text feature representations corresponding to the real-world image, perform diffusion denoising on the noise latent features, the text feature representations and the image latent representations through a diffusion transformer network, and reconstruct a super-resolution image; Among them, the diffusion transformer network uses a multimodal diffusion transformer control module to realize bidirectional information interaction between the noise latent features, the text feature representation and the image latent representation, and injects low-resolution local information from the image latent representation into the noise latent stream through a cross-stream convolutional layer.

8. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the real-world image super-resolution reconstruction method based on the diffusion transformer according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instruction is executed by a processor, the steps of the real-world image super-resolution reconstruction method based on a diffusion transformer as described in any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the real-world image super-resolution reconstruction method based on a diffusion transformer as described in any one of claims 1 to 6 are implemented.