Image generation method and system based on step-by-step super-resolution and regional differentiation

By adopting a step-by-step super-resolution and region differentiation method in image generation, the problem of inaccurate and low quality of the target area of ​​the diffusion model during high-resolution image generation is solved, and image generation with higher quality and accuracy is achieved.

CN119313565BActive Publication Date: 2025-05-06SHANDONG JIANZHU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411855449.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-17
Publication Date
2025-05-06
Estimated Expiration
2044-12-17

AI Technical Summary

Technical Problem

When generating high-resolution images, the diffusion model faces the problem of inaccurate target areas and low quality, making it difficult to effectively capture complex scenes and fine details of the image.

Method used

Using an image generation method based on stepwise super-resolution and regional differentiation, through a layer-by-layer generation strategy, low-resolution target images are created, and using them as a condition, the generation of higher-resolution images is gradually guided. The method includes acquiring the image to be processed and pre-processing, dividing the image area for noise addition, inputting feature extraction branches and U-NET network step by step for processing, generating conditional fusion features and performing reverse diffusion, and finally generating a high-quality target image.

Benefits of technology

Through the step-by-step super-resolution generation method, the global structure and local features of the image can be effectively captured, and the quality and accuracy of the generated image can be improved, solving the problems of inaccurate target areas and low quality in the traditional diffusion model when generating high-resolution images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119313565B_ABST
    Figure CN119313565B_ABST
Patent Text Reader

Abstract

The present invention discloses an image generation method and system based on step-by-step super-resolution and regional differentiation, belonging to the field of image generation technology. It includes: dividing images at different resolution levels into regions and adding noise based on the regional division results to obtain sub-regional noisy images; according to the resolution level, inputting the images into the feature extraction branch step by step, and generating conditional fusion features guided by the final target image generated at the previous resolution level; combining the diffusion time step, inputting the sub-regional noisy images into the U-NET network, splicing them with the conditional fusion features of the same resolution during the processing, and obtaining the noise prediction value; performing inverse diffusion based on the noise prediction value and the sub-regional noisy images to obtain the target image at the corresponding resolution level and diffusion time step. It can effectively capture the global structure and local features of the image, improve the quality of the generated results; and solve the shortcomings of the prior art in capturing the overall structure and fine details of complex scenes in the image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image generation, and in particular to an image generation method and system based on gradual super-resolution and regional differentiation. Background Art

[0002] The statements in this section merely mention background art related to the present invention and do not necessarily constitute prior art.

[0003] Image generation is a fundamental problem in computer vision. Its purpose is to generate highly realistic images by training network models. Image generation plays an important role in many fields, such as medical imaging, autonomous driving, and virtual reality. With the explosive growth of large-scale datasets and the rapid development of deep learning-related technologies, image generation has made significant progress, promoting the rise of multiple generative models, including generative adversarial networks (GANs), variational autoencoders (VAEs), and diffusion models. Among them, the diffusion model has demonstrated excellent high-quality image synthesis capabilities due to its outstanding performance in generation tasks, and has surpassed other generation methods in multiple tasks, gradually becoming a cutting-edge technology in the current field of image generation.

[0004] The diffusion model can not only generate images unconditionally, but also generate specific target images through conditional guidance. This conditional generation expands the application scenarios of image generation and promotes the concept of "image-to-image translation". Image-to-image translation is an important branch of image generation, and its goal is to map an input image to another output image in the target domain while maintaining certain key features. The diffusion model first gradually converts the original image into noise, and then through the inverse denoising process, the model gradually generates the target image from the noise. This process allows the model to capture the complex mapping relationship between the input image and the target image as much as possible, thereby ensuring the diversity and fidelity of the generated results; but there are still the following problems:

[0005] (1) The denoising process of the diffusion model faces the problem of difficulty in establishing a deterministic relationship between the pixels of the input image and the output image.

[0006] (2) Traditional diffusion models mainly process image pixels at a specific resolution. High-resolution images are often more complex and contain fine textures, complex structures or details. This method is insufficient in capturing the overall structure and fine details of complex scenes in images. Summary of the invention

[0007] In order to address the deficiencies of the prior art, the present invention provides an image generation method, system, electronic device, computer storage medium and computer program product based on progressive super-resolution and regional differentiation, which can effectively extract low-resolution global features and high-resolution local features, thereby improving the quality of generated images.

[0008] In a first aspect, the present invention provides an image generation method based on step-by-step super-resolution and regional differentiation;

[0009] An image generation method based on step-by-step super-resolution and regional differentiation, comprising:

[0010] Obtain the image to be processed and perform preprocessing to generate images at different resolution levels;

[0011] Divide images at different resolution levels into regions, and add noise to the images based on the region division results to obtain region-divided noisy images;

[0012] According to the resolution level, the image is input into the feature extraction branch step by step for processing, and the final target image generated by the previous resolution level is used as a guide to generate conditional fusion features; at the same time, combined with the diffusion time step, the sub-regional noisy image is input into the U-NET network step by step for processing, and during the processing, it is spliced ​​with the conditional fusion features of the same resolution to obtain the noise prediction value;

[0013] Based on the noise prediction value and the noisy image in different regions, inverse diffusion is performed to obtain the target image at the corresponding resolution level and diffusion time step until the final target image is generated.

[0014] In some implementations, the adding noise to the image based on the region division result specifically includes: replacing the background area of ​​the image with Gaussian noise.

[0015] In some implementations, the step of inputting the image into the feature extraction branch for processing according to the resolution level, and using the final target image generated at the previous resolution level as a guide, generating the conditional fusion feature specifically includes:

[0016] The final target image generated by the previous resolution level is input into the first feature extraction branch for processing to obtain a plurality of global feature maps of different resolutions; the image is input into the second feature extraction branch for processing, and a plurality of local feature maps of different resolutions are obtained with the global feature map as a guide;

[0017] Channel attention is used to fuse local feature maps and global feature maps of equal resolution to generate multiple conditional fusion features.

[0018] In some embodiments, the first feature extraction branch includes a plurality of first resolution feature layers connected in sequence, the first resolution feature layer including a patch merging block and a plurality of first resolution feature blocks;

[0019] The second feature extraction branch includes a plurality of second resolution feature layers connected in sequence, and the second resolution feature layer includes a plurality of second resolution feature blocks.

[0020] In some embodiments, the conditional fusion features are fused with resolution levels and diffusion process time steps through sinusoidal coding.

[0021] In some implementations, preprocessing the image to be processed specifically includes: processing the image to be processed by a bilinear interpolation method to obtain images at different resolution levels.

[0022] In a second aspect, the present invention provides an image generation system based on progressive super-resolution and regional differentiation;

[0023] An image generation system based on step-by-step super-resolution and regional differentiation, comprising:

[0024] The super-resolution generation module is configured to: obtain the image to be processed and perform pre-processing to generate images of different resolution levels;

[0025] The regional differentiation module is configured to: divide the images at different resolution levels into regions, and add noise to the images based on the region division results to obtain the region-by-region noisy images; according to the resolution levels, input the images into the feature extraction branches step by step for processing, and generate conditional fusion features guided by the final target image generated at the previous resolution level;

[0026] The inverse diffusion module is configured as follows: in combination with the diffusion time step, the sub-regional noisy image is input into the U-NET network step by step for processing, and during the processing, it is spliced ​​with the conditional fusion features of the same resolution to obtain the noise prediction value; based on the noise prediction value and the sub-regional noisy image, inverse diffusion is performed to obtain the target image at the corresponding resolution level and diffusion time step, until the final target image is generated.

[0027] In a third aspect, the present invention provides an electronic device;

[0028] An electronic device comprises a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the above-mentioned image generation method based on step-by-step super-resolution and regional differentiation.

[0029] In a fourth aspect, the present invention provides a computer-readable storage medium;

[0030] A computer-readable storage medium stores a computer program / instruction thereon, which, when executed by a processor, implements the steps of the above-mentioned image generation method based on gradual super-resolution and regional differentiation.

[0031] In a fifth aspect, the present invention provides a computer program product;

[0032] A computer program product comprises a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned image generation method based on step-by-step super-resolution and regional differentiation.

[0033] Compared with the prior art, the present invention has the following beneficial effects:

[0034] 1. The technical solution provided by the present invention proposes a strategy for generating images by dividing the regions according to the conflict between diversity and certainty when generating images using a diffusion model. Different generation strategies are formulated for different regions to ensure the accuracy of the generated images while improving the generalization ability.

[0035] 2. The technical solution provided by the present invention aims at the problem that the target area is inaccurate and of low quality when the traditional conditional diffusion model generates high-resolution images. A step-by-step super-resolution generation method is proposed. This method first generates a relatively low-resolution target image through a layer-by-layer generation strategy, and uses it as a condition to gradually guide the generation of higher-resolution images, and finally enriches and improves the details. This step-by-step generation method ensures that the global structure and local features of the image can be effectively captured at each stage, thereby improving the quality of the generated results.

[0036] 3. The technical solution provided by the present invention ensures that low-resolution global features and high-resolution local features can be effectively extracted, and more effective features can be extracted through efficient mixed conditional feature enhancement. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.

[0038] Figure 1 A schematic diagram of a process of an image generation method based on step-by-step super-resolution and regional differentiation provided in an embodiment of the present invention;

[0039] Figure 2 A schematic diagram of a network architecture of an image generation method based on progressive super-resolution and regional differentiation provided by an embodiment of the present invention;

[0040] Figure 3A schematic diagram of the framework of an image generation system based on progressive super-resolution and regional differentiation provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0041] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meanings as those commonly understood by those skilled in the art to which the present invention belongs.

[0042] It should be noted that the terms used herein are only for describing specific embodiments, and are not intended to limit exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that the terms "include" and "have" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0043] In the absence of conflict, the embodiments of the present invention and the features of the embodiments may be combined with each other.

[0044] Embodiment 1

[0045] The existing image generation through diffusion model is difficult to establish a deterministic relationship between the pixels of the input image and the output image, and is insufficient in capturing the overall structure and fine details of complex image scenes. Therefore, the present invention provides an image generation method based on gradual super-resolution and regional differentiation, which improves the quality of the generated image by combining the regional image denoising strategy, the gradual super-resolution generation strategy and the efficient mixing conditions.

[0046] Next, combine Figure 1-Figure 2 , a method for generating an image based on step-by-step super-resolution and regional differentiation disclosed in this embodiment is described in detail. The method for generating an image based on step-by-step super-resolution and regional differentiation comprises the following steps:

[0047] S1. Obtain the image to be processed and perform preprocessing to generate images at different resolution levels.

[0048] Exemplarily, the image to be processed may be an image of a virtual character. In the style transfer of the virtual character image, it is necessary to generate an image of the virtual character image.

[0049] Virtual character images often contain complex scenes. It is difficult to directly generate high-resolution images due to problems with detail capture and global structural coherence in complex scenes. To overcome this problem, this embodiment uses a step-by-step generation method to pre-process the image to be processed, and gradually generates a high-resolution image from a low-resolution image. The low-resolution image is used to provide the global structure of the image to avoid image incoherence due to the lack of global structure at high resolution; the image details of the high-resolution image are gradually enriched, thereby improving the integrity and details of the image.

[0050] Specifically, the image to be processed Images resized to different resolution levels using bilinear interpolation ,in, Indicates different resolution levels. The larger the number, the lower the resolution level. The highest resolution level, with stronger details. It is the lowest resolution level, with better global structure.

[0051] S2, divide the images at different resolution levels into regions, determine the target region and background region of the image, and add noise to the image based on the region division results to obtain the region-by-region noisy image. Specifically including:

[0052] S201, input images of different resolution levels into the PSPNet network for processing, and output an intermediate image containing the target area, and the part other than the target area is the background area.

[0053] Exemplarily, the target area is a person area, and the PSPNet network is a classic image segmentation model, which is not improved in this embodiment and will not be described in detail here.

[0054] To facilitate subsequent processing, in this embodiment, by acquiring the coordinates of all boundary points of the irregular target area (the boundary points are usually the outline or edge points of the area), a minimum rectangular area containing the target area is constructed.

[0055] Specifically, by calculating the minimum and maximum values ​​of x and y in the boundary point coordinates, we get , , , , then the four vertices of the rectangular area are: the upper left corner , upper right corner , lower left corner , lower right corner , this rectangular area is regarded as the target area, and the rest is regarded as the background area.

[0056] S202, adding random Gaussian noise to the background area, and not processing the target area, to obtain a noisy image for each region.

[0057] S3. According to the resolution level, the image is input into the feature extraction branch step by step for processing, and the final target image generated by the previous resolution level is used as a guide to generate conditional fusion features.

[0058] In order to capture broader contextual information and global features from low-resolution images, in this embodiment, global feature extraction is performed on the final target image generated by the previous resolution level through the first feature extraction branch to capture broad contextual information and global features; local feature extraction is performed on the image through the second feature extraction branch, and at the same time, the processing relies on the global information extracted by the first feature extraction branch, and the local features are refined through complex convolution.

[0059] In this embodiment, the feature extraction branch includes a first feature extraction branch and a second feature extraction branch, the first feature extraction branch includes 4 first resolution feature layers connected in sequence, the first resolution feature layer includes 1 patch merging block and 2 first resolution feature blocks; the second feature extraction branch includes 4 second resolution feature layers connected in sequence, the first second resolution feature layer includes 3 second resolution feature blocks, and the remaining second resolution feature layers include 2 second resolution feature blocks.

[0060] In this step, the resolution levels are arranged in ascending order, and the original image of each resolution level and the final target image generated by the previous resolution level are processed in sequence from small to large resolution levels.

[0061] As an implementation method, S3 specifically includes:

[0062] S301: The target image obtained at the previous resolution level Input the first feature extraction branch for processing to obtain multiple global feature maps with different resolutions; set the resolution level to Image The second feature extraction branch is input for processing to obtain multiple local feature maps with different resolutions.

[0063] Specifically, the target image obtained by the four first resolution feature layers on the previous resolution level Process them sequentially to obtain four global feature maps with different resolutions , which is the output of the second first-resolution feature block in each first-resolution feature layer. In each first-resolution feature layer, the input feature map is processed sequentially by the patch merging block and the two first-resolution feature blocks. The calculation process of each first-resolution feature block is expressed as follows:

[0064] ;

[0065] In the formula, Represents the output features of the previous first-resolution feature block. For the first first-resolution feature block, After the patch merging layer is processed Image; GELU represents the activation function for nonlinear mapping; Conv represents the convolution operation, BlkAttn represents the block-level attention mechanism, and LN represents the layer normalization. The patch merging layer means that the feature map is divided into small blocks (patches) and merged, thereby achieving resolution reduction and feature aggregation.

[0066] In this way, the first feature extraction branch can better capture the global characteristics of the image.

[0067] Correspondingly, the second feature extraction branch is also divided into four stages to obtain four local feature maps with different resolutions , which is the output of the last first-resolution feature block in each second-resolution feature layer.

[0068] Specifically, the resolution level is obtained by connecting four second-resolution feature layers in sequence. Image Processing is performed, and the global feature map output by the first feature extraction branch at the same stage is combined to obtain four local feature maps of different resolutions. In the first second resolution feature layer, the input feature map is processed sequentially by three second resolution feature blocks, and in the remaining second resolution feature layers, the input feature map is processed sequentially by two second resolution feature blocks. The calculation process of each second resolution feature is expressed as follows:

[0069] ;

[0070] In the formula, represents the global features extracted at the same stage in the first feature extraction branch, The output features of the previous second resolution feature block are used to guide the high-resolution local feature extraction by utilizing the global information extracted at low resolution.

[0071] S302, using channel attention to fuse the local feature map and the global feature map with equal resolution to generate four conditional fusion features. The conditional fusion feature is expressed as:

[0072] ;

[0073] In the formula, represents global average pooling, represents the Sigmoid activation function, FC represents the fully connected layer, i Represents the stage.

[0074] S4, combining the resolution level and diffusion time step, the sub-regional noise image with a resolution level of n and a time step of t , and input into the U-NET network step by step for processing. During the processing, it is spliced ​​with the conditional fusion features of the same resolution to obtain the noise prediction value.

[0075] The U-NET network consists of 4 encoders, a bottleneck layer and 4 decoders connected in sequence. Each encoder has 3 encoding blocks and each decoder has 3 decoding blocks. The encoder sizes are 64×64, 32×32, 16×16, and 8×8, the bottleneck layer size is 8×8, and the decoder sizes are 8×8, 16×16, 32×32, and 64×64. At the same time, the output of the encoder is concatenated with the conditional fusion features of the same resolution as the input of the decoder of the same size.

[0076] In addition, in order to make the U-Net network stage and time aware, before concatenating with the conditional fusion feature, as an implementation method, the resolution level n and the diffusion process time step t are fused into the conditional fusion feature through sinusoidal position encoding, which can be expressed as:

[0077] ;

[0078] Among them, t represents the time step of the diffusion process, d represents the dimension of the encoding, and j represents the dimension index. Based on this, the conditional fusion feature after embedding the sinusoidal position encoding is expressed as:

[0079] ;

[0080] Finally, the four resolutions The output of the decoder with the same resolution is concatenated for inverse diffusion of image generation to obtain the noise prediction value. .

[0081] S5. Noise-based prediction value And the region-by-region noise image Perform inverse diffusion to obtain the target image at the corresponding resolution level and diffusion time step .

[0082] Among them, the inverse diffusion process (denoising process) of the target area is expressed as:

[0083] ;

[0084] In the formula, represents the parameter that changes with time step t, and as t increases, It is also decreasing.

[0085] The inverse diffusion process (denoising process) of the background area is expressed as:

[0086] ;

[0087] Finally, the overall inverse diffusion process is expressed as:

[0088] ;

[0089] In the formula, , It is used to indicate which region each pixel i belongs to (target region or background region). If pixel i belongs to the target region, then =1, otherwise =0.

[0090] If the time step of the next diffusion process is 0, set n=n-1 and execute S3 until the resolution level is 1; otherwise, set t=t-1 and execute S4 until the final target image at the resolution level is generated.

[0091] In the reverse diffusion process, … , … , Step-by-step denoising, for , the model is generated first Effect diagram as conditional guidance Image generation: realizes image generation from low resolution to high resolution, from coarse to fine, and ensures the consistency of image generation.

[0092] As an implementation method, before executing S2, the above algorithm is trained, and the specific process is as follows:

[0093] Step 1: Obtain the existing data set and preprocess it, and divide the preprocessed data set into a test set, a validation set, and a training set according to a pre-planned ratio.

[0094] Since the size of image samples in the original dataset may be inconsistent, which will affect the feature extraction and subsequent learning process of the deep network model; therefore, it is necessary to normalize the size of the existing dataset. In addition, in order to maintain the correspondence between the target image and the original image in the image translation task, the original image and the target image should be stored in different folders, while ensuring that their file names remain consistent.

[0095] Step 2: Based on the preset loss function, perform the process described in S1-S4 on the training set for training until the preset number of iterations is reached; test with the test set and verify with the validation set.

[0096] The target area loss function is expressed as:

[0097] ;

[0098] The background area loss function is expressed as:

[0099] ;

[0100] In the formula, is the noise added to the background area at time step t during the forward diffusion process, is the noise predicted by the model.

[0101] The total loss function during network training can be defined as:

[0102] ;

[0103] in, and are the losses of the target area and background area at the nth resolution level, respectively. is a weighting factor used to balance the loss of the target area and background area loss During the learning process, the network will repeatedly perform back propagation training based on the loss function L, and the loss value will slowly decrease as the number of training rounds increases. When the loss value reaches the minimum value, the obtained network model is the best training result.

[0104] Embodiment 2

[0105] Combination Figure 3 , this embodiment discloses an image generation system based on step-by-step super-resolution and regional differentiation, including:

[0106] The super-resolution generation module is configured to: obtain the image to be processed and perform pre-processing to generate images of different resolution levels;

[0107] The regional differentiation module is configured to: divide the images at different resolution levels into regions, and add noise to the images based on the region division results to obtain the region-by-region noisy images; according to the resolution levels, input the images into the feature extraction branches step by step for processing, and generate conditional fusion features guided by the final target image generated at the previous resolution level;

[0108] The inverse diffusion module is configured as follows: in combination with the diffusion time step, the sub-regional noisy image is input into the U-NET network step by step for processing, and during the processing, it is spliced ​​with the conditional fusion features of the same resolution to obtain the noise prediction value; based on the noise prediction value and the sub-regional noisy image, inverse diffusion is performed to obtain the target image at the corresponding resolution level and diffusion time step, until the final target image is generated.

[0109] It should be noted that the super-resolution generation module, regional differentiation module and inverse diffusion module correspond to the steps in Embodiment 1, and the examples and application scenarios implemented by the modules and the corresponding steps are the same, but are not limited to the contents disclosed in Embodiment 1. It should be noted that the modules as part of the system can be executed in a computer system such as a set of computer executable instructions.

[0110] Embodiment 3

[0111] Embodiment 3 of the present invention provides an electronic device, including a memory and a processor, and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps of the above-mentioned image generation method based on gradual super-resolution and regional differentiation are completed.

[0112] Embodiment 4

[0113] Embodiment 4 of the present invention provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by a processor, the steps of the above-mentioned image generation method based on gradual super-resolution and regional differentiation are completed.

[0114] Embodiment 5

[0115] Embodiment 5 of the present invention provides a computer program product, including a computer program / instruction, which, when executed by a processor, implements the steps of the above-mentioned image generation method based on gradual super-resolution and regional differentiation.

[0116] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0117] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0118] These computer program instructions can also be loaded onto a computer or other programmable data processing device, and a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0119] The description of each embodiment in the above embodiments has different emphases. For parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0120] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An image generation method based on step-by-step super-resolution and regional differentiation, characterized in that: include: Obtain the image to be processed and perform preprocessing to generate images at different resolution levels; Divide images at different resolution levels into regions, and add noise to the images based on the region division results to obtain region-divided noisy images; According to the resolution level, the image is input into the feature extraction branch step by step for processing, and the final target image generated by the previous resolution level is used as a guide to generate conditional fusion features; Combined with the diffusion time step, the sub-regional noisy images are input into the U-NET network step by step for processing. During the processing, they are spliced ​​with the conditional fusion features of the same resolution to obtain the noise prediction value; Based on the noise prediction value and the noise-added image in each region, reverse diffusion is performed to obtain the target image at the corresponding resolution level and diffusion time step until the final target image is generated; According to the resolution level, the image is input into the feature extraction branch step by step for processing, and the final target image generated by the previous resolution level is used as a guide to generate conditional fusion features, which specifically includes: The final target image generated at the previous resolution level is input into the first feature extraction branch for processing to obtain a plurality of global feature maps of different resolutions; the image at the current resolution level is input into the second feature extraction branch for processing to obtain a plurality of local feature maps of different resolutions with the global feature map as a guide; Channel attention is used to fuse local feature maps with equal resolution and global feature maps to generate multiple conditional fusion features; The first feature extraction branch includes a plurality of first resolution feature layers connected in sequence, the first resolution feature layer includes a patch merging block and a plurality of first resolution feature blocks; The second feature extraction branch includes a plurality of second resolution feature layers connected in sequence, and the second resolution feature layer includes a plurality of second resolution feature blocks.

2. The image generation method based on step-by-step super-resolution and regional differentiation according to claim 1, characterized in that: The step of adding noise to the image based on the region division result specifically includes: replacing the background region of the image with Gaussian noise.

3. The image generation method based on step-by-step super-resolution and regional differentiation according to claim 1, characterized in that: The conditional fusion features are fused with resolution levels and diffusion process time steps through sinusoidal coding.

4. The image generation method based on step-by-step super-resolution and regional differentiation according to claim 1, characterized in that: The preprocessing of the image to be processed is specifically: processing the image to be processed by a bilinear interpolation method to obtain images of different resolution levels.

5. Image generation system based on step-by-step super-resolution and regional differentiation, characterized in that: include: The super-resolution generation module is configured to: obtain the image to be processed and perform pre-processing to generate images of different resolution levels; The regional differentiation module is configured to: divide the images at different resolution levels into regions, and add noise to the images based on the region division results to obtain the region-by-region noisy images; according to the resolution levels, input the images into the feature extraction branches step by step for processing, and generate conditional fusion features guided by the final target image generated at the previous resolution level; The inverse diffusion module is configured as follows: combining the diffusion time step, the sub-region noise-added image is input into the U-NET network step by step for processing, and during the processing, it is spliced ​​with the conditional fusion features of the same resolution to obtain the noise prediction value; Based on the noise prediction value and the noise-added image in each region, reverse diffusion is performed to obtain the target image at the corresponding resolution level and diffusion time step until the final target image is generated; According to the resolution level, the image is input into the feature extraction branch step by step for processing, and the final target image generated by the previous resolution level is used as a guide to generate conditional fusion features, which specifically includes: The final target image generated at the previous resolution level is input into the first feature extraction branch for processing to obtain a plurality of global feature maps of different resolutions; the image at the current resolution level is input into the second feature extraction branch for processing to obtain a plurality of local feature maps of different resolutions with the global feature map as a guide; Channel attention is used to fuse local feature maps with equal resolution and global feature maps to generate multiple conditional fusion features; The first feature extraction branch includes a plurality of first resolution feature layers connected in sequence, the first resolution feature layer includes a patch merging block and a plurality of first resolution feature blocks; The second feature extraction branch includes a plurality of second resolution feature layers connected in sequence, and the second resolution feature layer includes a plurality of second resolution feature blocks.

6. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the image generation method based on step-by-step super-resolution and regional differentiation according to any one of claims 1 to 4.

7. A computer-readable storage medium having a computer program / instruction stored thereon, characterized in that: When the computer program / instruction is executed by a processor, the steps of the image generation method based on step-by-step super-resolution and regional differentiation described in any one of claims 1 to 4 are implemented.

8. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the image generation method based on step-by-step super-resolution and regional differentiation described in any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Remote sensing image super-resolution reconstruction method and system based on diffusion model

    CN118735785A

  • Pathological section image super-resolution method based on diffusion model

    CN118967449A