Image generation method, image generation network, electronic device and medium

By generating images based on the text content input by the user and the diffusion model, combined with the lighting texture joint perception module, the problem that image generation in the prior art does not meet user needs is solved, and image generation of rich lighting and texture features is achieved.

CN120599095APending Publication Date: 2025-09-05ACADEMY OF BROADCASTING SCI STATE ADMINISTATION OF PRESS PUBLICATION RADIO FILM & TELEVISION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510585605.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

The existing image generation technology cannot meet users' needs for image texture features and lighting features, and the generated images are not rich enough and do not meet user needs.

Method used

By generating an initial image based on the text content input by the user, and generating a candidate image set using the diffusion model, combining the lighting texture joint perception module and a multi-layer image generation network, the images are gradually adjusted to meet user needs.

Benefits of technology

It realizes the generation of images that meet user needs, with rich lighting characteristics and texture characteristics, and meets the user's image generation needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120599095A_ABST
    Figure CN120599095A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image generation, in particular to an image generation method, an image generation network, electronic equipment and a medium. The method comprises the steps of generating an initial image based on text content which is input by a user and is used for indicating an image generation requirement; inputting the initial image into a diffusion model to generate a first candidate image set; inputting each first candidate image and the random seed graph into a diffusion model to generate a second candidate image set in response to feedback information that the first candidate images in the first candidate image set fed back by the user do not meet the image generation requirement; wherein the number of the images in the second candidate image set is greater than the number of the images in the first candidate image set; and outputting a target image according to the second candidate image set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image generation technology, and more specifically, to an image generation method, an image generation network, an electronic device, and a readable storage medium. Background Art

[0002] In recent years, image generation technology has developed rapidly, and its applications have penetrated into multiple fields. Traditional image generation techniques typically rely on physical models, such as light tracing and global illumination algorithms. However, these technologies cannot meet user requirements for image generation. For example, the generated images lack rich texture features, and the lighting characteristics do not meet user requirements.

[0003] Therefore, there is an urgent need to provide a new image generation method to meet the image generation needs of users. Summary of the Invention

[0004] In view of this, the present disclosure proposes an image generation method, an image generation network, an electronic device, and a readable storage medium, which can generate images that meet user needs.

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an image generation method, comprising:

[0006] generating an initial image based on text content input by a user for indicating image generation requirements;

[0007] Inputting the initial image into a diffusion model to generate a first candidate image set;

[0008] In response to user feedback that none of the first candidate images in the first candidate image set meets the image generation requirement, inputting each of the first candidate images and a random seed graph into the diffusion model to generate a second candidate image set; wherein the number of images in the second candidate image set is greater than the number of images in the first candidate image set;

[0009] Output a target image based on the second candidate image set.

[0010] Optionally, the diffusion model includes an encoder, a noise image generation module, a bottleneck structure and a decoder.

[0011] The encoder is used to encode an input image into a first feature vector of the input image, where the first feature vector is used to represent semantic information of the input image;

[0012] The bottleneck structure is used to receive the first feature vector and extract the illumination feature and texture feature in the first feature vector to obtain a target feature vector;

[0013] The decoder is configured to receive the target feature vector and the noise image generated by the noise image generation module at a current time step, predict the noise added to the input image at each time step before the current time step, and generate the candidate image corresponding to the input image by removing the noise corresponding to each time step from the noise image according to the target feature vector;

[0014] The input image includes an initial image, and the candidate images include the first candidate image; or the input image includes the first candidate image and the random seed image, and the candidate images include the second candidate image.

[0015] Optionally, the bottleneck structure includes an illumination and texture joint perception module, which includes an illumination feature extraction submodule, a texture feature extraction submodule and a fusion submodule. The illumination feature extraction submodule is used to receive the first feature vector, extract illumination features from the first feature vector to obtain an illumination feature vector, the texture feature extraction submodule is used to receive the first feature vector, extract texture features from the first feature vector to obtain a texture feature vector, and the fusion submodule is used to fuse the illumination feature vector and the texture feature vector to obtain a fused feature vector; wherein, the fused feature vector is used to generate the target feature vector.

[0016] Optionally, the bottleneck structure includes multiple illumination and texture joint perception modules, and the illumination and texture joint perception modules are connected sequentially.

[0017] Optionally, the encoder is further configured to downsample the input image to obtain a first feature map of the input image, and obtain the first feature vector based on the first feature map; the decoder is further configured to upsample the received noisy image to obtain a second feature map of the noisy image, and obtain the candidate image based on the second feature map;

[0018] The spatial dimension of the first feature map is lower than the spatial dimension of the input image, and the spatial dimension of the second feature map is higher than the spatial dimension of the noise image.

[0019] Optionally, the number of downsampling times of the encoder and the number of upsampling times of the decoder are determined according to the resolution of the initial image, and the resolution of the initial image is determined according to the text content required for image generation.

[0020] Optionally, generating an initial image based on the text content includes:

[0021] Based on the text content, obtaining a text encoding vector corresponding to the text content;

[0022] The text encoding vector is input into the image generation model to generate an initial image.

[0023] According to a second aspect of an embodiment of the present disclosure, there is provided an image generation network, comprising:

[0024] A first image generation module is configured to receive text content input by a user indicating an image generation requirement, generate an initial image based on the text content, and input the initial image into a diffusion model to generate a first candidate image set;

[0025] a second image generation module, configured to, in response to user feedback indicating that none of the first candidate images in the first candidate image set meets the image generation requirement, input each of the first candidate images and a random seed graph into the diffusion model to generate a second candidate image set; wherein the number of images in the second candidate image set is greater than the number of images in the first candidate image set;

[0026] The target image output module is configured to output a target image according to the second candidate image set.

[0027] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device including a memory and a processor.

[0028] The memory is used to store computer instructions, and the processor is used to call the computer instructions from the memory to execute the method as described in any one of the first aspects.

[0029] According to a fourth aspect of an embodiment of the present disclosure, a non-volatile computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method as described in any one of the first aspects is implemented.

[0030] According to an embodiment of the present disclosure, an initial image is first generated based on the text content input by the user to indicate the image generation requirements. The initial image is input into a diffusion model to generate a first set of candidate images, allowing the user to determine whether there is a first candidate image that meets their image generation requirements in the first set of candidate images. If the user determines that there is no first candidate image that meets their image generation requirements in the first set of candidate images, in response to user feedback, each first candidate image and a random seed graph are input into the diffusion model to generate a second set of candidate images, and a target image is obtained based on the second set of candidate images. Through the multi-layer image generation network architecture of this embodiment, it is possible to gradually obtain a target image that meets the requirements from the text content of the image generation requirements input by the user.

[0031] Other features and advantages of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0033] Figure 1 It is a flowchart of an image generation method provided by an embodiment of the present disclosure.

[0034] Figure 2 It is a schematic diagram of the composition structure of a diffusion model provided by an embodiment of the present disclosure.

[0035] Figure 3 It is a schematic diagram of the composition structure of an image generation network provided by an embodiment of the present disclosure.

[0036] Figure 4 It is a schematic diagram of the composition structure of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION

[0037] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that unless otherwise specifically stated, the relative arrangement of components and steps, numerical expressions and numerical values ​​set forth in these embodiments do not limit the scope of the present disclosure.

[0038] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the present disclosure, its application, or uses.

[0039] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.

[0040] In all examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.

[0041] It should be noted that like reference numerals and letters refer to like items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.

[0042] It should be noted that the collection, storage, use, processing, transmission, provision, disclosure, and deletion of data involved in this disclosure are all carried out in compliance with the relevant data protection laws and policies of the country or region where they are located, and with the full authorization of the corresponding data owner.

[0043] In recent years, image generation technology has developed rapidly, and its applications have penetrated into multiple fields. Traditional image generation techniques typically rely on physical models, such as light tracing and global illumination algorithms. However, these technologies cannot meet user requirements for image generation. For example, the generated images lack rich texture features, and the lighting characteristics do not meet user requirements.

[0044] Based on this, the embodiments of the present disclosure provide a new image generation solution, which can gradually obtain a target image that meets the requirements based on the text content of the image generation requirements input by the user.

[0045] Figure 1 The image generation method provided by the embodiment of the present disclosure is shown. Figure 1 As shown, the method may include steps S110 to S140.

[0046] Step S110 : generating an initial image based on the text content input by the user for indicating image generation requirements.

[0047] In this embodiment, the text content entered by the user to indicate the image generation requirement may vary depending on the type of image generation task. The text content may be descriptive text content. As an example, the text content may only include the main content of the target image, such as "a tiger walking in the forest." The text content may also include the background environment of the target image, such as "on a sunny beach." The text content may also include the style of the target image, such as "realistic style." As another example, the text content may also include a detailed description of the target image, such as "blue parasols and beach chairs on the beach." The text content may also include the emotion or atmosphere that the target image is intended to convey, such as "a warm family scene." As another example, the text content may also include constraints, such as "generate a 1028*1028 resolution image." The text content may also be a collection of several keywords, such as "forest scene, sunlight filtering through leaves, fallen leaves and small animals on the ground, realistic style."

[0048] In some examples, step S110 may include steps S111 to S112.

[0049] Step S111 : obtaining a text encoding vector based on the text content input by the user for indicating image generation requirements.

[0050] As an example, the text content input by the user to indicate the image generation requirement can be preprocessed to extract keywords (such as theme, scene, style, details, etc.) from the text content. The preprocessed keywords are encoded to obtain a text encoding vector.

[0051] The method for encoding the preprocessed keywords to obtain a text encoding vector may include: using a pretrained word embedding model to map the keywords into a high-dimensional vector space to obtain a keyword sequence; and encoding the keyword sequence using a text encoder to generate a text encoding vector. As an example, the word embedding model may be a BERT model. The text encoder may be a recurrent neural network (RNN).

[0052] Step S112: input the text encoding vector into the image generation model to generate an initial image.

[0053] As an example, the image generation model can be any one of a generative adversarial network, a variational molecular encoder, and a diffusion model.

[0054] Step S120: input the initial image into the diffusion model to generate a first candidate image set.

[0055] In this embodiment, the diffusion model may include an encoder, a noise image generation module, and a decoder. The encoder is configured to encode an initial image into a first feature vector of the initial image, which is used to represent the semantic information of the initial image. The decoder is configured to receive the first feature vector and the noise image generated by the noise image generation module at the current time step, predict the noise added to the initial image at each time step prior to the current time step, remove the noise corresponding to each time step from the noise image through an inverse diffusion process, and generate multiple first candidate images corresponding to the initial image. These first candidate images constitute a first candidate image set.

[0056] In some examples, only the initial image may be input into the diffusion model, and a system random source (such as clock time) during the operation of the diffusion model may be used to generate noise, thereby generating a first set of candidate images.

[0057] In other examples, the initial image and a first random seed map can be input into the diffusion model to generate a first set of candidate images. The random seed map is an image generated based on a specific random seed. By introducing the random seed, controllable randomness is added to the image generation process.

[0058] Step S130 , in response to user feedback that none of the first candidate images in the first candidate image set meets the image generation requirement, each first candidate image and the random seed graph are input into a diffusion model to generate a second candidate image set.

[0059] The number of images in the second candidate image set is greater than the number of images in the first candidate image set.

[0060] In this embodiment, it is possible that none of the first candidate images in the first candidate image set meet the image generation requirements. For example, the lighting characteristics or texture characteristics of each first candidate image may not meet the image generation requirements. After determining that none of the first candidate images in the first candidate image set meet the image generation requirements, the user may provide feedback through the interactive interface indicating that the image generation result does not meet the image generation requirements.

[0061] In this embodiment, in response to user feedback that none of the first candidate images in the first candidate image set meets the image generation requirements, each first candidate image and a random seed graph are input into the diffusion model to generate a second candidate image set containing multiple second candidate images.

[0062] In some examples, the random seed map input to the diffusion model together with the first candidate image may be a second random seed map. The second random seed map and the first random seed map may have different random seeds, which can increase the diversity of generated images.

[0063] Step S140: outputting a target image according to the second candidate image set.

[0064] In this embodiment, it is possible that none of the second candidate images in the second candidate image set meet the image generation requirements. After determining that none of the second candidate images in the second candidate image set meet the image generation requirements, the user can provide feedback through the interactive interface indicating that the image generation result does not meet the image generation requirements.

[0065] In this embodiment, step S140 may include: in response to user feedback that a second candidate image in the second candidate image set meets the image generation requirement, taking the second candidate image selected by the user as the target image and outputting the target image.

[0066] Step S140 may also include: in response to user feedback that none of the second candidate images in the second candidate image set meet the image generation requirements, inputting each second candidate image and the third random seed graph into the diffusion model to generate a third candidate image set containing multiple third candidate images; and outputting the target image based on the third candidate image set.

[0067] The image generation method of this embodiment can be performed by at least two layers of image generation modules and a target image output module. The first layer of the image generation module receives text content input by a user to indicate image generation requirements, generates an initial image based on the text content, and inputs the initial image into a diffusion model to generate a first set of candidate images. The second layer of the image generation module, in response to user feedback indicating that none of the first candidate images in the first set of candidate images meet the image generation requirements, inputs each first candidate image and a random seed map into the diffusion model to generate a second set of candidate images. The target image output module outputs a target image based on the second set of candidate images.

[0068] The embodiment of the present disclosure first generates an initial image based on the text content input by the user to indicate the image generation requirements, inputs the initial image into a diffusion model, and generates a first candidate image set for the user to determine whether there is a first candidate image that meets the image generation requirements in the first candidate image set; if the user determines that there is no first candidate image that meets the image generation requirements in the first candidate image set, in response to the user's feedback information, each first candidate image and a random seed map are input into the diffusion model to generate a second candidate image set, and a target image is obtained based on the second candidate image set. Through the multi-layer image generation network architecture of this embodiment, it is possible to gradually obtain a target image that meets the requirements from the text content of the image generation requirements input by the user.

[0069] In some embodiments, the user needs to generate a system with rich lighting features and texture features. Figure 2 As shown, the diffusion model 200 in this embodiment may include an encoder 210 , a noise image generation module 220 , a bottleneck structure 230 , and a decoder 240 .

[0070] The encoder 210 is used to encode the input image into a first feature vector of the input image, wherein the first feature vector is used to represent semantic information of the input image.

[0071] The noise image generation module 220 is configured to add noise to an input image to generate a noise image. In one example, random noise may be added to the input image to generate the noise image. In another example, a preset random seed map may be received and the random seed map may be integrated into the input image to generate the noise image.

[0072] The bottleneck structure 230 is used to receive the first feature vector and extract the illumination feature and texture feature in the first feature vector to obtain the target feature vector.

[0073] In some examples, the bottleneck structure 230 may include an illumination and texture joint perception module, which may include an illumination feature extraction submodule, a texture feature extraction submodule, and a fusion submodule.

[0074] The illumination feature extraction submodule is used to receive the first feature vector and extract illumination features from the first feature vector to obtain an illumination feature vector. The illumination feature extraction submodule can be, for example, a spherical harmonic illumination feature extraction block that implements fast illumination calculation and rendering based on spherical harmonics.

[0075] The texture feature extraction submodule is configured to receive the first feature vector and extract texture features from the first feature vector to obtain a texture feature vector. The texture feature extraction submodule may, for example, be a Gaussian texture feature extraction block that smooths the first feature vector using a Gaussian filter to extract the texture feature vector. By selecting different Gaussian kernel parameters, texture information at different scales can be extracted.

[0076] The fusion submodule is used to fuse the illumination feature vector and the texture feature vector to obtain a fused feature vector. The fused feature vector is used to generate the target feature vector. As an example, fusing the illumination feature vector and the texture feature vector may include performing a weighted operation on the illumination feature vector and the texture feature vector.

[0077] The decoder 240 is used to receive the target feature vector and the noise image generated by the noise image generation module 220 at the current time step, and predict the noise added to the input image at each time step before the current time step, remove the noise corresponding to each time step from the noise image according to the target feature vector, and generate a candidate image corresponding to the input image.

[0078] In this embodiment, the input image includes an initial image, and the candidate images include a first candidate image; or, the input image includes the first candidate image and a random seed image, and the candidate images include a second candidate image.

[0079] In some examples, the target feature vectors can be passed directly as conditional inputs to each layer of the decoder, and the decoder incorporates these target feature vectors both when predicting the noise added to the input image at each time step and when generating the image.

[0080] The diffusion model of this embodiment can generate an image with rich illumination features and texture features that meets the user's image generation requirements.

[0081] In some embodiments, the bottleneck structure 230 may include multiple illumination and texture joint perception modules, and each illumination and texture joint perception module is connected in sequence. In some examples, the bottleneck structure 230 may include four illumination and texture joint perception modules, namely a first illumination and texture joint perception module, a second illumination and texture joint perception module, a third illumination and texture joint perception module, and a fourth illumination and texture joint perception module. Among them, the first illumination and texture joint perception module is used to receive the first feature vector and output a first fused feature vector. The second illumination and texture joint perception module is used to receive the first fused feature vector output by the first illumination and texture joint perception module and output a second fused feature vector. The third illumination and texture joint perception module is used to receive the second fused feature vector output by the second illumination and texture joint perception module and output a third fused feature vector. The fourth illumination and texture joint perception module is used to receive the third fused feature vector output by the third illumination and texture joint perception module and output a target feature vector.

[0082] By sequentially connecting multiple illumination and texture joint perception modules, a balance can be achieved between extracting complete and detailed illumination and texture features and achieving feature extraction efficiency. Furthermore, the illumination and texture joint perception module can be used to perform an implicit rendering process to achieve the purpose of controlling illumination.

[0083] In some embodiments, the encoder 210 may further be configured to downsample the input image to obtain a first feature map of the input image, and to obtain a first feature vector based on the first feature map. The decoder 240 may further be configured to upsample the received noisy image to obtain a second feature map of the noisy image, and to obtain a candidate image based on the second feature map. The spatial dimension of the first feature map is lower than that of the input image, and the spatial dimension of the second feature map is higher than that of the noisy image.

[0084] In some examples, the encoder 210 may be further configured to perform bicubic downsampling on the input image to obtain a first feature map of the input image, and to obtain a first feature vector based on the first feature map. The decoder 240 may be further configured to perform bicubic upsampling on the received noisy image to obtain a second feature map of the noisy image, and to obtain a candidate image based on the second feature map.

[0085] In this embodiment, the encoder downsamples the input image to reduce the spatial dimension of the image, extracting high-level semantic features from the low-dimensional image, thereby reducing computational complexity. The decoder upsamples the received noisy image to gradually add details to the low-dimensional image, ultimately generating a high-resolution candidate image.

[0086] The number of downsampling cycles for the encoder and upsampling cycles for the decoder can be determined based on the resolution of the original image. In some examples, the number of downsampling cycles for the encoder and upsampling cycles for the decoder can both be two. By appropriately downsampling and upsampling, a good balance can be achieved between the richness of semantic features and computational complexity.

[0087] In some examples, an image generation network may include a first-layer image generation module, a second-layer image generation module, a third-layer image generation module, and a target image output module. In the second-layer image generation module, the first candidate image and the first random seed map output by the decoder in the first-layer image generation module are used as inputs to the encoder. The fused image, formed by fusing the first candidate image and the random seed map, is bicubic downsampled twice. The decoder then bicubic upsamples the fused image twice, and the second candidate image output by the decoder is used as input to the third-layer image generation module. In this example, since the first-layer image generation module does not have an available bicubic upsampled output image as input, only four bicubic downsampling steps are used as input.

[0088] The present disclosure also provides an image generation network, such as Figure 3 As shown, the image generation network 300 may include:

[0089] The first image generation module 310 is configured to receive text content input by a user indicating an image generation requirement, generate an initial image based on the text content, and input the initial image into a diffusion model to generate a first candidate image set;

[0090] A second image generation module 320 is configured to, in response to user feedback indicating that none of the first candidate images in the first candidate image set meet the image generation requirements, input each first candidate image and the random seed graph into the diffusion model to generate a second candidate image set; wherein the number of images in the second candidate image set is greater than the number of images in the first candidate image set;

[0091] The target image output module 330 is configured to output a target image according to the second candidate image set.

[0092] For the specific implementation of each module in the image generation network in this embodiment, reference can be made to the description in the aforementioned embodiments of the present disclosure, which will not be repeated here.

[0093] According to an embodiment of the present disclosure, an initial image is first generated based on the text content input by the user to indicate the image generation requirements. The initial image is input into a diffusion model to generate a first set of candidate images, allowing the user to determine whether there is a first candidate image that meets their image generation requirements in the first set of candidate images. If the user determines that there is no first candidate image that meets their image generation requirements in the first set of candidate images, in response to user feedback, each first candidate image and a random seed graph are input into the diffusion model to generate a second set of candidate images, and a target image is obtained based on the second set of candidate images. Through the multi-layer image generation network architecture of this embodiment, it is possible to gradually obtain a target image that meets the requirements from the text content of the image generation requirements input by the user.

[0094] The present disclosure also provides an electronic device. Figure 4 As shown, the electronic device 1000 may include a memory 1010 and a processor 1020. The memory 1010 may be used to store computer instructions, and the processor 1020 may be used to call computer instructions from the memory 1010 to perform all or part of the steps of any of the methods described in the aforementioned embodiments of the present disclosure. It should be noted that the processor 1020 may include one or more processors to execute instructions, and the memory 1010 may also include one or more memories to store computer instructions. In one example, the electronic device 1000 may be a cloud server.

[0095] The present disclosure also provides a non-volatile computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the method described in any of the aforementioned embodiments of the present disclosure. Optionally, the computer-readable storage medium may be a non-transitory storage medium, but is not limited thereto and may also be a transient storage medium.

[0096] An embodiment of the present disclosure further provides a computer program product, which may include a computer program. When the computer program is executed by a processor, any method in the aforementioned embodiments of the present disclosure may be implemented.

[0097] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement any of the methods in the aforementioned embodiments of the present disclosure.

[0098] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.

[0099] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.

[0100] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, and conventional procedural programming languages ​​such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.

[0101] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.

[0102] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0103] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.

[0104] The flowcharts and block diagrams in the accompanying drawings show the possible implementation architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment or part of an instruction, and the module, program segment or part of the instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that implementation by hardware, implementation by software, and implementation by a combination of software and hardware are all equivalent.

[0105] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terms used herein are chosen to best explain the principles of the embodiments, practical applications, or technical improvements to technologies in the marketplace, or to enable other persons skilled in the art to understand the embodiments disclosed herein. The scope of the present disclosure is defined by the appended claims.

Claims

1. An image generation method, characterized in that: include: generating an initial image based on text content input by a user for indicating image generation requirements; Inputting the initial image into a diffusion model to generate a first candidate image set; In response to user feedback that none of the first candidate images in the first candidate image set meets the image generation requirement, inputting each of the first candidate images and a random seed graph into the diffusion model to generate a second candidate image set; wherein the number of images in the second candidate image set is greater than the number of images in the first candidate image set; Output a target image based on the second candidate image set.

2. The method according to claim 1, characterized in that The diffusion model includes an encoder, a noise image generation module, a bottleneck structure and a decoder. The encoder is used to encode an input image into a first feature vector of the input image, where the first feature vector is used to represent semantic information of the input image; The bottleneck structure is used to receive the first feature vector and extract the illumination feature and texture feature in the first feature vector to obtain a target feature vector; The decoder is configured to receive the target feature vector and the noise image generated by the noise image generation module at a current time step, predict the noise added to the input image at each time step before the current time step, and generate the candidate image corresponding to the input image by removing the noise corresponding to each time step from the noise image according to the target feature vector; The input image includes an initial image, and the candidate images include the first candidate image; or the input image includes the first candidate image and the random seed image, and the candidate images include the second candidate image.

3. The method according to claim 2, characterized in that The bottleneck structure includes an illumination and texture joint perception module, which includes an illumination feature extraction submodule, a texture feature extraction submodule and a fusion submodule. The illumination feature extraction submodule is used to receive the first feature vector, extract illumination features from the first feature vector to obtain an illumination feature vector, the texture feature extraction submodule is used to receive the first feature vector, extract texture features from the first feature vector to obtain a texture feature vector, and the fusion submodule is used to fuse the illumination feature vector and the texture feature vector to obtain a fused feature vector; wherein, the fused feature vector is used to generate the target feature vector.

4. The method according to claim 3, characterized in that The bottleneck structure includes multiple illumination and texture joint perception modules, and each of the illumination and texture joint perception modules is connected in sequence.

5. The method according to claim 2, characterized in that The encoder is further configured to downsample the input image to obtain a first feature map of the input image, and obtain the first feature vector based on the first feature map. The decoder is further configured to upsample the received noisy image to obtain a second feature map of the noisy image, and obtain the candidate image based on the second feature map. The spatial dimension of the first feature map is lower than the spatial dimension of the input image, and the spatial dimension of the second feature map is higher than the spatial dimension of the noise image.

6. The method according to claim 5, characterized in that The number of downsampling times of the encoder and the number of upsampling times of the decoder are determined according to the resolution of the initial image, and the resolution of the initial image is determined according to the text content required for image generation.

7. The method according to claim 1, characterized in that Generating an initial image based on the text content includes: Based on the text content, obtaining a text encoding vector corresponding to the text content; The text encoding vector is input into the image generation model to generate an initial image.

8. An image generation network, characterized in that include: A first image generation module is configured to receive text content input by a user indicating an image generation requirement, generate an initial image based on the text content, and input the initial image into a diffusion model to generate a first candidate image set; a second image generation module, configured to, in response to user feedback indicating that none of the first candidate images in the first candidate image set meets the image generation requirement, input each of the first candidate images and a random seed graph into the diffusion model to generate a second candidate image set; wherein the number of images in the second candidate image set is greater than the number of images in the first candidate image set; The target image output module is configured to output a target image according to the second candidate image set.

9. An electronic device, characterized in that: including memory and processor, The memory is used to store computer instructions, and the processor is used to call the computer instructions from the memory to execute the method according to any one of claims 1 to 7.

10. A non-volatile computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.