Image generation method and apparatus thereof

The trained diffusion model processes noise images with weighted text descriptions to generate images that align with user preferences, addressing the lack of effective methods for content emphasis or avoidance in image generation.

CN116245749BActive Publication Date: 2025-07-15BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211691456.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-27
Publication Date
2025-07-15
Estimated Expiration
2042-12-27

AI Technical Summary

Technical Problem

When the prior art lacks effective means to generate images, it cannot meet the user's need to strengthen or avoid different description contents.

Method used

By acquiring text description information, its weights and noise images, denoising processing is performed using the diffusion model to generate a target image related to text description information.

Benefits of technology

This achieves the image generation of specific descriptive content based on user needs, and improves the image content generation effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116245749B_ABST
    Figure CN116245749B_ABST
Patent Text Reader

Abstract

The present disclosure provides an image generation method and apparatus, which relate to the field of artificial intelligence, and specifically relate to natural language processing and deep learning technologies. The specific implementation solution is as follows: obtain at least one text description information and the weight of each text description information; obtain a noise image to be processed; perform denoising processing on the noise image according to the noise image, at least one text description information, and the weight of each text description information, and generate a target image corresponding to the at least one text description information. The present disclosure supports inputting one or more text descriptions and assigning different weights, and generating an image by integrating multiple text descriptions with different weights, so as to meet the user's needs for strengthening or avoiding different description contents, thereby improving the image content generation effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence, specifically to natural language processing and deep learning technologies, and particularly to an image generation method, apparatus, electronic device, and storage medium. Background Art

[0002] The text-to-image task refers to generating an image based on an input sentence. When a user needs to generate an image with complex content, they often need to focus on or avoid certain parts in the text description. For example, in the text description "When the sun is setting, there are huge clouds on the horizon, the sea is rough, and a whale jumps out of the sea", the user hopes to highlight the part "a whale jumps out of the sea"; another example is when the user wants to generate "a huge whale flying in the sky", they do not want to see images related to "sea water".

[0003] However, there is currently a lack of effective means for generating images to meet the user's needs for enhancing or avoiding different description contents. Summary of the Invention

[0004] The present disclosure provides an image generation method, apparatus, electronic device, and storage medium.

[0005] According to a first aspect of the present disclosure, there is provided an image generation method, including:

[0006] Obtaining at least one text description information and the weight of each said text description information;

[0007] Obtaining a noise image to be processed;

[0008] Performing denoising processing on the noise image according to the noise image, the at least one text description information, and the weight of each said text description information to generate a target image corresponding to the at least one text description information.

[0009] According to a second aspect of the present disclosure, there is provided an image generation apparatus, including:

[0010] A first obtaining module, configured to obtain at least one text description information and the weight of each said text description information;

[0011] A second obtaining module, configured to obtain a noise image to be processed;

[0012] A generating module, configured to perform denoising processing on the noise image according to the noise image, the at least one text description information, and the weight of each said text description information to generate a target image corresponding to the at least one text description information.

[0013] According to a third aspect of the present disclosure, there is provided an electronic device, including:

[0014] at least one processor; and

[0015] a memory communicatively connected to the at least one processor; wherein

[0016] the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, enable the at least one processor to execute the method described in the foregoing first aspect.

[0017] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are for causing the computer to execute the method described in the foregoing first aspect.

[0018] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in the foregoing first aspect.

[0019] The technology according to the present disclosure can support inputting one or more text descriptions and assigning different positive or negative weights, and generating images by integrating multiple text descriptions with different weights, so as to meet the needs of users to strengthen or avoid different description contents, thereby improving the image content generation effect.

[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0022] Figure 1 is a flowchart of an image generation method provided by an embodiment of the present disclosure;

[0023] Figure 2 is an example diagram of a diffusion model process provided by an embodiment of the present disclosure;

[0024] Figure 3 is a flowchart of another image generation method provided by an embodiment of the present disclosure;

[0025] Figure 4 is an example diagram of yet another image generation method provided by an embodiment of the present disclosure;

[0026] Figure 5 is a structural block diagram of an image generation device provided by an embodiment of the present disclosure;

[0027] Figure 6 Block diagram of another image generation device provided by an embodiment of the present disclosure;

[0028] Figure 7 Block diagram of an electronic device for implementing an image generation method provided by an embodiment of the present disclosure. Detailed implementation manners

[0029] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0030] In the description of the present disclosure, unless otherwise specified, "and / or" is merely an association relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. These three situations.

[0031] The terms used in the embodiments of the present disclosure are merely for the purpose of describing specific embodiments, and are not intended to limit the embodiments of the present disclosure. The singular forms "a" and "the" used in the embodiments of the present disclosure and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0032] It should be understood that although the terms first, second, third, etc. may be used in the embodiments of the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the embodiments of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "when" as used herein may be interpreted as "when" or "when" or "in response to a determination".

[0033] Next, an image generation method, device, electronic device, and storage medium according to an embodiment of the present disclosure are described with reference to the accompanying drawings.

[0034] Figure 1 Flowchart of an image generation method provided by an embodiment of the present disclosure. As Figure 1 shown, the method may include, but is not limited to, the following steps.

[0035] In step 101, obtain at least one text description information and the weight of each text description information.

[0036] In a possible implementation, the present disclosure provides a text input interface for a user, through which the user can input at least one text description information and the weight set by the user for each text description information. In an embodiment of the present disclosure, at least one text description information and the weight of each text description information can be obtained through this text input interface. Herein, the "at least one" can be understood as one or more, and multiple can be understood as two or more, such as two, three, etc.

[0037] In another possible implementation, the present disclosure can provide a voice input interface for a user, through which the user can input voice description information. The present disclosure can obtain the voice description information through this voice input interface, perform speech recognition on the voice description information to obtain the corresponding text description information, and obtain the weight corresponding to each text description information by the user.

[0038] It should be noted that the present disclosure can also obtain text description information and the weight of each text description information in other ways, which will not be elaborated herein.

[0039] In an embodiment of the present disclosure, the above weight α i represents the degree of emphasis of the user on the corresponding text description y i When the weight α i is a positive value, it means that the user hopes to strengthen y i this description; when the weight α i is a negative value, it means that the user needs to avoid y i this description.

[0040] In step 102, a noise image to be processed is obtained.

[0041] Herein, in an embodiment of the present disclosure, the size of the noise image to be processed is the same as the size of the target image.

[0042] In a possible implementation, a Gaussian noise image can be generated and used as the noise image to be processed.

[0043] In step 103, based on the noise image, at least one text description information, and the weight of each text description information, the noise image is denoised to generate a target image corresponding to at least one text description information.

[0044] In an embodiment of the present disclosure, according to a noise image, at least one text description information, and the weight of each text description information, a diffusion model may be used to perform denoising processing on the noise image to generate a target image corresponding to the at least one text description information. Wherein, the correlation between the image content of the target image and the text content of the at least one text description information is greater than a certain threshold. Wherein, the diffusion model is a model that has been pre-trained, and the diffusion model may be an existing diffusion model. The present disclosure does not specifically limit the network structure of the diffusion model and will not elaborate further.

[0045] It should be noted that, in an embodiment of the present disclosure, the diffusion model is a type of generative model that converts Gaussian noise into samples of a known data distribution through an iterative denoising process, and the generated pictures have good diversity and realism. As Figure 2 shown, two processes of the diffusion model are presented. Among them, from right to left (from X0 to X T ) represents the forward process or the diffusion process, and from left to right (from X T to X0) represents the reverse process. The diffusion process gradually adds Gaussian noise to the original image and is a fixed Markov chain process. Finally, the image is gradually transformed into a Gaussian noise. The reverse process, on the other hand, restores the original image step by step through denoising, thereby realizing image generation.

[0046] By implementing the embodiments of the present disclosure, it is possible to support inputting one or more text descriptions and assigning different weights, and generating an image by integrating multiple text descriptions with different weights, thereby meeting the user's needs for strengthening or avoiding different description contents, and thus improving the image content generation effect.

[0047] Figure 3 The flowchart of another image generation method provided by the embodiment of the present disclosure is shown as Figure 3 follows. The method may include, but is not limited to, the following steps.

[0048] In step 301, obtain at least one text description information and the weight of each text description information.

[0049] In an embodiment of the present disclosure, the implementation manner of step 301 may be implemented by any one of the embodiments of the present disclosure respectively, and no limitation is made here and no further elaboration is provided.

[0050] In step 302, obtain the noise image to be processed.

[0051] In an embodiment of the present disclosure, the implementation manner of step 301 may be implemented by any one of the embodiments of the present disclosure respectively, and no limitation is made here and no further elaboration is provided.

[0052] In step 303, the noise image, at least one text description information, and the weight of each text description information are input into the diffusion model.

[0053] In step 304, starting from the noise image through the diffusion model, denoising the noise image with each text description information as a condition to obtain the first predicted image of each text description information, and denoising the noise image with an empty text description information as a condition to obtain the second predicted image.

[0054] In step 305, a new noise image is obtained according to the first predicted image of each text description information, the weight of each text description information, and the second predicted image, and the above process is iterated for T steps.

[0055] In a possible implementation manner, the first predicted image of each text description information and the weight of each text description information can be subjected to weighted summation processing to obtain an intermediate predicted image, and based on the first formula, the difference between the intermediate predicted image and the second predicted image is determined as the new noise image.

[0056] Wherein, in the embodiments of the present disclosure, the first formula is expressed as follows:

[0057]

[0058] Wherein, x t-1 is the new noise image, is the intermediate predicted image, is the second predicted image, s is an adjustable hyperparameter, and the value range of s is 1 to 8.

[0059] Optionally, in some embodiments of the present disclosure, the weight of each text description information can be normalized. The first predicted image of each text description information and the corresponding normalized weight are subjected to weighted summation processing to obtain an intermediate predicted image. Based on the above formula (1), the difference between the intermediate predicted image and the second predicted image is determined as the new noise image, and the above process is iterated for T steps, that is, the new noise image, at least one text description information, and the weight of each text description information are input into the diffusion model for a new round of prediction until it ends after iterating for T steps.

[0060] In step 306, the noise image obtained after iterating for T steps is determined as the target image corresponding to at least one text description information.

[0061] That is to say, the noise image obtained after iterating for T steps can be determined as the target image generated by integrating at least one text description information.

[0062] For example, taking the number of text description information as multiple. For the target image to be generated, obtain multiple text description information y1, y2, ..., y1 input by the user. n , and the corresponding weights α1, α2, …, α n Among them, the weight α i Represents the user's description of the text y i The degree of emphasis. i A positive value indicates that the user wants to strengthen y i This description; when α i A negative value indicates that the user needs to avoid y i This description. Normalize the weights so that their sum is one. The normalization formula is as follows:

[0063]

[0064] like Figure 4 As shown, in each iteration of the diffusion model, the empty text description information is used as a condition to obtain the t-1 Prediction And describe the information y in each text i As a condition, we get t-1 The prediction μ θ (x t |y i ), calculate the weighted sum of the prediction results to obtain the prediction value of multiple text descriptions in:

[0065]

[0066] The Classifier-free guidance method is combined with the diffusion model to strengthen the correlation between the generated image and the text description information. The final x can be obtained by the following formula (4): t-1 Predicted value:

[0067]

[0068] The above process is iterated for T steps to obtain the target image content x0 generated by integrating multiple text description information.

[0069] By implementing the embodiments of the present disclosure, the noise image can be denoised using a diffusion model based on the noise image, at least one text description information and the weight of each text description information to generate a target image corresponding to the at least one text description information. This can support the input of one or more text descriptions and assign different weights, and generate an image by combining multiple text descriptions with different weights, thereby meeting the user's needs for strengthening or avoiding different description contents, thereby improving the image content generation effect.

[0070] Figure 5 The block diagram of an image generation device provided by an embodiment of the present disclosure. As Figure 5 shown, the image generation device may include: a first acquisition module 501, a second acquisition module 502, and a generation module 503.

[0071] Among them, the first acquisition module 501 is used to acquire at least one text description information and the weight of each text description information.

[0072] The second acquisition module 502 is used to acquire the noise image to be processed.

[0073] The generation module 503 is used to perform denoising processing on the noise image according to the noise image, at least one text description information, and the weight of each text description information, and generate a target image corresponding to at least one text description information.

[0074] In some embodiments of the present disclosure, the generation module 503 is specifically configured to: input the noise image, at least one text description information, and the weight of each text description information into a diffusion model; starting from the noise image through the diffusion model, condition on each text description information, perform denoising processing on the noise image to obtain a first predicted image of each text description information, and condition on an empty text description information, perform denoising processing on the noise image to obtain a second predicted image; according to the first predicted image of each text description information, the weight of each text description information, and the second predicted image, obtain a new noise image; iterate the above process for T steps, and determine the noise image obtained after iterating for T steps as the target image corresponding to at least one text description information; where T is a positive integer.

[0075] In a possible implementation manner, the generation module 503 is specifically configured to: perform weighted summation processing on the first predicted image of each text description information and the weight of each text description information to obtain an intermediate predicted image; based on the first formula, determine the difference between the intermediate predicted image and the second predicted image as the new noise image.

[0076] Among them, in the embodiments of the present disclosure, the first formula is expressed as follows:

[0077]

[0078] where, x t-1 is the new noise image, is the intermediate predicted image, is the second predicted image, s is an adjustable hyperparameter, and the value range of s is 1 to 8.

[0079] Optionally, in some embodiments of the present disclosure, as Figure 6As shown, the device may further include a normalization module 604. The normalization module 604 is configured to normalize the weights of each text description information. Among them, Figure 6 601-603 in Figure 5 501-503 in have the same functions and structures.

[0080] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated here.

[0081] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device and a readable storage medium.

[0082] As Figure 7 shown, it is a block diagram of an electronic device according to an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0083] As Figure 7 shown, the electronic device includes: one or more processors 701, a memory 702, and an interface for connecting the components, including a high-speed interface and a low-speed interface. Each component is interconnected using different buses and may be mounted on a common motherboard or otherwise mounted as needed. The processor may process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, multiple processors and / or multiple buses may be used in conjunction with multiple memories and multiple memories if desired. Similarly, multiple electronic devices may be connected, each device providing some necessary operations (such as, as a server array, a set of blade servers, or a multi-processor system). Figure 7 One processor 701 is taken as an example in

[0084] The memory 702 is the non-transitory computer-readable storage medium provided by the present disclosure. Wherein, the memory stores instructions executable by at least one processor, so that the at least one processor executes the image generation method provided by the present disclosure. The non-transitory computer-readable storage medium of the present disclosure stores computer instructions for causing a computer to execute the image generation method provided by the present disclosure.

[0085] As a non-transitory computer-readable storage medium, the memory 702 can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the image generation method in the embodiments of the present disclosure. The processor 701 executes various functional applications and data processing of the server by running the non-transitory software programs, instructions, and modules stored in the memory 702, that is, implements the image generation method in the above method embodiments.

[0086] The memory 702 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device. In addition, the memory 702 may include a high-speed random access memory, and may also include non-transitory memories, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 702 may optionally include a memory remotely disposed relative to the processor 701, and these remote memories may be connected to the electronic device through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0087] The electronic device may further include: an input device 703 and an output device 704. The processor 701, the memory 702, the input device 703, and the output device 704 may be connected through a bus or other means, Figure 7 Taking the connection through the bus as an example.

[0088] The input device 703 can receive input digital or character information, and generate key signal inputs related to the user settings and function controls of the electronic device, such as input devices like a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 704 may include a display device, an auxiliary lighting device (for example, an LED), and a haptic feedback device (for example, a vibration motor), etc. The display device may include, but is not limited to, a liquid crystal display (LCD), a light emitting diode (LED) display, and a plasma display. In some embodiments, the display device may be a touch screen.

[0089] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuitry, integrated circuit systems, off-the-shelf ASICs (application specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0090] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor, and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, apparatus, and / or device (e.g., a magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0091] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).

[0092] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), the Internet, and blockchain networks.

[0093] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with blockchain.

[0094] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0095] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. An image generation method, comprising: Obtaining at least one text description information and the weight of each said text description information; Obtaining a noise image to be processed; Performing denoising processing on the noise image according to the noise image, the at least one text description information, and the weight of each said text description information to generate a target image corresponding to the at least one text description information; Wherein, the performing denoising processing on the noise image according to the noise image, the at least one text description information, and the weight of each said text description information to generate a target image corresponding to the at least one text description information includes: Inputting the noise image, the at least one text description information, and the weight of each said text description information into a diffusion model; Starting from the noise image through the diffusion model, performing denoising processing on the noise image with each said text description information as a condition to obtain a first predicted image of each said text description information, and performing denoising processing on the noise image with an empty text description information as a condition to obtain a second predicted image; Obtaining a new noise image according to the first predicted image of each said text description information, the weight of each said text description information, and the second predicted image; Iterating the above process for T steps, and determining the noise image obtained after iterating for T steps as the target image corresponding to the at least one text description information; wherein, the T is a positive integer.

2. The method according to claim 1, further comprising: Normalizing the weight of each said text description information.

3. The method according to claim 1 or 2, wherein The obtaining a new noise image according to the first predicted image of each said text description information, the weight of each said text description information, and the second predicted image includes: Performing weighted summation processing on the first predicted image of each said text description information and the weight of each said text description information to obtain an intermediate predicted image; Based on a first formula, determining the difference between the intermediate predicted image and the second predicted image as the new noise image; Wherein, the first formula is expressed as follows: wherein, is the new noise image, is the intermediate prediction image, is the second prediction image, is an adjustable hyperparameter, has a value range of 1 to 8.

4. An image generation device, comprising: A first obtaining module, configured to obtain at least one text description information and the weight of each said text description information; A second obtaining module, configured to obtain a noise image to be processed; A generating module, configured to perform denoising processing on the noise image according to the noise image, the at least one text description information, and the weight of each said text description information to generate a target image corresponding to the at least one text description information; Wherein, the generating module is specifically configured to: Input the noise image, the at least one text description information, and the weight of each said text description information into a diffusion model; Starting from the noise image through the diffusion model, performing denoising processing on the noise image with each said text description information as a condition to obtain a first predicted image of each said text description information, and performing denoising processing on the noise image with an empty text description information as a condition to obtain a second predicted image; A new noise image is obtained based on the first predicted image of each of the text description information, the weight of each of the text description information, and the second predicted image; The above process is iterated for T steps, and the noise image obtained after iterating for T steps is determined as the target image corresponding to the at least one text description information; wherein, T is a positive integer.

5. The apparatus according to claim 4, further comprising: A normalization module, configured to perform normalization processing on the weights of each of the text description information.

6. The device according to claim 4 or 5, wherein, The generation module is specifically configured to: Perform weighted summation processing on the first predicted image of each of the text description information and the weight of each of the text description information to obtain an intermediate predicted image; Based on a first formula, determine the difference between the intermediate predicted image and the second predicted image as the new noise image; wherein, the first formula is expressed as follows: Among them, is the new noise image, is the intermediate prediction image, is the second prediction image, is an adjustable hyperparameter, The value range of is 1 to 8.

7. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 3.

8. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 3.

9. A computer program product, comprising a computer program, wherein, The computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Image generation system training method, image generation method, and image generation system

    CN115170825A