Image generation method and electronic equipment

By acquiring and combining the first attribute image, the second attribute image and reference text to generate the target image, the problem of single image generation conditions in the prior art is solved, and multiple controls on image content, layout and details are realized to meet the personalized needs of users.

CN120147442AActive Publication Date: 2025-06-13HONOR DEVICE CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202311665824.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-05
Publication Date
2025-06-13
Estimated Expiration
2043-12-05

AI Technical Summary

Technical Problem

The existing image generation model is relatively single in terms of the conditions for controlling image generation and cannot meet the users' growing personalized needs.

Method used

By acquiring the first attribute image, the second attribute image, and the reference text, a target image is generated. The first attribute image is used to indicate layout information of the first object in the target image, the second attribute image is used to indicate image information of the second object in the target image, and the reference text is used to indicate contents of the target image.

Benefits of technology

Image generation that meets user's personalized needs in multiple aspects is realized, including multiple controls of content, layout and image details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147442A_ABST
    Figure CN120147442A_ABST
Patent Text Reader

Abstract

The invention discloses an image generation method and electronic equipment, and the method comprises the steps: obtaining a first attribute image, a second attribute image and a reference text, the first attribute image being used for indicating layout information of a first object in a target image, the second attribute image being used for indicating image information of a second object included in the target image, the reference text is used for indicating contents in the target image; and generating a target image based on the reference text, the first attribute image and the second attribute image. By implementing the method and the device, the increasing individual requirements of the user can be met through various control image generation conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of computer technologies, and in particular, to an image generation method and an electronic device. Background Art

[0002] In recent years, with the continuous development of artificial intelligence technologies, artificial intelligence technologies have been widely applied to the field of image generation, and high-quality images can be automatically generated through control information input by users. For example, when a user inputs descriptive text into a text-to-image (T2I) model, an image that generally conforms to the content of the descriptive text can be obtained. However, the conditions for controlling image generation in the T2I model are relatively single and cannot meet the growing personalized needs of users. Summary of the Invention

[0003] The present application provides an image generation method and an electronic device, which can enable the generated image to meet the description of the reference text in terms of content, meet the layout indicated by the first attribute image in terms of the layout of the first object in the target image, and meet the image indicated by the second attribute image in terms of the image of the second object in the target image, so as to meet the growing personalized needs of users in multiple aspects.

[0004] In a first aspect, the present application provides an image generation method, which includes:

[0005] Obtain a first attribute image, a second attribute image, and a reference text, where the first attribute image is used to indicate layout information of a first object in a target image, the second attribute image is used to indicate image information of a second object included in the target image, and the reference text is used to indicate content in the target image; generate the target image based on the reference text, the first attribute image, and the second attribute image.

[0006] By implementing the method described in the first aspect, the content of the target image, the layout of the first object in the target image, and the image of the second object in the target image can be controlled simultaneously by obtaining the first attribute image, the second attribute image, and the reference text. In this way, the generation conditions of the image can be controlled in multiple ways to meet the growing personalized needs of users.

[0007] In a possible implementation manner, the above manner of generating the target image based on the reference text, the first attribute image, and the second attribute image specifically includes: obtaining text features of the reference text; obtaining first attribute features based on the first attribute image; obtaining second attribute features based on the second attribute image and the text features; generating the target image based on the first attribute features, the second attribute features, and the text features.

[0008] In this way, multiple features can be obtained according to one or more of the first attribute image, the second attribute image, and the reference text, and the generation of the target image can be controlled based on these multiple features.

[0009] In a possible implementation, the method for generating a target image based on the first attribute feature, the second attribute feature, and the text feature specifically includes: using the first attribute feature as the main attribute feature and the second attribute feature as the auxiliary attribute feature, and inputting them into the encoder of the control model to obtain the latent space features output by each layer in the encoder of the control model; the encoder in the control model has the same structure as the encoder in the diffusion model, and the control model is used to control the diffusion model to generate the target image; using the text feature and the latent space features output by each layer in the encoder of the control model as the control conditions for image generation and inputting them into the diffusion model to generate the target image.

[0010] In this way, based on the encoder of the control model that has the same structure as the encoder in the diffusion model, when extracting latent space features with the first attribute feature as the main attribute feature, the second attribute feature can be used as the auxiliary attribute feature for control, so as to fuse the first attribute feature and the second attribute feature to generate latent space features. Then, by using the latent space features and the text feature together as the control conditions for the diffusion model during image generation, the compatibility between the output of the encoder of the control model and the basic control input in the diffusion model can be achieved (the basic control conditions of the diffusion model include the text feature), thereby realizing the control of image generation from multiple aspects such as text, layout information, and image information.

[0011] In a possible implementation, the control model further includes a pre-encoder. There are multiple first attribute images. The method for obtaining the first attribute feature based on the first attribute images specifically includes: respectively inputting each of the multiple first attribute images into the downsampling processing module of the pre-encoder to obtain the low-dimensional image corresponding to each first attribute image; inputting the low-dimensional image corresponding to each first attribute image into the layout feature extraction model of the pre-encoder to obtain the attribute feature corresponding to each first attribute image; inputting the attribute features corresponding to the multiple first attribute images into the attention weighting module of the pre-encoder to obtain the first attribute feature.

[0012] In this way, when there are multiple first attribute images, that is, the number of first objects is multiple, by performing layout feature extraction and attention weighting processing on the multiple first attribute images through the pre-encoder, the first attribute feature that fuses the layout features in the layout information of the multiple first objects can be obtained. Moreover, by performing downsampling processing on the multiple first attribute images in the pre-encoder, the spatial dimensions of the multiple first attribute images can be compressed, thereby reducing the data processing volume during feature extraction.

[0013] In a possible implementation, the control model further includes an image feature extraction model and a dimension transformation model; there are multiple second attribute images, and the method for obtaining the second attribute features based on the second attribute images and the text features specifically includes: inputting each of the multiple second attribute images into the image feature extraction model to obtain the attribute features corresponding to each second attribute image; inputting the attribute features corresponding to each second attribute image into the dimension transformation model to obtain the features to be fused corresponding to each second attribute image, where the dimensions of the attribute features corresponding to each second attribute image are different from the dimensions of the text features, and the dimensions of the features to be fused corresponding to each second attribute image are the same as the dimensions of the text features; fusing the features to be fused corresponding to the multiple second attribute images with the text features to obtain the second attribute features.

[0014] In this way, when there are multiple second attribute images, the number of second objects is multiple. After feature extraction, dimension transformation, and fusion with text features of the second attribute images, the second attribute features that integrate the features in the image information of multiple second objects and the text features can be obtained. And since the encoder of the control model uses the same structure as the encoder in the diffusion model (the basic input of the diffusion model includes text features), when obtaining the second attribute features, the dimensions of the attribute features corresponding to each second attribute image are first transformed to the dimensions of the text features and then fused with the text features, rather than simply fusing the attribute features corresponding to multiple second attribute images, which can be compatible with the basic text feature input in the control model.

[0015] In a possible implementation, before generating the target image based on the reference text, the first attribute image, and the second attribute image, the method further includes: determining the type to which the second object belongs; determining the control model corresponding to the category to which the second object belongs.

[0016] In this way, different control models can correspond to different categories of second objects.

[0017] In a possible implementation, the above further includes: obtaining a training image and a diffusion model; based on the training image, determining a training text, a first training attribute image, and a second training attribute image; the first training attribute image is used to indicate the layout information of the first training object in the training image, the second training attribute image is used to indicate the image information of the second training object in the training image, and the training text is used to indicate the content in the training image; based on the training image, the training text, the first training attribute image, the second training attribute image, and the diffusion model, performing model training to obtain the control model.

[0018] In this way, during model training, the parameters in the diffusion model can be frozen (i.e., the diffusion model is not trained), and only the parameters in the control model are trained, thereby improving the training speed of the model.

[0019] In a possible implementation manner, the method of training the control model based on the training image, training text, first training attribute image, second training attribute image, and diffusion model specifically includes: controlling the diffusion model to add noise to and denoise and restore the training image based on the training text, first training attribute image, and second training attribute image to obtain a restored image and predicted noise; obtaining a noise loss value based on the noise during noise addition and the predicted noise; determining the loss value of the second training object based on the second training object and the object corresponding to the second training object in the restored image; and performing model training based on the noise loss value and the loss value of the second training object to obtain the control model.

[0020] Based on the model training with the noise loss value, by adding the loss value of the second training object for training, the quality of model training can be effectively improved, making the quality of the obtained control model higher and the control of generating the target image more accurate.

[0021] In a second aspect, the present application provides an electronic device, including one or more processors and one or more memories. The one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code. The computer program code includes computer instructions. When the one or more processors execute the computer instructions, the electronic device executes the camera switching method in the first aspect and any possible implementation manner thereof, or executes the camera switching method in the second aspect and any possible implementation manner thereof.

[0022] In a third aspect, the present application provides an image generation device, and the image generation device includes functions / units for executing the image generation method in the first aspect and any possible implementation manner thereof.

[0023] In a fourth aspect, the present application provides a chip system. The chip system is applied to an electronic device. The chip system includes at least one processor and an interface. The interface is used to receive computer instructions and transmit them to the at least one processor; the at least one processor runs the computer instructions to enable the electronic device to execute the image generation method in the first aspect and any possible implementation manner thereof.

[0024] In a fifth aspect, the present application provides a computer-readable storage medium. The computer-readable storage medium stores computer instructions. When the computer instructions run on an electronic device, the electronic device is enabled to execute the image generation method in the first aspect and any possible implementation manner thereof.

[0025] In a sixth aspect, the present application provides a computer program product. When the computer program product runs on a computer, it causes the computer to execute the image generation method in the first aspect and any possible implementation manner thereof described above.

[0026] It can be understood that for the beneficial effects that can be achieved by the above-provided electronic device, image generation device, chip system, computer-readable storage medium, and computer program product, reference may be made to the beneficial effects in the first aspect and any possible implementation manner thereof, which will not be elaborated herein. Description of the Drawings

[0027] Figure 1A It is a schematic flowchart of T2I image generation;

[0028] Figure 1B It is a schematic flowchart of I2I image generation;

[0029] Figure 2 It is a schematic hardware structure diagram of an electronic device provided by an embodiment of the present application;

[0030] Figure 3 It is a schematic software structure diagram of an electronic device provided by an embodiment of the present application;

[0031] Figure 4 It is a schematic flowchart of an image generation method provided by an embodiment of the present application;

[0032] Figure 5 It is a schematic structure diagram of a control model and a diffusion model provided by an embodiment of the present application;

[0033] Figure 6 It is a schematic structure diagram of a pre-encoder provided by an embodiment of the present application;

[0034] Figure 7 It is a schematic diagram of vector replacement provided by an embodiment of the present application;

[0035] Figure 8 It is a schematic flowchart of obtaining a training image data set provided by an embodiment of the present application;

[0036] Figure 9 It is a schematic flowchart of obtaining a training control input provided by an embodiment of the present application;

[0037] Figure 10 It is a schematic flowchart of training a control model provided by an embodiment of the present application. Detailed Embodiments

[0038] The technical solutions in the embodiments of the present application will be clearly and elaborately described below with reference to the accompanying drawings. Among them, in the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may mean A or B. The "and / or" in the text is only a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B may mean: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.

[0039] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying or suggesting relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" is two or more.

[0040] The term "user interface (UI)" in the following embodiments of the present application is a media interface for interaction and information exchange between an application program or an operating system and a user, which realizes the conversion between the internal form of information and the form acceptable to the user. The user interface is source code written in a specific computer language such as Java and Extensible Markup Language (XML). The interface source code is parsed and rendered on an electronic device and finally presented as content recognizable by the user. The common manifestation form of the user interface is the graphical user interface (GUI), which refers to the user interface related to computer operations displayed in a graphical manner. It can be visible interface elements such as time, date, text, icons, buttons, menus, tabs, text boxes, dialog boxes, status bars, navigation bars, and Widgets displayed on the display screen of an electronic device.

[0041] To facilitate understanding of the solutions provided in the embodiments of the present application, the relevant terms involved in the embodiments of the present application are introduced below:

[0042] The latent diffusion model (LDM), simply referred to as the diffusion model, is a generative model that can be used to generate multimedia data such as high-quality images. In the process of generating an image by the diffusion model, it includes a forward process and a reverse process. Among them, in the forward process, the diffusion model continuously adds noise to the input data until a noisy data with a stable distribution is obtained. In the reverse process, the diffusion model predicts the noise added to the noisy data and removes the noise from the noisy data according to the predicted noise to obtain the desired generated image.

[0043] Optionally, generating an image based on a diffusion model includes generating an image based on text control (text to image, T2I), or generating an image by adjusting a reference image based on text (image to image, I2I).

[0044] Exemplarily, Figure 1A A process of generating an image by T2I is shown. Among them, the modules used by the diffusion model when generating an image include a text encoder, a Unet model, and a decoder of a variational auto-encoder (VAE) (simply referred to as a VAE decoder). Based on Figure 1A the shown process, the user can input text into the diffusion model. The text encoder in the diffusion model extracts features from the text to obtain text features. Then, the text encoder inputs the text features into the Unet model. The Unet model performs forward processing of adding noise and backward processing of denoising to obtain the features of the generated image. When the Unet model performs forward processing of adding noise, noise is randomly selected, and then the text features control the Unet model to perform denoising processing on the randomly selected noise. The features of the generated image obtained after denoising conform to the description of the user input text. Finally, the features of the generated image are input into the VAE decoder, and the VAE decoder decodes the features of the generated image into a pixel-level image, that is, the generated image.

[0045] Exemplarily, Figure 1B A process of generating an image by I2I is shown. Among them, the modules used by the diffusion model when generating an image include a VAE encoder, a text encoder, a Unet model, and a decoder of a variational auto-encoder (VAE). Based on Figure 1B the shown process, the user can input text and a reference image into the diffusion model. The text encoder in the diffusion model extracts features from the text to obtain text features, and the VAE encoder in the diffusion model extracts features from the reference image to obtain reference image features. Then, the VAE encoder inputs the reference image features into the Unet model. The Unet model randomly selects noise and adds noise to the reference image features according to the randomly selected noise to obtain noisy data. Then, the text features control the Unet model to perform denoising processing on the noisy data so that the features of the generated image obtained after denoising conform to the description of the user input text. Finally, the features of the generated image are input into the VAE decoder, and the VAE decoder decodes the features of the generated image into a pixel-level image, that is, the generated image.

[0046] The above T2I method supports text to control image generation, and the above I2I method supports text and image to control image generation. However, the essence of the I2I method is to support text to simply modify the reference image and cannot support more complex image generation control conditions. Therefore, the control conditions of these two methods are relatively single and cannot support the growing personalized needs of users.

[0047] To support various control conditions for image generation and thus meet the growing personalized needs of users, the embodiments of the present application provide an image generation method and an electronic device. The electronic device can be a terminal device, such as a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, a smart vehicle, etc., but is not limited thereto. The electronic device can also be a server. For example, it can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The embodiments of the present application do not limit the type of the electronic device. Exemplarily, the electronic device can be the electronic device 200.

[0048] The following introduces the hardware structure of the electronic device 200: As Figure 2 shown, Figure 2 FIG. shows a schematic diagram of the hardware structure of the electronic device 200. It should be understood that the electronic device 200 may have more or fewer components than those shown in the figure, may combine two or more components, or may have different component configurations. The various components shown in the figure can be implemented in hardware, software, or a combination of hardware and software, including one or more signal processing and / or application specific integrated circuits.

[0049] The electronic device 200 may include: a processor 110, a memory 120, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, and a display screen 194.

[0050] It can be understood that the structure schematically shown in the embodiments of the present invention does not constitute a specific limitation on the electronic device 200. In other embodiments of the present application, the electronic device 200 may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure can be implemented in hardware, software, or a combination of software and hardware.

[0051] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0052] Among them, the controller may be the nerve center and command center of the electronic device 200. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.

[0053] A memory 120 may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory may save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0054] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0055] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc. In some embodiments, the electronic device 200 may include one or N display screens 194, where N is a positive integer greater than 1.

[0056] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the terminal device can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example: Antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.

[0057] The mobile communication module 150 can provide solutions for wireless communications such as 2G / 3G / 4G / 5G applied to the terminal device. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves by Antenna 1, filter, amplify, and process the received electromagnetic waves, and then transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through Antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 150 can be disposed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 can be disposed in the same device.

[0058] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device, or displays an image or video. In some embodiments, the modulation and demodulation processor may be an independent device. In other embodiments, the modulation and demodulation processor may be independent of the processor 110 and be provided in the same device as the mobile communication module 150 or other functional modules.

[0059] The wireless communication module 160 may provide wireless communication solutions applied to terminal devices, including wireless local area networks (WLAN) (such as Wi-Fi networks), Bluetooth (BT), BLE broadcasts, global navigation satellite systems (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 160 may be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signals, and transmits the processed signals to the processor 110. The wireless communication module 160 may also receive the signal to be transmitted from the processor 110, perform frequency modulation on it, amplify it, and convert it into electromagnetic waves through the antenna 2 for radiation.

[0060] In some embodiments, antenna 1 of the terminal device is coupled to the mobile communication module 150, and antenna 2 is coupled to the wireless communication module 160, enabling the terminal device to communicate with the network and other devices through wireless communication technologies. The wireless communication technologies may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), Beidou Navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).

[0061] The following is an introduction to the software structure of the electronic device 200: As Figure 3 shown, Figure 3 This is the software architecture diagram of the electronic device 200 provided by the embodiments of this application. Among them, the software structure adopts a layered architecture, which divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. As Figure 3 shown, taking the Android system running on the AP as an example, in some embodiments, the Android system is divided into five layers, from top to bottom, namely the application layer, the application framework layer (framework), the system library, the hardware abstraction layer (HAL), and the kernel layer (kernel).

[0062] Among them, the application layer may include a series of application packages. The application packages may include apps such as camera, gallery, calendar, call, map, WLAN, Bluetooth, music, video, short message, etc. The application layer may also include systemUI (system user interface), which is used to display the interface of the electronic device, such as displaying the signal icon corresponding to the SIM card, displaying the call interface, etc. The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions. For example, the application framework layer may include window manager, content provider, view system, telephony manager, resource manager, notification manager, etc. The telephony manager is used to provide the call function of the electronic device, such as the management of call status (including answering, hanging up, etc.). The application framework layer may also include a radio interface layer (RIL), and the modem can interact with the telephony through the RIL.

[0063] The system libraries include Android runtime, surface manager, 3D graphics processing library, 2D graphics processing library, and media library, etc. The hardware abstraction layer (HAL) includes display HAL, camera HAL, audio HAL, and sensor HAL, etc. The kernel layer is the layer between hardware and software. The kernel layer at least includes display driver, camera driver, audio driver, sensor driver, etc. In this application, each model used in the image generation method is stored in the 3D / 2D graphics processing library of the system libraries. The user can provide the first attribute image, the second attribute image, and the reference text, etc. to the content provider through the applications in the application layer. The diffusion model and the control model stored in the system libraries call the first attribute image, the second attribute image, and the reference text from the content provider to generate the target image, and then hand it over to the display HAL, and the display HAL further calls the display driver for display.

[0064] It can be understood that the above software architecture is only an example. In specific implementation, more functional modules may be included in the above layers, which will not be elaborated in this application.

[0065] The following further describes in detail the image generation method provided by the embodiments of this application:

[0066] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of an image generation method provided by the embodiments of this application. Figure 4 The subject of execution of the method shown can be an electronic device, or the subject can be a chip in the electronic device. Figure 4Taking an electronic device as the execution subject of the method as an example for illustration. The same applies to the execution subject of the image generation method shown in other drawings of the embodiments of the present application, which will not be elaborated hereinafter. Figure 4 The shown image generation method includes steps 401 to 402. Among them:

[0067] 401. The electronic device obtains a first attribute image, a second attribute image, and a reference text. The first attribute image is used to indicate the layout information of a first object in the target image, the second attribute image is used to indicate the image information of a second object included in the target image, and the reference text is used to indicate the content in the target image.

[0068] In the embodiments of the present application, the target image is an image that the user expects to generate. Both the first object and the second object are objects in the target image, and the first object and the second object are the same or different.

[0069] The first attribute image is used to indicate the layout information of the first object. The layout information of the first object is attribute information related to the position of the first object, that is, it indicates the general structure of the first object in the target image. The layout information of the first object includes, but is not limited to, information such as the edge, depth, pose, contour, and position in the target image of the first object.

[0070] Optionally, the number of the first objects can be multiple, and these multiple first objects can be used to cover the main body and background in the target image. In this way, the general structure of the entire target image can be obtained based on the layout information of the foreground and the background. For example, when the main body of the target image to be generated is a portrait and the background is a landscape, the first attribute image corresponding to the portrait and the first attribute image corresponding to the background can be obtained, and the first attribute image corresponding to the portrait and the first attribute image corresponding to the background can jointly cover the general structure of the entire target image.

[0071] Optionally, for multiple different first objects, the multiple first attribute images can be different or the same. For example, for different first objects, the first attribute images can all be depth maps; or, when the first object is a portrait, the first attribute image can be a skeletal image of the portrait, and the skeletal image is used to indicate the pose of the portrait in the target image. When the first object is a landscape, the first attribute image can be a sketch of the landscape (such as a line drawing, etc.), which is used to describe the edge, contour, etc. of the landscape.

[0072] The second attribute image is used to indicate the image information of the second object. The image information of the second object is some personalized attribute information of the second object, and this personalized attribute information is more detailed attribute information than the general structure (i.e., layout information). The image information of the second object can include, but is not limited to, information such as the texture, color, and unique attribute information of the second object.

[0073] Optionally, the number of second objects can be multiple, and these multiple second objects can be respectively parts (or referred to as a portion) of the main object in the first object, so that the details of the main object in the target image can be further refined based on the multiple second attribute images of the multiple second objects. For example, when the target image to be generated includes a landscape and a portrait in front of the landscape, and the portrait is the main object (for example, it is required that the proportion of the portrait in the target image is greater than the proportion of the landscape in the target image), the second objects can be the face and the body in the portrait, and the two second attribute images are the face image and the clothing image respectively. The face image can indicate the texture, color, and unique attribute information of the face (such as the shape and size of facial features, etc.), and the clothing image can indicate the texture, color, and unique attribute information of the clothing (such as the style and fabric of the clothing, etc.). Based on the face image and the landscape image, the details of the portrait in the target image can be refined.

[0074] The reference text can be used to describe the content in the target image. Optionally, the reference text can assist the first attribute image and the second attribute image to further indicate the content of the first object and the second object. For example, the reference text can describe the positional relationship between multiple first objects, such as the portrait being in front of the landscape, etc. Or, the reference text can also indicate the additional content in the target image other than the first object and the second object. For example, the reference text can be that there is also a flying bird in the landscape, which is used to add a flying bird to the landscape in the target image.

[0075] Optionally, the user can directly input one or more of the first attribute image, the second attribute image, and the reference text. Exemplarily, the user can respectively input the skeletal image of the portrait and the landscape image as the first attribute images; the user inputs the face image and the clothing image as the second attribute images; the user inputs the prompt word of the target image as the reference text.

[0076] Optionally, the user can input one or more images to be processed, and the electronic device preprocesses the one or more images to be processed to obtain one or more of the first attribute image, the second attribute image, and the reference text. Exemplarily, the user inputs two pictures and a prompt word. Picture A includes a portrait and a landscape, and Picture B includes a face and clothing. The portrait in Picture A is different from the corresponding portrait in Picture B. Then the electronic device can obtain the skeletal image corresponding to the portrait in Picture A and the landscape image without the portrait based on Picture A, and obtain the face image and the clothing image based on Picture B.

[0077] Optionally, when the electronic device performs preprocessing based on the one or more images to be processed, the electronic device can identify the objects included in the one or more images to be processed and display the identifiers of these objects. The user selects the identifier of the first object and the identifier of the second object from the displayed identifiers of the multiple objects, and then the electronic device generates the first attribute image and the second attribute image according to the identifiers of the first object and the second object selected by the user.

[0078] This application does not limit the sources of the first attribute image, the second attribute image, and the reference text.

[0079] 402. The electronic device generates a target image based on the reference text, the first attribute image, and the second attribute image.

[0080] In this application, the electronic device can process the reference text, the first attribute image, and the second attribute image based on the control model and the diffusion model to generate a target image.

[0081] Next, through Figure 5 First, introduce the structures of the diffusion model and the control model used in this application when generating the target image:

[0082] As Figure 5 shown, among them, the diffusion model includes a text encoder, a Unet model, and a VAE decoder. The Unet model includes a Unet encoder, a Unet decoder, and an intermediate layer for connecting the Unet encoder and the Unet decoder. The structure introduction of the diffusion model can refer to the above Figure 1A related description.

[0083] The control model includes an encoder, a pre-encoder, an image feature extraction model, a dimension transformation model, and a feature fusion model. Among them, the structure of the encoder in the control model is the same as the structure of the Unet encoder in the diffusion model, and the outputs of each layer of the encoder in the control model are connected to each layer of the Unet encoder in the diffusion model.

[0084] Optionally, the outputs of each layer of the encoder in the control model can also be connected to each layer of the Unet decoder in the diffusion model. This application does not make a limitation on this.

[0085] Based on the structures of the control model and the diffusion model, in a possible implementation manner, the manner in which the electronic device generates a target image based on the reference text, the first attribute image, and the second attribute image specifically includes: The electronic device first obtains the text features of the reference text, obtains the first attribute features based on the first attribute image, then obtains the second attribute features based on the second attribute image and the text features, and finally generates a target image based on the first attribute features, the second attribute features, and the text features.

[0086] TakeFigure 5 For example, the specific process of this method includes:

[0087] ①. The electronic device inputs the reference text into the text encoder of the diffusion model for feature extraction, and the text features of the reference text can be obtained. The obtained text features are input into the Unet encoder and Unet decoder of the diffusion model for generating image control. At the same time, since the structure of the encoder in the control model is the same as that of the Unet encoder in the diffusion model, and text features need to be input into each layer in the Unet encoder, in order to be compatible with the encoder in the control model, the text features also need to be input into the control model for processing. As Figure 5 shown, the text features are input into the feature fusion model in the control model, and the output of the feature fusion model is connected to the input of the encoder in the control model.

[0088] ②. The electronic device inputs the first attribute image into the pre-encoder of the control model to obtain the first attribute features, and the obtained first attribute features are input into the encoder of the control model for processing.

[0089] Optionally, when there are multiple first attribute images (equivalent to multiple first objects), the electronic device can input multiple first attribute images into the pre-encoder of the control model together to obtain the first attribute features. The specific method includes: inputting each first attribute image in the multiple first attribute images into the downsampling processing module of the pre-encoder respectively to obtain the low-dimensional image corresponding to each first attribute image; inputting the low-dimensional image corresponding to each first attribute image into the layout feature extraction model of the pre-encoder to obtain the attribute feature corresponding to each first attribute image; inputting the attribute features corresponding to the multiple first attribute images into the attention weighting module of the pre-encoder to obtain the first attribute features. In this way, the obtained first attribute features can indicate the layout information of multiple first objects, and through the downsampling processing by the downsampling processing module, the spatial dimensions of the multiple first attribute images can be compressed, thereby reducing the data processing amount during layout feature extraction.

[0090] Optionally, after inputting the attribute features corresponding to the multiple first attribute images into the attention weighting module of the pre-encoder, channel changes can also be performed on the output of the attention weighting module, so that the obtained first attribute features can be compatible with the channel number requirements of the input of the encoder in the control model.

[0091] Exemplarily, taking the case where there are two first attribute images, and the two first attribute images are the background image and the bone image respectively, Figure 6A structure of a pre - encoder is shown. Among them, the skeletal image is input into the down - sampling processing module and the layout feature extraction module to obtain the attribute feature f1 corresponding to the skeletal image; the background image is input into the down - sampling processing module and the layout feature extraction module to obtain the attribute feature f0 corresponding to the background image. Exemplarily, the down - sampling processing module can perform three 1 / 2 down - samplings on the spatial dimension of the skeletal image or the background image, that is, reduce the spatial dimension to 1 / 8 of the input. Then, f1 is input into the attention module of the attention - weighted module to obtain two attention matrices (or called weights) a and 1 - a; calculate the Hadamard product fa1 of f1 and a in the spatial dimension, and the Hadamard product fa0 of f0 and 1 - a in the spatial dimension; then, stack these two Hadamard products to get fa = fa1+fa0; finally, perform channel transformation on fa through a convolutional layer to obtain the first attribute feature. Exemplarily, Figure 6 The attention module in it can be a channel attention mechanism module (convolutional block attention module, CBAM), etc., and this application does not limit it.

[0092] ③. The electronic device inputs the second attribute image into the image feature extraction model, and the image feature extraction model processes the second attribute image to obtain the attribute feature corresponding to the second attribute image; then, the attribute feature corresponding to the second attribute image undergoes dimensional transformation through the dimensional transformation model to obtain the feature to be fused corresponding to the second attribute image; the feature fusion model will fuse the text feature output by the text encoder in ① and the feature to be fused corresponding to the second attribute image, so as to obtain the second attribute feature, and the obtained second attribute feature is input into the encoder of the control model for processing.

[0093] Optionally, the image feature extraction model can adopt a general image feature extraction model, such as a model for extracting general features such as the color, brightness, and texture of an image; or the image feature extraction model can adopt a preset image feature extraction model corresponding to the second attribute image, such as an image feature extraction model for extracting face features, an image feature extraction model for extracting clothing features, and so on.

[0094] Optionally, when there are multiple second attribute images (equivalent to multiple first objects), if a general image feature extraction model is adopted, multiple second attribute images can be input into the image feature extraction model for processing; if a preset image feature extraction model corresponding to the second attribute image is adopted, one second attribute image can respectively correspond to an image feature extraction model and a dimensional transformation model. The outputs of multiple dimensional transformation models are all input into the feature fusion model to be fused with the text feature.

[0095] Optionally, when the feature fusion model fuses the text features and the features to be fused corresponding to the second attribute image, vector splicing or vector replacement methods can be used for fusion, etc. This application does not limit this. For example, as Figure 7 shown, when fusing the text features, the features to be fused corresponding to the face, and the features to be fused corresponding to the clothes, the slanted part in the text features is the padding part, which has no actual meaning. Then, the padding part can be replaced with the features to be fused corresponding to the face and the features to be fused corresponding to the clothes.

[0096] ④. After the encoder of the control model obtains the first attribute feature and the second attribute feature, the first attribute feature can be used as the main attribute feature, and the second attribute feature can be used as the auxiliary attribute feature and input into the encoder of the control model to obtain the hidden space features output by each layer in the encoder of the control model. Specifically, cross-attention modules can be set in each layer of the encoder of the control model. The cross-attention module is used to analyze the first attribute feature as the main attribute feature and the second attribute feature as the auxiliary attribute feature to obtain the hidden space features of each layer. For example, when the second attribute feature also contains some features indicating layout information, instead of using the layout information in the second attribute feature, the layout information indicated by the first attribute feature is used for layout analysis. For attributes other than layout, the first attribute feature is used for analysis to obtain the hidden space features of each layer. The electronic device inputs the hidden space features output by each layer of the encoder in the control model into each layer of the Unet encoder in the diffusion model.

[0097] ⑤. The diffusion model obtains the target image through the Unet model and the VAE decoder. Among them, the Unet model generates the target image in the way of adding noise and denoising and restoring, and uses the hidden space features and text features for control when generating the target image, so that the finally obtained target image meets the description of the reference text in terms of content, meets the layout indicated by the first attribute image in the layout of the first object, and meets the image indicated by the second attribute image (such as personalized attributes) in the image of the second object, thereby meeting the growing personalized needs of users in multiple aspects.

[0098] Optionally, before generating the target image based on the reference text, the first attribute image, and the second attribute image, it further includes: determining the type to which the second object belongs; determining the control model corresponding to the category to which the second object belongs.

[0099] As can be seen from the above description, the first attribute image can be used to control the general structure of the target image when generating the target image, while the second attribute image can be used to control the local attributes of the main object in the target image (i.e., when the second object is a part of the main object). When the first object is different, the same depth map can be obtained to indicate the general structure. When the main object is different, the second attribute images may be significantly different. Therefore, an appropriate control model can be selected for image generation according to the type of the second attribute image. Since the type of the second attribute image is determined by the second object, if the second object is a human face or clothes, the type to which the second object belongs is a portrait, and a control model that can control the local attributes in the portrait can be selected; if the second object is a house or a street, the type to which the second object belongs is a building, and a control model that can control the local attributes in the building can be selected; if the second object is an animal, a control model that can control the local attributes in the animal can be selected.

[0100] Based on Figure 4 the described embodiments, the content of the target image, the layout of the first object in the target image, and the image of the second object in the target image can be controlled simultaneously by obtaining the first attribute image, the second attribute image, and the reference text. In this way, the generation conditions of the control image can be diversified to meet the growing personalized needs of users.

[0101] For the control model and the diffusion model used in the above embodiments, during model training, only the control model needs to be trained, and there is no need to train the diffusion model. The trained diffusion model can be used to train the untrained control model, so as to obtain the control model and the diffusion model used in the above embodiments.

[0102] The training method of the control model will be further described in detail below:

[0103] In a possible implementation, the training method of the control model is as follows: First, obtain the training images and the diffusion model, then determine the training text, the first training attribute image, and the second training attribute image based on the training images, and finally perform model training based on the training images, the training text, the first training attribute image, the second training attribute image, and the diffusion model to obtain the control model.

[0104] Among them, when training the control model, it is necessary to first select appropriate training images, then obtain training texts, first training attribute images, and second training attribute images based on the training images, and then use the training texts, first training attribute images, and second training attribute images as the training inputs of the control model, so that the control model can control the diffusion model to generate images according to these training inputs. Among them, the first training attribute image is used to indicate the layout information of the first training object in the training image, the second training attribute image is used to indicate the image information of the second training object in the training image, and the training text is used to indicate the content in the training image.

[0105] The following introduces how to select appropriate training images by way of example:

[0106] Suppose that the images generated by the control model for controlling the diffusion model all include a portrait and a background, and the portrait is the main object in the generated image. Such a control model is called a portrait control model. Then, in order to train the portrait control model, the training image dataset can be obtained by first filtering various image datasets, and then the training images in the training image dataset can be used for model training. Figure 8 A process of obtaining the training image dataset is shown. Among them, various image datasets are sequentially filtered by image size, grayscale, text, number of people, face size, face angle, face occlusion, and face sharpness to obtain the training dataset. Exemplarily, images in various image datasets with a size smaller than the preset image size (such as the long side being less than 350 pixels), monochromatic, the proportion of the area occupied by text being greater than the preset threshold (such as 5%), the face size in the image being smaller than the preset face size (such as the long side being less than 75 pixels), the face angle in the image being deflected by more than the preset angle, the face being occluded, and the face being blurred are deleted, and the remaining images constitute the training image dataset.

[0107] The following introduces how to obtain the first training attribute image, the second training attribute image, and the training text based on the training images in the training image dataset:

[0108] Suppose that in the portrait control model, it is required that the first training attribute image includes a training skeleton image and a training background image, the second training attribute image includes a training face image and a training clothing image, and the training text describes the content in the training image. Then, the first training attribute image, the second training image, and the training text can be obtained according to the following process:

[0109] In the embodiments of the present application, since the training bone image and the training background image are images indicating layout information, and the training face image and the training clothing image are images indicating local attributes more detailed than the layout information, therefore, the training face image and the training clothing image can be obtained based on the training images of the original size, so that the training face image and the training clothing image will not lose feature information due to size compression. The training bone image and the training background image are obtained based on the training images of a small size, reducing the amount of data processing of the training bone image and the training background image in subsequent processes.

[0110] Exemplarily, as Figure 9 shown, the face and clothing in the training image can be detected first. According to the face detection result, the training image is cropped to obtain the training face image. According to the clothing detection result, the training clothing image is obtained. Then, according to the face detection result and the clothing detection result, the size of the training image is transformed to reduce the size of the training image so that the reduced training image includes the face, the clothing, and a part of the remaining background (for example, the training image is cropped and scaled to 512 pixels * 512 pixels). Then, a training text is matched for the training image with reduced size to obtain the training text (for example, the BLIP model can be used to obtain the text label of the image as the training text); the bone points of the training image with reduced size are detected to obtain the training bone image (for example, the trained bone key point detection model can be used to obtain the information of the bone key points in the portrait, so as to obtain the training bone image); the foreground is deducted from the training image with reduced size (such as the face and clothing included in the portrait are deducted from the training image with reduced size), and according to the analysis of the background part of the remaining image after deduction, the deducted part is predicted and repaired (for example, the trained image repair model can be used to repair the deducted area), obtaining a complete training background image.

[0111] It can be understood that when the training image is of other types, corresponding methods can also be adopted to obtain the first training attribute image and the second training attribute image corresponding to the training image, which is not limited in the present application.

[0112] After obtaining the first training attribute image, the second training attribute image, and the training text, the diffusion model can be frozen to train the control model. Among them, during the model training, for Figure 5 the control model structure in, the parameters in the pre-encoder, the encoder in the control model, and the dimension transformation model can be updated. Figure 5 The image feature extraction model in

[0113] In a possible implementation, the electronic device first controls the diffusion model to add noise to and denoise and restore the training image based on the training text, the first training attribute image, and the second training attribute image, to obtain a restored image and predicted noise; then, based on the noise during noise addition and the predicted noise, a noise loss value is obtained; based on the second training object and the object corresponding to the second training object in the restored image, a loss value of the second training object is determined; finally, based on the noise loss value and the loss value of the second training object, model training is performed to obtain a control model.

[0114] Exemplarily, as Figure 10 shown, Figure 10 FIG. is a schematic diagram of a training process of a control model provided by an embodiment of the present application.

[0115] Wherein:

[0116] First, the training image is encoded by a VAE encoder to obtain the latent space feature of the training image. Then, the Unet model randomly selects noise to add noise to the latent space feature of the training image. Exemplarily, the latent space feature of the training image is x 0 , the selected noise is Gaussian noise z, and then x 0 can be added with noise according to the following formula 1 to obtain the noise-added feature x t :

[0117]

[0118] where z~N(0,I), a t =1-β t , β t varies uniformly between β s (0.00085) and β s (0.012), and t is a random integer uniformly sampled from 0 to T. For example, T can take a value of 1000.

[0119] At the same time, the training background image, the training skeleton image, the training face image, and the training clothing image are input into the Unet model after being processed by the control model, and the training text is input into the Unet model after being processed by the text encoder.

[0120] Then, the Unet model predicts the noise in the noise-added feature x t according to the inputs of the control model and the text encoder, so as to restore the training image as much as possible, to obtain the restored latent space feature x 0 ', and finally the restored latent space feature x 0 ' is decoded by a VAE decoder to obtain a restored image. It should be noted that the number of sampling times for noise addition and noise restoration by the Unet model can be multiple times, and the present application does not limit this.

[0121] If the noise predicted by the Unet model during noise prediction based on the inputs of the control model and the text encoder is z', then the noise loss value L can be obtained based on z and z'. noise , where the method for obtaining the noise loss value is the same as that in the diffusion model and will not be elaborated here. Meanwhile, the loss value L of the training face can be obtained based on the differences between the training face and training clothes in the training image and the restored image. sim-p and the loss value L of the training clothes sim-p , where L can be obtained according to the following formula 2 sim-p :

[0122] L sim-p = 1 - sim(F(I p ), F(I' p )) (Formula 2)

[0123] where I p is the image of the training face / training clothes in the training image, I' p is the image of the training face / training clothes in the restored image, F() is used to extract the attribute features corresponding to the image, and sim() is used to calculate the cosine similarity between the two extracted attribute features. Therefore, L sim-p can indicate the difference between the training face / training clothes in the training image and the restored image.

[0124] Furthermore, the Unet model obtains the loss value L for training the control model according to the following formula 3:

[0125]

[0126] where is the similarity loss value coefficient. For Figure 10 , it is necessary to obtain the loss value L sim-p of the training face and the loss value L sim-p of the training clothes, then sum these two losses and multiply by the coefficient , and then multiply by L noise to obtain L. Since in L, the control error of the image generation by the reference text and the first training attribute image (including the training bone image and the training background image) can be indicated by L noise , and the control error of the image generation by the second training attribute image (including the training face image / training clothes image) can be indicated by L sim-p , therefore, training the model based on L can effectively improve the quality of model training compared to only using the noise loss value for model training, making the quality of the obtained control model higher and the control of generating the target image more accurate.

[0127] Optionally, during the model training process, the training images can also be randomly flipped left and right, randomly cropped, and the input of each control condition can be randomly emptied (such as inputting the second training attribute image and the reference text but not inputting the first training attribute image, etc.). This can help decouple each control input, so that the trained control model can also control the diffusion model to generate a target image with better quality when a control condition is missing.

[0128] It can be understood that if it is necessary to control the model to generate for other objects except the main object being a portrait, then Figure 5 the pre-encoder, image feature extraction model, and dimension transformation model, etc. in the shown control model can be adaptively extended, and the control model can be retrained by changing each control condition input to the control model during training.

[0129] This application embodiment also provides an electronic device, which may include: one or more processors and one or more memories. The one or more memories are coupled to the one or more processors, and the one or more memories are used to store computer program code. The computer program code includes computer instructions. When the one or more processors execute the computer instructions, the electronic device executes each function or step that the electronic device executes in the above method embodiment.

[0130] This application embodiment also provides an image generation device, which includes units for executing the functions in the electronic device in the above embodiment.

[0131] This embodiment also provides a computer-readable storage medium, in which computer instructions are stored. When the computer instructions run on an electronic device, the electronic device executes each function or step that the mobile phone executes in the above method embodiment.

[0132] This embodiment also provides a computer program product. When the computer program product runs on a computer, the computer executes each function or step that the mobile phone executes in the above method embodiment.

[0133] In addition, the embodiment of this application also provides a device, which may specifically be a chip, component, or module. The device may include a processor and a memory connected to each other; wherein, the memory is used to store computer execution instructions. When the device runs, the processor can execute the computer execution instructions stored in the memory, so that the chip executes each function or step that the mobile phone executes in the above method embodiment.

[0134] Among them, the electronic device, communication system, computer-readable storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here.

[0135] Through the description of the above embodiments, those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0136] In several embodiments provided in this application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the module or unit is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.

[0137] The unit described as a separated component may or may not be physically separated. The component displayed as a unit may be a physical unit or multiple physical units, that is, it may be located in one place, or it may be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0138] In addition, each functional unit in various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0139] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit them. Although the present application has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. An image generation method, characterized in that, the method includes: obtaining a first attribute image, a second attribute image, and a reference text, where the first attribute image is used to indicate the layout information of a first object in the target image, the second attribute image is used to indicate the image information of a second object included in the target image, and the reference text is used to indicate the content in the target image; generating the target image based on the reference text, the first attribute image, and the second attribute image.

2. The method according to claim 1, characterized in that, the generating the target image based on the reference text, the first attribute image, and the second attribute image includes: obtaining the text features of the reference text; obtaining first attribute features based on the first attribute image; obtaining second attribute features based on the second attribute image and the text features; generating the target image based on the first attribute features, the second attribute features, and the text features.

3. The method according to claim 2, characterized in that, the generating the target image based on the first attribute features, the second attribute features, and the text features includes: using the first attribute features as the main attribute features and the second attribute features as the auxiliary attribute features, and inputting them into the encoder of the control model to obtain the hidden space features output by each layer in the encoder of the control model; the encoder in the control model has the same structure as the encoder in the diffusion model, and the control model is used to control the diffusion model to generate the target image; inputting the text features and the hidden space features output by each layer in the encoder of the control model as the control conditions for image generation into the diffusion model to generate the target image.

4. The method according to claim 3, characterized in that, the control model further includes a pre-encoder, there are multiple first attribute images, and the obtaining first attribute features based on the first attribute image includes: inputting each first attribute image in the multiple first attribute images into the downsampling processing module of the pre-encoder to obtain the low-dimensional image corresponding to each first attribute image; inputting the low-dimensional image corresponding to each first attribute image into the layout feature extraction model of the pre-encoder to obtain the attribute features corresponding to each first attribute image; inputting the attribute features corresponding to the multiple first attribute images into the attention weighting module of the pre-encoder to obtain the first attribute features.

5. The method according to claim 3 or 4, characterized in that, the control model further includes an image feature extraction model and a dimension transformation model; there are multiple second attribute images, and the obtaining second attribute features based on the second attribute image and the text features includes: inputting each second attribute image in the multiple second attribute images into the image feature extraction model to obtain the attribute features corresponding to each second attribute image; Input the attribute features corresponding to each of the second attribute images into the dimension transformation model to obtain the features to be fused corresponding to each of the second attribute images. The dimensions of the attribute features corresponding to each of the second attribute images are different from the dimensions of the text features, and the dimensions of the features to be fused corresponding to each of the second attribute images are the same as the dimensions of the text features; Fuse the features to be fused corresponding to the multiple second attribute images with the text features to obtain second attribute features.

6. The method according to claim 3 or 4, wherein, before generating the target image based on the reference text, the first attribute image, and the second attribute image, further comprising: determining the type to which the second object belongs; determining the control model corresponding to the category to which the second object belongs.

7. The method according to claim 3 or 4, wherein, the method further comprises: obtaining a training image and the diffusion model; based on the training image, determining a training text, a first training attribute image, and a second training attribute image; the first training attribute image is used to indicate the layout information of a first training object in the training image, the second training attribute image is used to indicate the image information of a second training object in the training image, and the training text is used to indicate the content in the training image; performing model training based on the training image, the training text, the first training attribute image, the second training attribute image, and the diffusion model to obtain the control model.

8. The method according to claim 7, wherein, the performing model training based on the training image, the training text, the first training attribute image, the second training attribute image, and the diffusion model to obtain the control model includes: controlling the diffusion model to perform noise addition and denoising restoration on the training image based on the training text, the first training attribute image, and the second training attribute image to obtain a restored image and predicted noise; obtaining a noise loss value based on the noise during noise addition and the predicted noise; determining the loss value of the second training object based on the second training object and the object corresponding to the second training object in the restored image; performing model training based on the noise loss value and the loss value of the second training object to obtain the control model.

9. An electronic device, wherein, comprising: one or more processors, one or more memories; wherein, one or more memories are coupled to one or more processors, the one or more memories are used to store computer program code, the computer program code includes computer instructions, and when the one or more processors execute the computer instructions, the electronic device executes the method according to any one of claims 1-8.

10. A chip system, applied to an electronic device, wherein, The chip system includes at least one processor and an interface, where the interface is configured to receive computer instructions and transmit them to the at least one processor; the at least one processor runs the computer instructions to cause the electronic device to execute the method according to any one of claims 1-8.

11. A computer-readable storage medium, characterized in that it stores computer instructions, and when the computer instructions run on an electronic device, the electronic device is caused to execute the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Image generation method and device, electronic equipment and storage medium

    CN115861747A

  • Multi-path parallel text-to-image generation method and system

    CN116128998A

  • Virtual fitting model training method, virtual fitting method and electronic equipment

    CN116416416A

  • Image generation and diffusion model training method, electronic equipment and storage medium

    CN116450873A

  • Image generation method and device, electronic equipment and storage medium

    CN116797684A