Text-to-image method and apparatus, and electronic device

By receiving image modification prompts, the description information of specific objects in images generated by electronic devices can be modified, solving the problem of inaccurate image modification in existing technologies and achieving efficient and accurate local image editing.

WO2025261302A1PCT designated stage Publication Date: 2025-12-26VIVO MOBILE COMM CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/101211
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-21
Filing Date
2025-06-16
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

In existing technologies, when users use the image editing function of electronic devices, they cannot effectively modify part of the image without affecting other areas, resulting in the generated image not meeting the user's needs.

Method used

By receiving image modification prompts from the user, the description information of at least one object in the first image is modified, and a second image is generated, so that the content of the image retention area is the same as that of the first image, and only the content of the image editing area is modified.

Benefits of technology

This allows users to make partial modifications to images without having to re-enter the complete text prompts, generating images that conform to human-computer interaction logic, thus improving image generation efficiency and meeting users' actual needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025101211_26122025_PF_FP_ABST
    Figure CN2025101211_26122025_PF_FP_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of artificial intelligence. Disclosed are a text-to-image method and apparatus, and an electronic device. The method comprises: receiving an image modification prompt inputted by a user, the image modification prompt comprising descriptive information indicating modification of at least one object in a first image, and the first image being an image outputted on the basis of a first text-to-image prompt inputted by the user; and on the basis of the image modification prompt and the first image, outputting a second image, image content of a preserved image area of the first image being the same as that of a preserved image area of the second image, the preserved image area being an image area in the first image except for an image editing area, and the image editing area being an object image area of an object to be edited which is indicated by the image modification prompt.
Need to check novelty before this filing date? Find Prior Art

Description

Methods, apparatus and electronic equipment for creating Wensheng diagrams

[0001] Cross-reference to related applications

[0002] This application claims priority to Chinese Patent Application No. 202410812000.5, filed on June 21, 2024, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of artificial intelligence technology, specifically to a text-to-image method, apparatus, and electronic device, and more specifically, to a text-to-image method, apparatus, and electronic device. Background Technology

[0004] Currently, users can utilize the image generation function of electronic devices to create images. For example, a user can input an image generation prompt on the image generation interface of the electronic device. After receiving the input prompt, the electronic device generates a first image corresponding to the prompt and displays it on the image generation interface for the user to view. Furthermore, if the user is not satisfied with the image effect of the first image and wants to modify its content to obtain a new image, the user needs to re-enter the image generation prompt in the same format as the previous prompt to obtain a newly generated second image.

[0005] In related technologies, the text-to-image (TPI) function of electronic devices is model-based, and the model generates images based on the text-to-image prompts input by the current user. Furthermore, the model generates a completely new image for each text-to-image prompt input by the user. Therefore, the second image is a completely new image relative to the first image.

[0006] Thus, when a user wants to modify part of the content in the first image, the generated second image differs from the first image not only in the content of the image editing area but also in the content of the area outside the image editing area, resulting in the generated second image not being the image the user wants. Summary of the Invention

[0007] The purpose of this application is to provide a text-to-image method, apparatus, and electronic device capable of generating images that meet user modification needs.

[0008] In a first aspect, embodiments of this application provide a text-to-image method, including:

[0009] The system receives image modification prompts input by the user; wherein the image modification prompts include descriptive information indicating that at least one object in a first image should be modified, and the first image is an image output based on the first text-based image prompt input by the user;

[0010] Based on the image modification prompt and the first image, output the second image;

[0011] The image content of the image retention area of ​​the first image is the same as that of the image retention area of ​​the second image. The image retention area is the image area in the first image other than the image editing area, and the image editing area is the object image area of ​​the object to be edited indicated by the image modification prompt.

[0012] Secondly, embodiments of this application provide a text-based image processing device, comprising:

[0013] A receiving module is configured to receive image modification prompts input by a user; wherein the image modification prompts include descriptive information indicating modification of at least one object in a first image, and the first image is an image output based on the first text-based image prompt input by the user;

[0014] The processing module is used to output a second image based on the image modification prompt word received by the receiving module and the first image;

[0015] The image content of the image retention area of ​​the first image is the same as that of the image retention area of ​​the second image. The image retention area is the image area in the first image excluding the image editing area, and the image editing area is the object image area of ​​the object to be modified indicated by the image modification prompt.

[0016] Thirdly, embodiments of this application provide an electronic device including a processor and a memory, wherein the memory stores a program or instructions executable on the processor, and the program or instructions, when executed by the processor, implement the steps of the text-to-image method as described in the first aspect.

[0017] Fourthly, embodiments of this application provide a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method described in the first aspect.

[0018] Fifthly, embodiments of this application provide a chip, the chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run programs or instructions to implement the method as described in the first aspect.

[0019] In a sixth aspect, embodiments of this application provide a computer program product stored in a storage medium, which is executed by at least one processor to implement the method described in the first aspect.

[0020] In this embodiment, an image modification prompt is received from the user. The prompt includes descriptive information indicating the modification of at least one object in a first image. The first image is an image output based on the user-inputted first text-based image prompt. A second image is output based on the image modification prompt and the first image. The image content of the image retention area of ​​the first image is the same as that of the image retention area of ​​the second image. The image retention area of ​​the first image is the image region excluding the image editing area, and the image editing area is the object image region of the object to be edited, as indicated by the image modification prompt. In this solution, on the one hand, the user does not need to re-input the complete text-based image prompt; they only need to input the image modification prompt indicating which content in the image needs modification to generate the second image. Therefore, the generation process of the second image is more in line with human-computer interaction logic and more closely resembles the natural language interaction method of image modification between people. Furthermore, since the image modification prompts include descriptive information indicating the modification of at least one object in the first image, this application can directly edit the object image region of each object indicated by the image modification prompts without repeatedly inputting the prompts, thus achieving simultaneous editing of multiple regions of the first image and effectively improving image generation efficiency. On the other hand, since the second image only modifies the image content of the image editing area in the first image compared to the first image—that is, only modifying the image content mentioned by the user's image modification prompts—image regions not mentioned by the prompts are retained in the second image. In this way, the image content of the image regions that the user is satisfied with in the previously generated image is retained, and the image content of the image regions that the user wants to modify is updated, thereby obtaining an image that meets the user's actual needs. Attached Figure Description

[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 is a schematic diagram of a text-based graphical interface provided in some embodiments of this application;

[0023] Figure 2 is a flowchart illustrating the text generation method provided in some embodiments of this application;

[0024] Figure 3 is a schematic diagram of a text-based graphical interface provided in some embodiments of this application;

[0025] Figure 4 is a schematic diagram of a text-based graphical interface provided in some embodiments of this application;

[0026] Figure 5 is a schematic diagram of a text-based graphical interface provided in some embodiments of this application;

[0027] Figure 6 is a schematic diagram of the editing area provided in some embodiments of this application;

[0028] Figure 7 is a schematic diagram of the mask images provided in some embodiments of this application;

[0029] Figure 8 is a schematic diagram of noise images provided in some embodiments of this application;

[0030] Figure 9 is a schematic diagram of a text-based graphical interface provided in some embodiments of this application;

[0031] Figure 10 is a schematic diagram of the weight images provided in some embodiments of this application;

[0032] Figure 11 is a schematic diagram of the Q2 update process provided in some embodiments of this application;

[0033] Figure 12 is a schematic diagram of the K2 update process provided in some embodiments of this application;

[0034] Figure 13 is a schematic diagram of the structure of a text-based image processing device provided in some embodiments of this application;

[0035] Figure 14 is a schematic diagram of the structure of an electronic device provided in some embodiments of this application;

[0036] Figure 15 is a schematic diagram of the hardware structure of an electronic device provided in some embodiments of this application. Detailed Implementation

[0037] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0038] The terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such terms can be used interchangeably where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0039] The terms "at least one," "at least one of," etc., used in the specification and claims of this application refer to any one, any two, or a combination of two or more of the included items. For example, at least one of a, b, and c can mean: "a," "b," "c," "a and b," "a and c," "b and c," and "a, b, and c," where a, b, and c can be single or multiple. Similarly, "at least two" refers to two or more items, and its meaning is similar to that of "at least one."

[0040] The following will explain the terminology used in the embodiments of this application.

[0041] Stable Diffusion (SD) is an artificial intelligence model for text-to-image processing. Generally, an SD model consists of three main modules: a character encoder, a UNet network, and a Variational Auto-Encoder (VAE). UNet (also known as U-Net) is a Convolutional Neural Network (CNN) architecture for image segmentation. The UNet structure consists of an encoder and a decoder, and its shape resembles a U, hence the name UNet.

[0042] It should be noted that the text-to-image method provided in this application can be executed by electronic devices such as mobile phones, tablets, laptops, PDAs, and in-vehicle electronic devices. Some embodiments of this application use electronic devices as the executing entity to illustrate the text-to-image method provided in this application.

[0043] The text-based image generation method provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.

[0044] It should be noted that the text-to-image method provided in this application can be executed by electronic devices such as mobile phones, tablets, laptops, PDAs, and in-vehicle electronic devices. Some embodiments of this application use electronic devices as the executing entity to illustrate the text-to-image method provided in this application.

[0045] The text-based image generation method, apparatus, and electronic device provided in this application can be applied to image drawing scenarios. Specifically, the text-based image generation method provided in this application can be applied to scenarios where, after inputting text-based image prompts into an image drawing model for image drawing, the user wants to modify the drawn image.

[0046] In a specific application scenario, the text-based image method provided in this application embodiment can be applied to a scenario where the "dog" in an image of "a dog running on the grass" is replaced with a "cat", for example, scenario 1.

[0047] Scenario 1, as shown in Figure 1, when a user wants to draw an image of "a dog running on the grass", they can input the text image prompt 101 "draw a dog running on the grass on a sunny day" on the text image interface 10 of the electronic device. After receiving the text image prompt "draw a dog running on the grass on a sunny day", the electronic device inputs the text image prompt "draw a dog running on the grass on a sunny day" into the image drawing model to draw the image, so as to output an image A of "a dog running on the grass on a sunny day", and then display the image A on the text image interface 10. Furthermore, when a user wants to replace the "dog" in the drawn image of "a dog running on a sunny day in the grass" with a "cat", the user needs to modify the text-based image prompt, for example, changing the text-based image prompt "draw a dog running on a sunny day in the grass" to "draw a cat running on a sunny day in the grass". The user can continue to input the modified text-based image prompt 103 "draw a cat running on a sunny day in the grass" on the text-based image interface 10 of the electronic device, and input the modified text-based image prompt 103 "draw a cat running on a sunny day in the grass" into the image drawing model to draw the image, and then output an image of "a cat running on a sunny day in the grass". Then, the image 104 of "a cat running on a sunny day in the grass" is displayed on the text-based image interface 10. However, since the above image modification process involves modifying the original text-based image prompts and then inputting the modified prompts into the image rendering model for redrawing, not only the editable areas of the output image may change, but the non-editable areas may also undergo some changes. This can result in the output image not meeting the user's needs. For example, while the image of "a cat running on a grassy field on a sunny day" obtained through this method contains the word "cat," the background elements may have changed compared to the background elements in the initial image of "a dog running on a grassy field," such as the "grass" itself.

[0048] Another specific application scenario is that the text-based image method provided in this application embodiment can be applied to a scenario where the hair of the "beautiful woman" in the image of "a beautiful woman with long hair" is replaced with "short hair", for example, scenario 2.

[0049] Scenario 2: When a user wants to draw an image of "a beautiful woman with long hair," they can input the text-based image prompt "Draw a beautiful woman with long hair" into the image drawing model to generate an image of "a beautiful woman with long hair." Further, if the user wants to change the hair in the drawn image of "a beautiful woman with long hair" to "short hair," the user needs to modify the text-based image prompt, for example, changing "Draw a beautiful woman with long hair" to "Draw a beautiful woman with short hair," and then inputting the modified prompt "Draw a beautiful woman with short hair" into the image drawing model to generate an image B of "a beautiful woman with short hair." However, because the above image modification process involves modifying the original text-based image prompt and then inputting the modified prompt into the image drawing model for redrawing, not only will the editable areas change in the output image, but the non-editable areas will also change to some extent, resulting in an output image that does not meet the user's needs. For example, in the image of "a beautiful woman with short hair" obtained through the above method, although the woman's hair has changed from long hair to short hair, her facial features may have changed compared to the facial features of the woman in the initial image of "a beautiful woman with long hair". For example, the woman's face shape may have changed from "round face" to "oval face".

[0050] In the text-based image generation method, apparatus, and electronic device provided in this application embodiment, an image modification prompt is received from a user. The image modification prompt includes descriptive information indicating the modification of at least one object in a first image. The first image is an image output based on the first text-based image prompt input by the user. Then, based on the image modification prompt and the first image, a second image can be output. The image content of the image retention area of ​​the first image is the same as the image content of the image retention area of ​​the second image. The image retention area is the image region in the first image excluding the image editing area, and the image editing area is the object image region of the object to be edited indicated by the image modification prompt. In this solution, on the one hand, the user does not need to re-input the complete text-based image prompt; they only need to input the image modification prompt indicating which content in the image needs modification to generate the second image. Therefore, the generation process of the second image is more in line with human-computer interaction logic and is closer to the natural language interaction method of image modification between people. Furthermore, since the image modification prompts include descriptive information indicating the modification of at least one object in the first image, this application can directly edit the object image region of each object indicated by the image modification prompts without repeatedly inputting the prompts, thus achieving simultaneous editing of multiple regions of the first image and effectively improving image generation efficiency. On the other hand, since the second image only modifies the image content of the image editing area in the first image compared to the first image—that is, only modifying the image content mentioned by the user's image modification prompts—image regions not mentioned by the prompts are retained in the second image. In this way, the image content of the image regions that the user is satisfied with in the previously generated image is retained, and the image content of the image regions that the user wants to modify is updated, thereby obtaining an image that meets the user's actual modification needs.

[0051] Figure 2 is a flowchart illustrating the text-to-image method provided in this embodiment of the application. As shown in Figure 2, the text-to-image method provided in this embodiment of the application may include the following steps:

[0052] Step 201: The electronic device receives the image modification prompt text input by the user.

[0053] The image modification prompt includes descriptive information indicating that at least one object in the first image should be modified, and the first image is an image output based on the first text-based image prompt input by the user.

[0054] In some embodiments of this application, the above-mentioned image modification prompts may include, but are not limited to, text, images, audio, video, and other content.

[0055] For example, the first text-based image prompt includes the descriptive text of the first image, which is used to describe the first image.

[0056] For example, in scenario 1, the first text-based image prompt could be "draw a dog running on the grass on a sunny day," and "a dog running on the grass on a sunny day" would be the descriptive text for the first image, which would then be image A of "a dog running on the grass on a sunny day."

[0057] For example, in scenario 2, the first text prompt could be "draw a beautiful woman with long hair", and "a beautiful woman with long hair" would be the descriptive text for the first image, which would then be image B of "a beautiful woman with long hair".

[0058] In some embodiments of this application, the image modification prompts mentioned above include at least one of the following:

[0059] The instruction is to replace the description information of at least one object in the first image; wherein the object includes at least one of the following: people, animals, plants, scenery, buildings, and items;

[0060] The description information instructs modification of the object attribute information of at least one object in the first image; wherein the object attribute includes at least one of the following: the object's action, characteristics, and state.

[0061] In some embodiments of this application, the aforementioned persons include, but are not limited to, men, women, and children.

[0062] In some embodiments of this application, the animals mentioned above include, but are not limited to, livestock, wild animals, pets, birds, fish, insects, etc.

[0063] In some embodiments of this application, the above-mentioned plants include, but are not limited to, flowers, grasses, trees, wood, potted plants, etc.

[0064] In some embodiments of this application, objects of the type "scenery" may include: cultural landscapes and natural landscapes. Cultural landscapes may include: artificial hills, gardens, etc.; natural landscapes may include: rivers, lakes, seas, the sun, the moon, clouds, starry skies, the sky, snow-capped mountains, forests, or grasslands, etc.

[0065] In some embodiments of this application, objects of the building type include, but are not limited to, roads, houses, bridges, sculptures, water towers, etc.

[0066] In some embodiments of this application, objects of the type of article include, but are not limited to: clothing and accessories, daily necessities, office supplies, toys, tools, entertainment tools, and other inanimate objects.

[0067] For example, the object type of laptops, mobile phones, clothes, pants, shoes, socks, hats, chairs, plates, basins, vases, washing machines, etc.

[0068] In some embodiments of this application, the actions of the aforementioned object include, but are not limited to, the object's facial expressions and postures. Facial expressions may include: joy, anger, sorrow, and happiness; postures may include: running, sitting, standing, lying down, raising a hand, kicking a leg, etc.

[0069] In some embodiments of this application, the features of the above-mentioned objects include, but are not limited to: name, gender, age, shape, color, pattern, size, body shape such as height and weight, style, pattern, material and corresponding form, etc.

[0070] In some embodiments of this application, when the image modification prompt includes descriptive information indicating the replacement of at least one object in the first image, the image modification prompt includes descriptive information of at least one object and descriptive information of a replacement object corresponding to each of the at least one object.

[0071] Example 1: Taking Scenario 1 as an example, the image modification prompt could be "Replace the dog with a cat". Here, "dog" is an object in the first image, and "cat" is the object to replace "dog".

[0072] In some embodiments of this application, when the image modification prompt includes descriptive information for modifying the object attribute information of at least one object in the first image, the image modification prompt includes descriptive information for the object attribute information of each object in the at least one object, and descriptive information for the target object attribute information corresponding to each object. The image modification prompt is used to prompt that the object attribute information of each object be modified to the target object attribute information.

[0073] Example 2: Taking Scenario 1 as an example, in addition to "replace the dog with a cat", the image modification prompt can also include "change sunny day to rainy day". Here, "sunny day" is the object attribute information of the "sky" object in the first image, and "rainy day" is the target object attribute information of the "sky" object.

[0074] Example 3: Taking Scenario 1 as an example, the image modification prompt could be "Change long hair to short hair". Here, "long hair" is the object attribute information of the "hair" object in the first image, and "short hair" is the target object attribute information of the "hair" object.

[0075] Thus, the electronic device, through image modification prompts, can, on the one hand, replace at least one object in the first image indicated by the image modification prompts; and on the other hand, modify the object attribute information of at least one object in the first image indicated by the image modification prompts. Therefore, this application can directly edit the object image region of each object to be edited indicated by the image modification prompts, without repeatedly inputting text-based image prompts, thereby achieving simultaneous editing of multiple regions of the first image and effectively improving image generation efficiency.

[0076] Step 202: The electronic device modifies the prompt word and the first image based on the image and outputs the second image.

[0077] The image content of the image retention area of ​​the first image is the same as the image content of the image retention area of ​​the second image. The image retention area is the image area in the first image other than the image editing area, and the image editing area is the object image area of ​​the object to be edited indicated by the image modification prompt.

[0078] In some embodiments of this application, the aforementioned image editing area is the object image region of all objects to be edited, as indicated by the aforementioned image modification prompt.

[0079] In some embodiments of this application, before step 201 described above, the text-to-image method provided in this application further includes steps 205a and 205b. Furthermore, after step 202 described above, the text-to-image method provided in this application further includes step 206.

[0080] Step 205a: The electronic device receives the first text-based image prompt input by the user.

[0081] In some embodiments of this application, users can input the first text-based image prompt word by voice, text, or other means; this embodiment does not impose specific limitations here.

[0082] Step 205b: The electronic device displays a first message on the text-to-image interactive interface, the first message including a first image.

[0083] In some embodiments of this application, the electronic device inputs the first text-based image prompt into an image drawing model, outputs the first image, and then displays a first message on the text-based image interactive interface, the first message including the first image.

[0084] For example, in conjunction with Example 1, as shown in Figure 3, after the electronic device receives the first text-based image prompt "draw a dog running on a grassy field on a sunny day" input by the user, it displays the first message 102 on the text-based image interactive interface 10. The first message 102 includes the image A of "a dog running on a grassy field on a sunny day".

[0085] For example, in conjunction with Example 3, after the electronic device receives the first text-based image prompt "Draw a short-haired beauty" input by the user, it displays a first message on the text-based image interactive interface. The first message includes an image B of "a short-haired beauty".

[0086] Step 206: The electronic device displays a second message on the above-mentioned text-to-image interactive interface, the second message including a second image.

[0087] In some embodiments of this application, after the electronic device generates a second image based on the image modification prompt and the first image, it displays a second message on the text-to-image interactive interface, the second message including the second image.

[0088] For example, in conjunction with Example 1, as shown in Figure 5, after the electronic device generates an image C of "a cat running on a sunny day in the grass" based on the image modification prompt "replace the dog with a cat" and image A, the second message 103 is displayed on the text-to-image interactive interface 10, and the second message 103 includes image C.

[0089] For example, in conjunction with Example 2, after the electronic device generates an image of "a cat running in the grass on a rainy day" based on the image modification prompts "replace the dog with a cat and change the sunny day to a rainy day" and Image A, it displays a second message on the text-to-image interactive interface, the second message including the image of "a cat running in the grass on a rainy day".

[0090] For example, in conjunction with Example 3, after the electronic device receives the first text-based image prompt "Draw a short-haired beauty" input by the user, it displays a first message on the text-based image interactive interface. The first message includes an image D of "a short-haired beauty".

[0091] In this way, the electronic device displays a second message on the text-to-image interactive interface, allowing the user to view the second image on the text-to-image interactive interface.

[0092] In some embodiments of this application, after step 205a, the text-to-image method provided in this application further includes step 205c. Furthermore, after step 201, the text-to-image method provided in this application further includes step 207.

[0093] Step 205c: The electronic device displays a third message on the text-to-image interactive interface, the third message including the first text-to-image prompt.

[0094] In some embodiments of this application, after receiving a first text-based image prompt word input by a user, the electronic device displays a third message on the text-based image interaction interface, the third message including the first text-based image prompt word.

[0095] For example, in conjunction with Example 1, as shown in Figure 3, after the electronic device receives the first text-based image prompt "Draw a dog running on a grassy field on a sunny day" input by the user, it displays a third message 101 on the text-based image interactive interface 10. The third message 101 includes the first text-based image prompt "Draw a dog running on a grassy field on a sunny day".

[0096] For example, in conjunction with Example 3, after the electronic device receives the first text-based image prompt "Draw a short-haired beauty" input by the user, it displays a third message on the text-based image interaction interface. The third message includes the first text-based image prompt "Draw a short-haired beauty".

[0097] Step 207: The electronic device displays a fourth message on the image-to-text interactive interface, which includes an image modification prompt.

[0098] In some embodiments of this application, after receiving an image modification prompt word input by a user, the electronic device displays a fourth message on the text-to-image interactive interface, the fourth message including the image modification prompt word.

[0099] For example, in conjunction with Example 1, as shown in Figure 4, after the electronic device receives the image modification prompt "replace the dog with a cat" input by the user, it displays a fourth message 104 on the text-to-image interactive interface 10. The fourth message 104 includes the image modification prompt "replace the dog with a cat".

[0100] For example, in conjunction with Example 2, after the electronic device receives the user's input image modification prompt "replace the dog with a cat and change the sunny day to a rainy day", it displays a fourth message on the text-to-image interactive interface. The fourth message includes the image modification prompt "replace the dog with a cat and change the sunny day to a rainy day".

[0101] For example, in conjunction with Example 3, after the electronic device receives the image modification prompt "Change long hair to short hair" input by the user, it displays a fourth message on the text-to-image interaction interface. The fourth message includes the image modification prompt "Change long hair to short hair".

[0102] In this way, the electronic device displays a fourth message on the image editing interface, allowing the user to view image modification prompts on the image editing interface.

[0103] For example, taking Scenario 1 as an example, when the image modification prompt includes "replace the dog with a cat," the second image output based on the image modification prompt "replace the dog with a cat" and image A is shown in image C in Figure 5. This second image is an image of "a cat running on the grass on a sunny day." Combining Figures 1 and 5, the image content of the image retention area of ​​image A is the same as the image content of the image retention area of ​​image C. The image retention area is the image area in image A excluding the image editing area, and the image editing area is the object image area of ​​the "dog" to be edited, indicated by the image modification prompt "replace the dog with a cat."

[0104] For example, taking Scenario 1 as an example, combined with Example 2, when the image modification prompt includes "replace the dog with a cat, change sunny day to rainy day", the second image output by Image A, which is "a cat running on grass in the rain", is based on the image modification prompt "replace the dog with a cat, change sunny day to rainy day" and the second image of Image A. The second image is "a cat running on grass in the rain". Combining Figures 1 and 5, the image content of the image retention area of ​​Image A is the same as the image content of the image retention area of ​​the image of "a cat running on grass in the rain". The image retention area is the image area of ​​Image A excluding the image editing area, and the image editing area is the object image area of ​​the object to be edited, "dog", indicated by the image modification prompt "replace the dog with a cat, change sunny day to rainy day", and the object image area of ​​the object to be edited, "sky".

[0105] For example, taking scenario 2 as an example, combined with example 3, when the image modification prompt includes "change long hair to short hair", the second image output based on the image modification prompt "change long hair to short hair" and the image of "a beautiful woman with long hair" is the image of "a beautiful woman with short hair". The image content of the image retention area of ​​"a beautiful woman with long hair" is the same as the image content of the image retention area of ​​"a beautiful woman with short hair". The image retention area is the image area of ​​"a beautiful woman with long hair" excluding the image editing area, and the image editing area is the object image area of ​​the "hair" object to be edited, indicated by the image modification prompt "change long hair to short hair".

[0106] In the text-based image generation method provided in this application, an electronic device receives image modification prompts input by a user. These prompts include descriptive information indicating the modification of at least one object in a first image, which is an image output based on the user-inputted text-based image prompts. Then, based on the image modification prompts and the first image, a second image is output. The image content of the image retention area of ​​the first image is the same as that of the image retention area of ​​the second image. The image retention area is the image region in the first image excluding the image editing area, and the image editing area is the object image region of the object to be edited, as indicated by the image modification prompts. In this solution, on the one hand, the user does not need to re-input the complete text-based image prompts; they only need to input the image modification prompts indicating which content in the image needs modification to generate the second image. Therefore, the generation process of the second image is more in line with human-computer interaction logic and more closely resembles the natural language interaction method of image modification between people. Furthermore, since the image modification prompts include descriptive information indicating the modification of at least one object in the first image, this application can directly edit the object image region of each object indicated by the image modification prompts without repeatedly inputting the prompts, thus achieving simultaneous editing of multiple regions of the first image and effectively improving image generation efficiency. On the other hand, since the second image only modifies the image content of the image editing area in the first image compared to the first image—that is, only modifying the image content mentioned by the user's image modification prompts—image regions not mentioned by the prompts are retained in the second image. In this way, the image content of the image regions that the user is satisfied with in the previously generated image is retained, and the image content of the image regions that the user wants to modify is updated, thereby obtaining an image that meets the user's actual modification needs.

[0107] In some embodiments of this application, step 202 described above can be implemented by the following steps 202a and 202b:

[0108] Step 202a: The electronic device generates a second textual image prompt and a mask image for each of at least one object based on the image modification prompt and the first image.

[0109] In some embodiments of this application, the electronic device can input an image modification prompt and a first image into a multimodal large model, recognize the first image, and understand the image modification prompt to output a second text-based image prompt.

[0110] For example, taking Scenario 1 as an example, combined with Example 1, assuming the image modification prompt is "replace the dog with a cat", the second text image prompt generated by the electronic device based on image A and the prompt "replace the dog with a cat" is "draw a cat running on the grass on a sunny day".

[0111] For example, taking scenario 1 as an example, combined with example 2, assuming the image modification prompt is "replace the dog with a cat, change the sunny day to a rainy day", then the second text-based image prompt generated by the electronic device based on image A and the prompt "replace the dog with a cat, change the sunny day to a rainy day" is "draw a cat running on the grass in the rain".

[0112] For example, taking scenario 2 as an example, combined with example 3, assuming the image modification prompt is "change long hair to short hair", the second text image prompt generated by the electronic device based on image B and the image modification prompt "change long hair to short hair" is "draw a beautiful woman with short hair".

[0113] In some embodiments of this application, the image modification prompt includes modifying the description information of X objects in the first image, where X is a positive integer; the step 202a above, "generating a mask image for each object in at least one object based on the image modification prompt and the first image", can be implemented through the following steps 202a1 to 202a4:

[0114] Step 202a1: Based on the image modification prompt, the electronic device determines X image editing areas from the first image. Each image editing area is the object image region of one of the X objects.

[0115] In some embodiments of this application, the electronic device can input an image modification prompt and a first image into a multimodal large model, recognize the first image, and understand the image modification prompt, so as to output it to the image editing area of ​​each of the X objects.

[0116] For example, when the aforementioned multimodal large model outputs the image editing area for X objects, the image editing area can be a rectangular region. The multimodal large model can specifically output the coordinates of the upper left corner and the lower right corner of each rectangular region. For example, taking scenario 1 as an example, combined with example 1, when the image modification prompt is "replace the dog with a cat", that is, the image modification prompt only includes the description information of one object "dog", as shown in Figure 6, the upper left corner coordinates of the image editing area of ​​"dog" in image A are (224, 117), and the lower right corner coordinates are (758, 817).

[0117] Step 202a2: The electronic device acquires the mask image corresponding to each of the X image editing areas.

[0118] In some embodiments of this application, the electronic device calculates the center point of each of the X image editing areas based on the coordinates of the upper left corner and the lower right corner of each of the X rectangular areas. Then, it uses an instance segmentation algorithm to extract the object image of each of the X objects from the first image based on the center point of each image editing area, and obtains the mask image corresponding to each image editing area based on the object image of each object.

[0119] For example, referring to Example 1, assuming the image modification prompt is "replace the dog with a cat," meaning the prompt only includes descriptions of the object to be edited, "dog," the electronic device can use the top-left corner coordinates (224, 117) and bottom-right corner coordinates (758, 817) of the rectangle containing the "dog" in image A to calculate the center point of the rectangle. Then, using an instance segmentation algorithm, it extracts the "dog" image from image A based on the center point of the "dog" editing area and obtains a mask image of the "dog." As shown in Figure 7, the electronic device extracts the "dog" image from image A's image editing area 701 and obtains a mask image 702 based on it. In the mask image 702, the area 7021 corresponding to the "dog" is white, indicating that the area containing the "dog" in image A is the area to be edited. In the mask image of the "dog," the pixel value of each pixel in the area corresponding to the "dog" can be a first preset pixel threshold, such as 0. In the mask image for the "dog," all areas except the region corresponding to the "dog" are black, indicating that the areas in image A other than the "dog" do not need to be edited. The pixel value of each pixel in the other regions is a second preset pixel threshold, for example, 1.

[0120] For example, referring to Example 2, assuming the image modification prompt is "replace the dog with a cat, change sunny day to rainy day," meaning the prompt includes descriptions of both the "dog" and "sky" objects to be edited, the electronic device can use the coordinates of the top-left corner (224, 117) and bottom-right corner (758, 817) of the rectangle containing the "dog" in image A to calculate the center point of the rectangle. Then, using an instance segmentation algorithm, the "dog" image is extracted from image A based on the center point of the "dog" editing area. Based on this image, a mask image of the "dog" is obtained. As shown in Figure 7, the region 7021 corresponding to the "dog" in the mask image 702 is white, indicating that the region containing the "dog" in image A is the area to be edited. In the mask image of the "dog," the pixel value of each pixel in the region corresponding to the "dog" can be a first preset pixel threshold, such as 0. In the mask image for the "dog," all areas except the region corresponding to the "dog" are black, indicating that the areas in image A other than the "dog" do not need to be edited. The pixel value of each pixel in the other regions is a second preset pixel threshold, for example, 1.

[0121] Similarly, the electronic device uses the coordinates of the top-left and bottom-right corners of the rectangle containing the "sky" in image A to calculate the center point of the rectangle. Then, using an instance segmentation algorithm, it extracts the "sky" image from image A based on the center point of the "sky" editing region. Based on the "sky" image, it obtains a mask image of the "sky." In the mask image, the area corresponding to the "sky" is white, indicating that the area containing the "sky" in image A is the area that needs editing. In the mask image of the "sky," the pixel value of each pixel in the area corresponding to the "sky" can be a first preset pixel threshold, such as 0. The other areas in the mask image of the "sky," excluding the area corresponding to the "sky," are black, indicating that the other areas in image A, excluding the "sky," do not need editing. The pixel value of each pixel in the other areas is a second preset pixel threshold, such as 1.

[0122] For example, referring to Example 3, in the first mask image corresponding to "long hair," the area corresponding to "hair" is white, indicating that the area containing "hair" in image B is the region that needs to be edited. The pixel value of each pixel in the area corresponding to "hair" can be a first preset pixel threshold, such as 0. The other areas in the first mask image, excluding the area corresponding to "hair," are black, indicating that the other areas in image B, excluding "hair," do not need to be edited. The pixel value of each pixel in the other areas is a second preset pixel threshold, such as 1.

[0123] Step 202a3: When X is greater than 1, the electronic device superimposes the mask images corresponding to the X image editing areas to generate mask images of X objects.

[0124] In some embodiments of this application, the image sizes of the mask images corresponding to the X image editing areas are the same, and the image sizes of the mask images corresponding to the X image editing areas are consistent with the first image. Step 202a3 can be implemented through steps A1 and A2 as follows.

[0125] Step A1: The electronic device adds up the pixel values ​​of pixels at the same pixel position in the mask image corresponding to the X image editing areas to obtain the cumulative pixel value corresponding to the pixel position, and calculates the ratio between the cumulative pixel value corresponding to the pixel position and X.

[0126] For example, let's take the top-left corner of the mask image corresponding to X image editing areas as an example. The pixel values ​​of the pixels at the top-left corner of the mask image corresponding to the X image editing areas are added together to obtain the accumulated pixel value corresponding to the top-left corner point. Then, the ratio between the accumulated pixel value corresponding to the pixel position and X is calculated. The processing method for pixels at other positions in the mask image corresponding to the X image editing areas is the same as that for the top-left corner pixel, and will not be repeated here.

[0127] For example, referring to Example 2, if the image modification prompt is "replace the dog with a cat, change sunny day to rainy day," then the image modification prompt includes descriptive information for two objects: "dog" and "sky." The electronic device then adds the pixel values ​​of pixels at the same pixel location in the mask image corresponding to the "dog" image editing area and the mask image corresponding to the "sky" image editing area, obtaining the accumulated pixel value at that pixel location, and calculates the ratio between the accumulated pixel value at that pixel location and 2.

[0128] Step A2: The electronic device generates a mask image of X objects based on the ratio between the accumulated pixel value corresponding to the above pixel position and X.

[0129] In some embodiments of this application, the ratio between the accumulated pixel value corresponding to the above pixel position and X is the pixel value of the pixel point at the above pixel position in the mask image of X objects.

[0130] For example, taking scenario 1 as an example, combined with example 2, the pixel values ​​of the pixels at the same pixel position in the mask image corresponding to the image editing area of ​​"dog" and the mask image corresponding to the image editing area of ​​"sky" are added together, and the ratio between the resulting cumulative pixel value and 2 is the mask image of the two objects "dog" and "sky".

[0131] In some embodiments of this application, the electronic device can apply Gaussian blur to the superimposed mask image to obtain a mask image of X objects. By applying Gaussian blur to the superimposed mask image, the edges of the mask image of X objects can be smoothed.

[0132] In this way, the electronic device can accurately generate mask images of X objects by calculating the accumulated pixel value corresponding to the same pixel position in the mask image corresponding to X image editing areas, calculating the ratio between the accumulated pixel value corresponding to the pixel position and X, and then generating mask images of X objects based on the ratio between the accumulated pixel value corresponding to the pixel position and X.

[0133] Step 202a4: When X equals 1, the electronic device uses the mask image corresponding to the image editing area of ​​an object as the mask image.

[0134] In some embodiments of this application, when the image modification prompt only includes the description information of one object to be edited, the mask image corresponding to the image editing area of ​​that object to be edited is the mask image.

[0135] For example, taking scenario 1 as an example, combined with example 1, if the image modification prompt is "replace the dog with a cat", that is, the image modification prompt only includes the description information of the object to be edited, "dog", then the mask image corresponding to the image editing area of ​​"dog" is the mask image.

[0136] For example, referring to Example 3, if the image modification prompt is "change long hair to short hair", that is, if the image modification prompt only includes the description information of the object to be edited, "hair", then the mask image corresponding to the image editing area of ​​"hair" is the mask image.

[0137] In this way, the electronic device generates a mask image of X objects by superimposing the mask images corresponding to X image editing areas, and can quickly determine the X image editing areas in the first image through the mask images of X objects.

[0138] Step 202b: The electronic device outputs a second image based on the first text feature information of the second text-based image prompt, the first text feature information of the first text-based image prompt, and a mask image of at least one object.

[0139] In some embodiments of this application, the aforementioned first text feature information is used to characterize the textual semantics of the first textual image prompt word.

[0140] In some embodiments of this application, the aforementioned first text feature information may be: the text feature information of the first text corresponding to the aforementioned first text-based image prompt. The first text is used to describe the text content of the first text-based image prompt.

[0141] In this way, the electronic device can generate a second text-based image prompt and a mask image for each of at least one object based on the image modification prompt and the first image. Therefore, the user does not need to manually modify the first text-based image prompt to generate the second, improving image generation efficiency. Then, based on the second text-based image prompt, the first text feature information of the first text-based image prompt, and the mask image of at least one object, the second image is output. For the user, there is no need to re-enter the complete text-based image prompt; they only need to input the image modification prompt indicating which parts of the image to modify to generate the second image. Therefore, the generation process of the second image is more in line with human-computer interaction logic and more closely resembles the natural language interaction method of image modification between people.

[0142] In some embodiments of this application, before step 202b above, the text-to-image method provided in this application further includes step 203, and step 202b above can be implemented by step 202b1.

[0143] Step 203: The electronic device inputs the first text prompt word into the image drawing model and outputs the first image, the first text feature information, and the first feature map.

[0144] In some embodiments of this application, the first feature map described above is used to characterize the image features of the image formed after the first text is converted from image to text.

[0145] It is understood that the aforementioned first image is an image drawn by the image rendering model based on the first text-based image prompt. In other words, the aforementioned first image is an image obtained by the image rendering model in the electronic device through parsing the first feature map.

[0146] In some embodiments of this application, step 203 described above can be implemented by the following steps 203a to 203e:

[0147] Step 203a: The electronic device inputs the first textual image prompt and the first noise image into the image drawing model.

[0148] In some embodiments of this application, the first noise image described above is the original base map used for image rendering.

[0149] In some embodiments of this application, the size of the first noise image is less than or equal to the size of the first image. That is, when rendering the image, the final rendered image has a size greater than or equal to that of the noise image used.

[0150] In some embodiments of this application, the electronic device may generate a first noise image using a matrix filled with random numbers of a Gaussian noise distribution.

[0151] For example, assuming the matrix size corresponding to the first noisy image is [1,4,128,128], the image output by the image rendering model can be 1024x1024 in size. Here, [1,4,128,128] indicates that the first noisy image is four one-dimensional noisy images of size 128x128.

[0152] Step 203b: The electronic device extracts the first text feature information of the first textual prompt word through the image drawing model.

[0153] In some embodiments of this application, when the content contained in the aforementioned first text-based image prompt is non-text content, the image rendering model needs to convert the non-text content into first text and then extract the first text feature information of the first text. Alternatively, when the aforementioned first text-based image prompt contains first text, the image rendering model can directly extract the first text feature information of the first text.

[0154] In some embodiments of this application, the image rendering model can be an SD model. The electronic device can input the first text corresponding to the first text image prompt into the character encoder in the image rendering model, and encode the first text through the character encoder to obtain the first text feature information.

[0155] Example 4: Taking Scenario 1 as an example, combined with Example 1, the character encoder in the image drawing model can encode the text "draw a dog running on the grass" corresponding to the first text-based image prompt, and obtain the text feature information A1 corresponding to the text "draw a dog running on the grass", that is, the first text feature information mentioned above.

[0156] Example 5: Taking Scenario 2 as an example, combined with Example 3, the character encoder in the image drawing model can perform text encoding on the text "draw a beautiful woman with long hair" corresponding to the first text-based image prompt, and obtain the text feature information B1 corresponding to the text "draw a beautiful woman with long hair", which is the first text feature information mentioned above.

[0157] Step 203c: The electronic device generates a first feature map based on the first text feature information and the third feature map of the first noisy image using an image drawing model.

[0158] In some embodiments of this application, the third feature map described above is used to characterize the image features of the first noisy image.

[0159] In some embodiments of this application, an electronic device can extract features from a first noisy image using a UNet network in an image rendering model to obtain a third feature map of the first noisy image.

[0160] In some embodiments of this application, after the electronic device obtains the third feature map of the first noisy image through the UNet network in the image rendering model, it can fuse the first text feature information and the third feature map through the UNet network to obtain the first feature map.

[0161] For example, in conjunction with Example 4, an electronic device can fuse text feature information A1 and the third feature map through the UNet network in the image rendering model to obtain feature map A2, which is the first feature map mentioned above.

[0162] For example, in conjunction with Example 5, an electronic device can fuse text feature information B1 and the third feature map through the UNet network in the image rendering model to obtain feature map B2, which is the first feature map mentioned above.

[0163] Step 203d: The electronic device generates a first image based on the first feature map using an image drawing model.

[0164] In some embodiments of this application, the electronic device decodes the first feature map using a VAE in an image rendering model to generate a first image.

[0165] For example, an electronic device can decode feature map A2 through the VAE in the image rendering model to obtain image A, which is the first image mentioned above.

[0166] For example, an electronic device can decode feature map B2 through the VAE in the image rendering model to obtain image B, which is the first image mentioned above.

[0167] Step 203e: The electronic device outputs a first image, first text feature information, and a first feature map through the image drawing model.

[0168] In some embodiments of this application, after the electronic device generates a first image, first text feature information and a first feature map through an image rendering model, it outputs the first image, first text feature information and first feature map through the image rendering model.

[0169] For example, in conjunction with Example 1, the electronic device outputs image A, text feature information A1, and feature map A2 through an image rendering model.

[0170] For example, in conjunction with Example 3, the electronic device outputs image B, text feature information B1, and feature map B2 through an image rendering model.

[0171] In this way, while the electronic device inputs the first text-based image prompt into the image rendering model, it simultaneously inputs a first noisy image. This allows the image rendering model to render the image on an image base according to the rendering requirements of the first text-based image prompt, thus ensuring the accuracy of the image rendering.

[0172] In some embodiments of this application, the image rendering model includes a UNet network, which comprises N cascaded convolutional layers, where N is an integer greater than 1. Based on this, step 203c can be implemented through steps 203c1 and 203c2.

[0173] Step 203c1: The electronic device fuses the first text feature information and the third feature map through the first convolutional layer in the UNet network to obtain the first first feature map.

[0174] In some embodiments of this application, the electronic device fuses the first text feature information and the third feature map using the following formula:

[0175] In formula (1) Let Q1 represent a first feature map, Q1 represent a third feature map, V1 represent the first key text feature after feature extraction from the first convolutional layer in the first convolutional layer, K1 represent the first key text feature after feature extraction from the second convolutional layer in the first convolutional layer, T represent the transpose of Q1K1, and d1 represent Q1K1. T The size.

[0176] In this way, the electronic device can effectively reduce the noise in the third feature map by fusing the first text feature information and the third feature map through the first convolutional layer, thereby improving the image quality of the first image when generating the first image based on the third feature map.

[0177] Step 203c2: The electronic device fuses the first text feature information and the first feature map output by the (j-1)th convolutional layer in the UNet network to obtain the jth first feature map, and outputs N first feature maps.

[0178] Where j∈[2,N]; the first feature map includes N first feature maps.

[0179] In some embodiments of this application, the first text feature information and the first feature map output by the (j-1)th convolutional layer are fused:

[0180] Among them, attention in formula (2) 1mapj This represents the j-th first feature map. V represents the first feature map output by the (j-1)th convolutional layer. j K represents the j-th key text feature after the first convolutional layer in the j-th convolutional layer extracts the first text feature information. jThis represents the j-th second key text feature after the first text feature information is extracted through the second convolutional layer in the j-th convolutional layer, where T represents the second key text feature information extracted through the first text feature information. Transpose, d j express The size.

[0181] In this way, the electronic device can effectively reduce the noise in the first feature map output by the first feature map output by the first feature map output by the first feature map output by the first feature map output by the first feature map output by the first feature map output by the first feature map output by the first feature map output by the first feature map output by the first feature map, thereby improving the image quality of the first image.

[0182] It should be noted here that the first feature map output by the last convolutional layer is input into the VAE, and the first image is obtained by decoding the first feature map output by the last convolutional layer through the VAE.

[0183] For example, the image rendering model may include L cascaded UNet networks, where L is an integer greater than 1. The first feature map output from the last convolutional layer of the UNet and the first text feature information are used as the input to the next UNet layer, thus obtaining N*L first feature maps.

[0184] It should be noted here that the first feature map output by the last convolutional layer of the last UNet network is input into the VAE. The first image is obtained by decoding the first feature map output by the last convolutional layer of the last UNet network through the VAE.

[0185] Thus, when generating the first feature map through N convolutional layers in the UNet network, the first text feature information and the first feature map are fused layer by layer through the N convolutional layers, which can effectively reduce the noise in the first feature map, thereby improving the image quality of the first image when generating the first image based on the first feature map.

[0186] Step 202b1: The electronic device inputs the above-mentioned second text image prompt, the above-mentioned first text feature information, the above-mentioned first feature map and the mask image of at least one object into the above-mentioned image drawing model, and outputs the second image.

[0187] In some embodiments of this application, step 202b1 described above can be implemented by the following steps B1 to B5:

[0188] Step B1: The electronic device inputs the second textual image prompt, the first textual feature information, the first feature map, and the mask image of at least one object into the image drawing model.

[0189] Step B2: The electronic device extracts the second text feature information of the second textual prompt word through the image drawing model.

[0190] In some embodiments of this application, the aforementioned second text feature information is used to characterize the textual semantics of the second text image prompt.

[0191] In some embodiments of this application, the aforementioned second text feature information may be: the text feature information of the second text corresponding to the aforementioned second text-based image prompt. Wherein, the second text is used to describe the text content of the second text-based image prompt.

[0192] In some embodiments of this application, when the content contained in the aforementioned second text-based image prompt is non-text content, the electronic device needs the image rendering model to convert the non-text content into second text, and then extract the second text feature information of the second text. Alternatively, when the aforementioned second text-based image prompt contains second text, the image rendering model can directly extract the second text feature information of the second text.

[0193] In some embodiments of this application, the image rendering model can be an SD model. The electronic device can input the second text corresponding to the second text image prompt into the character encoder in the image rendering model, and encode the second text through the character encoder to obtain the feature information of the second text.

[0194] Example 6: Taking Scenario 1 as an example, combined with Example 1, the character encoder in the image drawing model can encode the text "draw a cat running on the grass" corresponding to the second text image prompt, and obtain the text feature information C1 corresponding to the text "draw a dog running on the grass", which is the second text feature information mentioned above.

[0195] Example 7: Taking Scenario 1 as an example, combined with Example 2, the character encoder in the image drawing model can encode the text "draw a cat running on the grass on a rainy day" corresponding to the second text-based image prompt, and obtain the text feature information C2 corresponding to the text "draw a cat running on the grass on a rainy day", which is the second text feature information mentioned above.

[0196] Example 8: Taking Scenario 2 as an example, combined with Example 3, the character encoder in the image drawing model can perform text encoding on the text "draw a short-haired beauty" corresponding to the second text image prompt, and obtain the text feature information D1 corresponding to the text "draw a short-haired beauty", which is the second text feature information mentioned above.

[0197] Step B3: The electronic device generates a second feature map based on the second text feature information using an image drawing model.

[0198] In some embodiments of this application, step B1 can be implemented by step a, and step B3 can be implemented by step b.

[0199] Step a: The electronic device inputs the second text image prompt, the first text feature information, the first feature map, the mask image of at least one object, and the second noise image into the image drawing model.

[0200] In some embodiments of this application, the second noise image may be the same as or different from the first noise image. This embodiment does not impose specific limitations here.

[0201] In this way, while the electronic device inputs the second text-based image prompt into the image rendering model, it simultaneously inputs a second noisy image. This allows the image rendering model to render the image on an image base according to the rendering requirements of the second text-based image prompt, thus ensuring the accuracy of the image rendering.

[0202] Step b: The electronic device generates a second feature map based on the second text feature information and the fourth feature map of the second noise image using an image drawing model.

[0203] In some embodiments of this application, an electronic device can extract features from a second noisy image using a UNet network in an image rendering model to obtain a fourth feature map of the second noisy image.

[0204] In some embodiments of this application, after the electronic device obtains the fourth feature map of the second noisy image through the UNet network in the image rendering model, it can fuse the second text feature information and the fourth feature map through the UNet network to obtain the second feature map.

[0205] For example, in conjunction with Example 7, an electronic device can fuse text feature information C1 and the fourth feature map through the UNet network in the image rendering model to obtain feature map C3, which is the second feature map mentioned above.

[0206] For example, in conjunction with Example 8, an electronic device can fuse text feature information C2 and the fourth feature map through the UNet network in the image rendering model to obtain feature map C4, which is the second feature map mentioned above.

[0207] For example, in conjunction with Example 9, an electronic device can fuse text feature information D1 and the fourth feature map through the UNet network in the image rendering model to obtain feature map D2, which is the second feature map mentioned above.

[0208] Step B4: The electronic device fuses the second text feature information and the first text feature information through an image rendering model to generate fused text feature information.

[0209] In some embodiments of the present application, the electronic device may fuse the first text feature information and the second text feature information through the UNet network in the image drawing model to obtain fused text feature information.

[0210] In some embodiments of the present application, since the above-mentioned fused text feature information is obtained by fusing the first text feature information and the second text feature information, this fused text feature information can represent the fused text semantics of the first text and the second text.

[0211] In some embodiments of the present application, the first text feature information includes sub-text feature information corresponding to each word in the first text-to-image prompt; the second text feature information includes sub-text feature information corresponding to each word in the second text-to-image prompt; the above step B4 can be implemented through the following steps C1 to C3:

[0212] Step C1, the electronic device determines the first sub-text feature information of each object in at least one object in the first text feature information.

[0213] In some embodiments of the present application, each character of the first text in the first text-to-image prompt is mapped to the corresponding position of the first text feature information.

[0214] For example, when the maximum length of the first text is 77, that is, the first text includes at most 77 characters, the first text feature information is 77-dimensional, including the feature information corresponding to each word in the first text, and the arrangement order of the feature information corresponding to each word in the first text feature information is the same as the arrangement order of the characters in the first text.

[0215] For example, in combination with Example 1, the first 13 feature information in the first text feature information are the feature information corresponding to "draw", "a", "only", "sunny", "day", "on", "the", "grass", "land", "running", "dog", respectively, and the last 64 feature information are the feature information corresponding to empty characters.

[0216] Further, in combination with Example 1, assuming that the object to be edited is a dog, the electronic device determines the position of the character "dog" in the first text "draw a dog running on the sunny grassland", and then determines the first sub-text feature information corresponding to "dog" from the same position of the first text feature information. In combination with Example 1, "dog" is the 13th character of the first text, so the 13th feature information in the first text feature information is used as the first sub-text feature information corresponding to "dog".

[0217] Combined with Example 2, assuming the objects to be edited are a dog and the sky, the electronic device determines the position of the character "dog" in the first text "Draw a dog running on the grass", and then determines the first sub-text feature information corresponding to "dog" from the positions with the same first text feature information. Combined with Example 2, "dog" is the 13th character of the first text, so the 13th feature information in the first text feature information is used as the first sub-text feature information of "dog", and "sunny day" is the 4th and 5th characters of the first text, so the 4th and 5th feature information in the first text feature information are used as the first sub-text feature information of "weather".

[0218] For example, combined with Example 3, the first 8 feature information in the first text feature information are the feature information corresponding to "draw", "one", "lady", "long", "hair", "of", "beautiful", "woman" respectively, and the latter 69 feature information are the feature information corresponding to empty characters.

[0219] Furthermore, the electronic device determines the position of the character "long hair" in the first text "Draw a beautiful woman with long hair", and then determines the first sub-text feature information corresponding to "long hair" from the positions with the same first text feature information. Combined with Example 1, "long hair" is the 4th and 5th characters of the first text, so the 4th and 5th feature information in the first text feature information are used as the first sub-text feature information corresponding to "long hair".

[0220] Step C2: The electronic device determines the second sub-text feature information corresponding to the description information of each object in the second text feature information.

[0221] It should be noted that the method for determining the second sub-text feature information corresponding to the description information of each object in the second text feature information is similar to the method for determining the first sub-text feature information. To avoid repetition, it will not be elaborated here in this embodiment.

[0222] Step C3: The electronic device replaces the first sub-text feature information of each object in the first text feature information with the second sub-text feature information corresponding to the description information of each object, and obtains the fused text feature information.

[0223] In some embodiments of the present application, the electronic device keeps the other feature information in the first text feature information unchanged, and only replaces the first sub-text feature information of each object in the first text feature information with the second sub-text feature information corresponding to the description information of each object, then the fused text feature information can be obtained.

[0224] For example, referring to Example 1, assuming the object to be edited is "dog" and the replacement object corresponding to "dog" is "cat", the other feature information in the first text feature information remains unchanged, and only the first sub-text feature information corresponding to "dog" in the first text feature information is replaced with the second sub-text feature information corresponding to "cat" to obtain the fused text feature information.

[0225] For example, referring to Example 2, assuming the objects to be edited are "dog" and "weather", the replacement object corresponding to "dog" is "cat", and the target object attribute information corresponding to "sunny day" is "rainy day", the other feature information in the first text feature information remains unchanged. Only the first sub-text feature information corresponding to "dog" in the first text feature information is replaced with the second sub-text feature information corresponding to "cat", and the first sub-text feature information corresponding to "sunny day" is replaced with the second sub-text feature information corresponding to "rainy day", so as to obtain the fused text feature information.

[0226] For example, in conjunction with Example 3, assuming the object to be edited is "hair" and the target object attribute information corresponding to "long hair" is "short hair", the other feature information in the first text feature information remains unchanged, and only the first sub-text feature information corresponding to "long hair" in the first text feature information is replaced with the second sub-text feature information corresponding to "short hair" to obtain the fused text feature information.

[0227] In this way, the electronic device obtains fused text feature information by replacing the first sub-text feature information of each object in the first text feature information with the second sub-text feature information corresponding to the description information of each object. This makes the feature information in the fused text feature information consistent with the first text feature information, except for the second sub-text feature information corresponding to the description information of each object. Therefore, when the second image is generated using the fused text feature information, the region in the second image except for at least one object is consistent with the first image, thereby obtaining an image that meets the user's actual needs.

[0228] Step B5: The electronic device outputs a second image based on the fused text feature information, a mask image of at least one object, a first feature map, and a second feature map using an image drawing model.

[0229] In some embodiments of this application, an electronic device may use a mask image of at least one object to determine a first weighted image of a first feature map and a second weighted image of a second feature map, and then use the first weighted image and the second weighted image to weight and fuse the first feature map and the second feature map to obtain a fused feature map.

[0230] For example, for scenario 1, combining examples 1 and 3, the electronic device can fuse feature map A2 and feature map C2 through the UNet network in the image rendering model to obtain a first fused feature map. Then, the fused text feature information obtained by fusing the first text feature information with the second text feature information is fused with the first fused feature map to obtain a second fused feature map. Finally, the second fused feature map is decoded through the VAE in the image rendering model to obtain image A.

[0231] For example, for scenario 1, combined with examples 4 and 6, the electronic device can fuse feature map A2 and feature map C3 through the UNet network in the image rendering model to obtain a first fused feature map. Then, the fused text feature information obtained by fusing text feature information A1 with text feature information C1 is fused with the first fused feature map to obtain a second fused feature map. Then, the second fused feature map is decoded through the VAE in the image rendering model to obtain image C.

[0232] For example, for scenario 1, combined with examples 4 and 7, the electronic device can fuse feature map A2 and feature map C4 through the UNet network in the image rendering model to obtain a first fused feature map. Then, the fused text feature information obtained by fusing text feature information A1 with text feature information C2 is fused with the first fused feature map to obtain a second fused feature map. Then, the VAE in the image rendering model decodes the second fused feature map to obtain an image of "a cat running on the grass on a rainy day".

[0233] For example, for scenario 2, combined with examples 5 and 8, the electronic device can fuse feature map B2 and feature map D2 through the UNet network in the image rendering model to obtain a first fused feature map. Then, the fused text feature information obtained by fusing text feature information B1 with text feature information D1 is fused with the first fused feature map to obtain a second fused feature map. Then, the second fused feature map is decoded through the VAE in the image rendering model to obtain image D.

[0234] Thus, the electronic device extracts the second textual feature information of the second text-based image prompt through the image rendering model, thereby extracting the textual semantics of the second text-based image prompt. Then, based on the second textual feature information, the image rendering model generates a second feature map that conforms to the textual semantics of the second text-based image prompt. The image rendering model fuses the second textual feature information and the first textual feature information to generate fused textual feature information that combines the textual semantics of the first text-based image prompt and the textual semantics of the second text-based image prompt. Then, based on the fused textual feature information, the mask image of at least one object, the first feature map, and the second feature map, when the image rendering model outputs the second image, the area covered by the mask image of at least one object in the second image conforms to the textual semantics of the second text-based image prompt, and the area not covered by the mask image of at least one object in the second image conforms to the textual semantics of the first text-based image prompt. This ensures that the area covered by the mask image of one object in the second image is consistent with the first image, and the area not covered by the mask image of one object in the second image changes according to the image modification prompt. As a result, the second image is an image in which only at least one object in the first image is changed, while other areas remain unchanged, thus obtaining an image that meets the user's needs.

[0235] Thus, the electronic device inputs the first text-based image prompt into the image rendering model and outputs a first image, first text feature information, and a first feature map. It then inputs the second text-based image prompt, the first text feature information, the first feature map, and a mask image of at least one object into the image rendering model and outputs a second image. This ensures that the area not covered by the mask image of at least one object in the second image conforms to the textual semantics of the first text-based image prompt, and that the area covered by the mask image of one object in the second image is consistent with the first image. As a result, the second image is an image in which only at least one object in the first image is changed, while other areas remain unchanged, thereby obtaining an image that meets the user's needs.

[0236] In some embodiments of this application, step B5 can be implemented by the following steps D1 and D2:

[0237] Step D1: The electronic device obtains a fused feature map based on the mask image of at least one object, the first feature map, and the second feature map.

[0238] In some embodiments of this application, step D1 can be implemented by the following steps E1 to E3:

[0239] Step E1: The electronic device uses a mask image of at least one object as a first weighted image.

[0240] In some embodiments of this application, the size of the mask image of at least one object can be adjusted to the size of the first feature image based on the size of the first feature image to obtain a first weight image of the first feature image, such that the pixels of the first weight image correspond one-to-one with the pixels of the first feature image.

[0241] Step E2: The electronic device performs a difference operation between the preset pixel threshold and the pixel value of each pixel in the mask image of at least one object to obtain a second weighted image.

[0242] In some embodiments of this application, an electronic device may use a preset pixel threshold to perform a difference operation on the pixel values ​​of the mask image of at least one object to obtain the pixel value of each pixel in the second weight image, thereby determining the second weight image. The second weight image has the same size as the first mask image, and the pixels of the second weight image correspond one-to-one with the pixels of the mask image of at least one object.

[0243] In some embodiments of this application, after the electronic device determines the second weight image, it can adjust the second weight image according to the size of the second feature image, adjusting the size of the second weight image to match the size of the second feature image, so that the pixels of the second weight image correspond one-to-one with the pixels of the second feature image.

[0244] For example, when the size of the second weight image is 1024*1024 and the size of the second feature map is 4096*4096, the electronic device can upsample the second weight image and adjust its size to 4096*4096.

[0245] Step E3: The electronic device uses the first weighted image and the second weighted image to perform weighted fusion of the first feature map and the second feature map to obtain a fused feature map.

[0246] In some embodiments of this application, the electronic device can use the pixel value of a first pixel in a first weighted image as the weight of a second pixel in a first feature map, and the pixel value of a third pixel in a second weighted image as the weight of a fourth pixel in the second feature map. The pixel values ​​of the second and fourth pixels are then weighted to obtain the pixel value of a pixel in the fused feature map, thus obtaining the fused feature map. Here, the first pixel is a pixel in the first weighted image, the second pixel is a pixel in the first feature map corresponding to the first pixel, the third pixel is a pixel in the second weighted image corresponding to the first pixel, and the fourth pixel is a pixel in the fused feature map corresponding to the first pixel.

[0247] Thus, by using the mask image of at least one object as the first weight image, and weighting and fusing the first feature map and the second feature map with the first weight image and the second weight image, a fused feature map is obtained. This ensures that the regions in the fused feature map other than the region where the first object is located are consistent with the first feature map, and the other regions in the fused feature map are consistent with the second feature map. Therefore, when the second image is generated using the fused feature map, the regions in the second image other than the region where at least one object is located are consistent with the first feature map, and the region where at least one object is located will change according to the image modification prompt, thereby improving the image editing effect.

[0248] Step D2: The electronic device outputs a second image based on the fused feature map and fused text feature information.

[0249] In some embodiments of this application, an electronic device can fuse fused feature maps and fused text feature information through a UNet network to obtain a target image.

[0250] In some embodiments of this application, the image rendering model includes a UNet network, which comprises N cascaded convolutional layers. The first feature map includes N first feature maps, each corresponding one-to-one with one of the N convolutional layers, where N is an integer greater than 1. Based on this, step D2 can be implemented through steps F1 to F3 as follows.

[0251] Step F1: The electronic device fuses the fused text feature information and the fused feature map through the first convolutional layer in the UNet network to obtain the first first fused feature map, and obtains the first second fused feature map based on the mask image of at least one object, the first first feature map corresponding to the first convolutional layer and the first first fused feature map.

[0252] In some embodiments of this application, the electronic device fuses the fused text feature information and the fused feature map using the following formula:

[0253] Among them, in formula (3) P1 represents the first fused feature map, O1 represents the first key fused text feature after feature extraction by the first convolutional layer in the first convolutional layer, G1 represents the first key fused text feature after feature extraction by the second convolutional layer in the first convolutional layer, T represents the transpose of P1O1, and l1 represents P1O1. T The size.

[0254] It should be noted that the generation process of the first second fusion feature map can refer to steps E1 to E3 above, and will not be repeated here in this embodiment.

[0255] Step F2: The electronic device fuses the fused text feature information and the (i-1)th second fused feature map output by the (i-1)th convolutional layer in the UNet network to obtain the i-th first fused feature map. Based on the mask image of at least one object, the i-th first feature map corresponding to the i-th convolutional layer, and the i-th first fused feature map, the i-th second fused feature map is obtained; i∈[2,N].

[0256] In some embodiments of this application, the electronic device fuses the fused text feature information and the (i-1)th second fused feature map output by the (i-1)th convolutional layer using the following formula:

[0257] Among them, in formula (4) This represents the i-th first fused feature map. P1 represents the (i-1)th second fused feature map, and O represents the fused text feature information. i G represents the first key fused text feature information after feature extraction of the fused text feature information by the first convolutional layer in the i-th convolutional layer. i T represents the second key fused text feature information after feature extraction of the fused text feature information by the second convolutional layer in the i-th convolutional layer, and T represents the second key fused text feature information after feature extraction of the fused text feature information by the second convolutional layer in the i-th convolutional layer. i O i Transpose, l i P represents i O i T The size.

[0258] It should be noted that the generation process of the i-th second fusion feature map can refer to steps E1 to E3 above, and will not be repeated here in this embodiment.

[0259] Step F2: The electronic device outputs the second image by passing the Nth second fusion feature map output by the Nth convolutional layer in the UNet network.

[0260] In some embodiments of this application, the electronic device can input the Nth second fusion feature map output from the Nth convolutional layer in the UNet network into the VAE, and decode the Nth second fusion feature map through the VAE to output a second image.

[0261] For example, the second fused feature map output by the last convolutional layer of the UNet network, the fused text feature information, and the first feature map output by each convolutional layer in the next level of the UNet network are used as inputs to the next level of the UNet network. By repeating this process, the second fused feature map output by the last convolutional layer of the last level of the UNet network can be obtained. Then, the Nth second fused feature map output by the Nth convolutional layer of the last level of the UNet network is input into the VAE. By decoding the Nth second fused feature map output by the Nth convolutional layer of the last level of the UNet network through the VAE, the second image can be output.

[0262] Thus, when generating the second image through the N convolutional layers in the UNet network, the fused text feature information and the fused feature map can be fused layer by layer through the N convolutional layers, thereby effectively reducing the noise in the fused feature map. As a result, when generating the second image based on the fused feature map, the image quality of the second image can be improved.

[0263] Thus, when an electronic device generates a second image based on a mask image of at least one object, a first feature image, a second feature image, and fused text feature information, it can use the mask image of at least one object to mask the areas in the first feature image other than those corresponding to the X objects to be edited, and fuse the feature data of the second feature image and the fused text feature information into the first feature image and the X image editing areas to obtain the second image. This makes the areas in the second image other than those corresponding to the X objects to be edited consistent with the first image, and only the areas corresponding to the X objects to be edited are changed according to the image editing prompts. Therefore, the second image is an image in which only the X objects to be edited in the first image are changed, while the other areas remain unchanged, thus obtaining an image that meets the user's needs.

[0264] In some embodiments of this application, prior to step a above, the text-based image method provided in this application further includes steps 204a and 204b:

[0265] Step 204a: The electronic device fills the first region of the first image with random noise M times based on the first quantity M, generating M third noise images.

[0266] Wherein, the first region is the region where at least one object is located in the first image; M is an integer greater than 1.

[0267] In some embodiments of this application, the first quantity M is the number of second images that the user wants to obtain after modifying the first image.

[0268] For example, in conjunction with Example 1, if a user wants to obtain four second images of "a cat running on the grass on a sunny day", the user can send the first quantity M to the electronic device by inputting the first quantity M (4) on the screen of the electronic device.

[0269] For example, in conjunction with Example 1, if a user wants to obtain three second images of "a beautiful woman with short hair", the user can send the first quantity M to the electronic device by inputting the first quantity M (3) on the screen of the electronic device.

[0270] In some embodiments of this application, after receiving the first quantity M, the electronic device can overlay a mask image of at least one object onto the first image to block other areas in the first image except for the first region. Then, based on the first quantity M, the first region of the first image is filled with random noise M times to generate M third noise images.

[0271] For example, referring to Example 1, as shown in Figure 8, the electronic device can overlay the mask image 702 of the "dog" onto image A, fill the first region where the "dog" is located with Gaussian noise, and generate the first third noise image 703. The first region 7031 where the "dog" is located in the third noise image is filled with Gaussian noise, while the other regions 7032 are consistent with image A. The generation process of the remaining M-1 third noise images is similar to that of the first third noise image, and will not be described in detail here.

[0272] For example, in conjunction with Example 2, in some embodiments of this application, the electronic device can overlay a mask image of at least one object onto image A, fill the first region containing the "dog" and the "sky" with Gaussian noise, and generate a first third noise image. The first region containing the "dog" and the "sky" in the third noise image is filled with Gaussian noise, while the other regions besides the "dog" remain consistent with the first image. The generation process of the remaining M-1 third noise images is similar to that of the first third noise image, and will not be described again here.

[0273] For example, in conjunction with Example 3, in some embodiments of this application, the electronic device can overlay a mask image of at least one object onto image B, fill the first region where the "hair" is located with Gaussian noise, and generate a first third noise image. The first region where the "hair" is located in the third noise image is filled with Gaussian noise, while the other regions are consistent with the first image. The generation process of the remaining M-1 third noise images is similar to that of the first third noise image, and will not be described again here.

[0274] Step 204b: The electronic device performs reverse processing on the M third noise images to obtain M fourth noise images.

[0275] The second noise image includes M fourth noise images.

[0276] In some embodiments of this application, the electronic device stitches together the above M third noise images to obtain an M-channel stitched noise image, then inputs the stitched noise image into an image rendering model, and performs a reverse operation on the M third noise images to obtain M fourth noise images.

[0277] In some embodiments of this application, the electronic device can first input the VAE's encoder to encode the spliced ​​noisy image, generate an encoded feature map, and then input the encoded feature map output by the encoder into the UNet network to output M fourth noise maps.

[0278] In some embodiments of this application, referring to Example 2, the spliced ​​noisy image is input into a first-level UNet. The first-level UNet network extracts features from the spliced ​​noisy image to obtain a spliced ​​noisy image feature map. Then, N is calculated using the following formula. mask :

[0279] Among them, in formula (5) This represents the feature map of the spliced ​​noisy image. t indicates that the current UNet network is at level t, and a... t α t-1 It is a constant, representing the estimate of noise in the UNet network.

[0280] In some embodiments of this application, the electronic device can use N calculated by formula (5) mask The splicing noise image I is used as input to the next-level UNett network. latent This process is repeated until the N output by the final UNet network is reached. mask As the second noisy image, N mask For the fourth noise image of the M channel, N can be... mask One channel is used as a fourth noise map.

[0281] In some embodiments of this application, step b above can be implemented by the following step b1:

[0282] Step b1: The electronic device generates M second feature maps based on the second text feature information and M fourth feature maps.

[0283] In some embodiments of this application, an electronic device can obtain the fourth feature map of each of the M fourth noise images through a UNet network, and then obtain the second feature map corresponding to the fourth noise image based on the second text feature information and the fourth feature map of each fourth noise image, so as to obtain M second feature maps.

[0284] In some embodiments of this application, step B5 above can be implemented by the following step G:

[0285] Step G: The electronic device outputs M second images based on the fused text feature information, the first feature map, M second feature maps and the mask image of at least one object.

[0286] In some embodiments of this application, an electronic device can obtain a second image corresponding to each second feature map based on fused text feature information, a first feature map, each of the M second feature maps, and a mask image of at least one object, so as to output M second images.

[0287] It should be noted that the process of generating the second image corresponding to each second feature map can refer to steps D1 and D2 above, and will not be repeated here in this embodiment.

[0288] In some embodiments of this application, referring to Example 1, when M is 4, the electronic device can draw four second images using an image drawing model, and then display these four second images on the text-based image interface. As shown in Figure 9, after the electronic device displays the fourth message on the text-based image interface, the electronic device can display a first message 103 containing four images of "a cat running on the grass" on the text-based image interface. As shown in Figure 9, the background of the four images of "a cat running on the grass" is the same as that of image A.

[0289] In this way, by using fused text feature information, a first feature map, M second feature maps, and a first mask image, M second images can be output at once, increasing the diversity of the output second images and further improving the image editing effect.

[0290] Taking a first noisy image as an example to generate the second image, the text-based image generation method provided in this application embodiment will be described exemplarily. Exemplarily, the text-based image generation method may include steps 21 to 25.

[0291] Step 21: The electronic device inputs the first textual image prompt word P and the first noisy image N into the image drawing model, and outputs the first image and the attention map corresponding to the first textual image prompt word P (i.e., the first feature map mentioned above).

[0292] The first noisy image N is a matrix filled with random numbers distributed by Gaussian noise. Taking the generation of a 1024x1024 first image as an example, the matrix size of the first noisy image N is generally [1,4,128,128].

[0293] For example, an electronic device can encode the first text image prompt word P1 using a character encoder in an image rendering model to obtain first text feature information. Then, the first text feature information and the input noise map N are input into the denoising module UNet network in the image rendering model to obtain the first feature map.

[0294] For example, the UNet network has a multi-layered network structure, and each layer performs an attention operation on its input. The result of the attention is called an attention map. Attention represents the mutual stimulation response of multiple features, i.e., feature fusion.

[0295] For example, the attention operation can be represented by the following formula:

[0296] In formula (6), Q is related to the first noisy image N, K and V are related to the first text image prompt word P, T represents the transpose of QK, and d represents QK. T The size.

[0297] For example, suppose the UNet network includes N cascaded convolutional layers, where N is an integer greater than 1. Then, the first noisy image features can be extracted through the first convolutional layer of the UNet network to obtain Q corresponding to the first convolutional layer. The first text feature information can be extracted through the first convolutional layer in the first convolutional layer of the UNet network to obtain K corresponding to the first convolutional layer. The first text feature information can be extracted through the second convolutional layer in the first convolutional layer to obtain V corresponding to the first convolutional layer. Then, the attention map output by the first convolutional layer is obtained using the above formula (6).

[0298] Furthermore, the attention map output by the (j-1)th convolutional layer is used as Q corresponding to the jth convolutional layer. The first convolutional layer in the jth convolutional layer of the UNet network is used to extract features from the first text feature information to obtain K corresponding to the first convolutional layer. The second convolutional layer in the jth convolutional layer is used to extract features from the first text feature information to obtain V corresponding to the jth convolutional layer. Then, the above formula is used to obtain the attention map output by the jth convolutional layer, j∈[2,N].

[0299] For example, the UNet network needs to be executed multiple times, with the output of the UNet network serving as the input for the next UNet network, and so on iteratively, for example, 50 times; finally, the output attention map of the UNet is passed to the decoder of the VAE to decode and output the first image.

[0300] Step 22: The electronic device inputs the first image and the image modification prompt into the large modality model, and outputs the rectangular box of the editing area in the first image, as well as the modified text-to-image prompt (i.e., the second text-to-image prompt mentioned above) P2.

[0301] For example, a multimodal large model can be used to perform image recognition on the first image and understand image modification prompts. After processing the first image and image modification prompts by the multimodal large model, the multimodal large model will output two pieces of information:

[0302] 1) The rectangle of the editing area. The format is: the coordinates of the top left corner (x, y) and the coordinates of the bottom right corner (x, y). Referring to Example 1, the coordinates of the top left corner of the rectangle are (224, 117), and the coordinates of the bottom right corner are (758, 817).

[0303] 2) Revised prompt for the text-to-image drawing. Based on Example 1, the revised prompt for the text-to-image drawing is: Draw a cat running on the grass.

[0304] For example, the final output of the multimodal large model uses a lightweight data exchange format (JavaScript Object Notation, JSON) to unify the above two pieces of information.

[0305] For example, referring to Example 1, assuming it contains an edit region, the output of the above multimodal large model is: {Edit region: [224,117,758,817], Drawing instruction: "Draw a cat running on the grass"}.

[0306] For example, referring to Example 1, assuming the user needs to edit multiple regions, the output will also be multiple editable regions. The output of the above multimodal large model is: {Editable regions: [224,117,758,817]; [0,0,250,1024], text prompt: "Draw a cat running on the grass in the rain"}.

[0307] Step 23: The electronic device determines the mask image of at least one object in the image modification prompt based on the rectangle.

[0308] For example, if there is only one object to be edited in the image modification prompt, then there is only one editing area for the object to be edited in the first image. The electronic device calculates the center point of the rectangle of the editing area, and then uses an instance segmentation algorithm to extract an object image of the object to be edited from the first image based on the center point of the image editing area. Based on the object image of the object to be edited, a first mask image of at least one object is obtained.

[0309] For example, if there are X objects to be edited in the image modification prompt, then there is an editing area for each of the X objects to be edited in the first image. The electronic device calculates the center point of the rectangular area of ​​each editing area based on the coordinates of the upper left and lower right corners of the rectangular area. Then, it uses an instance segmentation algorithm to extract the object image of each object to be edited from the first image based on the center point of each editing area. Based on the object image of each object to be edited, it obtains the mask image corresponding to each object to be edited. Then, it adds up the mask images of the X "editing areas" and takes the average to obtain the mask image of at least one object.

[0310] For example, after acquiring a mask image of at least one object, the electronic device can further apply a Gaussian blur to the mask image of the at least one object to smooth the edges of the mask image of the at least one object. Specifically, applying a Gaussian blur to the mask image of the at least one object can smooth the edges of the mask image of the at least one object, making the boundary between the edited area and the non-edited area in the target image transition more naturally and smoothly.

[0311] Step 24: The electronic device determines the first weight image corresponding to each convolutional layer of the UNet network in the image rendering model based on the mask image of at least one object.

[0312] For example, the size of the attention map in the UNet network is represented as [c, h, w]. Here, c represents the number of channels in the attention map, h represents the height of the attention map, and w represents the width of the attention map. The size of the attention map in each convolutional layer of the UNet network is not necessarily the same.

[0313] For example, the resolution of the mask image of at least one object is the same as that of the first image. In order to perform subsequent processing on each layer of attention map in the UNet network, the electronic device needs to obtain the weight image corresponding to each layer of attention map based on the mask image of at least one object, and the resolution of the weight image is consistent with the resolution of the attention map.

[0314] For example, an electronic device can resize the mask image of at least one object according to the size of the attention map of each layer of the UNet network, adjusting the mask image of at least one object to a resolution of [h, w], thus obtaining the first weight image of the attention map. This first weight image can be represented by a mask. LThis indicates that L represents the number of layers in UNet, which can be 1, 2, 3, ... Referring to Example 1, the mask image of at least one object as shown in Figure 6 is resized to obtain the first weight image mask corresponding to each convolutional layer. L As shown in Figure 10, the size of the attention map in the first convolutional layer is [4, 4096, 4096]. Therefore, the size of the first weight image mask1 corresponding to the first convolutional layer is [4, 4096, 4096], which is a 4-channel image of 4096*4096. The size of the attention map in the second convolutional layer is [4, 2048, 2048]. Therefore, the size of the first weight image mask2 corresponding to the second convolutional layer is [4, 2048, 2048], which is a 4-channel image of 2048*2048. The size of the attention map in the third convolutional layer is [4, 1024, 1024]. Therefore, the size of the first weight image mask3 corresponding to the third convolutional layer is [4, 1024, 1024], which is a 4-channel image of 1024*1024. The size of the attention map in the fourth convolutional layer is [4, 2048, 2048]. Therefore, the size of the first weight image mask4 corresponding to the fourth convolutional layer is [4, 2048, 2048], which is a 2048*2048 image with 4 channels. The size of the attention map in the fifth convolutional layer is [4, 4096, 4096]. Therefore, the size of the first weight image mask5 corresponding to the fifth convolutional layer is [4, 4096, 4096], which is a 4096*4096 image with 4 channels. As shown in Figure 10, the first weight image mask corresponding to each convolutional layer... L The only difference is the size, which is different from the first mask image, but the image content remains the same.

[0315] Step 25: The electronic device inputs the first weighted image, the second raw image prompt, the first noisy image N, and the attation map corresponding to the first raw image prompt into the image drawing model, and outputs the second image.

[0316] For example, the parameters required to calculate the attention map corresponding to the first text-based image prompt, i.e., Q, K, and V in formula (6), are denoted as Q1, K1, and V1. In some embodiments of this application, if the second text-based image prompt and the first noisy image are input into the image rendering model separately, the attention map corresponding to the second text-based image prompt can also be obtained. The parameters required to calculate the attention map corresponding to the first text-based image prompt, i.e., Q, K, and V in formula (6), are denoted as Q2, K2, and V2.

[0317] For example, the electronic device can update Q2 using the following formula to obtain Q′2: Q′2=Q1*mask L +Q2*(1-mask L (7)

[0318] For example, in conjunction with Example 1, the update process of Q2 is shown in Figure 11, mask L After adjusting the size of the mask image for the dog, the area corresponding to the dog in Q′2 remains consistent with Q2, and the other areas in Q′2, excluding the area corresponding to the dog, remain consistent with Q1.

[0319] For example, both K1 and V1 are influenced by the first text-based image prompt P1. K1 and V1 are both feature information matrices, typically of size [n, 77, 64], where n varies in different convolutional layers of the UNet network; 77 represents the maximum text length that the image rendering model can process, and each character / word in the first text-based image prompt P is mapped to a corresponding position in these 77 dimensions; 64 is a predefined parameter of the image rendering model. Each character / word in the first text-based image prompt P has corresponding feature information in K1 and V1, with the size of the feature information corresponding to one character / word being n × 64. Similarly, K2 and V2 are both feature information matrices, with the same size as K1 and V1. Each character / word in the modified text-based image prompt P2 has corresponding feature information in K2 and V2. Compared to the first text-based image prompt P1, to determine which words have been replaced, simply replace the feature information corresponding to the words before replacement in K1 and V1 with the feature information corresponding to the words after replacement in K2 and V2 (size [n, 64]), and then update K2 and V2 to obtain K′2 and V′2.

[0320] Furthermore, referring to Example 1, the update process of K2 is shown in Figure 12. Compared with the first text-based image prompt P1, the modified text-based image prompt P2 replaces "dog" with "cat". Simply replace the feature information corresponding to "dog" in K1 with the feature information corresponding to "cat" in K2 (size [n, 64]) to update K2, resulting in K′2. As shown in Figure 12, each word corresponds to a feature information of size n×64.

[0321] It should be noted that the process of updating V2 to obtain V′2 is similar to the process of updating K2 to obtain K′2, and will not be described again in this embodiment.

[0322] For example, after updating Q2, K2, and V2 to obtain Q′2, K′2, and V′2, the attention map2 corresponding to the second drawing instruction is recalculated using the following formula:

[0323] For example, the UNet output attention map2, along with K′2 and V′2, is used as the input to the UNet network for the next iteration, and this process is repeated iteratively, for example, 50 times. Finally, the attention map2 output by the UNet network is input into the VAE's Decoder module, and the Decoder module decodes it to generate the second image.

[0324] In summary, the text-based image method provided in this application can accurately determine the editing area without manually cutting out the editing area of ​​the first image, keeping the non-editing area unchanged, and the boundary between the editing and non-editing areas is natural and smoother. Furthermore, the text-based image method provided in this application only requires a one-time input of modification requirements and can edit multiple objects in the image simultaneously, eliminating the need to repeatedly input image modification prompts.

[0325] Taking the generation of the second image as an example where the noise map consists of M third noise images, the text-based image generation method provided in this application will be described exemplarily. Exemplarily, the text-based image generation method may include steps 31 to 36.

[0326] Step 31: The electronic device draws a model from the first textual image prompt word P1 and the first noisy image N input by the user, and outputs the first image and the attention map (i.e. the first feature map mentioned above) corresponding to the first textual image prompt word P1.

[0327] Step 32: The electronic device inputs the first image and the image modification prompt into the large modality model, and outputs the rectangular box of the editing area in the first image, as well as the modified text-to-image prompt (i.e., the second text-to-image prompt mentioned above) P2.

[0328] Step 33: The electronic device determines the mask image of at least one object in the image modification prompt based on the rectangle.

[0329] It should be noted that steps 31 to 33 in this embodiment are similar to steps 21 to 23 in terms of implementation process. To avoid repetition, they will not be described again here.

[0330] Step 34: Fill the first region of the first image with random noise M times based on the first quantity M to generate M third noise images Imask, and perform the reverse operation on the M third noise images Imask to obtain the second noise image Nma.

[0331] It should be noted that the implementation process of step 34 is similar to that of steps 204a and 204b. To avoid repetition, this embodiment will not describe it again here.

[0332] Step 35: The electronic device determines the first weight image corresponding to each convolutional layer of the UNet network in the image rendering model based on the first mask image.

[0333] Step 36: The electronic device inputs the first weighted image, the second raw image prompt, the second noisy image Nma, and the attention map corresponding to the first raw image prompt into the image drawing model, and outputs the second image.

[0334] It should be noted that steps 35 and 36 in this embodiment are similar to the implementation process of steps 24 to 25. To avoid repetition, they will not be described again here.

[0335] In summary, the text-based image generation method provided in this application can accurately determine the editing area without manually cutting out the editing area of ​​the first image, preserving the non-editing area unchanged, and resulting in a smoother and more natural transition between the editing and non-editing areas. Furthermore, the text-based image generation method provided in this application only requires a one-time input of editing requirements, allowing simultaneous editing of multiple objects in the image without repeatedly entering image modification prompts. Simultaneously, it can generate multiple different images at once for the user to choose from, reducing user operations and increasing the diversity of edited image outputs.

[0336] The above-described method embodiments, or various possible implementations of the method embodiments, can be executed individually, or, provided there are no contradictions, they can be combined with each other. The specific implementation can be determined according to actual usage requirements, and this application embodiment does not impose any restrictions on this.

[0337] The text-to-image method provided in this application can be executed by a text-to-image device. This application uses a text-to-image device executing the method as an example to illustrate the text-to-image device provided in this application.

[0338] Figure 13 is a schematic diagram of the structure of the image processing device provided in the embodiment of this application. As shown in Figure 13, the image processing device 1300 includes: a receiving module 1301, used to receive an image modification prompt word input by a user; wherein, the image modification prompt word includes descriptive information indicating modification of at least one object in a first image, and the first image is an image output based on the first image modification prompt word input by the user; a processing module 1302, used to output a second image based on the image modification prompt word received by the receiving module 1301 and the first image; wherein, the image content of the image retention area of ​​the first image and the image content of the image retention area of ​​the second image are the same, the image retention area is the image area in the first image excluding the image editing area, and the image editing area is the object image area of ​​the object to be modified indicated by the image modification prompt word.

[0339] In some embodiments of this application, the image modification prompt includes at least one of the following: descriptive information indicating the replacement of at least one object in the first image; wherein the object includes at least one of the following: people, animals, plants, scenery, buildings, and items; descriptive information indicating the modification of object attribute information of at least one object in the first image; wherein the object attribute includes at least one of the following: the object's action, features, and state.

[0340] In some embodiments of this application, the processing module 1302 is specifically used for:

[0341] Based on the image modification prompt and the first image, generate a second text-based image prompt and a mask image for each of the at least one object; based on the second text-based image prompt, the first text feature information of the first text-based image prompt, and the mask image of the at least one object, output the second image.

[0342] In some embodiments of this application, the image modification prompt includes modifying the description information of X objects in the first image, where X is a positive integer; the processing module 1302 is specifically used for:

[0343] Based on the image modification prompt, X image editing areas are determined from the first image, where each image editing area is the object image region of one of the X objects; a mask image is obtained for each of the X image editing areas; if X is greater than 1, the mask images corresponding to the X image editing areas are superimposed to generate a mask image for the X objects; if X equals 1, the mask image corresponding to the image editing area of ​​one object is used as the mask image.

[0344] In some embodiments of this application, the image sizes of the mask images corresponding to the X image editing areas are the same; the processing module 1302 is specifically used to: add the pixel values ​​of pixels at the same pixel position in the mask images corresponding to the X image editing areas to obtain the accumulated pixel value corresponding to the pixel position, and calculate the ratio between the accumulated pixel value corresponding to the pixel position and X; and generate the mask images of the X objects based on the ratio between the accumulated pixel value corresponding to the pixel position and X.

[0345] In some embodiments of this application, the processing module 1302 is further configured to input the first text-based image prompt into an image rendering model and output the first image, the first text feature information, and the first feature map before outputting the second image based on the second text-based image prompt, the first text feature information of the first text-based image prompt, and the mask image of the at least one object; specifically, the processing module 1302 is configured to input the second text-based image prompt, the first text feature information, the first feature map, and the mask image of the at least one object into the image rendering model and output the second image.

[0346] In some embodiments of this application, the processing module 1302 is specifically configured to: input the second text-based image prompt, the first text feature information, the first feature map, and the mask image of the at least one object into the image rendering model; the image rendering model extracts the second text feature information of the second text-based image prompt; the image rendering model generates a second feature map based on the second text feature information; the image rendering model fuses the second text feature information and the first text feature information to generate fused text feature information; and the image rendering model outputs the second image based on the fused text feature information, the mask image of the at least one object, the first feature map, and the second feature map.

[0347] In some embodiments of this application, the first text feature information includes sub-text feature information corresponding to each word in the first text-to-image prompt; the second text feature information includes sub-text feature information corresponding to each word in the second text-to-image prompt; the processing module 1302 is specifically used to: determine the first sub-text feature information of each object in the at least one object in the first text feature information; determine the second sub-text feature information corresponding to the description information of each object in the second text feature information; replace the first sub-text feature information of each object in the first text feature information with the second sub-text feature information corresponding to the description information of each object, to obtain the fused text feature information.

[0348] In some embodiments of this application, the processing module 1302 is specifically used to: obtain a fused feature map based on the mask image of the at least one object, the first feature map, and the second feature map; and output the second image based on the fused feature map and the fused text feature information.

[0349] In some embodiments of this application, the processing module 1302 is specifically used to: use the mask image of the at least one object as a first weight image; perform a difference operation between a preset pixel threshold and the pixel value of each pixel in the mask image of the at least one object to obtain a second weight image; and use the first weight image and the second weight image to perform weighted fusion of the first feature map and the second feature map to obtain the fused feature map.

[0350] In some embodiments of this application, the image rendering model includes a UNet network, which includes N cascaded convolutional layers. The first feature map includes N first feature maps, each corresponding to one of the N convolutional layers, where N is an integer greater than 1. The processing module 1302 is specifically used to: fuse the fused text feature information and the fused feature map through the first convolutional layer in the UNet network to obtain a first first fused feature map, and based on the mask image of the at least one object, the first first feature map corresponding to the first convolutional layer, and the first... The first fusion feature map is used to obtain the first second fusion feature map; the fused text feature information and the (i-1)th second fusion feature map output by the (i-1)th convolutional layer in the UNet network are fused to obtain the ith first fusion feature map, and the ith second fusion feature map is obtained based on the mask image of the at least one object, the ith first feature map corresponding to the ith convolutional layer, and the ith first fusion feature map; i∈[2,N]; the ith second image is output by the ith second fusion feature map output by the ith convolutional layer in the UNet network.

[0351] In some embodiments of this application, the processing module 1302 is specifically configured to: input the first text-based image prompt and the first noise image into the image rendering model; the image rendering model extracts the first text feature information of the first text-based image prompt; the image rendering model generates the first feature map based on the first text feature information and the third feature map of the first noise image; the image rendering model generates the first image based on the first feature map; and the image rendering model outputs the first image, the first text feature information, and the first feature map.

[0352] In some embodiments of this application, the image rendering model includes a UNet network, which includes N cascaded convolutional layers, where N is an integer greater than 1; the processing module 1302 is specifically used to: fuse the first text feature information and the third feature map through the first convolutional layer in the UNet network to obtain the first first feature map; fuse the first text feature information and the first feature map output by the (j-1)th convolutional layer in the UNet network to obtain the jth first feature map, so as to output N first feature maps; where j∈[2, N]; the first feature map includes the N first feature maps.

[0353] In some embodiments of this application, the processing module 1302 is specifically used to: input the second text image prompt, the first text feature information, the first feature map, the mask image of the at least one object, and the second noise image into the image drawing model; and generate the second feature map based on the second text feature information and the fourth feature map of the second noise image.

[0354] In some embodiments of this application, the above-mentioned device 1300 further includes: a noise-adding module 1303 and a reverse processing module 1304; the noise-adding module 1303 is used to fill a first region of the first image with random noise M times based on a first quantity M before inputting the second text-based image prompt, the first text feature information, the first feature map, the mask image of the at least one object and the second noise image into the image drawing model, thereby generating M third noise images, wherein the first region is the region where the at least one object in the first image is located; M is an integer greater than 1; the reverse processing module 1304 is used to perform reverse processing on the M third noise images to obtain M fourth noise images, wherein the second noise images include the M fourth noise images; the processing module 1302 is specifically used to generate M second feature maps based on the second text feature information and the M fourth feature maps; and output M second images based on the fused text feature information, the first feature map, the M second feature maps and the mask image of the at least one object.

[0355] In some embodiments of this application, the receiving module 1301 is further configured to receive a first text-to-image prompt word input by the user before receiving the image modification prompt word input by the user; the device 1300 further includes a display module 1305, configured to display a first message on the text-to-image interactive interface, the first message including the first image; the display module 1305 is further configured to display a second message on the text-to-image interactive interface, the second message including the second image.

[0356] In some embodiments of this application, the display module 1305 is further configured to display a third message on the text-to-image interactive interface after receiving a first text-to-image prompt word input by the user, the third message including the first text-to-image prompt word; the display module 1305 is further configured to display a fourth message on the text-to-image interactive interface after receiving an image modification prompt word input by the user, the fourth message including the image modification prompt word.

[0357] In the image editing device provided in this application, the device receives image modification prompts input by a user; wherein the image modification prompts include descriptive information indicating the modification of at least one object in a first image, the first image being an image output based on the first image editing prompt input by the user; and a second image is output based on the image modification prompts and the first image; wherein the image content of the image retention area of ​​the first image and the image content of the image retention area of ​​the second image are the same, the image retention area being the image area in the first image excluding the image editing area, and the image editing area being the object image area of ​​the object to be modified indicated by the image modification prompts. In this solution, on the one hand, for the user, the user does not need to re-input the complete image editing prompts; the user only needs to input the image modification prompts indicating which content in the image to modify to generate the second image. Therefore, the generation process of the second image is more in line with human-computer interaction logic and is closer to the natural language interaction method of image modification between people. Furthermore, since the image modification prompts include descriptive information indicating the modification of at least one object in the first image, this application can directly edit the object image region of each object indicated by the image modification prompts without repeatedly inputting the prompts, thus achieving simultaneous editing of multiple regions of the first image and effectively improving image generation efficiency. On the other hand, since the second image only modifies the image content of the image editing area in the first image compared to the first image—that is, only modifying the image content mentioned by the user's image modification prompts—image regions not mentioned by the prompts are retained in the second image. In this way, the image content of the image regions that the user is satisfied with in the previously generated image is retained, and the image content of the image regions that the user wants to modify is updated, thereby obtaining an image that meets the user's actual modification needs.

[0358] The texturing device 1300 in this application embodiment can be an electronic device or a component within an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices besides a terminal. For example, the electronic device can be a mobile phone, tablet computer, laptop computer, PDA, in-vehicle electronic device, mobile internet device (MID), augmented reality (AR) / virtual reality (VR) device, robot, wearable device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc. It can also be a server, network attached storage (NAS), personal computer (PC), television (TV), ATM, or self-service machine, etc. This application embodiment does not specifically limit the specific type of device.

[0359] The image processing device in this application embodiment can be a device with an operating system. The operating system can be Android, iOS, or other possible operating systems; this application embodiment does not specifically limit the specific operating system.

[0360] The image processing device provided in this application can implement the various processes implemented in the method embodiments of Figures 1 to 12 and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0361] Optionally, as shown in FIG14, this application embodiment also provides an electronic device 1400, including a processor 1401 and a memory 1402. The memory 1402 stores a program or instructions that can run on the processor 1401. When the program or instructions are executed by the processor 1401, they implement the various steps of the above-described text-to-image method embodiment and can achieve the same technical effect. To avoid repetition, they will not be described again here.

[0362] It should be noted that the electronic devices in the embodiments of this application include the mobile electronic devices and non-mobile electronic devices described above.

[0363] Figure 15 is a schematic diagram of the hardware structure of an electronic device that implements an embodiment of this application.

[0364] The electronic device 1500 includes, but is not limited to, components such as: radio frequency unit 1501, network module 1502, audio output unit 1503, input unit 1504, sensor 1505, display unit 1506, user input unit 1507, interface unit 1508, memory 1509, and processor 810.

[0365] Those skilled in the art will understand that the electronic device 1500 may also include a power supply (such as a battery) for powering various components. The power supply may be logically connected to the processor 1510 through a power management system, thereby enabling functions such as managing charging, discharging, and power consumption through the power management system. The electronic device structure shown in Figure 15 does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements, which will not be elaborated here.

[0366] The processor 1510 is configured to receive image modification prompts input by a user; wherein the image modification prompts include descriptive information indicating modification of at least one object in a first image, the first image being an image output based on the first text-based image prompt input by the user; and output a second image based on the image modification prompts and the first image; wherein the image content of the image retention area of ​​the first image and the image content of the image retention area of ​​the second image are the same, the image retention area being the image region in the first image excluding the image editing area, and the image editing area being the object image region of the object to be edited indicated by the image modification prompts.

[0367] In some embodiments of this application, the image modification prompts mentioned above include at least one of the following:

[0368] The instruction is to replace the description information of at least one object in the first image; wherein the object includes at least one of the following: people, animals, plants, scenery, buildings, and items;

[0369] Description information indicating modification of object attribute information of at least one object in the first image; wherein the object attributes include at least one of the following: object action, features, and state.

[0370] In some embodiments of this application, the processor 1510 is specifically configured to generate a second text-based image prompt and a mask image of each of the at least one object based on the image modification prompt and the first image; and output a second image based on the second text-based image prompt, the first text feature information of the first text-based image prompt, and the mask image of the at least one object.

[0371] In some embodiments of this application, the image modification prompt includes modifying the description information of X objects in the first image, where X is a positive integer; the processor 1510 is specifically configured to, based on the image modification prompt, determine X image editing areas from the first image, where each image editing area is an object image region of one of the X objects; obtain a mask image corresponding to each of the X image editing areas; when X is greater than 1, superimpose the mask images corresponding to the X image editing areas to generate a mask image of the X objects; when X equals 1, use the mask image corresponding to the image editing area of ​​one object as the mask image.

[0372] In some embodiments of this application, the image sizes of the mask images corresponding to the X image editing areas are the same; the processor 1510 is specifically used to add the pixel values ​​of pixels at the same pixel position in the mask images corresponding to the X image editing areas to obtain the accumulated pixel value corresponding to the pixel position, and calculate the ratio between the accumulated pixel value corresponding to the pixel position and X; based on the ratio between the accumulated pixel value corresponding to the pixel position and X, the mask images of the X objects are generated.

[0373] In some embodiments of this application, processor 1510 is further configured to input the first text-based image prompt into an image rendering model before outputting the second image based on the second text-based image prompt, the first text feature information of the first text-based image prompt, and the mask image of the at least one object, and output the first image, the first text feature information, and the first feature map; processor 1510 is specifically configured to input the second text-based image prompt, the first text feature information, the first feature map, and the mask image of the at least one object into the image rendering model and output the second image.

[0374] In some embodiments of this application, the processor 1510 is specifically configured to input the second text-based image prompt, the first text feature information, the first feature map, and the mask image of the at least one object into the image rendering model; the image rendering model extracts the second text feature information of the second text-based image prompt; the image rendering model generates a second feature map based on the second text feature information; the image rendering model fuses the second text feature information and the first text feature information to generate fused text feature information; and the image rendering model outputs a second image based on the fused text feature information, the mask image of the at least one object, the first feature map, and the second feature map.

[0375] In some embodiments of this application, the first text feature information includes sub-text feature information corresponding to each word in the first text-to-image prompt; the second text feature information includes sub-text feature information corresponding to each word in the second text-to-image prompt; the processor 1510 is specifically configured to determine the first sub-text feature information of each object in the at least one object in the first text feature information; determine the second sub-text feature information corresponding to the description information of each object in the second text feature information; and replace the first sub-text feature information of each object in the first text feature information with the second sub-text feature information corresponding to the description information of each object to obtain fused text feature information.

[0376] In some embodiments of this application, the processor 1510 is specifically configured to obtain a fused feature map based on the mask image of the at least one object, the first feature map, and the second feature map; and output a second image based on the fused feature map and the fused text feature information.

[0377] In some embodiments of this application, the processor 1510 is specifically configured to use the mask image of the at least one object as a first weight image; perform a difference operation between a preset pixel threshold and the pixel value of each pixel in the mask image of the at least one object to obtain a second weight image; and use the first weight image and the second weight image to perform weighted fusion of the first feature map and the second feature map to obtain a fused feature map.

[0378] In some embodiments of this application, the image rendering model includes a UNet network, which includes N cascaded convolutional layers. The first feature map includes N first feature maps, each corresponding to one of the N convolutional layers, where N is an integer greater than 1. The processor 1510 is specifically configured to fuse the fused text feature information and the fused feature map through the first convolutional layer in the UNet network to obtain a first first fused feature map, and based on the mask image of the at least one object, the first first feature map corresponding to the first convolutional layer, and the first... The first fusion feature map is used to obtain the first second fusion feature map; the fused text feature information and the (i-1)th second fusion feature map output by the (i-1)th convolutional layer in the UNet network are fused to obtain the i-th first fusion feature map, and the i-th second fusion feature map is obtained based on the mask image of the at least one object, the i-th first feature map corresponding to the i-th convolutional layer and the i-th first fusion feature map; i∈[2,N]; the second image is output through the Nth second fusion feature map output by the Nth convolutional layer in the UNet network.

[0379] In some embodiments of this application, the processor 1510 is specifically configured to input the first text-based image prompt and the first noisy image into the image rendering model; the image rendering model extracts the first text feature information of the first text-based image prompt; the image rendering model generates a first feature map based on the first text feature information and the third feature map of the first noisy image; the image rendering model generates a first image based on the first feature map; and the image rendering model outputs the first image, the first text feature information, and the first feature map.

[0380] In some embodiments of this application, the image rendering model includes a UNet network, which includes N cascaded convolutional layers, where N is an integer greater than 1; the processor 1510 is specifically used to fuse the first text feature information and the third feature map through the first convolutional layer in the UNet network to obtain a first first feature map; and to fuse the first text feature information and the first feature map output by the (j-1)th convolutional layer in the UNet network to obtain a j-th first feature map, so as to output N first feature maps; where j∈[2, N]; the first feature map includes the N first feature maps.

[0381] In some embodiments of this application, the processor 1510 is specifically configured to input the second text image prompt, the first text feature information, the first feature map, the mask image of the at least one object, and the second noise image into the image rendering model; and generate a second feature map based on the second text feature information and the fourth feature map of the second noise image.

[0382] In some embodiments of this application, the processor 1510 is further configured to, before inputting the second text-based image prompt, the first text feature information, the first feature map, the mask image of the at least one object, and the second noise image into the image rendering model, fill a first region of the first image with random noise M times based on a first quantity M to generate M third noise images, wherein the first region is the region where the at least one object in the first image is located; M is an integer greater than 1; perform reverse processing on the M third noise images to obtain M fourth noise images, wherein the second noise images include the M fourth noise images; the processor 1510 is specifically configured to generate M second feature maps based on the second text feature information and the M fourth feature maps; and output M second images based on the fused text feature information, the first feature map, the M second feature maps, and the mask image of the at least one object.

[0383] In some embodiments of this application, the processor 1510 is further configured to receive a first text-to-image prompt word input by the user before receiving the image modification prompt word input by the user; display a first message on the text-to-image interactive interface, the first message including a first image; and the processor 1510 is specifically configured to display a second message on the text-to-image interactive interface, the second message including a second image.

[0384] In some embodiments of this application, the processor 1510 is further configured to, after receiving a first text-to-image prompt word input by a user, display a third message on the text-to-image interactive interface, the third message including the first text-to-image prompt word; and after receiving an image modification prompt word input by a user, display a fourth message on the text-to-image interactive interface, the fourth message including the image modification prompt word.

[0385] In the electronic device provided in this application, the electronic device receives image modification prompts input by a user; wherein the image modification prompts include descriptive information indicating the modification of at least one object in a first image, the first image being an image output based on the first text-based image prompt input by the user; and a second image is output based on the image modification prompts and the first image; wherein the image content of the image retention area of ​​the first image and the image content of the image retention area of ​​the second image are the same, the image retention area being the image region in the first image excluding the image editing area, and the image editing area being the object image region of the object to be edited indicated by the image modification prompts. In this solution, on the one hand, for the user, there is no need to re-input the complete text-based image prompts; the user only needs to input the image modification prompts indicating which content in the image to modify to generate the second image. Therefore, the generation process of the second image is more in line with human-computer interaction logic and is closer to the natural language interaction method of image modification between people. Furthermore, since the image modification prompts include descriptive information indicating the modification of at least one object in the first image, this application can directly edit the object image region of each object indicated by the image modification prompts without repeatedly inputting the prompts, thus achieving simultaneous editing of multiple regions of the first image and effectively improving image generation efficiency. On the other hand, since the second image only modifies the image content of the image editing area in the first image compared to the first image—that is, only modifying the image content mentioned by the user's image modification prompts—image regions not mentioned by the prompts are retained in the second image. In this way, the image content of the image regions that the user is satisfied with in the previously generated image is retained, and the image content of the image regions that the user wants to modify is updated, thereby obtaining an image that meets the user's actual modification needs.

[0386] It should be understood that in some embodiments of this application, the input unit 1504 may include a graphics processing unit (GPU) 15041 and a microphone 15042. The GPU 15041 processes image data of still images or videos obtained by an image capture device (such as a camera) in video capture mode or image capture mode. The display unit 1506 may include a display panel 15061, which may be configured in the form of a liquid crystal display, an organic light-emitting diode, or the like. The user input unit 1507 includes a touch panel 15071 and at least one of other input devices 15072. The touch panel 15071 is also called a touch screen. The touch panel 15071 may include a touch detection device and a touch controller. Other input devices 15072 may include, but are not limited to, physical keyboards, function keys (such as volume control buttons, power buttons, etc.), trackballs, mice, and joysticks, which will not be described in detail here.

[0387] The memory 1509 can be used to store software programs and various data. The memory 1509 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback, image playback, etc.). Furthermore, the memory 1509 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory 1509 in this embodiment includes, but is not limited to, these and any other suitable types of memory.

[0388] Processor 1510 may include one or more processing units; optionally, processor 1510 integrates an application processor and a modem processor, wherein the application processor mainly handles operations involving the operating system, user interface, and applications, and the modem processor mainly handles wireless communication signals, such as a baseband processor. It is understood that the aforementioned modem processor may also not be integrated into processor 1510.

[0389] This application also provides a readable storage medium storing a program or instructions. When the program or instructions are executed by a processor, they implement the various processes of the above-described text-to-image method embodiments and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0390] The processor is the processor in the electronic device described in the above embodiments. The readable storage medium includes computer-readable storage media, such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0391] This application embodiment also provides a chip, which includes a processor and a communication interface. The communication interface is coupled to the processor. The processor is used to run programs or instructions to implement the various processes of the above-described text-to-image method embodiment and achieve the same technical effect. To avoid repetition, it will not be described again here.

[0392] It should be understood that the chip mentioned in the embodiments of this application may also be referred to as a system-on-a-chip, system chip, chip system, or system-on-a-chip, etc.

[0393] This application provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the various processes of the above-described text-to-image method embodiments, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0394] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of this application is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0395] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a computer software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0396] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A method for generating images from text, comprising: The system receives image modification prompts input by the user; wherein the image modification prompts include descriptive information indicating that at least one object in a first image should be modified, and the first image is an image output based on the first text-based image prompt input by the user; Based on the image-modified prompt and the first image, output the second image; The image content of the image retention area of ​​the first image is the same as that of the image retention area of ​​the second image. The image retention area is the image region in the first image excluding the image editing area, and the image editing area is the object image region of the object to be edited indicated by the image modification prompt.

2. The method according to claim 1, wherein, The image modification prompt includes at least one of the following: The instruction is to replace the description information of at least one object in the first image; wherein the object includes at least one of the following: people, animals, plants, scenery, buildings, and items; Description information indicating modification of object attribute information of at least one object in the first image; wherein the object attributes include at least one of the following: object action, features, and state.

3. The method according to claim 1, wherein, The step of modifying the prompt word based on the image and the first image, and outputting the second image, includes: Based on the image modification prompt and the first image, a second textual image prompt and a mask image for each of the at least one object are generated; Based on the second text-based image prompt, the first text feature information of the first text-based image prompt, and the mask image of the at least one object, a second image is output.

4. The method according to claim 3, wherein, The image modification prompt includes modifying the description information of X objects in the first image, where X is a positive integer; The step of generating a mask image for each of the at least one objects based on the image-modified prompt and the first image includes: Based on the image modification prompts, X image editing areas are determined from the first image, and each image editing area is the object image region of one of the X objects; Obtain the mask image corresponding to each of the X image editing areas; When X is greater than 1, the mask images corresponding to the X image editing areas are superimposed to generate the mask images of the X objects; When X equals 1, the mask image corresponding to the image editing area of ​​an object is used as the mask image.

5. The method according to claim 4, wherein, The mask images corresponding to the X image editing areas have the same image size; The step of superimposing the mask images corresponding to the X image editing areas to obtain the mask images of the X objects includes: Add the pixel values ​​of pixels at the same pixel position in the mask image corresponding to the X image editing areas to obtain the cumulative pixel value corresponding to the pixel position, and calculate the ratio between the cumulative pixel value corresponding to the pixel position and X; Based on the ratio between the accumulated pixel value corresponding to the pixel position and X, a mask image of the X objects is generated.

6. The method according to claim 3 or 4, wherein, Before outputting the second image based on the second text-based image prompt, the first text feature information of the first text-based image prompt, and the mask image of the at least one object, the method further includes: Input the first textual image prompt into the image rendering model, and output the first image, the first text feature information, and the first feature map; The step of outputting a second image based on the second text-based image prompt, the first text feature information of the first text-based image prompt, and the mask image of the at least one object includes: The second textual image prompt, the first textual feature information, the first feature map, and the mask image of the at least one object are input into the image drawing model, and the second image is output.

7. The method according to claim 6, wherein, The step of inputting the second text-based image prompt, the first text feature information, the first feature map, and the mask image of the at least one object into the image rendering model, and outputting the second image, includes: The second textual image prompt, the first text feature information, the first feature map, and the mask image of the at least one object are input into the image drawing model; The image rendering model extracts the second text feature information of the second textual image prompt; The image rendering model generates a second feature map based on the second text feature information; The image rendering model fuses the second text feature information and the first text feature information to generate fused text feature information; The image rendering model outputs a second image based on the fused text feature information, the mask image of the at least one object, the first feature map, and the second feature map.

8. The method according to claim 7, wherein, The first text feature information includes sub-text feature information corresponding to each word in the first text-generated image prompt; the second text feature information includes sub-text feature information corresponding to each word in the second text-generated image prompt. The step of fusing the second text feature information and the first text feature information to obtain fused text feature information includes: Determine the first sub-text feature information of each object in at least one of the objects in the first text feature information; Determine the second sub-text feature information corresponding to the description information of each object in the second text feature information; Replace the first sub-text feature information of each object in the first text feature information with the second sub-text feature information corresponding to the description information of each object to obtain the fused text feature information.

9. The method according to claim 7 or 8, wherein, The step of outputting the second image based on the fused text feature information, the mask image of the at least one object, the first feature map, and the second feature map includes: A fused feature map is obtained based on the mask image of the at least one object, the first feature map, and the second feature map; Based on the fused feature map and the fused text feature information, a second image is output.

10. The method according to claim 9, wherein, The process of obtaining a fused feature map based on the mask image of the at least one object, the first feature map, and the second feature map includes: Use the mask image of the at least one object as the first weight image; A second weighted image is obtained by performing a difference operation between a preset pixel threshold and the pixel value of each pixel in the mask image of the at least one object; The first feature map and the second feature map are weighted and fused using the first weighted image and the second weighted image to obtain a fused feature map.

11. The method according to claim 9 or 10, wherein, The image rendering model includes a UNet network, which includes N cascaded convolutional layers. The first feature map includes N first feature maps, which correspond one-to-one with the N convolutional layers, where N is an integer greater than 1. The step of outputting a second image based on the fused feature map and the fused text feature information includes: The fused text feature information and the fused feature map are fused through the first convolutional layer in the UNet network to obtain the first first fused feature map. Based on the mask image of the at least one object, the first first feature map corresponding to the first convolutional layer, and the first first fused feature map, the first second fused feature map is obtained. The fused text feature information and the (i-1)th second fused feature map output by the (i-1)th convolutional layer in the UNet network are fused to obtain the ith first fused feature map. Based on the mask image of the at least one object, the ith first feature map corresponding to the ith convolutional layer, and the ith first fused feature map, the ith second fused feature map is obtained; i∈[2,N]; The second image is output through the Nth second fusion feature map output from the Nth convolutional layer in the UNet network.

12. The method according to claim 6, wherein, The step of inputting the first text-based image prompt into the image rendering model and outputting the first image, the first text feature information, and the first feature map includes: The first textual image prompt and the first noisy image are input into the image rendering model; The image rendering model extracts the first text feature information of the first textual image prompt word; The image rendering model generates a first feature map based on the first text feature information and the third feature map of the first noisy image; The image rendering model generates a first image based on the first feature map; The image rendering model outputs the first image, the first text feature information, and the first feature map.

13. The method according to claim 12, wherein, The image rendering model includes a UNet network, which consists of N cascaded convolutional layers, where N is an integer greater than 1. The step of outputting the first feature map based on the first text feature information and the third feature map includes: The first text feature information and the third feature map are fused through the first convolutional layer in the UNet network to obtain the first first feature map; The first text feature information and the first feature map output by the (j-1)th convolutional layer in the UNet network are fused to obtain the jth first feature map, so as to output N first feature maps. Where j∈[2,N]; the first feature map includes the N first feature maps.

14. The method according to claim 7, wherein, The step of inputting the second text-based image prompt, the first text feature information, the first feature map, and the mask image of the at least one object into the image rendering model includes: The second textual image prompt, the first text feature information, the first feature map, the mask image of the at least one object, and the second noise image are input into the image rendering model; The step of generating a second feature map based on the second text feature information includes: A second feature map is generated based on the second text feature information and the fourth feature map of the second noisy image.

15. The method according to claim 14, wherein, Before inputting the second text-based image prompt, the first text feature information, the first feature map, the mask image of the at least one object, and the second noise image into the image rendering model, the method further includes: Based on a first quantity M, the first region of the first image is filled with random noise M times to generate M third noise images, where the first region is the region where at least one object in the first image is located; M is an integer greater than 1. The M third noise images are reverse-processed to obtain M fourth noise images, and the second noise image includes the M fourth noise images; The process of generating a second feature map based on the second text feature information and the fourth feature map of the second noise image includes: Based on the second text feature information and the M fourth feature maps, M second feature maps are generated; The step of outputting a second image based on the fused text feature information, the first feature map, the second feature map, and the mask image of the at least one object includes: Based on the fused text feature information, the first feature map, the M second feature maps, and the mask image of the at least one object, M second images are output.

16. The method according to claim 1, wherein, Before receiving the image modification prompt input by the user, the method further includes: Receive the first text-based image prompt input by the user; The first message is displayed on the Wenshengtu interactive interface, and the first message includes the first image; After modifying the prompt word based on the image and the first image, and outputting the second image, the method further includes: The text-to-image interactive interface displays a second message, which includes the second image.

17. The method according to claim 16, wherein, After receiving the first text-based image prompt word input by the user, the method further includes: On the text-to-image interactive interface, a third message is displayed, which includes the first text-to-image prompt. After receiving the image modification prompt input by the user, the method further includes: The image editing interface displays a fourth message, which includes the image modification prompt.

18. A text-based image processing device, comprising: A receiving module is configured to receive image modification prompts input by a user; wherein the image modification prompts include descriptive information indicating modification of at least one object in a first image, and the first image is an image output based on the first text-based image prompt input by the user; The processing module is used to output a second image based on the image modification prompt word received by the receiving module and the first image; The image content of the image retention area of ​​the first image is the same as that of the image retention area of ​​the second image. The image retention area is the image area in the first image excluding the image editing area, and the image editing area is the object image area of ​​the object to be modified indicated by the image modification prompt.

19. The apparatus according to claim 18, wherein, The image modification prompt includes at least one of the following: The instruction is to replace the description information of at least one object in the first image; wherein the object includes at least one of the following: people, animals, plants, scenery, buildings, and items; Description information indicating modification of object attribute information of at least one object in the first image; wherein the object attributes include at least one of the following: object action, features, and state.

20. The apparatus according to claim 18, wherein, The processing module is specifically used for: Based on the image modification prompt and the first image, a second textual image prompt and a mask image for each of the at least one object are generated; Based on the second text-based image prompt, the first text feature information of the first text-based image prompt, and the mask image of the at least one object, a second image is output.

21. The apparatus according to claim 20, wherein, The image modification prompt includes modifying the description information of X objects in the first image, where X is a positive integer; the processing module is specifically used for: Based on the image modification prompts, X image editing areas are determined from the first image, and each image editing area is the object image region of one of the X objects; Obtain the mask image corresponding to each of the X image editing areas; When X is greater than 1, the mask images corresponding to the X image editing areas are superimposed to generate the mask images of the X objects; When X equals 1, the mask image corresponding to the image editing area of ​​an object is used as the mask image.

22. The apparatus according to claim 21, wherein, The mask images corresponding to the X image editing areas have the same image size; the processing module is specifically used for: Add the pixel values ​​of pixels at the same pixel position in the mask image corresponding to the X image editing areas to obtain the cumulative pixel value corresponding to the pixel position, and calculate the ratio between the cumulative pixel value corresponding to the pixel position and X; Based on the ratio between the accumulated pixel value corresponding to the pixel position and X, a mask image of the X objects is generated.

23. The apparatus according to claim 20 or 21, wherein, The processing module is further configured to input the first text-based image prompt into an image drawing model and output the first image, the first text feature information, and the first feature map before outputting the second image based on the second text-based image prompt, the first text feature information, and the mask image of the at least one object; The processing module is specifically used to input the second text image prompt, the first text feature information, the first feature map, and the mask image of the at least one object into the image drawing model, and output the second image.

24. The apparatus according to claim 23, wherein, The processing module is specifically used for: The second textual image prompt, the first text feature information, the first feature map, and the mask image of the at least one object are input into the image drawing model; The image rendering model extracts the second text feature information of the second textual image prompt; The image rendering model generates a second feature map based on the second text feature information; The image rendering model fuses the second text feature information and the first text feature information to generate fused text feature information; The image rendering model outputs the second image based on the fused text feature information, the mask image of the at least one object, the first feature map, and the second feature map.

25. The apparatus according to claim 24, wherein, The first text feature information includes sub-text feature information corresponding to each word in the first text-generated image prompt; the second text feature information includes sub-text feature information corresponding to each word in the second text-generated image prompt. The processing module is specifically used for: Determine the first sub-text feature information of each object in at least one of the objects in the first text feature information; Determine the second sub-text feature information corresponding to the description information of each object in the second text feature information; The first sub-text feature information of each object in the first text feature information is replaced with the second sub-text feature information corresponding to the description information of each object to obtain the fused text feature information.

26. The apparatus according to claim 24 or 25, wherein, The processing module is specifically used for: A fused feature map is obtained based on the mask image of the at least one object, the first feature map, and the second feature map; Based on the fused feature map and the fused text feature information, the second image is output.

27. The apparatus according to claim 26, wherein, The processing module is specifically used for: Use the mask image of the at least one object as the first weight image; A second weighted image is obtained by performing a difference operation between a preset pixel threshold and the pixel value of each pixel in the mask image of the at least one object; The first feature map and the second feature map are weighted and fused using a first weighted image and a second weighted image to obtain the fused feature map.

28. The apparatus according to claim 26 or 27, wherein, The image rendering model includes a UNet network, which includes N cascaded convolutional layers. The first feature map includes N first feature maps, which correspond one-to-one with the N convolutional layers, where N is an integer greater than 1. The processing module is specifically used for: The fused text feature information and the fused feature map are fused through the first convolutional layer in the UNet network to obtain the first first fused feature map. Based on the mask image of the at least one object, the first first feature map corresponding to the first convolutional layer, and the first first fused feature map, the first second fused feature map is obtained. The fused text feature information and the (i-1)th second fused feature map output by the (i-1)th convolutional layer in the UNet network are fused to obtain the ith first fused feature map. Based on the mask image of the at least one object, the ith first feature map corresponding to the ith convolutional layer, and the ith first fused feature map, the ith second fused feature map is obtained. i∈[2,N]; The second image is output by the Nth second fusion feature map output from the Nth convolutional layer in the UNet network.

29. The apparatus according to claim 23, wherein, The processing module is specifically used to: input the first text-based image prompt and the first noise image into the image drawing model; The image rendering model extracts the first text feature information of the first text-based image prompt word; The image rendering model generates the first feature map based on the first text feature information and the third feature map of the first noisy image; The image rendering model generates the first image based on the first feature map; The image rendering model outputs the first image, the first text feature information, and the first feature map.

30. The apparatus according to claim 29, wherein, The image rendering model includes a UNet network, which consists of N cascaded convolutional layers, where N is an integer greater than 1. The processing module is specifically used for: The first text feature information and the third feature map are fused through the first convolutional layer in the UNet network to obtain the first first feature map; The first text feature information and the first feature map output by the (j-1)th convolutional layer in the UNet network are fused to obtain the jth first feature map, so as to output N first feature maps. Where j∈[2,N]; the first feature map includes the N first feature maps.

31. The apparatus according to claim 24, wherein, The processing module is specifically used for: The second textual image prompt, the first text feature information, the first feature map, the mask image of the at least one object, and the second noise image are input into the image rendering model; The second feature map is generated based on the second text feature information and the fourth feature map of the second noisy image.

32. The apparatus according to claim 31, wherein, The device further includes: a noise-adding module and a reverse processing module; The noise-adding module is used to fill the first region of the first image with random noise M times based on a first quantity M before inputting the second text image prompt, the first text feature information, the first feature map, the mask image of the at least one object and the second noise image into the image rendering model, thereby generating M third noise images. The first region is the region where the at least one object in the first image is located; M is an integer greater than 1. The reverse processing module is used to reverse process the M third noise images to obtain M fourth noise images, wherein the second noise image includes the M fourth noise images; The processing module is specifically used to generate M second feature maps based on the second text feature information and the M fourth feature maps; Based on the fused text feature information, the first feature map, the M second feature maps, and the mask image of the at least one object, M second images are output.

33. The apparatus according to claim 18, wherein, The receiving module is also configured to receive the first text image prompt word input by the user before receiving the image modification prompt word input by the user; The device further includes a display module for displaying a first message on the image-text interactive interface, the first message including the first image; The display module is further configured to display a second message on the text-to-image interactive interface, the second message including the second image.

34. The apparatus according to claim 33, wherein, The display module is also used to display a third message on the text-to-image interactive interface after receiving the first text-to-image prompt word input by the user, the third message including the first text-to-image prompt word; The display module is also configured to display a fourth message on the text-to-image interactive interface after receiving the image modification prompt word input by the user, the fourth message including the image modification prompt word.

35. An electronic device comprising a processor and a memory, the memory storing a program or instructions executable on the processor, the program or instructions, when executed by the processor, implementing the steps of the text-based image method as claimed in any one of claims 1 to 17.

Citation Information

Patent Citations

  • Image content regeneration method and device, equipment and storage medium

    CN116740210A

  • Semantic-based image editing method and system and medium

    CN117274436A

  • Video generation method and device, electronic equipment and readable storage medium

    CN118042246A

  • Figure generation method and device and electronic equipment

    CN118691700A

  • Computer-assisted text and visual styling for images

    US10049477B1