Image generation method and apparatus, electronic device, and storage medium

By obtaining reference text information and reference images, determining image generation guidance information, and generating target images using the target text image model, the problem of large randomness in image generation in the prior art is solved, and the user's personalized image generation control and similarity requirements are realized.

WO2025146025A1PCT designated stage expired Publication Date: 2025-07-10BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/143995
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-03
Filing Date
2024-12-30
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

The images generated by the existing literary and artistic graphics model are relatively random, and the image content specified by the user is difficult to accurately reflect in the generated images, resulting in a large difference between the generated images and the user's expectations and cannot meet the actual needs of the users.

Method used

By acquiring the reference text information and the reference image, the image generation guidance information is determined, and the influence weight of the reference text information and the reference image to different regions of the target image to be generated is indicated. The target image is generated using the target text image model, so that the second subject in the generated image is similar to the first subject.

Benefits of technology

It realizes the accurate definition of the appearance of things in the image when generating images, meets the user's personalized image generation control needs, ensures the similarity between the generated image and the reference image, and maintains the generation control ability of the background part.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024143995_10072025_PF_FP_ABST
    Figure CN2024143995_10072025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to an image generation method and apparatus, an electronic device, and a storage medium. The method comprises: acquiring reference text information and a reference image, wherein the reference text information comprises a target phrase, the reference image comprises a first subject, and the target phrase is related to the first subject; determining image generation guidance information, wherein the image generation guidance information is used for indicating influence weights of the reference text information and the reference image on different areas of a target image to be generated, and the influence weights of the reference image on different areas of said target image are different; and inputting the reference text information, the reference image, and the image generation guidance information into a target text-to-image model to generate a target image, wherein the target image comprises a second subject, and the second subject is similar to the first subject.
Need to check novelty before this filing date? Find Prior Art

Description

Image generation method, device, electronic device, and storage medium

[0001] This application claims priority to the Chinese invention patent application entitled “Image Generation Method, Device, Electronic Device and Storage Medium” filed on January 3, 2024, with application number 202410010715.9. The entire contents of that application are incorporated herein by reference. Technical Field

[0002] The present disclosure relates to the field of artificial intelligence technology, and in particular to an image generation method, device, electronic device, and storage medium. Background Art

[0003] With the continuous development of deep learning technology, pre-trained text-based graph models have become a research hotspot in the field of artificial intelligence. These models can automatically generate images that match a given text description. This technology has not only attracted widespread attention in academia but has also been widely applied in industries such as virtual reality, game design, and advertising creativity.

[0004] However, the images generated by existing text-based image models are highly random. Even if the user inputs a specified image, the objects in the user-specified image cannot be optimally reflected in the generated image. This makes the generated image often significantly different from the image the user expects and cannot meet the user's actual needs. Summary of the Invention

[0005] In order to solve the above technical problems or at least partially solve the above technical problems, the present disclosure provides an image generation method, an apparatus, an electronic device, and a storage medium.

[0006] In a first aspect, the present disclosure provides an image generation method, including: obtaining reference text information and a reference image; the reference text information includes a target phrase; the reference image includes a first subject; the target phrase is related to the first subject; determining image generation guidance information, the image generation guidance information is used to indicate the influence weights of the reference text information and the reference image on different areas of the target image to be generated; the reference image has different influence weights on different areas of the target image; inputting the reference text information, the reference image, and the image generation guidance information into a target text image model to generate a target image, the target image including a second subject, and the second subject is similar to the first subject.

[0007] In a second aspect, the present disclosure also provides an image generation device, including: an acquisition module, used to acquire reference text information and a reference image; the reference text information includes a target phrase; the reference image includes a first subject; the target phrase is related to the first subject; a generation guidance information determination module, used to determine image generation guidance information, the image generation guidance information is used to indicate the influence weights of the reference text information and the reference image on different areas of the target image to be generated; the reference image has different influence weights on different areas of the target image; an image generation module, used to input the reference text information, the reference image, and the image generation guidance information into a target text image model to generate a target image, the target image includes a second subject, and the second subject is similar to the first subject.

[0008] In a third aspect, the present disclosure also provides an electronic device, comprising: one or more processors; a storage device for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the image generation method as described above.

[0009] In a fourth aspect, the present disclosure further provides a computer-readable storage medium having a computer program stored thereon, which implements the image generation method as described above when the program is executed by a processor.

[0010] In a fifth aspect, the present disclosure further provides a computer program product, which includes computer-executable instructions, and which implements the above-mentioned image generation method when executed by a processor. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.

[0012] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0013] FIG1 is a flow chart of an image generation method provided by an embodiment of the present disclosure;

[0014] FIG2 is a flowchart of a method for implementing S120 provided in an embodiment of the present disclosure;

[0015] FIG3 is a flow chart of a method for implementing S123 provided in an embodiment of the present disclosure;

[0016] 4-6 are schematic diagrams of the image generation method provided by the embodiments of the present disclosure;

[0017] FIG7 is a schematic structural diagram of an image generating device according to an embodiment of the present disclosure;

[0018] FIG8 is a schematic structural diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0019] In order to more clearly understand the above-mentioned objectives, features and advantages of the present disclosure, the scheme of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features therein can be combined with each other in the absence of conflict.

[0020] In the following description, many specific details are set forth to facilitate a full understanding of the present disclosure, but the present disclosure may also be implemented in other ways different from those described herein; it is obvious that the embodiments in the specification are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0021] FIG1 is a flow chart of an image generation method provided by an embodiment of the present disclosure. This embodiment is applicable to situations where image generation is performed on a client or server. The method can be executed by an image generation device, which can be implemented in software and / or hardware. The device can be configured in an electronic device, such as a terminal or a server. If the electronic device is a terminal, it specifically includes but is not limited to a smartphone, a PDA, a tablet computer, a wearable device with a display screen, a desktop computer, a laptop computer, an all-in-one computer, a smart home device, etc.

[0022] As shown in FIG1 , the method may specifically include:

[0023] S110 , obtaining reference text information and a reference image; the reference text information includes a target phrase; the reference image includes a first subject; the target phrase is related to the first subject.

[0024] The reference text information may be, for example, text information serving as a basis for a text map, which may be text description information of an image that a user wishes to generate, that is, prompts in the text map technology.

[0025] The target phrase may be, for example, a phrase in the reference text information, which defines what the target image needs to include. The target image may specifically refer to a person, an animal, or an object.

[0026] It should be noted that after obtaining the reference text information, the electronic device may not be able to directly determine which phrase in the reference text information is the target phrase, and needs to use a text mask to determine it.

[0027] The text mask may be, for example, information used to describe the position of the target phrase in the reference text information, and is used to prompt the user what content in the reference image the user wishes to emphasize in the image the user wishes to generate.

[0028] There are many methods for obtaining a text mask, and this application does not impose any restrictions thereto. Optionally, a target phrase is determined by interacting with the user; the target phrase is compared with the reference text information to determine the position of the target phrase in the reference text information, thereby obtaining a text mask. Alternatively, the reference text information is segmented to obtain multiple phrases; the target phrase is determined by matching each phrase with an object in a reference image; the target phrase is compared with the reference text information to determine the position of the target phrase in the reference text information, thereby obtaining a text mask. The target phrase is determined by matching each phrase with an object in a reference image. For example, it may be determined which or which phrases specify objects that appear in the reference image, and the phrases corresponding to the objects that appear in the reference image are used as target phrases. The objects in the reference image are things in the reference image, such as people, animals, or objects.

[0029] The reference image can be, for example, a user-provided image that reflects the appearance of the subject that the user wishes the target phrase to refer to in the generated image (i.e., the final output image). The first subject is an object in the reference image, which can specifically be a person, animal, or object. The target phrase is related to the first subject, for example, the first subject can be an example of the target phrase.

[0030] For example, the reference image depicts a white Persian cat sleeping on a windowsill. The first subject is a kitten, and the reference text is "Kitten running on the grass." The target phrase is "kitten." The white Persian cat in the reference image is associated with the target phrase "kitten" because the target phrase "kitten" can be used to refer to a white Persian cat. The white Persian cat in the reference image can be understood as an example of the target phrase "kitten."

[0031] S120 , determining image generation guidance information, where the image generation guidance information indicates influence weights of the reference text information and the reference image on different regions of the target image to be generated; the reference image has different influence weights on different regions of the target image.

[0032] S130: Input the reference text information, the reference image, and the image generation guidance information into the target text-image model to generate a target image, where the target image includes a second subject, and the second subject is similar to the first subject.

[0033] The target image is the image generated by the target text graph model. The second subject is the object in the target image, which can be a person, an animal, or an object.

[0034] The second subject is related to the first subject. For example, the similarity between the second subject and the first subject may exceed a set similarity threshold. In other words, the second subject is generated based on the first subject, and the two subjects share a similar appearance. In some scenarios, the second subject and the first subject can be considered the same subject. That is, the subject in the target image and the reference image is the same.

[0035] For example, the content of the reference image is a white Persian cat sleeping on the windowsill. The first subject is the white Persian cat in the reference image (hereinafter referred to as white Persian cat A). The reference text information is "kitten running on the grass". The target phrase is "kitten". The image generation guidance information indicates which area or areas in the target image need to consider the influence of the reference image on image generation while generating the image based on the reference text information; and which areas only need to generate the image based on the reference text information. The content of the target image finally generated is another white Persian cat running on the grass. The second subject is the white Persian cat in the target image (hereinafter referred to as white Persian cat B). Comparing the reference image and the target image, the two subjects (white Persian cat A and white Persian cat B) are similar in appearance. To a certain extent, they can be considered to be the same white Persian cat; but white Persian cat A and white Persian cat B have different movements and different backgrounds.

[0036] The above technical solution obtains reference text information and a reference image by setting; the reference text information includes a target phrase; the reference image includes a first subject; the target phrase is related to the first subject; determines image generation guidance information, the image generation guidance information is used to indicate the influence weight of the reference text information and the reference image on different areas of the target image to be generated; the reference image has different influence weights on different areas of the target image; inputs the reference text information, the reference image, and the image generation guidance information into a target text-image model to generate a target image, the target image including a second subject that is similar to the first subject. This solution provides an image generation method that allows users to define the appearance of objects in the generated image, which can meet the user's needs for personalized image generation control.

[0037] In addition, during the process of generating the target image, the above-mentioned technical solution can accurately limit the influence range of the reference text information and the reference image on the target image through the image generation guidance information, so that the second subject is similar to the first subject, while ensuring that the reference text information's ability to control the generation of other parts of the target image (such as the background part of the image) is not affected.

[0038] In the above technical solution, there are multiple specific implementation methods for S120. This application does not limit this. In one embodiment, optionally, the implementation method of S120 may include:

[0039] S121: Input reference text information into a target text-based image model to obtain a first image.

[0040] It should be emphasized that, in this step, in the process of generating the first image, the reference text information is used but the reference image is not used, that is, the generation of the first image is completely controlled by the reference text information.

[0041] S122: Determine an image mask based on the first image and the target phrase.

[0042] The image mask is used to indicate the influence range of the reference image on the target image.

[0043] It's important to note that the image mask is determined based on the first image and the target phrase because, in some scenarios, it can be roughly assumed that the position of the subject remains unchanged in images generated using the same reference text information. By determining the image mask based on the first image and the target phrase, we can more accurately estimate the position of the second subject in the target image and use the image mask to indicate the extent to which the reference image affects the target image.

[0044] For example, assuming the reference text information is "kitten on the grass", the same text graph model is used to generate an image based on the reference text information. If multiple generations are performed, the kitten's movements, colors, and breeds may be different in the generated images, but the position of the kitten in each generated image tends to be consistent. The generated image is processed through this step, and the resulting image mask describes the position of the kitten in the image. It can be used to estimate the position of the kitten in the target image and indicate the position of the kitten in the target image with the help of the image mask. This position is the position where the reference image needs to act, that is, the range of influence of the reference image on the target image.

[0045] There are many ways to implement this step, and this application does not limit this. Exemplarily, the implementation method of this step includes: using a cross-attention mechanism to process the target phrase and the first image to obtain a first cross-attention map. The first cross-attention map includes multiple attention values; different attention values ​​correspond to different areas in the first image, and the size of the attention value reflects the correlation between the target phrase and the area corresponding to the attention value; based on the first cross-attention map and the first preset threshold range, the image mask is determined.

[0046] Optionally, determining the image mask based on the first cross-attention map and the first preset threshold range may include: retaining attention values ​​within the first preset threshold range in the first cross-attention map, setting attention values ​​outside the first preset threshold range in the first cross-attention map to 0, to obtain a second cross-attention map. Normalizing the second cross-attention map to obtain the image mask.

[0047] When normalizing the second cross-attention map, it can be normalized according to the maximum value of the second cross-attention map, that is, normalized according to the following formula:

[0048] in, is the attention value in the second cross attention map, For the attention value Normalized results.

[0049] As above, optionally, before S120 , the method further includes: acquiring a text mask; the text mask is used to indicate the position of the target phrase in the reference text information; and determining the target phrase in the reference text information based on the text mask.

[0050] S123. Determine image generation guidance information based on the image mask.

[0051] There are many ways to implement this step, and this application does not limit this. Optionally, the area in the target image that corresponds to the image mask is called the second target area, and the area in the target image that does not correspond to the image mask is called the second non-target area. In one embodiment, the influence weight of the reference image on the second target area can be simply set to a, and the influence weight of the reference text information on the second non-target area can be simply set to b, where a>b, and both a and b are positive numbers.

[0052] In another embodiment, optionally, referring to FIG3 , the implementation method of this step includes:

[0053] S310 : Determine a key area in the reference image based on the image features of the current image to be denoised, the image mask, the text features of the reference text information, and the text mask.

[0054] Key regions are areas within a reference image that require particular consideration or emphasis when generating the target image. Typically, these areas include all or part of the first subject. Whether or not these areas are given particular consideration or emphasis directly impacts the degree of similarity between the second subject and the first.

[0055] Optionally, the specific implementation method of S310 may be: dividing the reference image into multiple first sub-regions; based on the image features of the current image to be denoised and the image mask, respectively calculating the first correlation value between the image of each first sub-region and the image of the third target region; the third target region is the region corresponding to the current image to be denoised and the image mask; based on the text features of the reference text information and the text mask, respectively calculating the second correlation value between the image of each first sub-region and the target phrase; based on the first correlation value and the second correlation value, determining the final score of each first sub-region in the reference image; based on the final score of each first sub-region in the reference image, determining the key region from the reference image.

[0056] Furthermore, referring to FIG4 , the reference image is divided into n first subregions; the reference text information is encoded using an encoder to obtain text features of the reference text information; based on the text features and text mask of the reference text information, text features of the target phrase in the reference text information are obtained. The text features of these target phrases are aggregated using a weighted pooling method to obtain text context features; the text context features are replicated to obtain multiple text context features, thereby ensuring a one-to-one correspondence between the text context features and the first subregions; based on each pair of text context features and the image features of the image within the first subregion, a second correlation value between the image of the first subregion and the target phrase is obtained. The second correlation values ​​of all the first subregion images and the target phrase are organized into a second correlation graph. The current image to be denoised is encoded using an encoder to obtain image features of the current image to be denoised; image features of the third target region are obtained based on the image features of the current image to be denoised and an image mask; the image features of the third target region are aggregated using a weighted pooling method to obtain image context features; the image context features are replicated to obtain multiple image context features, thereby making the image context features correspond one-to-one with the first subregion; based on the image context features and the image features of the images within the first subregion, a first correlation value between the image of the first subregion and the image of the third target region is obtained. The first correlation values ​​of all the images of the first subregion and the images of the third target region are organized into a first correlation graph.

[0057] Alternatively, the text context feature C can be obtained by using the following formula (1) and formula (2): textual :

[0058] Using equations (3) and (4) below, we can get the image context feature C visual :

[0059] in, and is the trainable parameter matrix, f ctis the text feature of the target phrase, z t is the image feature of the third target area.

[0060] Optionally, the first correlation value S visual and the second correlation value S textual Substituting into formula (5), we can obtain the final score S of the first sub-region.

[0061] in, is a preset parameter and is a fixed value.

[0062] Continuing to refer to FIG4 , the key area can be determined from the reference image with the help of the second threshold range. For example, if the screening threshold is set to [0.05, 1], the result shown in FIG4 is obtained.

[0063] S320 : Determine image generation guidance information based on image features of the current image to be denoised, image features of the key area, text features of the reference text information, and the image mask.

[0064] There are multiple methods for implementing this step, which are not limited in this application. For example, the method for implementing this step may include: determining a first influence weight based on the image features of the current image to be denoised, the image features of the key area, and the image mask; determining a second influence weight based on the image features of the current image to be denoised and the text features of the reference text information; and determining image generation guidance information based on the first influence weight and the second influence weight.

[0065] Further, referring to Figure 5, the first influence weight is determined based on the image features of the current image to be denoised, the image features of the key area, and the image mask. Specifically, it can include: utilizing the cross-attention mechanism to determine the visual cross-attention map of the current image to be denoised and the key area based on the image features of the current image to be denoised and the image features of the key area; determining the first influence weight based on the visual cross-attention map of the current image to be denoised and the key area and the image mask.

[0066] Furthermore, determining the first influence weight based on the visual cross-attention map of the current image to be denoised and the key area and the image mask may include: taking the product of the visual cross-attention map of the current image to be denoised and the key area and the image mask as the first influence weight.

[0067] As mentioned above, since the key area is determined based on the image mask, the step of "using the cross-attention mechanism to determine the visual cross-attention map of the current image to be denoised and the key area based on the image features of the current image to be denoised and the image features of the key area" can achieve the purpose of limiting the influence range of the reference image.

[0068] By setting the first influence weight based on the visual cross-attention map of the current image to be denoised and the key area and the image mask, it is possible to accurately limit the influence range of the reference image and ensure the reference text information's ability to control image generation in areas outside the reference image's influence range.

[0069] The specific implementation method of "determining the second influence weight based on the image features of the current image to be denoised and the text features of the reference text information" may include: utilizing the cross-attention mechanism to determine the text cross-attention map of the current image to be denoised and the reference text information based on the image features of the current image to be denoised and the text features of the reference text information; and determining the second influence weight based on the cross-attention.

[0070] A specific implementation method of “determining the image generation guidance information based on the first influence weight and the second influence weight” may include: using the sum of the first influence weight and the second influence weight as the image generation guidance information.

[0071] By adopting the above-mentioned method for determining image generation guidance information, it is possible to limit the reference image to only affect the image generation in a local area of ​​the target image, rather than affecting the image generation in the entire area of ​​the target image. Those skilled in the art will understand that if the reference image and the reference text information simultaneously affect the image generation in the entire area of ​​the target image, if the influence weight of the reference image is relatively heavy, the control ability of the reference text information over the image generation will be weakened. If the influence weight of the reference text information is relatively heavy, the second subject in the generated target image will be dissimilar to the first subject. The above-mentioned method can be used to ensure that the reference text information has the ability to control the generation of the target image while ensuring that the second subject is similar to the first subject.

[0072] On the basis of the above technical solution, optionally, after S130 , the target image may be directly output as the final image, or the method may be configured to further include: fusing the first image and the target image to obtain the final image.

[0073] There are various specific implementation methods for "fusing the first image and the target image to obtain a final image", which are not limited in this application. For example, if the first image includes a first target area and a first non-target area, the first target area includes a subject corresponding to the target phrase; the first non-target area is an image in the first image other than the first target area; the target image includes a second target area, and the second target area includes a second subject; fusing the first image and the target image to obtain a final image includes: fusing the image of the first target area in the first image with the image of the second target area in the target image to obtain a third image; and fusing the third image with the image of the first non-target area in the first image to obtain the final image.

[0074] Exemplarily, the final image can be obtained using the following formula.

[0075] Among them, z is the final image, z TI is the target image, z T is the first image, z TI -z T For the third image, In the absence of reference text information and reference image input, the target text graph model is used to generate the result of denoising the image to be denoised.

[0076] Based on the above technical solution, optionally, the process of generating the first image using the target text graph model includes N rounds of image denoising; each round of image denoising includes: determining image generation guidance information; inputting reference text information, reference image, and image generation guidance information into the target text graph model to generate a target image.

[0077] For example, as shown in Figure 6, the text-to-image system includes an image mask determination unit, an adaptive scoring module, a guidance information determination unit, and a target text graph model. The image mask determination unit is configured to obtain an image mask based on a first image and a target phrase; the adaptive scoring module is configured to determine key regions in a reference image; the guidance information determination unit is configured to generate image generation guidance information; and the target text graph model is configured to perform denoising on the image to be denoised based on the image generation guidance information, reference text information, and the reference image.

[0078] Assume that the Wenshengtu process includes N rounds of image denoising, where N is a positive integer greater than or equal to 1. Each round of image denoising can be divided into two stages: the first stage is the reference image influence range determination stage, and the second stage is the reference image influence range injection stage.

[0079] Assuming that the current noise reduction process is the Pth round, where P is a positive integer less than or equal to N, in the first stage, no reference image is used, and the adaptive scoring module does not operate. Because no reference image is used, the first influence weight obtained by the guidance information determination unit is 0. The guidance information determination unit also determines a second influence weight based on the image features of the current image to be denoised and the text features of the reference text information; determines image generation guidance information 1 based on the first and second influence weights; and inputs image generation guidance information 1 into the target text graph model. The target text graph model performs noise reduction processing on the Pth round of noise reduction based on image generation guidance information 1, obtaining a first image. The image mask determination unit obtains an image mask based on the first image and the target phrase.

[0080] In the second stage, using the reference image, the adaptive scoring module works to determine the key area in the reference image based on the image features of the image to be denoised in the P-th round, the image mask, the text features of the reference text information, and the text mask. The guidance information determination unit determines the first influence weight based on the image features of the current image to be denoised, the image features of the key area, and the image mask; determines the second influence weight based on the image features of the current image to be denoised and the text features of the reference text information; and determines the image generation guidance information 2 based on the first influence weight and the second influence weight. The target text-to-image model performs denoising processing on the image to be denoised in the P-th round based on the image generation guidance information 2 to obtain the target image.

[0081] Then, the target image obtained by denoising in the P-th round can be used as the final image of this round of denoising, or the fusion result of the first image obtained by denoising in the P-th round and the target image obtained by denoising in the P-th round can be used as the final image.

[0082] Subsequently, if P = M, the final image of the P-th round of denoising is output. If P < M, the final image of the P-th round of denoising is used as the image to be denoised in the (P + 1)-th round, and the (P + 1)-th round of denoising continues. It should be noted that in the same round of denoising, optionally, the image to be denoised in the P-th round used in the first stage and the second stage is the same image. When performing the first round of denoising, the image to be denoised in the first round used in the first stage and the second stage is a pre-specified noise image.

[0083] It should also be noted that in FIG. 6, the adaptive scoring module also uses a time step when working, that is, based on the image features of the image to be denoised in the P-th round, the image mask, the time information in the P-th round, the text features of the reference text information, and the text mask, the key area is determined in the reference image. Here, the time step is used to control the fineness of the denoising of the current target text-to-image model.

[0084] In FIG. 6, for ease of understanding, the guidance information determination unit and the target text-to-image model are both drawn in the first stage and the second stage. Optionally, in practice, the text generation image system only includes one guidance information determination unit and one target text-to-image model. That is, in FIG. 6, the guidance information determination unit used in the first stage and the second stage is the same unit, and the target text-to-image model used in the first stage and the second stage is the same model.

[0085] Based on the above technical solution, optionally, the method further includes: obtaining a sample data pair, the sample data pair including a sample image, sample text information and a sample noise image; the sample image including a third subject, the sample text information including a target sample phrase, and the target sample phrase being related to the third subject; the sample noise image being the result of adding noise to the sample image; determining sample image generation guidance information, the sample image generation guidance information indicating the influence weights of the sample text information and the sample image on different areas of the fourth image to be generated; inputting the sample image, the sample text information and the sample image generation guidance information sample into the target text graph model to be trained to obtain the fourth image; and adjusting the parameters of the target text graph model and / or the model used in the process of determining the sample image generation guidance information based on the similarity between the fourth image and the sample image.

[0086] Among them, the sample image has a similar meaning to the reference image, the sample text information has a similar meaning to the reference sample information, and the third subject has a similar meaning to the first subject. The target sample phrase has a similar meaning to the target phrase. The sample image generation guidance information has a similar meaning to the image generation guidance information, and the fourth image has a similar meaning to the target image. The only difference is that the sample image, sample text information, third subject, target sample phrase, sample image generation guidance information, and fourth image are used in the stage of training the text-to-image system, while the reference image, reference sample information, first subject, target phrase, image generation guidance information, and target image are used in the stage of generating images using the text-to-image system model.

[0087] Based on the similarity between the fourth image and the sample image, parameters in the target text graph model and / or the model used in determining the guidance information for generating the sample image are adjusted. For example, based on the similarity between the fourth image and the sample image, one or more parameters in the image mask determination unit, the adaptive scoring module, the guidance information determination unit, and the target text graph model in the text-to-image system can be adjusted. This allows the text-to-image system to use images to define the appearance of objects in the generated image, thereby meeting the user's needs for personalized image generation control.

[0088] The technical solution provided by the embodiment of the present disclosure has the following advantages over the prior art: the technical solution provided by the embodiment of the present disclosure obtains reference text information and a reference image through settings; the reference text information includes a target phrase; the reference image includes a first subject; the target phrase is related to the first subject; image generation guidance information is determined, and the image generation guidance information is used to indicate the influence weights of the reference text information and the reference image on different areas of the target image to be generated; the reference image has different influence weights on different areas of the target image; the reference text information, the reference image, and the image generation guidance information are input into the target text image model to generate a target image, and the target image includes a second subject, and the second subject is similar to the first subject. It provides an image generation method that allows the user to limit the appearance of things in the generated image, which can meet the user's personalized image generation control needs.

[0089] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0090] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0091] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0092] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0093] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present disclosure.

[0094] FIG7 is a schematic diagram of the structure of an image generation device in an embodiment of the present disclosure. The image generation device provided in the embodiment of the present disclosure can be configured in a client or a server. Referring to FIG7 , the image generation device specifically includes:

[0095] An acquisition module 510 is configured to acquire reference text information and a reference image; the reference text information includes a target phrase; the reference image includes a first subject; and the target phrase is related to the first subject.

[0096] A generation guidance information determination module 520 is configured to determine image generation guidance information, wherein the image generation guidance information is configured to indicate the influence weights of the reference text information and the reference image on different regions of the target image to be generated; the reference image has different influence weights on different regions of the target image;

[0097] The image generation module 530 is configured to input the reference text information, the reference image, and the image generation guidance information into a target text graph model to generate a target image, where the target image includes a second subject that is similar to the first subject.

[0098] Furthermore, a guidance information determination module 520 is generated, which is used to:

[0099] Inputting the reference text information into the target text graph model to obtain a first image;

[0100] determining an image mask based on the first image and the target phrase;

[0101] Based on the image mask, the image generation guidance information is determined.

[0102] Furthermore, a guidance information determination module 520 is generated, which is used to:

[0103] Before determining an image mask based on the first image and the target phrase, obtaining a text mask; the text mask is used to indicate a position of the target phrase in the reference text information;

[0104] The target phrase is determined in the reference text information based on the text mask.

[0105] Furthermore, the process of generating an image using the target text graph model is a process of performing denoising on the image to be denoised;

[0106] Generate guidance information determination module 520, which is used to:

[0107] Determining a key area in the reference image based on image features of the current image to be denoised, the image mask, text features of the reference text information, and the text mask;

[0108] The image generation guidance information is determined based on the image features of the current image to be denoised, the image features of the key area, the text features of the reference text information, and the image mask.

[0109] Furthermore, a guidance information determination module 520 is generated, which is used to:

[0110] Dividing the reference image into a plurality of first sub-regions;

[0111] Based on the image features of the current image to be denoised and the image mask, respectively calculating first correlation values ​​between the images of each of the first sub-regions and the image of a third target region; the third target region is the region corresponding to the current image to be denoised and the image mask;

[0112] Calculating a second relevance value between the image of each first sub-region and the target phrase based on the text features of the reference text information and the text mask;

[0113] determining a final score of each of the first sub-regions in the reference image based on the first correlation value and the second correlation value;

[0114] A key region is determined from the reference image based on the final score of each of the first sub-regions in the reference image.

[0115] Furthermore, a guidance information determination module 520 is generated, which is used to:

[0116] determining a first influence weight based on the image features of the current image to be denoised, the image features of the key area, and the image mask;

[0117] determining a second influence weight based on image features of the current image to be denoised and text features of the reference text information;

[0118] The image generation guidance information is determined based on the first influence weight and the second influence weight.

[0119] Furthermore, the device also includes a synthesis module, which is used to:

[0120] The reference text information, the reference image, and the image generation guidance information are input into a target text-based graph model. After a target image is generated, the first image and the target image are fused to obtain a final image.

[0121] Furthermore, the first image includes a first target area and a first non-target area, the first target area includes a subject corresponding to the target phrase; the first non-target area is an image other than the first target area in the first image; the target image includes a second target area, and the second target area includes the second subject;

[0122] Synthesis modules for:

[0123] The fusing the first image and the target image to obtain a final image includes:

[0124] fusing an image of a first target area in the first image and an image of a second target area in the target image to obtain a third image;

[0125] The third image is fused with the image of the first non-target area in the first image to obtain a final image.

[0126] Furthermore, the Wensheng graph process includes N rounds of image denoising;

[0127] Each round of image denoising includes: determining image generation guidance information; inputting the reference text information, the reference image, and the image generation guidance information into the target text-image model to generate a target image.

[0128] Furthermore, the device also includes a training module for:

[0129] Acquire a sample data pair, the sample data pair comprising a sample image, sample text information, and a sample noise image; the sample image comprises a third subject, the sample text information comprises a target sample phrase, and the target sample phrase is related to the third subject; the sample noise image is a result of adding noise to the sample image;

[0130] determining sample image generation guidance information, where the sample image generation guidance information indicates influence weights of the sample text information and the sample image on different regions of a fourth image to be generated;

[0131] Inputting the sample image, sample text information, and sample image generation guidance information sample into the target text-based image model to be trained to obtain a fourth image;

[0132] Based on the similarity between the fourth image and the sample image, parameters of the target cultural graph model and / or the model used in the process of determining the sample image to generate the guidance information are adjusted.

[0133] The image generation device provided in the embodiment of the present disclosure can execute the steps executed by the client or server in the image generation method provided in the embodiment of the method of the present disclosure, and has the execution steps and beneficial effects, which will not be repeated here.

[0134] FIG8 is a schematic diagram of the structure of an electronic device in an embodiment of the present disclosure. Specific reference will be made to FIG8 below, which shows a schematic diagram of the structure of an electronic device 1000 suitable for implementing an embodiment of the present disclosure. The electronic device 1000 in the embodiment of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable electronic devices, and the like, as well as fixed terminals such as digital TVs, desktop computers, smart home devices, and the like. The electronic device shown in FIG8 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.

[0135] As shown in FIG8 , the electronic device 1000 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes to implement the image generation method of the embodiment described in the present disclosure according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1008 into a random access memory (RAM) 1003. Various programs and information required for the operation of the electronic device 1000 are also stored in the RAM 1003. The processing device 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0136] Typically, the following devices may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device 1000 to communicate with other devices wirelessly or by wire to exchange information. Although FIG8 shows the electronic device 1000 with various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may alternatively be implemented or have.

[0137] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart, thereby implementing the image generation method described above. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 1009, or installed from the storage device 1008, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0138] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include an information signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated information signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0139] In some embodiments, the client and server can communicate using any known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital information communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), an internet (e.g., the Internet), and a peer-to-peer network (e.g., an ad hoc peer-to-peer network), as well as any known or future developed network.

[0140] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0141] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device:

[0142] Acquire reference text information and a reference image; the reference text information includes a target phrase; the reference image includes a first subject; the target phrase is related to the first subject;

[0143] Determining image generation guidance information, wherein the image generation guidance information is used to indicate influence weights of the reference text information and the reference image on different regions of the target image to be generated; the reference image has different influence weights on different regions of the target image;

[0144] The reference text information, the reference image, and the image generation guidance information are input into a target text-graph model to generate a target image, where the target image includes a second subject, and the second subject is similar to the first subject.

[0145] Optionally, when the above one or more programs are executed by the electronic device, the electronic device may also execute other steps described in the above embodiments.

[0146] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).

[0147] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0148] The units involved in the embodiments described in this disclosure may be implemented in software or hardware, wherein the name of a unit does not necessarily limit the unit itself.

[0149] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0150] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0151] According to one or more embodiments of the present disclosure, the present disclosure provides an electronic device, including:

[0152] one or more processors;

[0153] a memory for storing one or more programs;

[0154] When the one or more programs are executed by the one or more processors, the one or more processors implement any image generating method provided in the present disclosure.

[0155] According to one or more embodiments of the present disclosure, the present disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any image generating method as provided in the present disclosure.

[0156] An embodiment of the present disclosure further provides a computer program product, which includes a computer program or instructions, and when the computer program or instructions are executed by a processor, the image generation method described above is implemented.

[0157] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

[0158] The foregoing description is intended only to provide specific embodiments of the present disclosure, intended to enable those skilled in the art to understand and implement the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments described herein, but rather to be construed in the broadest manner consistent with the principles and novel features disclosed herein.

Claims

1. An image generation method, characterized in that, Comprising: Obtaining reference text information and a reference image; The reference text information includes target phrases; The reference image includes a first subject; the target phrase is related to the first subject; Determining image generation guidance information, which is used to indicate the influence weights of the reference text information and the reference image on different regions of the target image to be generated; The influence weights of the reference image on different regions of the target image are different; Inputting the reference text information, the reference image, and the image generation guidance information into a target text-to-image model to generate a target image, where the target image includes a second subject, and the second subject is similar to the first subject.

2. The method according to claim 1, wherein The determining of the image generation guidance information includes: Inputting the reference text information into the target text-to-image model to obtain a first image; Determining an image mask based on the first image and the target phrase; Determining the image generation guidance information based on the image mask.

3. The method according to claim 2, characterized in that, Before determining the image mask based on the first image and the target phrase, it includes: Obtaining a text mask; the text mask is used to indicate the position of the target phrase in the reference text information; Determining the target phrase in the reference text information based on the text mask.

4. The method according to claim 3, wherein The process of generating an image using the target text-to-image model is a process of denoising a to-be-denoised image; The determining of the image generation guidance information based on the image mask includes: Determining a key region in the reference image based on the image features of the current to-be-denoised image, the image mask, the text features of the reference text information, and the text mask; Determining the image generation guidance information based on the image features of the current to-be-denoised image, the image features of the key region, the text features of the reference text information, and the image mask.

5. The method according to claim 4, characterized in that, The determining of the key region in the reference image based on the image features of the current to-be-denoised image, the image mask, the text features of the reference text information, and the text mask includes: Dividing the reference image into multiple first sub-regions; Based on the image features of the current to-be-denoised image and the image mask, respectively calculating the first correlation values between the images of the first sub-regions and the image of a third target region; the third target region is the region corresponding to the current to-be-denoised image and the image mask; Based on the text features of the reference text information and the text mask, respectively calculating the second correlation values between the images of the first sub-regions and the target phrase; Based on the first correlation values and the second correlation values, determining the final scores of the first sub-regions in the reference image; Based on the final scores of the first sub-regions in the reference image, determining the key region from the reference image.

6. The method according to claim 4, wherein The determining of the image generation guidance information based on the image features of the current to-be-denoised image, the image features of the key region, the text features of the reference text information, and the image mask includes: Determine a first influence weight based on the image features of the current image to be denoised, the image features of the key region, and the image mask; Determine a second influence weight based on the image features of the current image to be denoised and the text features of the reference text information; Determine the image generation guidance information based on the first influence weight and the second influence weight.

7. The method according to claim 2, characterized in that After inputting the reference text information, the reference image, and the image generation guidance information into the target text-to-image model to generate a target image, it further includes: Fuse the first image and the target image to obtain a final image.

8. The method according to claim 7, wherein The first image includes a first target region and a first non-target region. The first target region includes a main body corresponding to the target phrase; the first non-target region is the image in the first image except the first target region; the target image includes a second target region, and the second target region includes the second main body. The step of fusing the first image and the target image to obtain a final image includes: Fuse the image of the first target region in the first image and the image of the second target region in the target image to obtain a third image; Fuse the third image with the image of the first non-target region in the first image to obtain a final image.

9. The method according to claim 1, wherein The text-to-image process includes an image denoising process of N rounds; Each round of the image denoising process includes: determining image generation guidance information; Input the reference text information, the reference image, and the image generation guidance information into the target text-to-image model to generate a target image.

10. The method according to claim 1, characterized in that The method further includes: Obtain a sample data pair, where the sample data pair includes a sample image, sample text information, and a sample noise image; the sample image includes a third main body, the sample text information includes a target sample phrase, and the target sample phrase is related to the third main body; the sample noise image is the result of adding noise to the sample image; Determine sample image generation guidance information, where the sample image generation guidance information indicates the influence weights of the sample text information and the sample image on different regions of the to-be-generated fourth image; Input the sample image, the sample text information, and the sample image generation guidance information into the target text-to-image model to be trained to obtain a fourth image; Adjust the parameters in the target text-to-image model and / or the model used in the process of determining the sample image generation guidance information based on the similarity between the fourth image and the sample image.

11. An image generation device, characterized in that, It includes: An acquisition module, configured to acquire reference text information and a reference image; The reference text information includes a target phrase; The reference image includes a first main body; the target phrase is related to the first main body; A generation guidance information determination module, configured to determine image generation guidance information, where the image generation guidance information is used to indicate the influence weights of the reference text information and the reference image on different regions of the to-be-generated target image; The influence weights of the reference image on different regions of the target image are different; An image generation module, configured to input the reference text information, the reference image, and the image generation guidance information into a target text-to-image model to generate a target image, where the target image includes a second subject, and the second subject is similar to the first subject.

12. An electronic device, characterized in that, The electronic device includes: One or more processors; A storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-10.

13. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method according to any one of claims 1-10.

14. A computer program product, including computer-executable instructions, where the computer-executable instructions, when executed by a processor, implement the method according to any one of claims 1-10.

Citation Information

Patent Citations

  • Image generation method and server

    CN116597039A

  • Text-to-image generation model optimization method and device, equipment and storage medium

    CN116611496A

  • Data processing method and device based on AIGC, electronic equipment and storage medium

    CN116704062A

  • Pentograph model training method and device, equipment and storage medium

    CN117173504A

  • Image generation method and device, electronic equipment, storage medium and program product

    CN117437317A