Method and apparatus for generating image, device, medium and program product

By generating first and second text descriptions and utilizing a pre-trained machine learning model, the problem of the lack of controllability in image generation models is solved, achieving more accurate and efficient image generation and improving the user experience.

WO2026084648A1PCT designated stage Publication Date: 2026-04-23LEMON INC(GB)
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
LEMON INC(GB)
Filing Date
2025-10-13
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing image diffusion generation models lack controllability during the generation process, resulting in users often having to spend a lot of time adjusting the generated images. Furthermore, traditional image editing techniques fail to fully utilize the capabilities of text-to-image models, leading to poor image quality and efficiency.

Method used

By acquiring the original image and target instructions, first and second text descriptions are generated, and a target image is generated using a pre-trained machine learning model. By combining a visual language model and a diffusion model, the understanding of images and text is enhanced, thereby improving the accuracy and efficiency of image generation.

Benefits of technology

The system enhances the understanding of images and text during image editing, resulting in more accurate images, improved image generation efficiency, and a better user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SG2025050665_23042026_PF_FP_ABST
    Figure SG2025050665_23042026_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present disclosure relate to a method and apparatus for generating an image, a device, a medium and a program product. The method comprises: determining an original image and a target instruction for the original image, the target instruction indicating an adjustment for the original image. The method further comprises: on the basis of the original image and the target instruction, generating a first text description for describing the original image and a second text description for describing a target image to be obtained, the target image being an adjusted original image. The method further comprises: generating the target image on the basis of the original image, the first text description and the second text description.
Need to check novelty before this filing date? Find Prior Art

Description

[0001]Cross-Reference to Related Applications for Methods, Apparatus, Devices, Media, and Program Products for Generating Images This application claims priority to Chinese Patent Application No. 202411457553.X, filed on October 17, 2024, entitled "Methods, Apparatus, Devices, Media, and Program Products for Generating Images," the entire contents of which are incorporated herein by reference. Technical Field Embodiments of this disclosure generally relate to the field of image processing, and more specifically to methods, apparatuses, devices, media, and program products for generating images. Background Art With the advancement of technology, various neural network model technologies are developing rapidly. Among them, the diffusion model is a powerful generative model. Its components typically include a forward diffusion process, gradually destroying the structure of the data, and a reverse process, gradually restoring the data through learning. In the field of image generation, the diffusion model performs particularly well, capable of generating high-quality, realistic images. Therefore, the diffusion image generation model has a wide range of applications, including artistic creation, image restoration, and data augmentation. With the expansion of application scenarios, the fields requiring image processing are increasing. Image diffusion models are playing an increasingly important role in these different fields. Moreover, as the amount of image data processed continues to increase, the generation quality and speed of diffusion image generation models need to be continuously improved, bringing new opportunities and challenges to the field of image generation. Summary of the Invention Embodiments of this disclosure provide a method, apparatus, device, medium, and program product for generating images. According to a first aspect of this disclosure, a method for generating an image is provided. The method includes determining an original image and target instructions for the original image, the target instructions indicating adjustments to the original image. The method further includes generating, based on the original image and the target instructions, a first text description for describing the original image and a second text description for describing a target image to be obtained, the target image being the adjusted original image. The method further includes generating the target image based on the original image, the first text description, and the second text description. In a second aspect of this disclosure, an apparatus for generating an image is provided. The apparatus includes a target instruction determination module configured to determine an original image and a target instruction for the original image, the target instruction indicating adjustments to the original image; a text description generation module configured to generate a first text description for describing the original image and a second text description for describing the target image to be obtained, based on the original image and the target instruction, the target image being the adjusted original image; and a target image generation module configured to generate the target image based on the original image, the first text description, and the second text description.In a third aspect of this disclosure, an electronic device is provided, including at least one processor; and a storage device for storing at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the method according to the first aspect of this disclosure. In a fourth aspect of this disclosure, a computer-readable storage medium is provided, having stored thereon a computer program that, when executed by a processor, implements the method according to the first aspect of this disclosure. In a fifth aspect of this disclosure, a computer program product is provided. The computer program product includes a computer program that, when executed by a processor, implements the method according to the first aspect of this disclosure. It should be understood that the content described in this content section is not intended to limit the key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Brief Description of the Drawings The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of exemplary embodiments of this disclosure taken in conjunction with the accompanying drawings, in which like reference numerals generally represent like parts. Figure 1 is a schematic diagram illustrating an example environment in which the devices and / or methods of some embodiments of the present disclosure may be implemented; Figure 2 is a schematic diagram illustrating a method for generating an image according to some embodiments of the present disclosure; Figure 3 is a schematic diagram illustrating an example architecture of a system for generating an image according to some embodiments of the present disclosure; Figure 4 is a schematic diagram illustrating an example architecture of a machine learning model according to some embodiments of the present disclosure; Figure 5 is a schematic diagram illustrating another example architecture of a machine learning model according to some embodiments of the present disclosure; Figure 6 is a schematic diagram illustrating a method for acquiring sample images according to some embodiments of the present disclosure; Figure 7 is a schematic diagram illustrating an example of the performance of image editing according to some embodiments of the present disclosure; Figure 8 is a schematic diagram illustrating an example of the score generated for text-to-image pairs according to some embodiments of the present disclosure; Figure 9 is a schematic block diagram illustrating an apparatus for generating an image according to some embodiments of the present disclosure; Figure 10 is a schematic block diagram illustrating an example device suitable for implementing various embodiments of the present disclosure. In the various figures, the same or corresponding reference numerals denote the same or corresponding parts. It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of data) should comply with the requirements of relevant laws, regulations, and related provisions. It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operation of this disclosure, based on the prompt message. As an optional but non-limiting implementation, the prompt message sent to the user in response to receiving a user's active request can be, for example, a pop-up window, where the prompt message can be presented in text form. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device. It is understood that the above notification and user authorization process is merely illustrative and does not constitute a limitation on the implementation of this disclosure; other methods that comply with relevant laws and regulations can also be applied to the implementation of this disclosure. The embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the accompanying drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure. In the description of embodiments of this disclosure, the term "comprising" and similar terms should be understood as open-ended inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an embodiment" or "this embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below. Currently, image diffusion generation models can generate realistic and diverse images based solely on textual descriptions. However, these generated images often lack controllability; to some extent, the generation process is like rolling dice until a good output is obtained. To gain greater control over the generated content, image editing—the ability to generate images from images—has become an ideal function, allowing modification of the input image through additional instructions. This can be seen as the intersection of image generation and understanding, both of which are quite mature. However, until now, image editing technology itself has lagged far behind image generation and understanding. In traditional approaches, to obtain a model capable of image editing, two corresponding descriptive texts are typically used to generate image pairs. Then, suitable image pairs are selected from these pairs to train a text-to-image model, resulting in an editing model.However, the image data generated by the above methods is of low quality, either changing too much or remaining almost unchanged, failing to fully utilize the inherent capabilities of the text-to-image model. Therefore, the images generated by the trained model are inaccurate, and most do not meet user requirements, requiring users to spend a significant amount of time adjusting them. To address this, embodiments of this disclosure propose a method for generating images. In this method, a computing device first acquires an original image and determines target instructions for adjusting the original image. Next, the computing device uses the original image and target instructions to generate a first text description describing the original image and a second text description describing the target image to be obtained, where the target image is the image obtained by adjusting the original image. The computing device further utilizes the original image, the first text description, and the second text description to generate the target image. This method adds text descriptions corresponding to both the target image and the original image during the target image generation process, enhancing the understanding of the image and text during image editing, resulting in more accurate generated images, improved image generation efficiency, and a better user experience. Embodiments of this disclosure will now be described in further detail with reference to the accompanying drawings. Figure 1 illustrates an example environment in which the devices and / or methods of embodiments of this disclosure may be implemented. In environment 100, computing device 106 can be used to perform image editing operations, such as generating a target image 112 by editing an original image 102. Examples of computing device 106 include, but are not limited to, personal computers, server computers, handheld or laptop devices, mobile devices (such as mobile phones, personal digital assistants (PDAs), media players, etc.), multiprocessor systems, consumer electronics, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. Computing device 106 can acquire the original image 102. The original image 102 can be of various types. For example, the original image 102 may be an image captured by a user through a camera, or it may be an image obtained by the user through image generation software. Since the user may need to further modify or adjust the original image 102, the user can determine the target instruction 104 for adjusting the image as needed. In one example, the target instruction 104 is a text instruction. In another example, target instruction 104 can be a voice instruction. The above examples are merely for describing this disclosure and not for limiting it.In one example, image 104 shows a person standing in the sun. If the user wants a similar image, but taken at night under the moon, the user might need to input a target command to adjust the image, such as adding the moon. In another example, image 104 shows a person standing in the sun. If the user wants an image of that person dancing, the target command would be to make the person in the image dance. Therefore, users can provide various adjustment commands to the original image. Then, after receiving the original image 102 and the target instruction 104, the computing device 106 can use the original image 102 and the target instruction 104 to generate a first text description 108 describing the original image and a second text description 110 for the target image 112 to be adjusted. The first text description 108 can describe the scene content in the original image 102. For example, what objects are in the image, what these objects are doing, and the environment they are in. The second text description 110 can describe the scene content in the target image to be generated, such as what objects are in the image, what these objects are doing, and the environment they are in, including a description of the image after adjusting the original image. For example, the computing device 106 can use a Visual Language Model (VLM) to process the original image 102 and the target instruction 104 to obtain the first text description 108 and the second text description 110. Then, the computing device 106 can generate the target image 112 based on the first text description 108 and the second text description 110, and further combine them with the original image 104. In one example, When generating a target image, computing device 106 can utilize a trained machine learning model to process the first text description 108, the second text description 110, and the original image 102 to generate a target image 112. In some embodiments, computing device 106 can utilize a pre-established mapping relationship between the two text descriptions, the image, and the image to be generated to determine the target image 112 corresponding to the first text description 108, the second text description 110, and the original image 102. The above examples are merely for describing this disclosure and are not intended to limit this disclosure. Those skilled in the art can use any suitable method to obtain the target image 112 corresponding to the first text description 108, the second text description 110, and the original image 102. Through this method, text descriptions corresponding to the target image and the original image are added during the generation of the target image, which enhances the understanding of the image and text during image editing, resulting in more accurate generated images, improved image generation efficiency, and enhanced user experience.The above description, with reference to FIG1, illustrates an example environment in which some embodiments of the present disclosure's devices and / or methods can be implemented. The following description, with reference to FIG2, illustrates a method for generating an image according to some embodiments of the present disclosure. The method in FIG2 can be executed by the computing device 106 in FIG1 or any suitable computing device. In example method 200, at block 202, computing device 106 determines an original image 102 and a target instruction 104 for the original image, the target instruction 104 indicating adjustments to the original image 102. For example, due to different user needs, the acquired image may not be satisfactory, requiring certain adjustments to meet the user's needs. In this case, the user can determine the desired image state based on the information presented in the original image 102, thereby providing the target instruction 104 requiring adjustments to the original image 102. As mentioned above, if the original image 104 depicts a person standing under the sun, and the user wants to adjust it to a picture of a person under the moon at night, the user can provide a target instruction to adjust the original image. At box 204, based on the original image and the target instruction, computing device 106 generates a first text description 108 for describing the original image 102 and a second text description 110 for describing the target image 112 to be obtained, the target image 112 being an adjusted original image. In some embodiments, computing device 106 may utilize a pre-trained first machine learning model to generate the first text description 108 for describing the original image 102 and the second text description 110 for describing the target image 112 to be obtained. For example, the pre-trained first machine learning model is a VLM model. The computing device then applies the original image 102 and the target instruction 104 to the first machine learning model to generate the first text description 108 and the second text description 110. Additionally, the first machine learning model can be trained using sample images, sample instructions, and the two sample descriptions. In some embodiments, computing device 106 determines the first text description 108 and the second text description 110 corresponding to the original image 102 and the target instruction 104 using a pre-established mapping table of images, instructions, and two texts. For example, the mapping table includes four columns: the first column is the image, the second column is the target instruction, the third column is the description corresponding to the image in the first example, and the fourth column is the description corresponding to the image to be retrieved.Therefore, computing device 104 can look up the first text description 108 and the second text description 110 corresponding to the original image 102 and the target instruction 104 from the mapping table. The above examples are only for describing this disclosure and are not intended to specifically limit this disclosure. At block 206, computing device 106 generates target image 112 based on original image 102, first text description 108 and second text description 110. After obtaining original image 102, first text description 108 and second text description 110, computing device 106 can further process this information to generate target image 112. In some embodiments, in order to generate target image, computing device 106 can first use original image 102 and first text description 108 to generate a first intermediate representation for original image 102. In addition, computing device 106 can also process second text description 110 to generate a second intermediate representation for target image 112. Then computing device 104 can use the first intermediate representation and the second intermediate representation to generate target image 112. In one example, computing device 106 can combine a first intermediate representation and a second intermediate representation to generate a first combined intermediate representation. Then, computing device 106 generates a target image based on the first combined intermediate representation. For example, the target image 112 described above is generated by applying the original image 102, the first text description 108, and the second text description 110 to a second machine learning model. This second machine learning model includes a first part and a second part, wherein the first part includes a first self-attention layer, and the second part includes a second self-attention layer. When generating the target image using this second machine learning model, the first part of the second machine learning model is used to process the original image 104 and the first text description 108 to generate a first intermediate representation for the first self-attention layer of the first part; and the second part of the second machine learning model can be used to process the second text description 110 to generate a second intermediate representation for the second self-attention layer of the second part. Then, the computing device 106 can further generate a combined intermediate representation by inputting the first intermediate representation into the second self-attention layer to combine it with the second intermediate representation, for example, by concatenating corresponding parts of the two intermediate representations to generate the combined intermediate representation. Then, the second part of the second machine learning model uses the combined intermediate representation to generate the target image 1120. This process will be described below with reference to Figures 4 and 5.This method adds text descriptions corresponding to both the target image and the original image during the generation of the target image. This enhances the understanding of the image and text during image editing, resulting in more accurate generated images, improved image generation efficiency, and a better user experience. The above description, in conjunction with Figure 2, illustrates a schematic diagram of a method for generating images according to some embodiments of this disclosure. The following description, in conjunction with Figure 3, illustrates a schematic diagram of an example architecture for generating images according to some embodiments of this disclosure. In example 300, the original image 302 needs to be adjusted. The editing instruction 304 for adjusting the original image 302 is to add a moon. At this time, the original image 302 and the editing instruction 304 can be input into a pre-trained visual language model 306 to generate an input description 308 and an output description 310. oInput description 308 describes the original image 302, which is information obtained after understanding the original image 302, such as a person standing in the sun. Output description 310, generated after understanding the original image 302 and editing instructions 304, describes the target image 314 generated after adjusting the original image 302. Since the editing instruction is to add a moon, output description 310 describes a person standing in the night with a bright moon in the background. Then, the original image 302, input description 308, and output description 310 are input into diffusion model 312 to generate the target image. The diffusion model 312 described in Figure 3 above is an example of the second machine learning model in Figure 2. Figure 3 above, in conjunction with the schematic diagram of an example architecture of a system for generating images according to some embodiments of the present disclosure, is described. Below, in conjunction with Figures 4 and 5, two example architectures of machine learning models according to some embodiments of the present disclosure are described. Figure 4 shows an example architecture 400 of the second machine learning model. In this example architecture, the machine learning model is a U-shaped Network (UNet) model. It comprises two parts, a first part 424 and a second part 426. The first part determines the conditional time step 0418 and then combines it with the conditional image xc 412, such as the original image. After processing through a series of network layers, it is input to the cross-attention module 408. At this point, it can also be further combined with the input text yc 416, such as the first text description 108, for further processing. Then the generated Q, K', and P are input to the causal self-attention module 402 for further processing. Similarly, for the second part 426, which is used to generate the target image 112, in this part, time step t 422 is obtained and then combined with the noisy image 414 for further processing. After processing through a series of network layers, it is input to the cross-attention module 410. In addition, the cross-attention module 410 also receives the output text y 420, such as the second text description 110, for processing. Then the generated Q, K, and / or P are input to the causal self-attention module 404 for further processing. In addition, to ensure that the generated target image has sufficient information from the original image, the intermediate representation 406 of the original image, such as K' and f, from the causal self-attention module 402 is also provided to the causal self-attention module 404.Then, the causal self-attention module 404 combines K' and V with K and F respectively, for example, K' \ K, V' \ V. Then, the causal self-attention module 404 processes (0 K' + K V' + V~) to generate the target image. Figure 5 illustrates another example architecture 500 of a second machine learning model according to some embodiments of the present disclosure. In this example architecture, the machine learning model is an MM-DiT model, which includes two parts, a first part 516 and a second part 518. The first part is processed starting at conditional time step / c 512 and then combines it with the input text and image {yc + xc) 508, for example, a first text description 108 and the original image 102. After processing through a series of network layers, the input is a causal self-attention module 502, which has 0', K'', and IV, as well as 0', K'', and Fz'. Similarly, for the second part 518, which is used to generate the target image 112, in this part, time step / 514 is obtained, and then the output text and noisy image (.y+xi) 510 are combined for further processing, for example, the output text y can be the second text description 1 10. After processing through a series of network layers, the input is a causal self-attention module 504 for further processing. The causal self-attention module 504 has 0, K, and F, as well as Qt, Kt, and ''. In addition, in order to ensure that the generated target image has sufficient information from the original image, the intermediate representation 506 for the original image in the causal self-attention module 502, such as K' and Vi', as well as '' and IV, is also provided to the causal self-attention module 504. Then, the causal self-attention module 504... The self-attention module 504 combines the corresponding K and V values ​​respectively. Then, the causal self-attention module 404 processes the combined data to generate the target image 112o. As shown in Figures 4 and 5, a causal self-attention structure is introduced, allowing the output image branch to query the keys and values ​​of the self-attention in the input branch. No noise is added to the input image. To distinguish between the input branch and the output, a special fixed time step tc is given as the embedding. Since the features of the input branch are fixed at different time steps, a static encoder can be used during inference, and its KV cache is queried by the output branch. Therefore, the overall runtime is similar to that of a single branch. The inference process of the second neural network model is described above with two specific model structures.The second machine learning model described in Figure 2 can be obtained through training. This training process can be performed on other computing devices or on computing device 106. When training the model on computing device 106, it can first acquire the original sample image, the first sample text description, the second sample text description, and the sample target image corresponding to the original sample image. Then, computing device 106 uses the original sample image, the first sample text description, the second sample text description, and the sample target image to train the second machine learning model. Furthermore, in traditional schemes, the data selection for training the machine learning model is relatively coarse, using only two thresholds from Contrastive Language-Image Pre-training (CLIP), and the designed editing model is relatively simple, relying solely on diffusion for language and image understanding without utilizing a stronger modeling approach. Additionally, the traditional training process is singular, performing only a single training iteration and failing to adequately consider the hierarchical quality of the training editing data. The sample data processing and model training process of this disclosure, described below, solves the above problems. The process of obtaining sample images for training a second machine learning model is described below with reference to Figure 6, for example, obtaining original sample images and target sample images. This method can be performed by the computing device 106 described in Figure 1 or any suitable computing device. At box 602, the computing device 106 obtains first and second prompt texts for generating images. For example, the computing device 106 can receive the first and second prompt texts, which can be used to generate two images for training the second machine learning model, from a user or any suitable computing device. At box 604, the computing device 106 generates candidate original images and candidate target images based on the first and second prompt texts. After obtaining the first and second prompt texts, the computing device 106 can use these texts to obtain the corresponding images. In some implementations, when generating candidate original images and candidate target images, the computing device 106 can generate a third intermediate representation for the first prompt text based on the first prompt text. Furthermore, the computing device 106 can also generate a fourth intermediate representation for the second prompt text using the second prompt text. Next, the computing device 106 can generate a candidate original image using the third intermediate representation, and generate a candidate target image using both the third and fourth intermediate representations. Therefore, in the process of generating the target image, information from both the first and second prompt texts is utilized.For example, computing device 106 uses a pre-trained text-to-image (T2I) model to generate a target image and an original image; in this case, the T2I model is a weakly edited model. Additionally, the T2I model can be further trained using selected sample original images and sample target images; this trained model can be used as a second machine learning model. In some embodiments, when generating a candidate target image, computing device 106 can combine a third intermediate representation and a fourth intermediate representation to generate a combined intermediate representation, which, for ease of description, can also be referred to as a second combined intermediate representation. Then, computing device 106 uses the fourth intermediate representation and the second combined intermediate representation to generate a candidate target image. In one example, computing device 106 generates a target intermediate representation by performing a weighted operation on the fourth intermediate representation and the second combined intermediate representation. For example, different weights are assigned to the fourth intermediate representation and the second combined intermediate representation to calculate the intermediate representation. Then, the target intermediate representation is used to generate a candidate target image. For example, to fully utilize the capabilities of a pre-trained T2I model, a weighted mutual attention mechanism is used to create candidate image pairs using the T2I model. Specifically, mutual attention refers to a self-attention operation that mixes the generated input and output images. This process works by mimicking generating two images from a single image but using different text descriptions. However, simply applying mutual self-attention can lead to excessive similarity between images, which is close to a reconstruction effect. Therefore, a new parameter can be added by mixing the self-attention outputs to balance this "reconstruction" and "regeneration" (independent generation). At box 606, computing device 106 determines sample original images and sample target images based on candidate original images and candidate target images. Computing device 106 can generate multiple pairs of candidate original images and candidate target images using a large amount of cue text, and then computing device 106 selects the sample original image and sample target image from these candidate original image and candidate target image pairs. For example, Figure 7 illustrates a schematic diagram of an example of the performance of image editing according to some embodiments of the present disclosure. In Example 700, there are two main dimensions for evaluating image editing performance: "cue alignment" and "image similarity." The former requires the edited image to remain consistent with the input edited text. The latter requires preserving necessary details in the input image that should not be changed. The goal is to balance these two aspects and achieve an optimal point 706. For example, as shown at point 702, the highest image similarity indicates image reconstruction. As shown at point 704, the highest cue alignment value indicates image regeneration. It is assumed that there exists an inherent trade-off curve 708 for the editing capabilities of a given model.The goal of the alignment process is to push the compromise towards the optimal upper right corner region to achieve the optimal point. Therefore, according to the example in Figure 7, sample original images and sample target images for training the second machine learning model need to be selected from these candidate original image and candidate target image pairs. Figure 8 illustrates a schematic diagram of an example of the scores generated for text-to-image pairs according to some embodiments of this disclosure. In example 800, the CLIP score of the generated T2I pair. A good trade-off curve 804 can be created between reconstruction and regeneration to simulate the compromise curve 708 in Figure 7. Accordingly, the alignment editing model can achieve similar or higher orientation scores (instant alignment) and higher image similarity compared to regeneration. At point 802, the image operation becomes image regeneration without incorporating original image information for image editing. Therefore, when determining sample images from candidate images, computing device 106 can determine the similarity, such as CLIP image similarity, between the candidate original image and candidate target image for each pair of candidate original image and candidate target image. Furthermore, the computing device 106 can further determine the orientation score associated with the second prompt text and the candidate target image, such as the CUP orientation score. Next, the computing device uses the similarity and orientation scores to determine the original sample image and the target sample image. In one example, the computing device determines the target score for the candidate target image based on the similarity and orientation scores. In another example, the mapping relationship between the similarity and orientation scores and the final score can be obtained to determine the target score corresponding to the similarity and orientation scores. For example, the target score corresponding to the similarity and orientation scores can be determined using a preset function. The above examples are merely for describing this disclosure and are not intended to limit the specific scope of this disclosure. Any suitable method can be used to obtain the corresponding target score based on the similarity and orientation scores. If the target score is less than or equal to a first threshold score, the candidate original image and the candidate target image are not determined as the original sample image and the target image. If the target score is greater than the first threshold score, the candidate original image and the candidate target image are determined as the original sample image and the target sample image. Additionally, in some embodiments, when the target score is greater than a first threshold score, it is also necessary to further determine the face or other relevant similarity metrics between the candidate original image and the candidate target image. If the similarity is less than or equal to the threshold similarity, the candidate original image and candidate target image are not determined as the sample original image and sample target image. If the similarity is greater than the threshold similarity, the candidate original image and candidate target image are determined as the sample original image and sample target image.Additionally, in some embodiments, when the target score is greater than a first threshold score, it is also necessary to further determine the matching score of the candidate target image to the second prompt text. If the matching score is less than or equal to the second threshold score, the candidate original image and candidate target image are not determined as sample original images and sample target images. If the matching score is greater than the second threshold score, the candidate original image and candidate target image are determined as sample original images and sample target images. For example, to obtain better image pairs, when determining the sample original images and sample target images from the candidate original images and candidate target images, CLIP image similarity and CLIP orientation score can be combined as a scoring function to sort and select all output images. After obtaining the sample original images and sample target images, a second machine learning model can be trained. x c> x > y c > y} and a U (gives Hongzhenjing⑴ where Qin is the noise image from X and , E and % are random noise and diffusion model respectively. λ is a parameter controlling the ratio between edit data and T2I data. In some embodiments, when acquiring candidate original images and candidate target images, the first and second prompt texts can be applied to the trained second machine learning model to generate candidate original images and candidate target images, and then the model is iteratively adjusted. Based on the above information, the second machine learning model can be trained by repeating two simple steps: Step I: Given a weak edit model, adjust the parameter samples, diversifying the edit pairs to fall around the compromise curve in Figure 7 or Figure 8. For the first round, the T2I model can be treated as a weak model to generate pairs with weighted mutual attention. For subsequent iterations, sample different random noise and text context-free grammar (CFG) to generate diversified pairs. Step II: Sort and filter the sampled pairs, which should be located near the upper right edge of the sampled region. Then, fine-tune the weak generator using these sampled pairs. Repeat these two steps, Until the model converges to optimal performance. The training process of the second machine learning model (e.g., the diffusion model) is described above. As mentioned above, the second machine learning model takes image descriptions as input, not just instructions. Therefore, the content of the descriptions may have a significant impact on the final results. This may be due to the limited text understanding ability of the second machine learning model, which will favor certain descriptions even if they have the same meaning. Therefore, the first machine learning model can be used to automatically write these text descriptions for images. As described above, the first machine learning model, such as VLM, is used to generate text descriptions. The training process of the first machine learning model is described below. Similar to training the second machine learning model, the sampling, selection and adjustment strategies are also used to train the first machine learning model to produce the best input / output descriptions. In this training (2) where. is The input and output descriptions are defined in the diagram, where r represents the reward function that measures the editing quality given the instruction y and the corresponding input image x and output image (). 0 is a corresponding parameter. Specifically, for each input image and instruction pair in the training dataset, different input / output descriptions are first created. Then, the best description leading to successful editing is selected to fine-tune a pre-trained first machine learning model to learn the output description from the input image and instruction prompts. Figure 9 illustrates an apparatus for generating images according to some embodiments of the present disclosure. As shown in Figure 9, apparatus 900 includes a target instruction determination module 902 configured to determine an original image and a target instruction for the original image, the target instruction indicating adjustments to the original image; a text description generation module 904 configured to generate a first text description for describing the original image and a second text description for describing the target image to be obtained, the target image being the adjusted original image, based on the original image and the target instruction; and a target image generation module 906 configured to generate the target image based on the original image, the first text description, and the second text description. In some embodiments, the text description generation module 904 includes: a first application module configured to generate a first text description and a second text description by applying an original image and a target instruction to a first machine learning model. In some embodiments, the target image generation module 906 includes: a first intermediate representation generation module configured to generate a first intermediate representation for the original image based on the original image and the first text description; a second intermediate representation generation module configured to generate a second intermediate representation for the target image based on the second text description; and a first generation module configured to generate a target image based on the first and second intermediate representations. In some embodiments, the first generation module includes: a first combined intermediate representation generation module configured to generate a first combined intermediate representation by combining the first and second intermediate representations; and a second generation module configured to generate the target image based on the first combined intermediate representation. In some embodiments, the target image generation module 906 includes a second application module configured to generate the target image by applying the original image, the first text description, and the second text description to a second machine learning model.In some embodiments, training the second machine learning model includes: a sample data determination module configured to determine a sample original image, a first sample text description, a second sample text description, and a sample target image corresponding to the sample original image; and a training module configured to train the second machine learning model based on the sample original image, the first sample text description, the second sample text description, and the sample target image. In some embodiments, the apparatus 900 further includes: a prompt text acquisition module configured to acquire a first prompt text and a second prompt text for generating an image; a first candidate image generation module configured to generate a candidate original image and a candidate target image based on the first prompt text and the second prompt text; and a first sample image determination module configured to determine the sample original image and the sample target image based on the candidate original image and the candidate target image. In some embodiments, the first candidate image generation module includes: a third intermediate representation generation module configured to generate a third intermediate representation for the first prompt text based on the first prompt text; a fourth intermediate representation generation module configured to generate a fourth intermediate representation for the second prompt text based on the second prompt text; a candidate original image generation module configured to generate a candidate original image based on the third intermediate representation; and a first candidate target image generation module configured to generate a candidate target image based on the third and fourth intermediate representations. In some embodiments, the first candidate target image generation module includes: a second combined intermediate representation generation module configured to generate a second combined intermediate representation by combining the third and fourth intermediate representations; and a second candidate target image generation module configured to generate a candidate target image based on the fourth intermediate representation and the second combined intermediate representation. In some embodiments, the second candidate target image generation module includes: a target intermediate representation generation module configured to generate a target intermediate representation by performing a weighted operation on the fourth intermediate representation and the second combined intermediate representation; and a third candidate target image generation module configured to generate a candidate target image based on the target intermediate representation. In some embodiments, the first sample image determination module includes: a first similarity determination module configured to determine the similarity between a candidate original image and a candidate target image; a direction score determination module configured to determine a direction score associated with the second prompt text and the candidate target image; and a second sample image determination module configured to determine a sample original image and a sample target image based on the similarity and direction score.In some embodiments, the second sample image determination module includes: a target score determination module configured to determine a target score for a candidate target image based on similarity and orientation scores; and a third sample image determination module configured to determine a candidate original image and a candidate target image as sample original image and sample target image in response to a target score greater than a first threshold score. In some embodiments, the third sample image determination module includes: a second similarity determination module configured to determine the facial similarity between the candidate original image and the candidate target image in response to a target score greater than a first threshold score; and a second sample image determination module configured to determine the candidate original image and the candidate target image as sample original image and sample target image in response to a similarity greater than a threshold similarity. In some embodiments, the third sample image determination module includes: a matching score determination module configured to determine a matching score of the candidate target image that satisfies a second prompt text in response to a target score greater than a first threshold score; and a second sample image determination module configured to determine the candidate original image and the candidate target image as sample original image and sample target image in response to a matching score greater than a second threshold score. In some embodiments, the first candidate image generation module includes a second candidate image generation module configured to generate candidate original images and candidate target images by applying first and second prompt texts to a second machine learning model. FIG10 shows a schematic block diagram of an example device 1000 that can be used to implement embodiments of the present disclosure. The computing device 106 in FIG1 can be implemented using device 1000. As shown, device 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 1002 or loaded from storage unit 1008 into random access memory (RAM) 1003. Various programs and data required for the operation of device 1000 may also be stored in RAM 1003. CPU 1001, ROM 1002 and RAM 1003 are interconnected to each other via bus 804. Input / output (I / O) interface 1005 is also connected to bus 804. Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of display, speaker, etc.; storage page 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless communication transceiver, etc.Communication unit 1009 allows device 1000 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunications networks. The various processes and procedures described above, such as methods 200 and 600, may be executed by processing unit 1001. For example, in some embodiments, methods 200 and 600 may be implemented as computer software programs tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by CPU 1001, one or more actions of the example methods 200 and 600 described above may be performed. This disclosure may be a method, apparatus, system, and / or computer program product. A computer program product may include a computer-readable storage medium on which computer-readable program instructions for performing various aspects of this disclosure are loaded. Computer-readable storage media can be tangible devices capable of holding and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example—but not limited to—electrical storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of computer-readable storage media include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital multifunction disc (DVD), memory sticks, floppy disks, mechanical encoding devices, such as punch cards or recessed protrusions storing instructions thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not to be construed as a transient signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through fiber optic cables), or electrical signals transmitted through wires. The computer-readable program instructions described herein can be downloaded from the computer-readable storage medium to various computing / processing devices, or downloaded via a network, such as the Internet, local area network, wide area network, and / or wireless network, to an external computer or external storage device. The network may include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers.Each computing / processing device's network adapter card or network interface receives computer-readable program instructions from the network and forwards these instructions for storage on a computer-readable storage medium within the respective computing / processing device. The computer program instructions used to perform the operations of this disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, status setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Smalltalk, C++, etc., and conventional procedural programming languages ​​such as "C" or similar languages. The computer-readable program instructions may execute entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, electronic circuitry, such as programmable logic circuitry, field-programmable gate arrays (FPGAs), or programmable logic arrays (PLAs), is personalized using state information from computer-readable program instructions. This electronic circuitry can execute computer-readable program instructions to implement various aspects of this disclosure. Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions. These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.Computer-readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks of the flowchart and / or block diagram. The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, segment, or portion of an instruction containing one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the blocks may occur in a different order than that shown in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. Various embodiments of this disclosure have been described above. The above description is exemplary and not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is chosen to best explain the principles, practical applications, or technical improvements to the technology in the market, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

Claims 1. A method for generating an image, comprising: Determine the original image and the target instructions for the original image, the target instructions indicating adjustments for the original image; Based on the original image and the target instruction, generate a first text description for describing the original image and a second text description for describing the target image to be obtained, wherein the target image is an adjusted original image; and generate the target image based on the original image, the first text description and the second text description.

2. The method of claim 1, wherein generating a first text description describing the original image and a second text description describing a target image to be obtained comprises: The first text description and the second text description are generated by applying the original image and the target instruction to a first machine learning model.

3. The method of claim 1, wherein generating the target image comprises: Based on the original image and the first text description, a first intermediate representation for the original image is generated; based on the second text description, a second intermediate representation for the target image is generated; and based on the first intermediate representation and the second intermediate representation, the target image is generated.

4. The method of claim 3, wherein generating the target image based on the first intermediate representation and the second intermediate representation comprises: The first combined intermediate representation is generated by combining the first intermediate representation and the second intermediate representation; The target image is generated based on the intermediate representation of the first combination.

5. The method of claim 1, wherein generating the target image comprises: The target image is generated by applying the original image, the first text description, and the second text description to a second machine learning model.

6. The method of claim 5, wherein the training of the second machine learning model comprises: Determine the original sample image, the first sample text description, the second sample text description, and the sample target image corresponding to the original sample image; And based on the original image of the sample, the text description of the first sample, the text description of the second sample, and the target image of the sample, the second machine learning model is trained.

7. The method of claim 6, further comprising: Get the first and second prompt texts used to generate the image; Based on the first prompt text and the second prompt text, candidate original images and candidate target images are generated; and based on the candidate original images and the candidate target images, the sample original image and the sample image are determined. Target image.

8. The method of claim 7, wherein generating the candidate original image and the candidate target image comprises: Based on the first prompt text, generate a third intermediate representation for the first prompt text; Based on the second prompt text, a fourth intermediate representation for the second prompt text is generated; The candidate original image is generated based on the third intermediate representation; and the candidate target image is generated based on the third intermediate representation and the fourth intermediate representation.

9. The method of claim 8, wherein generating the candidate target image comprises: The second combined intermediate representation is generated by combining the third intermediate representation and the fourth intermediate representation; and the candidate target image is generated based on the fourth intermediate representation and the second combined intermediate representation.

10. The method of claim 9, wherein generating the candidate target image based on the fourth intermediate representation and the second combined intermediate representation comprises: The target intermediate representation is generated by weighting the fourth intermediate representation and the intermediate representation of the second combination. And based on the intermediate representation of the target, the candidate target image is generated.

1. The method of claim 10, wherein determining the original sample image and the target sample image comprises: Determine the similarity between the candidate original image and the candidate target image; Determine the orientation score associated with the second prompt text and the candidate target image; And based on the similarity and the orientation score, the original image of the sample and the target image of the sample are determined.

12. The method of claim 11, wherein determining the sample original image and the sample target image based on the similarity and the direction score comprises: Based on the similarity and the orientation score, a target score is determined for the candidate target image; And in response to the target score being greater than a first threshold score, the candidate original image and the candidate target image are determined as the sample original image and the sample target image.

13. The method of claim 12, wherein responsive to the target score being greater than a first threshold score, determining the candidate raw image and the candidate target image as the sample raw image and the sample target image comprises: In response to the target score being greater than a first threshold score, the similarity of the faces in the candidate original image and the candidate target image is determined; And in response to the similarity being greater than a threshold similarity, the candidate original image and the candidate target image are determined as the sample original image and the sample target image.

14. The method of claim 12, wherein in response to the target score being greater than a first threshold score, the candidate original image and the candidate target image are determined as the sample original image and the sample target image. Such as including: In response to the target score being greater than a first threshold score, the candidate target image is determined to satisfy the matching score of the second prompt text; And in response to the matching score being greater than the second threshold score, the candidate original image and the candidate target image are determined as the sample original image and the sample target image. 15.The method of claim 7, wherein generating the candidate original image and the candidate target image based on the first prompt text and the second prompt text comprises: The candidate original image and the candidate target image are generated by applying the first prompt text and the second prompt text to the second machine learning model.

16. An apparatus for generating an image, comprising: A target instruction determination module is configured to determine an original image and a target instruction for the original image, the target instruction indicating adjustments for the original image; A text description generation module is configured to generate a first text description for describing the original image and a second text description for describing the target image to be obtained, based on the original image and the target instruction, wherein the target image is an adjusted original image; and a target image generation module is configured to generate the target image based on the original image, the first text description and the second text description.

17. An electronic device, comprising: At least one processor; And a storage device for storing at least one program, which, when executed by the at least one processor, causes the at least one processor to implement the method according to any one of claims 1-15.

18. A computer-readable storage medium having a computer program stored thereon, the computer program implementing the method according to any one of claims 1-15 when executed by a processor.

19. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-15. 18

Citation Information

Patent Citations

  • Image generation method and device, equipment and storage medium

    CN117173284A

  • Method for editing face image through text, terminal and storage medium

    CN117576257A

  • Image data pair generation method and device, image editing method and device, equipment and medium

    CN118521677A

  • Neural compositing by embedding generative technologies into non-destructive document editing workflows

    US20240135611A1