Image generation method and device, medium, equipment and computer program product
By receiving description text and generation conditions, determining the target generation model, and generating images based on this information, the problem of difficulty in fine control in the image generation process in the prior art is solved, and the feature consistency and user experience are improved.
Patent Information
- Application Number
- CN202510273120.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-05-30
AI Technical Summary
The existing literary and biographical technology is difficult to perform fine control during the image generation process, resulting in possible deviations between the generated image and user needs.
By receiving the description text and generation conditions of the target image to be generated, the target generation model corresponding to the description text is determined, and the target image is generated based on the description text, the generation conditions and the target generation model. Generation conditions are used to constrain the features of the target object in the target image to ensure consistency of the generated image features.
It realizes fine control of features in the image generation process, ensures the continuity of different images generated, and provides technical support for image generation in different scenarios to improve user experience.
Smart Images

Figure CN120070672A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image generation, and in particular, to an image generation method, apparatus, medium, device, and computer program product. Background Art
[0002] Currently, text-to-image technology can generate diverse images based on text descriptions to provide users with the required images. However, during this process, the generated images may deviate from the images required by the users, and it is difficult to finely control the image generation process through text descriptions. Summary of the Invention
[0003] This Summary of the Invention section is provided to introduce concepts in a brief form that will be described in detail in the following Detailed Description section. This Summary of the Invention section is not intended to identify key features or essential features of the claimed technical solution, nor is it intended to be used to limit the scope of the claimed technical solution.
[0004] In a first aspect, the present disclosure provides an image generation method, the method comprising: Receiving a description text and generation conditions of a target image to be generated, wherein the generation conditions are used to constrain features of a target object in the target image; Determining a target generation model corresponding to the description text; Obtaining the target image based on the description text, the generation conditions, and the target generation model.
[0005] In a second aspect, the present disclosure provides an image generation apparatus, the apparatus comprising: A receiving module, configured to receive a description text and generation conditions of a target image to be generated, wherein the generation conditions are used to constrain features of a target object in the target image; A first determination module, configured to determine a target generation model corresponding to the description text; A first generation module, configured to obtain the target image based on the description text, the generation conditions, and the target generation model.
[0006] In a third aspect, the present disclosure provides a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processing device, the steps of the method in the first aspect are implemented.
[0007] In a fourth aspect, the present disclosure provides an electronic device, comprising: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of the method in the first aspect.
[0008] In a fifth aspect, the present disclosure provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the method described in the first aspect.
[0009] In the above technical solution, during the process of generating an image based on the description text, the generation conditions can be combined simultaneously, so that feature constraints can be performed during the image generation process of the target image based on the generation conditions, so that the features of the target object in the target image can match the generation conditions, and the features of the target object in different images determined based on the same generation conditions can be kept consistent, ensuring the continuity of the generated different images. It can not only achieve fine control of the features during the image generation process, but also provide technical support for the feature consistency in the image generation process under different scenarios, broaden the application scenarios of the method of the present disclosure, and improve the user experience.
[0010] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In combination with the drawings and with reference to the following specific implementation manners, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale. In the drawings: Figure 1 is a flowchart of an image generation method provided according to an embodiment of the present disclosure.
[0012] Figure 2 is a schematic diagram of the generation process of an image generation method provided according to an embodiment of the present disclosure.
[0013] Figure 3 is a block diagram of an image generation device provided according to an embodiment of the present disclosure.
[0014] Figure 4 shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0015] The embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0016] It should be understood that the various steps described in the method embodiments of the present disclosure may be executed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0017] As used herein, the term "comprising" and its variations are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0018] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependent relationships.
[0019] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly specified in the context, it should be understood as "one or more".
[0020] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0021] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0022] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.
[0023] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0024] It is understandable that the above notification and the process of obtaining user authorization are only illustrative and do not limit the implementation manner of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0025] Meanwhile, it is understandable that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.
[0026] As described in related technologies, those skilled in the art have found that when generating images, if images of the same person are generated in different scenarios, due to the uncertainty of the diffusion model, the generated images of the same person may have the same appearance or may have different appearances. Based on this, the following embodiments are provided in the present disclosure to ensure feature consistency during the image generation process.
[0027] Figure 1 Shown is a flowchart of an image generation method provided according to an embodiment of the present disclosure. As Figure 1 shown, the method may include: In step 11, receive a description text of a target image to be generated and generation conditions, where the generation conditions are used to constrain the features of a target object in the target image.
[0028] Among them, the description text may be text input by the user based on their own needs, and it may be in natural language mode. Also, for example, in the scenario of generating an animation based on a novel, the description text may be a storyboard script obtained after performing storyboard detection on the novel text. For example, the novel text may be input into a semantic storyboard model to perform storyboard detection on the novel text, and multiple storyboard scripts are obtained. Each storyboard script contains the features of the objects and the features of the scenes in the image of that storyboard. The semantic storyboard model may be trained based on the large language model LLM (Large Language Model) in the art.
[0029] Among them, the generation conditions may be image conditions or text conditions.
[0030] In some embodiments, the user can determine the image conditions based on their own needs. For example, if the target object in the target image to be generated by the user is person A and is desired to have the appearance shown in the picture of person A, then the picture of person A can be provided when generating the image, so that the image can be generated based on the picture of person A during the image generation process to constrain the features of the target object A, making the features of A in the generated target image consistent with the features in the picture of person A, and ensuring the feature consistency of this person A. The feature consistency may mean the same appearance, such as the character's hairstyle, clothing, etc.
[0031] In some other embodiments, the user can determine text conditions based on their own needs. For example, the user can describe in natural language the characteristics they hope the target object to exhibit. For example, the characteristics of person A can be described so that during the image generation process, an image can be generated based on the text of A's characteristics, thereby constraining the characteristics of person A, making the characteristics of A in the generated target image consistent with the characteristics described in the text of A's characteristics, and ensuring the consistency of the characteristics of person A.
[0032] In step 12, determine the target generation model corresponding to the description text.
[0033] Among them, the image generation process in different application scenarios may have different preferences. As an example, multiple image generation models can be preset in advance, and the application scenario information corresponding to the image generation model can be configured. Then, the description text can be matched with the application scenario information, and the image generation model that matches the description text can be used as the target generation model.
[0034] As another example, the determination of the target generation model corresponding to the description text may include: Perform theme recognition based on the description text to determine the theme type corresponding to the description text.
[0035] In one scenario, the description text is determined based on a novel text. In this embodiment, theme recognition can be performed based on the description text to determine the theme type corresponding to the description text. The image generation styles under different theme types are usually different. For example, the ancient style xianxia type and the modern romance type usually have different corresponding image generation styles. Therefore, in this embodiment, different image generation models can be configured in advance for different theme types. In this embodiment, a theme recognition model can be pre-trained based on a classification model in the art, and then the description text can be input into the theme recognition model to determine the theme type.
[0036] After that, the image generation model corresponding to the theme type is determined as the target generation model.
[0037] Thus, the target generation model applied in the image generation process can be made to better match the description text, improving the fineness of the target image generation to a certain extent, thereby enhancing the accuracy of the subsequent generated target image and meeting the user's usage requirements.
[0038] In step 13, based on the description text, generation conditions, and target generation model, obtain a target image, and the characteristics of the target object of the target image match the generation conditions.
[0039] Such as Figure 2As shown, the generation conditions may include Image C1 and Image C2. Image C1 is used to represent Character 1, and Image C2 is used to represent Character 2. The target image obtained based on the description text and the generation conditions is shown as Image C3. Among them, the features of Character 1 in Image C3 match those in Image C1, and the features of Character 2 in Image C3 match those in Image C2, so as to achieve the feature consistency constraint in the image generation process.
[0040] In the above technical solution, in the process of image generation based on the description text, the generation conditions can be combined at the same time, so that feature constraints can be performed in the image generation process of the target image based on the generation conditions, so that the features of the target object in the target image can match the generation conditions, and the features of the target object in different images determined based on the same generation conditions can be kept consistent, ensuring the continuity of the generated different images. It can not only achieve fine control of the features in the image generation process, but also provide technical support for the feature consistency in the image generation process under different scenarios, broaden the application scenarios of the method of the present disclosure, and improve the user experience.
[0041] In some embodiments, the generation conditions include image conditions, that is, the generation conditions are embodied in the form of images. For example, the user can directly upload a portrait of A, so that the feature representation of A in the generated target image is the same as that of A in the portrait, so as to achieve the consistency of the character features in the image generation process.
[0042] Correspondingly, obtaining the target image based on the description text, the generation conditions, and the target generation model includes: Inputting the description text and the image conditions into the target generation model, so that the target generation model performs multi-modal attention mechanism processing based on the description text and the image conditions to obtain the target image.
[0043] As an example, the target generation model can be a model implemented based on the OminiControl algorithm, and the target generation model can be obtained by training based on training data including training description texts and training image pairs.
[0044] In this step, the description text, the image conditions, and the input image can be input into the target generation model for normalization processing to obtain the normalized features corresponding to the description text, the image conditions, and the input image respectively, where the input image is an initial noise image or an intermediate image generated in the previous generation step of the target generation model.
[0045] Among them, in the process of the target generation model generating an image, it is obtained by gradually denoising a noise image. Then the input image is initially an initial noise image, and a denoised image can be obtained after the first generation step. In subsequent image generation steps, the input image can be the intermediate image generated in the previous generation step of the current step, that is, the denoised image, so that denoising can be achieved through each generation step to obtain the final target image.
[0046] As an example, in this step, the descriptive text can be transformed into a latent representation through a text encoder, and the image condition and the input image can be transformed into a latent representation through an image encoder. This latent representation can be a token. Then, standardization processing is performed based on each token. For example, the above standardization processing process can be implemented based on the Norm model. The Norm model is used to represent the normalization technique to ensure the stability of the data distribution and improve the training process and performance of the model.
[0047] After that, based on the target generation model, multi-modal attention feature processing is performed on the standardized features to obtain the target image.
[0048] For example, each standardized feature can be input into the multi-modal attention module of the target generation model for multi-modal attention feature processing to calculate the correlation between different types of tokens. By calculating the attention between the tokens of the text and the image through this process, it can be determined which text information or image information has an important impact on image generation. Further, the result of the attention mechanism processing can pass through a multi-layer perceptron and a residual connection module to obtain the processing result to complete the denoising process of the current generation step and obtain the denoised image. The processing result of the last generation step can be used as the target image through gradual denoising.
[0049] Thus, through the above technical solution, the descriptive text and the image condition are input into the target generation model for multi-modal attention mechanism processing, so as to be able to display the relationship between the modeled image condition, the input image, and the descriptive text, ensure the full utilization of the information of the image condition, and flexibly integrate the image condition during the image generation process to improve the accuracy of the obtained target image.
[0050] In some embodiments, the generation condition includes a text condition, where the text condition can be a natural language description of the features of the target object by the user based on their needs.
[0051] Correspondingly, obtaining the target image based on the descriptive text, the generation condition, and the target generation model may include: Generating an image corresponding to the text condition based on the text condition.
[0052] Among them, a text condition can be input into an image generation model to obtain an image corresponding to the text condition. The image generation model can be a general text-to-image model in the art, such as a diffusion model, etc., and the present disclosure does not limit this.
[0053] After that, the description text and the image corresponding to the text condition are input into the target generation model, so that the target generation model performs multi-modal attention mechanism processing based on the description text and the image corresponding to the text condition to obtain the target image. Among them, in this step, the image corresponding to the text condition can be used as an image condition, so that the target image can be obtained based on the method described above, which will not be elaborated here.
[0054] Thus, through the above technical solution, it is possible to support a user to input a text condition, without the user actively performing image processing or image retrieval, simplify the determination method of generation conditions, and further improve the user experience.
[0055] In some embodiments, the method may further include: Display the image corresponding to the text condition.
[0056] Among them, after generating the image corresponding to the text condition, the image can be displayed so that the user can determine whether the features in the image are the features that meet their needs based on the displayed image.
[0057] In response to a confirmation operation on the image corresponding to the text condition, execute the step of inputting the description text and the image corresponding to the text condition into the target generation model, so that the target generation model performs multi-modal attention mechanism processing based on the description text and the image corresponding to the text condition to obtain the target image.
[0058] As an example, if the user determines that the image corresponding to the generated text condition meets their needs, they can confirm it through a confirmation operation. For example, a confirmation control can be displayed in the display interface, and the user can confirm it by clicking the confirmation control. After the user confirms, the step of inputting the description text and the image corresponding to the text condition into the target generation model can be further executed, so that the target generation model performs multi-modal attention mechanism processing based on the description text and the image corresponding to the text condition to obtain the target image, to avoid deviation of the finally generated target image due to errors between the generated image corresponding to the text condition and the user's needs, improve the accuracy of the target image, simplify the determination method of the image condition, and increase the interaction with the user during the image generation process to improve user participation.
[0059] In response to an editing operation on the text condition, generate a new image based on the text obtained from the editing operation; Use the new image as the image corresponding to the text condition to return the step of displaying the image corresponding to the text condition.
[0060] As an example, if the user believes that the image corresponding to the generated text condition does not meet their requirements, they can re-edit the text condition through an editing operation. For example, an editing control can be displayed on the display interface, and the user can click on the editing control to perform an editing operation, such as modifying the text condition or adding a text condition. After the user finishes editing and obtains a new text, a new image can be generated based on the new text, and the newly generated image can be re-displayed for the user to confirm until an image confirmed by the user is obtained. Then, execute the step of inputting the description text and the image corresponding to the text condition into the target generation model, so that the target generation model performs multi-modal attention mechanism processing based on the description text and the image corresponding to the text condition to obtain the target image, thereby avoiding waste of computing resources caused by image generation based on biased image conditions, ensuring the accuracy and efficiency of image generation, increasing the interaction with the user during the image generation process, and enhancing user participation.
[0061] In some embodiments, the training data of the target generation model is determined in the following manner: Obtain training text, where the training text contains feature text of at least one training object.
[0062] Among them, the feature text of the training object and the training text containing the training object can be obtained by pre-annotation or based on the existing object feature text in the application scenario. The training text can also contain the scenario text of the training object to describe the scenario features where the training object is located.
[0063] Generate multiple candidate images based on the training text.
[0064] Among them, the training text can be input into an image generation model to generate multiple candidate images based on the same training text. The image generation model can be implemented based on text-to-image models in the art, such as diffusion models, etc., and the present disclosure does not limit this.
[0065] After that, match the multiple candidate images, and determine the training data based on the matching result and the training text.
[0066] Thus, through the above technical solution, a training data set for the target generation model can be constructed through image generation technology. While ensuring the accuracy of the training data, the process of constructing the training data is simplified, the influence of less paired image data with consistent object features on the training process of the target generation model is avoided, and the training efficiency of the target generation model is improved.
[0067] In some embodiments, the training text only contains the feature text of one training object, that is, the training text is a description of an image containing a single object. Among them, generating multiple candidate images based on the training text can generate an input prompt word based on the training text to limit the generation of a double-column graph, thereby obtaining multiple candidate images. For example, the feature text of the training object is as follows: A woman with long black hair, wearing a white Hanfu with gold embroidery wearing a string of buddhist beads on her wrist The training text is as follows: Two side-by-side anime image showcases the same person with the same appearance in different contexts: {A woman with long black hair, wearing a white Hanfu with gold embroidery wearing a string of buddhist beads on her wrist.}; Right Grid: the person is {in a street}; Left Grid: the person is {in a park} Thus, based on the training text and the image generation model, candidate images of the same person in two different scenarios can be generated.
[0068] Correspondingly, matching the multiple candidate images and determining the training data based on the matching result and the training text may include: Determining the matching result of the image pair, where each image pair contains two candidate images corresponding to the training text, and the matching result of the image pair is used to indicate whether the objects in the two candidate images match.
[0069] Among them, the candidate images are generated based on the same training text, but there is a certain probability that the image generation model generates consistent or inconsistent images during the generation process. Therefore, the candidate images are further judged and screened. For example, two candidate images generated based on the same training text can be used as an image pair.
[0070] As an example, determining the matching result of an image pair can be to determine the similarity between the objects in two of the image pairs. If the similarity is greater than or equal to a preset threshold, it can be considered that the matching result of the image pair is a match. If the similarity is less than the preset threshold, it is considered that the matching result of the image pair is a mismatch.
[0071] As another example, two candidate images in the image pair can be input into a multi-modal large language model for determination, so that the model can determine whether the objects in the two candidate images are the same object. If they are the same object, it can be considered that the matching result of the image pair is a match. If it is determined based on the model that the objects in the two candidate images are not the same object, it is considered that the matching result of the image pair is a mismatch.
[0072] After that, if the matching result of the image pair is a match, the image pair is used as a training image pair, and the training data is determined based on the training text and the training image pair.
[0073] As an example, the training data can be directly based on the training text and the training image pair. Another example is that the training image pair can be further annotated, and the training image pair with the annotation result of a match and its training text are used as the training data to avoid the influence of the model determination result error on the training data, further improve the accuracy and effectiveness of the training data, and improve the accuracy of the target generation model.
[0074] Among them, in the process of training the target generation model based on the training data, one image in the training image pair can be used as an image condition, the training text can be used as a description text and input into the model, and the other image in the training image pair can be used as a target image. Thus, based on the target image and the output image of the model, loss calculation is performed for model training. The loss calculation and parameter update methods in this model training process can be based on the implementation method of the OminiControl algorithm in the art, which is not limited herein.
[0075] In some embodiments, the training text contains feature texts of multiple training objects, that is, the training text is a description of an image containing multiple objects, and the training text is expressed as follows: Character U1 and character U2 are hugging each other. Character U1 has feature XX1, and character U2 has feature XX2.
[0076] Correspondingly, generating multiple candidate images based on the training text may include: Generating a first image based on the training text, and generating a second image of each training object based on the feature text of each training object in the training text. The candidate images include the first image and the second image.
[0077] Continuing with the above example, a first image P1 can be generated based on the above training text. The first image contains character U1 and character U2. A second image P2 of character U1 is generated based on the features XX1 of character U1, and a second image P3 of character U2 is generated based on the features XX2 of character U2, thereby obtaining the multiple candidate images.
[0078] Correspondingly, matching the multiple candidate images and determining the training data based on the matching result and the training text includes: Determining the matching result of the image pair. Each image pair contains the first image corresponding to the training text and each of the second images corresponding to the training text. The matching result of the image pair includes the sub-results of the matching of each of the second images in the first image.
[0079] As an example, the first image P1, the second image P2 of character U1, and the second image P3 of character U2 can be used as an image pair. Then, the first image P1 and the second image P2 of character U1 can be input into a multi-modal large language model for determination, so that the model determines whether there is an object in the first image P1, and the object is the same object as the object in the second image P2. If there is an object in the first image P1 that is consistent with the object in the second image P2, it can be considered that the sub-result of the matching of the second image in the first image is a match. If there is no object in the first image P1 that is consistent with the object in the second image P2, it can be considered that the sub-result of the matching of the second image in the first image is a mismatch.
[0080] Further, if all the sub-results in the matching result of the image pair are matches, the image pair is used as a training image pair, and the training data is determined based on the training text and the training image pair.
[0081] As an example, the training text and the training image pair can be directly used as training data. Another example is that the training image pair can be further annotated, and the training image pair with the annotation result of match and its training text are used as training data to avoid the influence of the model determination result error on the training data, further improve the accuracy and effectiveness of the training data, and improve the accuracy of the target generation model.
[0082] Among them, in the process of training the target generation model based on the training data, each of the second images in the training image pair can be used as an image condition, the training text can be used as a description text and input into the model, and the first image in the training image pair can be used as the target image, so as to calculate the loss based on the target image and the output image of the model for model training.
[0083] Thus, through the above technical solutions, training data pairs including multiple objects can be quickly constructed, enabling the target generation model to support image conditions with multiple objects during image generation, so as to be applicable to image generation in multi-object scenarios, broaden the application scenarios of the method of the present disclosure, and improve the user experience at the same time.
[0084] Based on the same inventive concept, the present disclosure also provides an image generation device, as Figure 3 shown, the device 10 includes: A receiving module 100, configured to receive a description text and generation conditions of a target image to be generated, where the generation conditions are used to constrain features of a target object in the target image; A first determination module 200, configured to determine a target generation model corresponding to the description text; A first generation module 300, configured to obtain the target image based on the description text, the generation conditions, and the target generation model.
[0085] Optionally, the generation conditions include image conditions; The first generation module includes: A first generation sub-module, configured to input the description text and the image conditions into the target generation model, so that the target generation model performs multi-modal attention mechanism processing based on the description text and the image conditions to obtain the target image.
[0086] Optionally, the generation conditions include text conditions; The first generation module includes: A second generation sub-module, configured to generate an image corresponding to the text conditions based on the text conditions; A third generation sub-module, configured to input the description text and the image corresponding to the text conditions into the target generation model, so that the target generation model performs multi-modal attention mechanism processing based on the description text and the image corresponding to the text conditions to obtain the target image.
[0087] Optionally, the device further includes: A display module, configured to display the image corresponding to the text conditions; A first processing module, configured to trigger the third generation sub-module to input the description text and the image corresponding to the text conditions into the target generation model in response to a confirmation operation on the image corresponding to the text conditions, so that the target generation model performs multi-modal attention mechanism processing based on the description text and the image corresponding to the text conditions to obtain the target image; A second generation module, configured to, in response to an editing operation on the text condition, generate a new image based on the text obtained from the editing operation; and use the new image as the image corresponding to the text condition to trigger the display module to display the image corresponding to the text condition.
[0088] Optionally, the training data of the target generation model is determined by a second determination module, and the second determination module includes: An acquisition sub-module, configured to acquire training text, where the training text includes feature text of at least one training object; A fourth generation sub-module, configured to generate a plurality of candidate images based on the training text; A first determination sub-module, configured to perform matching on the plurality of candidate images, and determine the training data based on the matching result and the training text.
[0089] Optionally, the training text includes only the feature text of one training object; The first determination sub-module includes: A second determination sub-module, configured to determine the matching result of an image pair, where each image pair includes two candidate images corresponding to the training text, and the matching result of the image pair is used to indicate whether the objects in the two candidate images match; A third determination sub-module, configured to, if the matching result of the image pair is a match, use the image pair as a training image pair, and determine the training data based on the training text and the training image pair.
[0090] Optionally, the training text includes feature text of a plurality of training objects; The fourth generation sub-module includes: A fifth generation sub-module, configured to generate a first image based on the training text, and generate a second image of the training object based on the feature text of each training object in the training text, where the candidate images include the first image and the second image; The first determination sub-module includes: A fourth determination sub-module, configured to determine the matching result of an image pair, where each image pair includes the first image corresponding to the training text and each second image corresponding to the training text, and the matching result of the image pair includes sub-results of the matching of each second image in the first image; A fifth determination sub-module, configured to, if the sub-results in the matching result of the image pair are all matches, use the image pair as a training image pair, and determine the training data based on the training text and the training image pair.
[0091] Optionally, the first determination module includes: An identification sub-module, configured to perform theme identification based on the description text and determine the theme type corresponding to the description text; A sixth determination sub-module, configured to determine the image generation model corresponding to the theme type as the target generation model.
[0092] Reference is now made to Figure 4 , which shows a schematic structural diagram of an electronic device (such as a terminal device or a server) 600 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The electronic device shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.
[0093] As Figure 4 shown, the electronic device 600 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 601, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage device 608 into a random access memory (RAM) 603. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0094] Generally, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device 600 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 4 the electronic device 600 with various devices is shown, it should be understood that it is not required to implement or include all the shown devices. More or fewer devices may be implemented or included alternatively.
[0095] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the storage device 608, or installed from the ROM 602. When the computer program is executed by the processing device 601, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.
[0096] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0097] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed network.
[0098] The above computer-readable medium can be included in the above electronic device; or can exist separately without being assembled into the electronic device.
[0099] The above computer-readable medium carries one or more programs, which when executed by the electronic device, cause the electronic device to: receive a description text and generation conditions of a target image to be generated, where the generation conditions are used to constrain the features of a target object in the target image; determine a target generation model corresponding to the description text; and obtain the target image based on the description text, the generation conditions, and the target generation model.
[0100] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0101] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0102] The modules described in the embodiments of the present disclosure can be implemented in software or in hardware. In some cases, the name of a module does not constitute a limitation on the module itself. For example, the receiving module can also be described as "the module that receives the description text and generation conditions of the target image to be generated".
[0103] The functions described above herein can be performed, at least in part, by one or more hardware logic components. By way of example and not limitation, the types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0104] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0105] According to one or more embodiments of the present disclosure, Example 1 provides an image generation method, the method comprising: Receiving a description text of a target image to be generated and generation conditions, wherein the generation conditions are used to constrain features of a target object in the target image; Determining a target generation model corresponding to the description text; Obtaining the target image based on the description text, the generation conditions, and the target generation model.
[0106] According to one or more embodiments of the present disclosure, Example 2 provides the method of Example 1, wherein the generation conditions include image conditions; The obtaining the target image based on the description text, the generation conditions, and the target generation model includes: Inputting the description text and the image conditions into the target generation model, so that the target generation model performs multi-modal attention mechanism processing based on the description text and the image conditions to obtain the target image.
[0107] According to one or more embodiments of the present disclosure, Example 3 provides the method of Example 1, wherein the generation conditions include text conditions; The obtaining the target image based on the description text, the generation conditions, and the target generation model includes: Generating an image corresponding to the text conditions based on the text conditions; Inputting the description text and the image corresponding to the text conditions into the target generation model, so that the target generation model performs multi-modal attention mechanism processing based on the description text and the image corresponding to the text conditions to obtain the target image.
[0108] According to one or more embodiments of the present disclosure, Example 4 provides the method of Example 3, the method further comprising: Displaying the image corresponding to the text conditions; In response to a confirmation operation on the image corresponding to the text conditions, performing the step of inputting the description text and the image corresponding to the text conditions into the target generation model, so that the target generation model performs multi-modal attention mechanism processing based on the description text and the image corresponding to the text conditions to obtain the target image; In response to an editing operation on the text conditions, generating a new image based on the text obtained from the editing operation; Using the new image as the image corresponding to the text conditions to return to perform the step of displaying the image corresponding to the text conditions.
[0109] According to one or more embodiments of the present disclosure, Example 5 provides the method of Example 1, and the training data of the target generation model is determined by the following method: Obtain training texts, where the training texts contain feature texts of at least one training object; Generate a plurality of candidate images based on the training texts; Match the plurality of candidate images, and determine the training data based on the matching result and the training texts.
[0110] According to one or more embodiments of the present disclosure, Example 6 provides the method of Example 5, and the training texts contain only feature texts of one training object; The matching of the plurality of candidate images and determining the training data based on the matching result and the training texts includes: Determine the matching result of an image pair, where each image pair contains two candidate images corresponding to the training text, and the matching result of the image pair is used to indicate whether the objects in the two candidate images match; If the matching result of the image pair is a match, use the image pair as a training image pair, and determine the training data based on the training text and the training image pair.
[0111] According to one or more embodiments of the present disclosure, Example 7 provides the method of Example 5, and the training texts contain feature texts of a plurality of training objects; The generating a plurality of candidate images based on the training texts includes: Generate a first image based on the training text, and generate a second image of each training object based on the feature text of each training object in the training text, and the candidate images include the first image and the second image; The matching of the plurality of candidate images and determining the training data based on the matching result and the training texts includes: Determine the matching result of an image pair, where each image pair contains the first image corresponding to the training text and each second image corresponding to the training text, and the matching result of the image pair includes sub-results of the matching of each second image in the first image; If the sub-results in the matching result of the image pair are all matches, use the image pair as a training image pair, and determine the training data based on the training text and the training image pair.
[0112] According to one or more embodiments of the present disclosure, Example 8 provides the method of Example 1, and determining the target generation model corresponding to the description text includes: Perform theme recognition based on the described text to determine the theme type corresponding to the described text; Determine the image generation model corresponding to the theme type as the target generation model.
[0113] According to one or more embodiments of the present disclosure, Example 9 provides an image generation device, which includes: A receiving module, configured to receive the description text of the target image to be generated and the generation conditions, where the generation conditions are used to constrain the features of the target object in the target image; A first determination module, configured to determine the target generation model corresponding to the description text; A first generation module, configured to obtain the target image based on the description text, the generation conditions, and the target generation model.
[0114] According to one or more embodiments of the present disclosure, Example 10 provides a computer-readable medium, on which a computer program is stored, and when the computer program is executed by a processing device, the steps of the method described in any one of Examples 1-8 are implemented.
[0115] According to one or more embodiments of the present disclosure, Example 11 provides an electronic device, including: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of the method described in any one of Examples 1-8.
[0116] According to one or more embodiments of the present disclosure, Example 12 provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method described in any one of Examples 1-8 are implemented.
[0117] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the technical features (but not limited to) having similar functions disclosed in the present disclosure.
[0118] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be construed as limitations on the scope of the present disclosure. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, the various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0119] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims. With regard to the apparatus in the above embodiments, the specific manner in which each module performs an operation has been described in detail in the embodiments related to the method, and will not be elaborated herein.
Claims
1. An image generation method, characterized in that: The method comprises: Receiving a description text and a generation condition of a target image to be generated, wherein the generation condition is used to constrain features of a target object in the target image; Determining a target generation model corresponding to the description text; Based on the description text, the generation condition and the target generation model, the target image is obtained.
2. The method according to claim 1, characterized in that: The generation conditions include image conditions; The step of obtaining the target image based on the description text, the generation condition and the target generation model includes: The description text and the image condition are input into the target generation model, so that the target generation model performs multimodal attention mechanism processing based on the description text and the image condition to obtain the target image.
3. The method according to claim 1, characterized in that The generation condition includes a text condition; The step of obtaining the target image based on the description text, the generation condition and the target generation model includes: Based on the text condition, generating an image corresponding to the text condition; The description text and the image corresponding to the text condition are input into the target generation model, so that the target generation model performs multimodal attention mechanism processing based on the description text and the image corresponding to the text condition to obtain the target image.
4. The method according to claim 3, characterized in that: The method further comprises: Displaying an image corresponding to the text condition; In response to a confirmation operation on the image corresponding to the text condition, the step of inputting the description text and the image corresponding to the text condition into the target generation model so that the target generation model performs multimodal attention mechanism processing based on the description text and the image corresponding to the text condition to obtain the target image is performed; In response to an editing operation on the text condition, generating a new image based on the text obtained by the editing operation; The new image is used as the image corresponding to the text condition, so as to return to the step of displaying the image corresponding to the text condition.
5. The method according to claim 1, characterized in that The training data of the target generation model is determined in the following manner: Acquire a training text, wherein the training text contains at least one feature text of a training object; generating a plurality of candidate images based on the training text; The plurality of candidate images are matched, and the training data is determined based on the matching result and the training text.
6. The method according to claim 5, characterized in that The training text contains only one feature text of the training object; The matching of the plurality of candidate images and determining the training data based on the matching results and the training text includes: Determining a matching result of an image pair, wherein each of the image pairs includes two candidate images corresponding to the training text, and the matching result of the image pair is used to indicate whether objects in the two candidate images match; If the matching result of the image pair is a match, the image pair is used as a training image pair, and the training data is determined based on the training text and the training image pair.
7. The method according to claim 5, characterized in that The training text includes a plurality of feature texts of the training objects; The step of generating a plurality of candidate images based on the training text comprises: Generate a first image based on the training text, and generate a second image of the training object based on the feature text of each training object in the training text, wherein the candidate images include the first image and the second image; The matching of the plurality of candidate images and determining the training data based on the matching results and the training text includes: Determining matching results of image pairs, wherein each of the image pairs includes a first image corresponding to the training text and each of the second images corresponding to the training text, and the matching results of the image pairs include sub-results of each of the second images being matched in the first image; If all sub-results in the matching result of the image pair are matched, the image pair is used as a training image pair, and the training data is determined based on the training text and the training image pair.
8. The method according to claim 1, characterized in that The determining of the target generation model corresponding to the description text includes: Performing subject matter identification based on the description text to determine the subject matter type corresponding to the description text; The image generation model corresponding to the subject matter type is determined as the target generation model.
9. An image generating device, characterized in that: The device comprises: A receiving module, used for receiving a description text and a generation condition of a target image to be generated, wherein the generation condition is used for constraining features of a target object in the target image; A first determination module, used to determine a target generation model corresponding to the description text; The first generation module is used to obtain the target image based on the description text, the generation condition and the target generation model.
10. A computer readable medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processing device, the steps of the method according to any one of claims 1 to 8 are implemented.
11. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 8.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.