Methods, apparatus, devices and media for generating images

By providing recommended labels and pre-trained tuning models in the image generation tool, users can directly adjust the style, background, or foreground attributes of the image, solving the problem that image generation in existing technologies does not meet user needs and achieving a simpler and more effective image generation process.

CN117593404BActive Publication Date: 2025-10-31BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202311743763.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-18
Publication Date
2025-10-31
Estimated Expiration
2043-12-18

AI Technical Summary

Technical Problem

Existing machine learning models for generating images from text often fail to accurately meet user needs, requiring users to constantly adjust the input text to obtain a satisfactory image, a process that is complex and unstable.

Method used

An image generation method and apparatus are provided, which generate an initial image by obtaining prompt words and provide recommended labels near the image, allowing users to directly click on the labels to adjust the style, background or foreground attributes of the image, and use a pre-trained adjustment model to achieve image refinement.

Benefits of technology

The image generation process is simplified, allowing users to quickly and effectively adjust image content to meet expectations. This avoids instability caused by readjusting prompts and improves the accuracy and efficiency of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117593404B_ABST
    Figure CN117593404B_ABST
Patent Text Reader

Abstract

Methods, apparatus, devices, and media for generating images are provided. In one method, a prompt word is obtained to specify the image to be generated. A first image is generated based on the prompt word. The first image and at least one recommended label for adjusting the first image are provided, wherein the at least one recommended label is used to adjust at least one attribute of the first image: style, background, and objects in the foreground. Using exemplary implementations of this disclosure, the recommended label can provide a user with suggestions on modifying image content, and the user can directly click the label to adjust the image content without having to modify the prompt word again. In this way, images that better meet the user's expectations can be generated in a simpler and more efficient manner.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Exemplary implementations of this disclosure generally relate to image generation processing, and more particularly to methods, apparatus, devices, and computer-readable storage media for generating images based on prompt words. Background Technology

[0002] Machine learning techniques have been widely used for visual tasks. For example, various machine learning models for generating images based on text have been proposed. Users can use text to describe the desired image; however, the images generated by machine learning models sometimes do not meet the user's needs, forcing the user to constantly adjust the input text. Therefore, there is a need for simpler and more efficient image generation techniques. Summary of the Invention

[0003] In a first aspect of this disclosure, a method for generating an image is provided. In this method, a cue word is obtained to specify the image to be generated. A first image is generated based on the cue word. The first image and at least one recommended tag for adjusting the first image are provided, wherein the at least one recommended tag is used to adjust at least one of the following attributes of the first image: style, background, and at least one attribute of an object in the foreground.

[0004] In a second aspect of this disclosure, an apparatus for generating an image is provided. The apparatus includes: an acquisition module configured to acquire a prompt word specifying an image to be generated; a generation module configured to generate a first image based on the prompt word; and a providing module configured to provide the first image and at least one recommended tag for adjusting the first image, wherein the at least one recommended tag is used to adjust at least one of the following attributes of the first image: style, background, and at least one attribute of an object in the foreground.

[0005] In a third aspect of this disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processing unit.

[0006] In a fourth aspect of this disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to a first aspect of this disclosure.

[0007] It should be understood that the content described in this content section is not intended to limit the key or essential features of the implementation of this disclosure, nor is it intended to restrict the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0008] In the following detailed description, the above and other features, advantages, and aspects of the various implementations of this disclosure will become more apparent, taken in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals denote the same or similar elements, wherein:

[0009] Figure 1 A block diagram of an application environment according to an exemplary implementation of this disclosure is shown;

[0010] Figure 2 A block diagram illustrating some implementations of this disclosure for generating images is shown;

[0011] Figure 3 A block diagram is shown illustrating a process for generating recommendation tags according to some implementations of this disclosure;

[0012] Figure 4 A block diagram of an image adjusted based on different recommendation tags according to some implementations of this disclosure is shown;

[0013] Figure 5 A block diagram is shown of a conversion model corresponding to multiple recommendation tags according to some implementations of this disclosure;

[0014] Figure 6 A block diagram is shown illustrating the process of obtaining user adjustment requirements according to some implementations of this disclosure;

[0015] Figure 7 A flowchart illustrating a template-based process for obtaining prompt words according to some implementations of this disclosure is shown;

[0016] Figure 8 A block diagram is shown illustrating a process for generating an image based on edited prompts according to some implementations of this disclosure;

[0017] Figure 9 A block diagram is shown illustrating a process for editing prompt words according to some implementations of this disclosure;

[0018] Figure 10 A flowchart is shown of a method for generating an image according to some implementations of this disclosure;

[0019] Figure 11 A block diagram of an apparatus for generating an image according to some implementations of the present disclosure is shown; and

[0020] Figure 12 A block diagram of a device capable of implementing various implementations of the present disclosure is shown. Detailed Implementation

[0021] Implementations of this disclosure will now be described in more detail with reference to the accompanying drawings. While some implementations of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the implementations set forth herein. Rather, these implementations are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and implementations of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0022] In the description of the implementation methods disclosed herein, the term "comprising" and similar terms should be understood as open inclusion, i.e., "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". Other explicit and implicit definitions may also be included below. As used herein, the term "model" can represent the relationships between various data. For example, the aforementioned relationships can be obtained based on various currently known and / or future-developed technical solutions.

[0023] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.

[0024] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure through appropriate means in accordance with relevant laws and regulations, and user authorization should be obtained.

[0025] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0026] As an optional but non-restrictive implementation, in response to a user's active request, a prompt message can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose whether to "agree" or "disagree" to provide personal information to the electronic device.

[0027] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0028] The term "in response to" as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of subsequent actions performed in response to such event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; while in others, they may be performed some time after the occurrence of the event or the fulfillment of the condition.

[0029] Example Environment

[0030] Machine learning techniques have been widely used for visual tasks. For example, several machine learning models have been proposed for generating images based on text. Users can use text to describe the image they wish to generate. Figure 1 A block diagram 100 of an application environment according to an exemplary implementation of this disclosure is shown. Figure 1 As shown, in the interface 130 of the image generation tool, the user can enter a prompt word 110, and the image generation tool will return an image 120 generated based on the prompt word 110.

[0031] Specifically, the prompt 110 might include "dining table, coffee cup," and the generated image 120 would include the content specified by the prompt, such as a cup of coffee on a dining table. However, images generated by machine learning models sometimes don't meet user needs, forcing users to constantly adjust the input text. For example, a user might not be satisfied with the background of image 120, or might want to generate an image with a different style, etc. In this case, the user has to manually adjust the prompt text; however, this adjustment process might introduce other unstable factors. For example, the user might want to modify the desktop texture or the color of the coffee cup in the new image, etc. Therefore, a simpler and more effective image generation technology solution is needed.

[0032] Summary of image generation

[0033] To at least partially address the shortcomings of the prior art, a method for generating images is proposed according to an exemplary implementation of this disclosure. See also Figure 2 A summary of an exemplary implementation according to this disclosure is provided. Figure 2 A block diagram 200 for generating images according to some implementations of this disclosure is shown. For example... Figure 2As shown, in the interface 210 of the image generation tool, a prompt word 110 for specifying the image to be generated can be received, and a first image (e.g., image 120) can be generated based on the prompt word 110. Furthermore, image 120 can be provided, and at least one recommendation label 220 for adjusting image 120 can be provided near image 120 (e.g., at a position immediately below image 120).

[0034] like Figure 2 As shown, at least one recommendation tag 220 may include one or more tags for adjusting at least one of the following attributes of image 120: style, background, and at least one attribute of objects in the foreground. Specifically, recommendation tags 220-1 and 220-2 can change the background of the image, recommendation tag 220-3 can modify the style of the image, and recommendation tag 220-4 can modify at least one attribute of objects in the foreground of the generated image 120, and so on. Users can interact 230 with the aforementioned recommendation tags to perform the corresponding adjustment process.

[0035] Using the exemplary implementation of this disclosure, recommendation tags can provide users with suggestions on modifying image content, and users can directly click on the tags to adjust the image content without having to modify the prompts. In this way, images that better meet user expectations can be generated in a simpler and more effective manner.

[0036] Detailed process of image generation

[0037] Having outlined an example implementation according to this disclosure, further details regarding the provision of recommendation tags will be described below. Recommendation tags can be generated in various ways according to the example implementation of this disclosure. See also... Figure 3 ,Should Figure 3 A block diagram 300 illustrates a process for generating recommendation tags according to some implementations of this disclosure. For example... Figure 3 As shown, the factors 330 that generate the recommendation label 220 may include at least one of the prompt words 110 and the image 120.

[0038] For example, semantic analysis can be performed on cue word 110 to determine individual keywords 310, and so on. Here, keyword 310 can represent an element to be included in the image, and thus keyword 310 can be used to determine recommendation tag 220. For example, if cue word 110 includes the keywords "dining table" and "coffee cup," the generated recommendation tag could indicate "adjust the color of the dining table," "adjust the style of the coffee cup," and so on.

[0039] According to an example implementation of this disclosure, image analysis can be performed on the generated image 120 to extract image-related style features 320, texture features 324, color features 322, and geometric features 326, etc. Furthermore, recommendation tags can be generated based on any of the aforementioned features; for example, recommendation tags corresponding to features different from those currently detected can be provided.

[0040] For example, assuming image 120 has a realistic style, recommended tags could be provided to adjust it to a "line painting style," "oil painting style," etc. Similarly, assuming texture feature 324 indicates the dining table has a wood grain texture, recommended tags could be provided to adjust the dining table in image 120 to a "glass texture," "tablecloth texture," etc. And assuming color feature 322 indicates image 120 has a "dark tone," recommended tags could be provided to adjust the color of image 120 to a "bright tone," etc.

[0041] Alternatively and / or additionally, assuming that geometric feature 326 indicates that image 120 includes a foreground (e.g., including a dining table and a coffee cup on the table) and a background (e.g., the wall behind the dining table), recommended labels for modifying the foreground and / or background of image 120 can be provided. Specifically, the recommended labels could prompt the user to adjust the background to a sea background, or a starry sky background, etc.

[0042] Using the example implementation of this disclosure, one or more candidate recommendation tags can be provided to the user based on multi-faceted analysis of the prompt word 110 and the image 120. This simplifies the operational complexity of image generation, thereby generating images that better meet user expectations in a simpler and more effective way.

[0043] According to an example implementation of this disclosure, a user can select a desired recommendation tag and then adjust the content in image 120 corresponding to that recommendation tag. Specifically, upon receiving an interaction (e.g., referred to as a first interaction) targeting a target recommendation tag among at least one recommendation tag, the image generation tool can receive an adjustment request corresponding to the target recommendation tag. Further, the image generation tool can provide a new image (e.g., referred to as a second image), which is determined based on the adjustment request and the first image.

[0044] See Figure 4 This describes more details about using recommended tags to adjust images. Figure 4 A block diagram 400 illustrates an image adjusted based on different recommendation tags according to some implementations of this disclosure. Assuming a user presses recommendation tags 220-3, image 120 will be transformed to... Figure 4The image 410 shown has a line-style design. Alternatively and / or additionally, assuming the user presses the recommendation tab 220-1, image 120 will be converted to... Figure 4 Image 420 is shown with a sea background. Alternatively and / or additionally, assuming the user presses recommendation tab 220-2, image 120 will be converted to... Figure 4 Image 430 with a starry background is shown.

[0045] According to an example implementation of this disclosure, if the generated image 120 does not meet the user's expectations, the user only needs to select a suitable recommended tag through simple operations such as clicking to perform subsequent image adjustment tasks. In this way, the user does not need to readjust the prompts, thus avoiding the problem of introducing more uncontrollable factors into the newly generated image due to prompt adjustment.

[0046] According to one example implementation of this disclosure, each recommendation tag can correspond to a dedicated image adjustment model, see [link to relevant documentation]. Figure 5 Describe more details. Figure 5 A block diagram 500 is shown illustrating a conversion model corresponding to multiple recommendation tags according to some implementations of this disclosure. For example... Figure 5 As shown, the recommended label 220-1 used to adjust the background of an image to a sea background can correspond to adjustment model 510. Here, adjustment model 510 can be a pre-trained machine learning model for modifying the background of an image. Specifically, adjustment model 510 can adjust the background of the input image to one or more pre-specified sea backgrounds.

[0047] According to an example implementation of this disclosure, an original image can be obtained, and the background portion of the original image can be replaced with an image of the ocean using methods such as manual processing. Furthermore, the original image and the replaced image can be used as training data. A large amount of training data can be obtained in a similar manner, and then the final adjusted model 510 can be obtained through an iterative training process.

[0048] It should be understood that although the above only illustrates the case where the image adjustment process is directly triggered by clicking the recommendation tag, alternatively and / or additionally, more interaction can be provided after the user clicks the recommendation tag. For example, the user can specify adjustment needs via text and / or images, thereby adjusting the image background to the specified content.

[0049] Figure 6 A block diagram 600 illustrates a process for obtaining user adjustment requirements according to some implementations of this disclosure. For example... Figure 6As shown, suppose a user presses label 220-1 to change the background of an image to a sea background. At this point, various prompts can be provided to the user to obtain their adjustment requests 612. For example, control 620 allows the user to input their adjustment requests via text. For instance, the user can enter keywords such as "sea" or "beach" in the text input control 622 to specify the exact content of the background image.

[0050] Alternatively and / or additionally, control 630 may allow the user to specify adjustment requirements graphically. For example, the user can click on a predefined image 632, etc., to specify the specific content of the background image. Alternatively and / or additionally, controls may be provided to the user for specifying the background image, allowing the user to select a background image from the client device's photo album or from other remote locations on the network.

[0051] It should be understood that the various recommended tags presented in the accompanying drawings are merely illustrative, and alternatively and / or additionally, one or more other recommended tags may be provided. For example, assuming the generated image is a person image, one or more other recommended tags such as face swapping, hairstyle swapping, clothing swapping, and jewelry swapping may be provided. If the user selects the face swapping tag, a control may be provided to allow the user to specify the image to replace the current person's face, and so on.

[0052] Using the example implementation of this disclosure, users can specify adjustment requirements in various ways. Furthermore, the adjustment model corresponding to the user's selected recommendation label can output an image that better matches the user's expectations based on those adjustment requirements.

[0053] For example, if a user selects a recommendation tag 220-1 (e.g., referred to as the target recommendation tag), the image generation tool will select an adjustment model 510 that matches the recommendation tag 220-1. Furthermore, this adjustment model 510 can be used to generate a new image (e.g., referred to as the second image) based on the adjustment requirements and image 120.

[0054] It should be understood that the adjustment model 510 described herein can be obtained based on a predefined training dataset. Compared to large-scale text-to-image machine learning models, adjustment models performing a specific adjustment task typically have a simpler structure and fewer parameters. Therefore, the workload involved in obtaining the adjustment model and using it to perform the inference process is generally smaller. Using the example implementation of this disclosure, images can be adjusted in a more accurate and efficient manner by pre-obtaining a dedicated, small-scale adjustment model.

[0055] Furthermore, such as Figure 5As shown, the recommended label 220-2 for adjusting the background of an image to a starry sky background corresponds to adjustment model 520; the recommended label 220-3 for adjusting the image to a line style corresponds to adjustment model 530; and the recommended label 220-4 for adjusting the attributes of the foreground in the image corresponds to adjustment model 540. The above adjustment models can be obtained using a predetermined training dataset in a similar manner.

[0056] According to an example implementation of this disclosure, after providing the user with the adjusted new image, at least one recommended tag for adjusting the new image can be further provided. For example, at least one recommended tag can be generated based on the analysis of the new image, in a manner similar to that described above. Alternatively and / or additionally, new recommended tags can be generated based on the recommended tags selected by the user for generating the new image, and so on. In this way, it is convenient for the user to further adjust the image to generate a final image that meets their own needs.

[0057] According to one example implementation of this disclosure, the prompt word 110 used to generate the image can be obtained in several ways. For example, the user can input the prompt word through text editing, or alternatively and / or additionally, the prompt word can be obtained from an image template. See also Figure 7 Describe more details about obtaining prompt words. Figure 7 A block diagram 700 illustrates a template-based process for retrieving prompt words according to some implementations of this disclosure. For example... Figure 7 As shown, the interface 710 of the image generation tool can provide at least one image template indicating the image to be generated. Each image template may have predetermined prompts, and the user can select a template to generate a similar image.

[0058] According to an example implementation of this disclosure, upon receiving an interaction (e.g., referred to as a second interaction) with respect to a target image template in at least one image template, a prompt word template corresponding to the target image template can be provided. Furthermore, the prompt word can be determined based on receiving an interaction (e.g., referred to as a third interaction) with respect to the prompt word template. Using the example implementation of this disclosure, the user does not need to input the prompt word themselves, but can obtain the prompt word from a predetermined template, thereby generating images in a faster and more efficient manner.

[0059] Specifically, template 720 can be used to generate coffee-themed images, and the prompt word for template 720 could be, for example, "A cup of coffee on the table...". A user can press control 722 to generate a similar image. Specifically, when the user presses control 722, they can directly input the prompt word "A cup of coffee on the table...". According to an example implementation of this disclosure, if the user interactively confirms the prompt word template, the prompt word template can be determined as the prompt word. In other words, with user confirmation, the prompt word can be submitted to the image generation tool, thereby generating the corresponding image.

[0060] For example, template 730 can be used to generate images of cartoon characters. The prompt for template 730 could be something like "sweet girl, cartoon style, wireframe...". Users can press control 732 to generate similar cartoon characters. Specifically, when a user presses control 732, they can directly input the prompt template "sweet girl, cartoon style, wireframe...". With user confirmation, the prompt can be submitted to the image generation tool, thereby generating the corresponding image.

[0061] Alternatively and / or additionally, users can edit this prompt template to generate custom prompts. See also Figure 8 To describe more details, the Figure 8 A block diagram 800 illustrates a process for generating an image based on edited prompts, according to some implementations of this disclosure. For example... Figure 8 As shown, when the user presses control 732, the prompt words from template 730 can be directly copied to control 820 for inputting prompt words. The user can perform editing operations in control 820; for example, as shown in box 822, a "half-body portrait" can be added to the prompt words. Upon submitting the modified prompt words, the generated image 830 can be a cartoon-style half-body portrait. In this way, on the one hand, prompt word templates are provided to users, reducing the complexity of user input operations; on the other hand, users can modify the specific content of the prompt word templates to generate images that better suit their needs.

[0062] According to one example implementation of this disclosure, the prompt word template can have a structured format. See also Figure 9 To describe more details, the Figure 9 A block diagram 900 illustrates a process for editing prompt words according to some implementations of this disclosure. For example... Figure 9 As shown, in interface 910, a control 920 can be provided to ask the user whether they want to edit the prompt word of the selected template. If the user presses control 924, an image can be generated directly based on the current prompt word template. Alternatively and / or additionally, if the user presses control 922, an editing control 930 can be further provided.

[0063] In the editing control 930, the prompt template can be presented in a structured format. In this case, the prompt template can include at least one descriptive word specifying the image to be generated, and the value of at least one descriptive word. For example... Figure 9 As shown, descriptive terms can include, for example, foreground, style, and multiple attributes of the character in the foreground, such as hair, clothing, earrings, etc. The "=" sign can be used to separate the descriptive term from its value; for example, "style = comic style, wireframe" specifies that a comic-style wireframe image will be generated.

[0064] It should be understood that Figure 9 This illustration only shows some examples of descriptive words in the prompt template. Alternatively and / or additionally, descriptive words may further include, but are not limited to, image style, foreground, background, hue, time, location, people, etc. For example, setting "hue = warm" can generate a warm-toned image, setting "time = night" can generate an image for a nighttime period, and so on. Furthermore, specific information such as the age, hair color, and clothing of people can be specified. Using the example implementations of this disclosure, various aspects of the generated image can be explicitly specified, thereby facilitating an image generation tool to more accurately understand the user's detailed needs.

[0065] According to an example implementation of this disclosure, the prompt template includes editable and non-editable portions. Specifically, box 932 shows the prompt for the non-editable portion. For example, it can be specified that the salient features of the prompt template related to the image template are non-editable, while other portions are editable. In the prompt template, "sweet girl" specifies a salient feature of the generated image (i.e., the person in the foreground is "sweet girl"), and "foreground = sweet girl" is a non-editable portion and is presented in box 932. As another example, "cartoon style, wireframe" in the prompt template specifies the presentation style of the generated image, which is also a salient feature specific to this image template; therefore, "style = cartoon style, wireframe" is a non-editable portion and is presented in box 932.

[0066] According to an example implementation of this disclosure, a user can enter additional information in box 934, which indicates the editable portion. For example, a user can specify the hair color, shirt color, earring color, etc., of a cartoon character. Using this example implementation, on the one hand, it prevents users from accidentally modifying the core content of the prompt template, thus preventing the generation of image content that does not conform to the image template; on the other hand, it allows users to add custom content based on the prompt template to refine their specific needs.

[0067] It should be understood that the above is merely an example, and alternatively and / or additionally, editable and non-editable parts can be defined in different ways. For example, the descriptive words (e.g., style, background, etc.) in the structured format may be non-editable, while the values ​​of the descriptive words may be editable. This approach prevents users from disrupting the structured format of the prompts during editing, thus ensuring that the prompts accurately reflect user needs.

[0068] According to an example implementation of this disclosure, a user can modify the value of a descriptor and update the prompt word template upon receiving an interaction (e.g., a fourth interaction) regarding the value of a target descriptor among at least one descriptor. For example, assuming the prompt word template includes "background = white", the user can modify the specific value of the background, such as "background = gray". In this case, the generated image 940 will have a gray background. In this way, users can adjust the prompt words in a simpler and more efficient manner to generate images that better suit their needs.

[0069] Using the exemplary implementation of this disclosure, recommendation tags can provide users with suggestions on modifying image content, and users can directly click on the tags to adjust the image content without having to modify the prompts. In this way, images that better meet user expectations can be generated in a simpler and more effective manner.

[0070] Example process

[0071] Figure 10 A flowchart of a method 1000 for generating an image according to some implementations of this disclosure is shown. At box 1010, a prompt word is obtained to specify the image to be generated. At box 1020, a first image is generated based on the prompt word. At box 1030, the first image and at least one recommended label for adjusting the first image are provided, wherein the at least one recommended label is used to adjust at least one of the following attributes of the first image: style, background, and at least one attribute of an object in the foreground.

[0072] According to one example implementation of this disclosure, at least one recommendation label is determined based on at least one of the following: keywords in the prompt, style features of the first image, color features, texture features, and geometric features.

[0073] According to an example implementation of this disclosure, the method further includes: in response to receiving a first interaction for a target recommendation tag among at least one recommendation tag, receiving an adjustment request corresponding to the target recommendation tag; and providing a second image, the second image being determined based on the adjustment request and the first image.

[0074] According to one example implementation of this disclosure, the second image is determined based on: selecting a machine learning model that matches the target recommended label; and using the machine learning model to generate the second image based on the adjustment requirements and the first image.

[0075] According to one example implementation of this disclosure, the adjustment requirement includes at least one of text and image, and the method further includes: providing at least one recommended label for adjusting the second image.

[0076] According to an example implementation of this disclosure, obtaining the prompt word includes: providing at least one image template indicating the image to be generated; in response to receiving a second interaction for a target image template in the at least one image template, providing a prompt word template corresponding to the target image template; and in response to receiving a third interaction for the prompt word template, determining the prompt word.

[0077] According to one example implementation of this disclosure, determining the prompt word includes at least one of the following: determining the prompt word template as a prompt word in response to determining that the third interaction instruction prompt word template has been confirmed; and determining the edited prompt word template as a prompt word in response to determining that the third interaction instruction prompt word template has been edited.

[0078] According to an example implementation of this disclosure, the prompt word template includes at least one descriptive word for specifying the image to be generated and the value of the at least one descriptive word, wherein the at least one descriptive word includes at least one of the following: style, foreground, background, hue, time, location, and people of the image to be generated.

[0079] According to one example implementation of this disclosure, the prompt word template includes an editable portion and a non-editable portion.

[0080] According to one example implementation of this disclosure, the method further includes: updating the prompt word template in response to receiving a fourth interaction for a value of a target descriptor in at least one descriptor.

[0081] Example devices and equipment

[0082] Figure 11 A block diagram of an apparatus 1100 for generating an image according to some implementations of the present disclosure is shown. The apparatus 1100 includes: an acquisition module 1110 configured to acquire a cue word specifying an image to be generated; a generation module 1120 configured to generate a first image based on the cue word; and a providing module 1130 configured to provide the first image and at least one recommended tag for adjusting the first image, wherein the at least one recommended tag is used to adjust at least one of the following attributes of the first image: style, background, and at least one attribute of an object in the foreground.

[0083] According to one example implementation of this disclosure, at least one recommendation label is determined based on at least one of the following: keywords in the prompt, style features of the first image, color features, texture features, and geometric features.

[0084] According to an example implementation of this disclosure, the apparatus further includes: a receiving module configured to receive an adjustment request corresponding to the target recommendation tag in response to receiving a first interaction for at least one target recommendation tag; and an image providing module configured to provide a second image, the second image being determined based on the adjustment request and the first image.

[0085] According to one example implementation of this disclosure, the second image is determined based on: a selection module configured to select a machine learning model that matches the target recommended label; and a calling module configured to utilize the machine learning model to generate the second image based on the adjustment requirements and the first image.

[0086] According to one example implementation of this disclosure, the adjustment requirements include at least one of text and image, and the apparatus further includes: a label providing module configured to provide at least one recommended label for adjusting the second image.

[0087] According to an example implementation of this disclosure, the acquisition module includes: an image template providing module configured to provide at least one image template indicating an image to be generated; a prompt word template providing module configured to provide a prompt word template corresponding to the target image template in response to receiving a second interaction for a target image template among the at least one image template; and a determination module configured to determine a prompt word in response to receiving a third interaction for the prompt word template.

[0088] According to an example implementation of this disclosure, the determining module includes at least one of the following: a first determining module configured to determine the prompt word template as a prompt word in response to determining that the third interaction indication prompt word template has been confirmed; and a second determining module configured to determine the edited prompt word template as a prompt word in response to determining that the third interaction indication prompt word template has been edited.

[0089] According to an example implementation of this disclosure, the prompt word template includes at least one descriptive word for specifying the image to be generated and the value of the at least one descriptive word, wherein the at least one descriptive word includes at least one of the following: style, foreground, background, hue, time, location, and people of the image to be generated.

[0090] According to one example implementation of this disclosure, the prompt word template includes an editable portion and a non-editable portion.

[0091] According to one example implementation of this disclosure, the apparatus further includes: an update module configured to update a prompt word template in response to receiving a fourth interaction for a value of a target descriptor in at least one descriptor.

[0092] Figure 12 A block diagram of a device 1200 capable of implementing various implementations of the present disclosure is shown. It should be understood that... Figure 12 The computing device 1200 shown is merely exemplary and should not be construed as limiting the functionality and scope of the implementation described herein. Figure 12 The computing device 1200 shown can be used to implement the method described above.

[0093] like Figure 12 As shown, computing device 1200 is in the form of a general-purpose computing device. Components of computing device 1200 may include, but are not limited to, one or more processors or processing units 1210, memory 1220, storage devices 1230, one or more communication units 1240, one or more input devices 1250, and one or more output devices 1260. Processing unit 1210 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 1220. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of computing device 1200.

[0094] Computing device 1200 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 1200, including but not limited to volatile and non-volatile media, removable and non-removable media. Memory 1220 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. Storage device 1230 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and / or data (e.g., training data for training) and can be accessed within computing device 1200.

[0095] The computing device 1200 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not explicitly stated... Figure 12As shown, disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks can be provided. In these cases, each drive can be connected to a bus (not shown) via one or more data media interfaces. Memory 1220 may include computer program product 1225 having one or more program modules configured to perform various methods or actions of various implementations of this disclosure.

[0096] The communication unit 1240 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 1200 can be implemented as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, the computing device 1200 can operate in a networked environment using logical connections to one or more other servers, networked personal computers (PCs), or another network node.

[0097] Input device 1250 can be one or more input devices, such as a mouse, keyboard, trackball, etc. Output device 1260 can be one or more output devices, such as a monitor, speaker, printer, etc. Computing device 1200 can also communicate with one or more external devices (not shown) via communication unit 1240 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with computing device 1200, or with any device (e.g., network card, modem, etc.) that enables computing device 1200 to communicate with one or more other computing devices. Such communication can be performed via input / output (I / O) interface (not shown).

[0098] According to exemplary implementations of this disclosure, a computer-readable storage medium is provided that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above. According to exemplary implementations of this disclosure, a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.

[0099] Various aspects of this disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to this disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0100] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions / actions specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and / or other device to operate in a particular manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions / actions specified in one or more blocks of the flowchart and / or block diagram.

[0101] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions / actions specified in one or more boxes of a flowchart and / or block diagram.

[0102] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0103] Various implementations of this disclosure have been described above. These descriptions are exemplary and not exhaustive, nor are they limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is chosen to best explain the principles, practical applications, or improvements to technology in the market, or to enable others skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for generating an image, comprising: In the dialog box, retrieve the prompt words used to specify the image to be generated, including: Provide at least one image template that indicates the image to be generated; In response to receiving a second interaction for a target image template in the at least one image template, a prompt word template corresponding to the target image template is provided, the prompt word template being editable; In response to receiving a third interaction for the prompt word template, the prompt word is determined; Generate a first image based on the prompt words; and The system provides the first image and at least one recommended label for adjusting the first image, the at least one recommended label being determined based at least on the prompt word and the first image, and the at least one recommended label corresponding to at least one machine learning model. Wherein, the at least one recommendation tag is used to invoke the at least one machine learning model to adjust at least one of the following attributes of the first image: style, background, and objects in the foreground, without adjusting the prompt words.

2. The method of claim 1, wherein the at least one recommendation tag is determined based on at least one of the following: keywords in the prompt words, style features, color features, texture features, and geometric features of the first image.

3. The method according to claim 1, further comprising: In response to receiving a first interaction for a target recommendation tag among the at least one recommendation tags, an adjustment request corresponding to the target recommendation tag is received; as well as A second image is provided, which is determined based on the adjustment requirements and the first image.

4. The method of claim 3, wherein the second image is determined based on: Select a machine learning model that matches the target recommendation label; and The second image is generated using the machine learning model, based on the adjustment requirements and the first image.

5. The method of claim 3, wherein the adjustment requirement includes at least one of text and images, and the method further comprises: Provide at least one recommended label for adjusting the second image.

6. The method of claim 1, wherein determining the prompt word comprises at least one of the following: In response to determining that the third interaction indicates the prompt word template has been acknowledged, the prompt word template is determined as the prompt word; and In response to determining that the third interaction indicates that the prompt word template has been edited, the edited prompt word template is identified as the prompt word.

7. The method of claim 1, wherein the prompt word template includes at least one descriptive word for specifying the image to be generated and the value of the at least one descriptive word, the at least one descriptive word including at least one of the following: style, foreground, background, hue, time, location, and people of the image to be generated.

8. The method of claim 7, wherein the prompt word template includes an editable portion and a non-editable portion.

9. The method of claim 8, further comprising: In response to receiving a fourth interaction for a value of a target descriptor among the at least one descriptor, the prompt word template is updated.

10. An apparatus for generating an image, comprising: The acquisition module is configured to retrieve, in a dialog box, prompts for specifying the image to be generated, including: An image template providing module is configured to provide at least one image template indicating the image to be generated; A prompt word template providing module is configured to, in response to receiving a second interaction for a target image template in the at least one image template, provide a prompt word template corresponding to the target image template, the prompt word template being editable; and A determination module is configured to determine the prompt word in response to receiving a third interaction for the prompt word template; A generation module is configured to generate a first image based on the prompt words; and A providing module is configured to provide the first image and at least one recommended label for adjusting the first image, the at least one recommended label being determined based at least on the prompt word and the first image, the at least one label corresponding to at least one machine learning model. Wherein, the at least one recommendation tag is used to invoke the at least one machine learning model to adjust at least one of the following attributes of the first image: style, background, and objects in the foreground, without adjusting the prompt words.

11. An electronic device, comprising: At least one processing unit; as well as At least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which, when executed by the at least one processing unit, cause the electronic device to perform the method according to any one of claims 1 to 9.

12. A computer-readable storage medium having a computer program stored thereon, the computer program causing the processor to implement the method according to any one of claims 1 to 9 when executed by a processor.

Citation Information

Patent Citations

  • Method and device for generating cartoon, equipment and medium

    CN116630453A

  • Image generation method and device, electronic equipment and storage medium

    CN117170559A

  • Content generation method and device, computer equipment and storage medium

    CN117171369A

  • Clothing try-on graph generation method, generation system and generation model training method

    CN117218226A