Method and apparatus for generating image, device, and medium
By obtaining the prompt word to generate images and providing recommendation tags to adjust the image content, the problem that image generation in the prior art does not meet user needs is solved, and a simpler and more effective image generation process is realized.
Patent Information
- Application Number
- PCT/CN2024/123097
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-09-30
- Publication Date
- 2025-06-26
AI Technical Summary
Existing text-based image generation models often do not meet user needs when generating images, resulting in users needing to continuously adjust the input text to achieve satisfactory results.
A method and apparatus are provided to generate an image by obtaining a prompt word and provide a recommended label to adjust the properties of the image's style, background or foreground object. Users can directly click on the recommendation tag to adjust the image content without re-adjusting the prompt words.
The image generation process is simplified, and the user can adjust the image content more intuitively, avoiding the unstable factors introduced due to the adjustment of prompt words, and improving the possibility that the generated image meets user expectations.
Smart Images

Figure CN2024123097_26062025_PF_FP_ABST
Abstract
Description
Method, apparatus, device and medium for generating an image
[0001] This application claims priority to the Chinese invention patent application entitled “Methods, devices, apparatus and media for generating images” and application number 2023117437630, filed on December 18, 2023, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] Exemplary implementations of the present disclosure generally relate to image generation processing, and more particularly, to methods, devices, apparatuses, and computer-readable storage media for generating images based on prompt words. Background Art
[0003] Machine learning technology has been widely used in vision tasks. For example, various machine learning models have been proposed for generating images based on text. Users can use text to describe the desired image, but the images generated by machine learning models sometimes do not meet user requirements, forcing users to constantly adjust the input text. Therefore, it is desirable to provide simpler and more effective image generation technology solutions.
[0004] Summary of the Invention
[0005] In a first aspect of the present disclosure, a method for generating an image is provided. In the method, a prompt word is obtained for specifying an image to be generated. A first image is generated based on the prompt word. The first image and at least one recommended tag for adjusting the first image are provided. The at least one recommended tag is used to adjust at least one of the following attributes of the first image: style, background, and at least one attribute of a foreground object.
[0006] In a second aspect of the present disclosure, a device for generating an image is provided. The device includes: an acquisition module configured to acquire a prompt word for specifying a to-be-generated image; a generation module configured to generate a first image based on the prompt word; and a provision module configured to provide the first image and at least one recommended tag for adjusting the first image, wherein the at least one recommended tag is used to adjust at least one of the following attributes of the first image: style, background, and at least one attribute of a foreground object.
[0007] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the present disclosure.
[0008] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the processor implements the method according to the first aspect of the present disclosure.
[0009] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the implementation of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other features, advantages and aspects of various implementations of the present disclosure will become more apparent hereinafter with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0011] FIG1 shows a block diagram of an application environment according to an exemplary implementation of the present disclosure;
[0012] FIG2 illustrates a block diagram for generating an image according to some implementations of the present disclosure;
[0013] FIG3 shows a block diagram of a process for generating recommendation tags according to some implementations of the present disclosure;
[0014] FIG4 illustrates a block diagram of an adjusted image based on different recommended tags according to some implementations of the present disclosure;
[0015] FIG5 shows a block diagram of conversion models corresponding to multiple recommendation tags respectively according to some implementations of the present disclosure;
[0016] FIG6 shows a block diagram of a process for obtaining a user's adjustment requirements according to some implementations of the present disclosure;
[0017] FIG7 shows a block diagram of a process of acquiring prompt words based on a template according to some implementations of the present disclosure;
[0018] FIG8 shows a block diagram of a process of generating an image based on edited prompt words according to some implementations of the present disclosure;
[0019] FIG9 illustrates a block diagram of a process for editing a prompt word according to some implementations of the present disclosure;
[0020] FIG10 shows a flowchart of a method for generating an image according to some implementations of the present disclosure;
[0021] FIG11 shows a block diagram of an apparatus for generating an image according to some implementations of the present disclosure; and
[0022] FIG12 illustrates a block diagram of a device capable of implementing various implementations of the present disclosure. DETAILED DESCRIPTION
[0023] The following describes implementations of the present disclosure in more detail with reference to the accompanying drawings. Although certain implementations of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the implementations described herein. Rather, these implementations are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and implementations of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0024] In the description of the implementation of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "an implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". The following may also include other explicit and implicit definitions. As used herein, the term "model" can represent the association relationship between various data. For example, the above-mentioned association relationship can be obtained based on a variety of technical solutions currently known and / or to be developed in the future.
[0025] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0026] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0027] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0028] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0029] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0030] As used herein, the term "in response to" refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of executing a subsequent action executed in response to the event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is satisfied. For example, in some cases, a subsequent action may be executed immediately upon the occurrence of the event or the satisfaction of the condition; in other cases, the subsequent action may be executed some time after the occurrence of the event or the satisfaction of the condition.
[0031] Sample Environment
[0032] Machine learning technology has been widely used in visual tasks. For example, a variety of machine learning models have been proposed for generating images based on text. A user can use text to describe the image they wish to generate. FIG1 shows a block diagram 100 of an application environment according to an exemplary implementation of the present disclosure. As shown in FIG1 , in an interface 130 of an image generation tool, a user can enter a prompt word 110, and the image generation tool will then return an image 120 generated based on the prompt word 110.
[0033] Specifically, the prompt word 110 may include "dining table, coffee cup", and the image 120 generated at this time will include the content specified by the prompt word, for example, a cup of coffee placed on the dining table. However, the image generated by the machine learning model sometimes does not meet the user's needs, which causes the user to have to constantly adjust the input text. For example, the user may not be satisfied with the background of the image 120, or want to generate an image of a different style, etc. At this time, the user has to manually adjust the text of the prompt word, however, other unstable factors may be introduced in the adjustment process. For example, the user may want to modify the desktop material or the color of the coffee cup in the new image, etc. At this time, it is expected that a simpler and more effective image generation technology solution can be provided.
[0034] Overview of Image Generation
[0035] In order to at least partially address the deficiencies in the prior art, according to an exemplary implementation of the present disclosure, a method for generating an image is proposed. Referring to FIG2 , an overview of an exemplary implementation of the present disclosure is described, which shows a block diagram 200 for generating an image according to some implementations of the present disclosure. As shown in FIG2 , in the interface 210 of the image generation tool, a prompt word 110 for specifying an image to be generated may be received, at which point a first image (e.g., image 120) may be generated based on the prompt word 110. Further, an image 120 may be provided, and at least one recommended tag 220 for adjusting the image 120 may be provided near the image 120 (e.g., immediately below the image 120).
[0036] As shown in Figure 2, at least one recommended tag 220 may include one or more tags for adjusting at least one of the following aspects of image 120: style, background, and at least one attribute of an object in the foreground. Specifically, recommended tags 220-1 and 220-2 may change the image background, recommended tag 220-3 may modify the style of an image, and recommended tag 220-4 may modify at least one attribute of an object in the foreground of the generated image 120, and so on. Users can interact 230 with these recommended tags to perform corresponding adjustments.
[0037] Using the exemplary implementations of this disclosure, recommended tags can provide users with suggestions for modifying image content, and users can directly click on the tags to adjust the image content without having to re-edit the prompt words. In this way, images that better meet user expectations can be generated in a simpler and more efficient manner.
[0038] Detailed process of image generation
[0039] Having described an overview of an example implementation of the present disclosure, further details regarding providing recommended tags will be described below. According to an example implementation of the present disclosure, recommended tags can be generated based on a variety of methods. Referring to FIG3 , FIG3 shows a block diagram 300 of a process for generating recommended tags according to some implementations of the present disclosure. As shown in FIG3 , the dependency factors 330 for generating the recommended tags 220 can include at least one of the prompt words 110 and the image 120.
[0040] For example, semantic analysis can be performed on the prompt word 110 to determine various keywords 310, etc. Here, the keywords 310 can represent elements to be included in the image, and thus the keywords 310 can be used to determine the recommended tags 220. For example, if the prompt word 110 includes the keywords "dining table" and "coffee cup", the generated recommended tags can indicate "adjust the color of the dining table", "adjust the style of the coffee cup", etc.
[0041] According to an exemplary implementation of the present disclosure, image analysis may be performed on the generated image 120 to extract style features 320, texture features 324, color features 322, and geometric features 326, etc., related to the image. Further, a recommended tag may be generated based on any of the aforementioned features. For example, a recommended tag corresponding to a feature different from the currently detected feature may be provided.
[0042] For example, assuming the style of image 120 is a realistic shot style, recommended tags may be provided for adjusting image 120 to a "line style," "oil painting style," etc. For another example, assuming texture feature 324 indicates that the dining table has a wood grain texture, recommended tags may be provided for adjusting the dining table in image 120 to a "glass texture," "tablecloth texture," etc. For another example, assuming color feature 322 indicates that the color of image 120 is a "dark gray tone," recommended tags may be provided for adjusting the color of image 120 to a "bright tone," etc.
[0043] Alternatively and / or additionally, assuming that the geometric feature 326 indicates that the image 120 includes a foreground (e.g., a dining table and a coffee cup on the dining table) and a background (e.g., a wall behind the dining table), a recommendation tag may be provided for modifying the foreground and / or background of the image 120. Specifically, the recommendation tag may prompt the user to adjust the background to a sea background, a starry sky background, or the like.
[0044] Using the exemplary implementation of the present disclosure, one or more candidate recommendation tags can be provided to the user based on a multi-faceted analysis of the prompt word 110 and the image 120. In this way, the complexity of image generation can be simplified, thereby generating an image that better meets the user's expectations in a simpler and more efficient manner.
[0045] According to an example implementation of the present disclosure, a user can select a desired recommended tag and then adjust the content corresponding to the recommended tag in the image 120. Specifically, when the image generation tool receives an interaction (e.g., referred to as a first interaction) with a target recommended tag among at least one recommended tag, it can receive an adjustment request corresponding to the target recommended tag. Furthermore, the image generation tool can provide a new image (e.g., referred to as a second image), and the second image is determined based on the adjustment request and the first image.
[0046] 4 for more details on adjusting an image using recommended tags, which shows a block diagram 400 of an adjusted image based on different recommended tags according to some implementations of the present disclosure. Assuming that the user presses the recommended tag 220-3, the image 120 will be converted to an image 410 with a line style as shown in FIG4. Alternatively and / or additionally, assuming that the user presses the recommended tag 220-1, the image 120 will be converted to an image 420 with a sea background as shown in FIG4. Alternatively and / or additionally, assuming that the user presses the recommended tag 220-2, the image 120 will be converted to an image 430 with a starry sky background as shown in FIG4.
[0047] According to an exemplary implementation of the present disclosure, if the generated image 120 does not meet the user's expectations, the user can simply select the appropriate recommended tag through a simple click to perform subsequent image adjustment tasks. In this way, the user does not need to readjust the prompt word, thus avoiding the problem of introducing more uncontrollable factors into the newly generated image due to prompt word adjustment.
[0048] According to an example implementation of the present disclosure, each recommended tag may correspond to a dedicated image adjustment model, as described in more detail in FIG5 . FIG5 shows a block diagram 500 of conversion models corresponding to multiple recommended tags, respectively, according to some implementations of the present disclosure. As shown in FIG5 , the recommended tag 220-1 for adjusting the background of an image to a sea background may correspond to an adjustment model 510. Here, the adjustment model 510 may be a pre-trained machine learning model for modifying the background of an image. Specifically, the adjustment model 510 may adjust the background of an input image to one or more pre-specified sea backgrounds.
[0049] According to an exemplary implementation of the present disclosure, an original image can be obtained, and the background portion of the original image can be replaced with an image of the sea through manual processing or other methods. Furthermore, the original image and the replaced image can be used as training data. A large amount of training data can be obtained in a similar manner, and the final adjustment model 510 can be obtained through an iterative training process.
[0050] It should be understood that although the above description only illustrates the case where the image adjustment process is directly triggered by clicking a recommended tag, alternatively and / or additionally, more interactions may be provided after the user clicks the recommended tag. For example, the user may specify adjustment requirements via text and / or images, thereby adjusting the image background to the specified content.
[0051] FIG6 illustrates a block diagram 600 of the process for obtaining a user's adjustment request according to some implementations of the present disclosure. As shown in FIG6 , assume that a user presses tab 220-1 to change the background of an image to an ocean background. In this case, various prompts can be provided to the user to obtain the user's adjustment request 612. For example, control 620 can allow the user to enter the adjustment request in text form. For example, the user can enter keywords such as "ocean" or "beach" in text input control 622 to specify the specific content of the background image.
[0052] Alternatively and / or additionally, control 630 may allow the user to specify adjustment requirements in an image format. For example, the user may click on a predefined image 632, etc., to specify the specific content of the background image. Alternatively and / or additionally, a control for specifying a background image may be provided to the user, allowing the user to select a background image from a photo album on the client device or from another remote location on the network.
[0053] It should be understood that the various recommended tags shown in the accompanying figures are merely illustrative, and alternatively and / or additionally, one or more other recommended tags may be provided. For example, if the generated image is of a person, one or more other recommended tags may be provided, such as face swap, hairstyle swap, clothing swap, jewelry swap, etc. If the user selects the face swap tag, a designation control may be provided to allow the user to specify an image to replace the current person's face, and so on.
[0054] By utilizing the exemplary implementation of the present disclosure, a user may be allowed to specify adjustment requirements in a variety of ways. Furthermore, an adjustment model corresponding to a recommended tag selected by the user may output an image that better meets the user's expectations based on the adjustment requirements.
[0055] For example, when the user selects the recommended tag 220-1 (e.g., the target recommended tag), the image generation tool selects the adjustment model 510 that matches the recommended tag 220-1. Further, the adjustment model 510 can be used to generate a new image (e.g., the second image) based on the adjustment requirement and the image 120.
[0056] It should be understood that the adjustment model 510 herein can be obtained based on a predefined training dataset. Compared to large-scale text-to-image machine learning models, an adjustment model for performing a specific adjustment task typically has a simpler structure and fewer parameters. As a result, the workload involved in obtaining the adjustment model and performing the inference process using the adjustment model is typically smaller. Using the example implementations of the present disclosure, by pre-acquiring a dedicated, small-scale adjustment model, images can be adjusted in a more accurate and efficient manner.
[0057] 5 , the recommended tag 220-2 for adjusting the background of an image to a starry sky background may correspond to an adjustment model 520; the recommended tag 220-3 for adjusting the image to a line style may correspond to an adjustment model 530; and the recommended tag 220-4 for adjusting the attributes of the foreground in the image may correspond to an adjustment model 540. The above-mentioned adjustment models may be obtained using a predetermined training dataset in a similar manner.
[0058] According to an exemplary implementation of the present disclosure, after providing the user with an adjusted new image, at least one recommended tag for adjusting the new image may be further provided. For example, at least one recommended tag may be generated based on analysis of the new image in a manner similar to that described above. Alternatively and / or additionally, new recommended tags may be generated based on the recommended tags selected by the user for generating the new image, and so on. In this manner, the user can further adjust the image to generate a final image that meets their needs.
[0059] According to an example implementation of the present disclosure, prompt words 110 for generating an image can be obtained in a variety of ways. For example, a user can enter prompt words through text editing. Alternatively and / or additionally, prompt words can be obtained from an image template. See Figure 7 for more details on obtaining prompt words, which shows a block diagram 700 of the process of obtaining prompt words based on a template according to some implementations of the present disclosure. As shown in Figure 7, an interface 710 of the image generation tool can provide at least one image template indicating the image to be generated. Each image template can have a predetermined prompt word, and the user can select a template to generate a similar image.
[0060] According to an example implementation of the present disclosure, upon receiving an interaction with a target image template among at least one image template (e.g., a second interaction), a prompt word template corresponding to the target image template may be provided. Furthermore, a prompt word may be determined based on receiving an interaction with the prompt word template (e.g., a third interaction). Using this example implementation, users do not need to enter prompt words themselves, but can instead obtain prompt words from a predetermined template, thereby generating images more quickly and efficiently.
[0061] Specifically, template 720 can be used to generate an image related to a coffee theme. The prompt word of template 720 can be, for example, "On the dining table, a cup of coffee..." The user can press control 722 to generate a similar image. Specifically, when the user presses control 722, the prompt word "On the dining table, a cup of coffee..." can be directly input. According to an example implementation of the present disclosure, if the user interactively confirms the prompt word template, the prompt word template can be determined as the prompt word. In other words, if the user confirms, the prompt word can be submitted to the image generation tool, and the corresponding image can be generated.
[0062] For another example, template 730 can be used to generate an image of a cartoon character. The prompt words for template 730 might be, for example, "sweet girl, cartoon style, wireframe..." The user can press control 732 to generate a similar cartoon character. Specifically, when the user presses control 732, they can directly enter the prompt word template "sweet girl, cartoon style, wireframe..." Upon user confirmation, the prompt words can be submitted to the image generation tool, which then generates the corresponding image.
[0063] Alternatively and / or additionally, the user can edit the prompt word template to generate a customized prompt word. For more details, see FIG8 , which shows a block diagram 800 illustrating the process of generating an image based on edited prompt words according to some implementations of the present disclosure. As shown in FIG8 , when the user presses control 732, the prompt words in template 730 can be directly copied to control 820 for inputting prompt words. The user can perform editing operations in control 820, for example, as shown in box 822, to add "bust" to the prompt words. Upon submitting the modified prompt words, the generated image 830 can be a comic-style bust. In this way, a prompt word template can be provided to the user, reducing the complexity of the user's input operations; on the other hand, the user can modify the specific content of the prompt word template to generate an image that better meets their needs.
[0064] According to an example implementation of the present disclosure, the prompt word template can have a structured format. For more details, see FIG9 , which shows a block diagram 900 of a process for editing prompt words according to some implementations of the present disclosure. As shown in FIG9 , in interface 910, a control 920 can be provided to inquire whether the user wants to edit the prompt word of the selected template. If the user presses control 924, an image can be generated directly based on the current prompt word template. Alternatively and / or additionally, if the user presses control 922, an edit control 930 can be further provided.
[0065] In the editing control 930, a prompt word template can be presented in a structured format. In this case, the prompt word template can include at least one descriptive word and at least one value for specifying the image to be generated. As shown in FIG9 , the descriptive words can include, for example, foreground, style, and multiple attributes of the person in the foreground, such as hair, clothing, earrings, etc. The descriptive word and its value can be separated by an "=" sign. For example, "style=manga-style, wireframe" can specify the generation of a manga-style wireframe image.
[0066] It should be understood that FIG9 merely schematically illustrates certain examples of descriptive words in the prompt word template. Alternatively and / or additionally, the descriptive words may further include, but are not limited to, the style, foreground, background, hue, time, location, characters, and the like of the image. For example, a warm-toned image may be generated by setting “hue = warm,” and an image of the night time period may be generated by setting “time = night,” and so on. Furthermore, specific information such as the character's age, hair color, clothing, and the like may be specified. Utilizing the exemplary implementation of the present disclosure, various aspects of the generated image may be explicitly specified, thereby facilitating the image generation tool to more accurately understand the user's detailed needs.
[0067] According to an example implementation of the present disclosure, the prompt word template includes an editable part and an uneditable part. Specifically, box 932 shows the prompt word of the uneditable part. For example, the salient feature part related to the image template in the prompt word template can be specified to be uneditable, while the other parts can be edited. "Sweet girl" in the prompt word template specifies the salient feature of the generated image (that is, the person in the foreground is a "sweet girl"). "Foreground = sweet girl" belongs to the uneditable part and is presented in box 932. For another example, "comic style, wireframe" in the prompt word template specifies the presentation style of the generated image, which is also a salient feature unique to the image template. Therefore, "style = comic style, wireframe" belongs to the uneditable part and is presented in box 932.
[0068] According to an example implementation of the present disclosure, a user can enter additional information in box 934 indicating an editable portion. For example, a user can specify a cartoon character's hair color, top color, earring color, and so on. This example implementation of the present disclosure prevents users from accidentally modifying the core content of the prompt word template, potentially generating images that do not conform to the image template. Furthermore, it allows users to add customized content to the prompt word template to meet their specific needs.
[0069] It should be understood that the above is merely an example. Alternatively and / or additionally, the editable and non-editable portions can be defined in different ways. For example, the descriptive words (e.g., style, background, etc.) in the structured format can be non-editable, while the values of the descriptive words can be editable. In this way, the user can avoid destroying the structured format of the prompt word while editing it, thereby ensuring that the prompt word accurately reflects the user's needs.
[0070] According to an example implementation of the present disclosure, a user can modify the value of a descriptor. Upon receiving an interaction (e.g., a fourth interaction) with the value of a target descriptor in at least one of the descriptors, the prompt word template is updated. For example, assuming the prompt word template includes "background = white," the user can modify the specific value of the background, e.g., "background = gray." In this case, the generated image 940 will have a gray background. In this way, users can adjust the prompt words in a simpler and more efficient manner, thereby generating an image that better meets their needs.
[0071] Using the exemplary implementations of this disclosure, recommended tags can provide users with suggestions for modifying image content, and users can directly click on the tags to adjust the image content without having to re-edit the prompt words. In this way, images that better meet user expectations can be generated in a simpler and more efficient manner.
[0072] Example Process
[0073] FIG10 illustrates a flow chart of a method 1000 for generating an image, according to some implementations of the present disclosure. At block 1010, a prompt word is obtained for specifying an image to be generated. At block 1020, a first image is generated based on the prompt word. At block 1030, the first image and at least one recommended tag for adjusting the first image are provided, wherein the at least one recommended tag is used to adjust at least one of the following attributes of the first image: style, background, and at least one attribute of a foreground object.
[0074] According to an example implementation of the present disclosure, at least one recommended tag is determined based on at least any one of the following: a keyword in the prompt word, a style feature, a color feature, a texture feature, and a geometric feature of the first image.
[0075] According to an example implementation of the present disclosure, the method further includes: in response to receiving a first interaction with a target recommended tag among at least one recommended tag, receiving an adjustment requirement corresponding to the target recommended tag; and providing a second image, the second image being determined based on the adjustment requirement and the first image.
[0076] According to an example implementation of the present disclosure, the second image is determined based on: selecting a machine learning model that matches the target recommendation tag; and generating the second image based on the adjustment requirement and the first image using the machine learning model.
[0077] According to an example implementation of the present disclosure, the adjustment requirement includes at least any one of text and an image, and the method further includes: providing at least one recommended tag for adjusting the second image.
[0078] According to an example implementation of the present disclosure, obtaining a prompt word includes: providing at least one image template indicating an image to be generated; in response to receiving a second interaction for a target image template among the at least one image template, providing a prompt word template corresponding to the target image template; and in response to receiving a third interaction for the prompt word template, determining a prompt word.
[0079] According to an example implementation of the present disclosure, determining the prompt word includes at least any one of the following: in response to determining that the third interaction indication prompt word template is confirmed, determining the prompt word template as the prompt word; and in response to determining that the third interaction indication prompt word template is edited, determining the edited prompt word template as the prompt word.
[0080] According to an example implementation of the present disclosure, the prompt word template includes at least one descriptive word for specifying the image to be generated and the value of at least one descriptive word, and the at least one descriptive word includes at least any one of the following: style, foreground, background, tone, time, place, and character of the image to be generated.
[0081] According to an exemplary implementation of the present disclosure, the prompt word template includes an editable portion and a non-editable portion.
[0082] According to an exemplary implementation of the present disclosure, the method further includes: in response to receiving a fourth interaction for a value of a target description word in the at least one description word, updating the prompt word template.
[0083] Example devices and equipment
[0084] FIG11 shows a block diagram of an apparatus 1100 for generating an image according to some implementations of the present disclosure. The apparatus 1100 includes: an acquisition module 1110 configured to acquire a prompt word for specifying an image to be generated; a generation module 1120 configured to generate a first image based on the prompt word; and a provision module 1130 configured to provide the first image and at least one recommended tag for adjusting the first image, wherein the at least one recommended tag is used to adjust at least one of the following attributes of the first image: style, background, and at least one attribute of a foreground object.
[0085] According to an example implementation of the present disclosure, at least one recommended tag is determined based on at least any one of the following: a keyword in the prompt word, a style feature, a color feature, a texture feature, and a geometric feature of the first image.
[0086] According to an example implementation of the present disclosure, the device further includes: a receiving module, configured to receive an adjustment requirement corresponding to a target recommended tag in response to receiving a first interaction with a target recommended tag among at least one recommended tag; and an image providing module, configured to provide a second image, the second image being determined based on the adjustment requirement and the first image.
[0087] According to an example implementation of the present disclosure, the second image is determined based on: a selection module configured to select a machine learning model that matches the target recommendation tag; and a calling module configured to utilize the machine learning model to generate the second image based on the adjustment requirements and the first image.
[0088] According to an exemplary implementation of the present disclosure, the adjustment requirement includes at least any one of text and an image, and the apparatus further includes: a tag providing module configured to provide at least one recommended tag for adjusting the second image.
[0089] According to an example implementation of the present disclosure, the acquisition module includes: an image template providing module, configured to provide at least one image template indicating an image to be generated; a prompt word template providing module, configured to provide a prompt word template corresponding to the target image template in response to receiving a second interaction for a target image template in at least one image template; and a determination module, configured to determine the prompt word in response to receiving a third interaction for the prompt word template.
[0090] According to an example implementation of the present disclosure, the determination module includes at least any one of the following: a first determination module, configured to determine the prompt word template as the prompt word in response to determining that the third interaction indication prompt word template is confirmed; and a second determination module, configured to determine the edited prompt word template as the prompt word in response to determining that the third interaction indication prompt word template is edited.
[0091] According to an example implementation of the present disclosure, the prompt word template includes at least one descriptive word for specifying the image to be generated and the value of at least one descriptive word, and the at least one descriptive word includes at least any one of the following: style, foreground, background, tone, time, place, and character of the image to be generated.
[0092] According to an exemplary implementation of the present disclosure, the prompt word template includes an editable portion and a non-editable portion.
[0093] According to an exemplary implementation of the present disclosure, the apparatus further includes: an updating module configured to update the prompt word template in response to receiving a fourth interaction for a value of a target description word in the at least one description word.
[0094] FIG12 shows a block diagram of a device 1200 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 1200 shown in FIG12 is merely exemplary and should not be construed as limiting the functionality and scope of the implementations described herein. The computing device 1200 shown in FIG12 can be used to implement the methods described above.
[0095] As shown in FIG12 , computing device 1200 is in the form of a general-purpose computing device. Components of computing device 1200 may include, but are not limited to, one or more processors or processing units 1210, memory 1220, storage device 1230, one or more communication units 1240, one or more input devices 1250, and one or more output devices 1260. Processing unit 1210 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 1220. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of computing device 1200.
[0096] The computing device 1200 typically includes a plurality of computer storage media. Such media can be any available media accessible to the computing device 1200, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 1220 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 1230 can be a removable or non-removable medium and can include a machine-readable medium such as a flash drive, a disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the computing device 1200.
[0097] The computing device 1200 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 12 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a “floppy disk”) and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 1220 may include a computer program product 1225 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.
[0098] The communication unit 1240 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 1200 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 1200 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or other network nodes.
[0099] Input device 1250 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 1260 may be one or more output devices, such as a display, a speaker, or a printer. Computing device 1200 may also communicate with one or more external devices (not shown) via communication unit 1240, as needed, such as storage devices, display devices, or the like, with one or more devices that allow a user to interact with computing device 1200, or with any device that allows computing device 1200 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0100] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.
[0101] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0102] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0103] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0104] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0105] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for generating an image, comprising: Get a prompt word for specifying an image to be generated; generating a first image based on the prompt word; as well as providing the first image and at least one recommended tag for adjusting the first image, The at least one recommended tag is used to adjust at least one of the following items of the first image: style, background, and at least one attribute of an object in the foreground.
2. The method according to claim 1, wherein the at least one recommended tag is determined based on at least any one of the following: keywords in the prompt word, style features, color features, texture features, and geometric features of the first image.
3. The method according to claim 1, further comprising: In response to receiving a first interaction with a target recommended tag among the at least one recommended tag, receiving an adjustment requirement corresponding to the target recommended tag; as well as A second image is provided, the second image being determined based on the adjustment requirement and the first image.
4. The method of claim 3, wherein the second image is determined based on: Selecting a machine learning model that matches the target recommendation label; and The second image is generated based on the adjustment requirements and the first image using the machine learning model.
5. The method according to claim 3, wherein the adjustment requirement includes at least any one of text and image, and the method further comprises: At least one recommended tag is provided for adjusting the second image.
6. The method according to claim 1, wherein obtaining the prompt word comprises: providing at least one image template indicative of an image to be generated; In response to receiving a second interaction with respect to a target image template among the at least one image template, providing a prompt word template corresponding to the target image template; as well as In response to receiving a third interaction with the cue word template, the cue word is determined.
7. The method according to claim 6, wherein determining the prompt word comprises at least one of the following: In response to determining that the third interaction indicates that the prompt word template is confirmed, determining the prompt word template as the prompt word; and In response to determining that the third interaction indicates that the cue word template is edited, determining the edited cue word template as the cue word.
8. The method according to claim 6, wherein the prompt word template includes at least one descriptive word for specifying the image to be generated and the value of the at least one descriptive word, and the at least one descriptive word includes at least any one of the following: style, foreground, background, tone, time, place, and person of the image to be generated.
9. The method according to claim 8, wherein the prompt word template comprises an editable part and a non-editable part.
10. The method according to claim 9, further comprising: In response to receiving a fourth interaction for a value of a target descriptive word in the at least one descriptive word, the prompt word template is updated.
11. A device for generating an image, comprising: An acquisition module configured to acquire a prompt word for specifying an image to be generated; A generating module, configured to generate a first image based on the prompt word; as well as a providing module configured to provide the first image and at least one recommended tag for adjusting the first image, The at least one recommended tag is used to adjust at least one of the following items of the first image: style, background, and at least one attribute of an object in the foreground.
12. An electronic device comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 10 when executed by the at least one processing unit.
13. A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor is caused to implement the method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Image transformation using interpretable transformation parameters
CN113498526A
Image generation method and device, computer equipment and storage medium
CN117112826A
Content generation method and device, computer equipment and storage medium
CN117171369A
Method and device for generating image, equipment and medium
CN117593404A
Image generating device for generating 3D images corresponding to user input sentences and operation method thereof
KR102597074B1
Cited By
Image generation method, image generation device, and computer readable storage medium
CN121600117A