Method and apparatus for generating image, device, and medium
By providing session messages in the image generation tool to guide users to adjust image elements, the problem that image generation in the prior art does not meet user needs is solved, and more efficient and accurate image generation is achieved.
Patent Information
- Application Number
- PCT/CN2024/123080
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-18
- Filing Date
- 2024-09-30
- Publication Date
- 2025-06-26
AI Technical Summary
Existing image generation technology based on text description is difficult to generate images that fully meet user needs, resulting in users needing to frequently adjust prompt words.
By providing a session message, the user is guided to specify the elements to be adjusted and their adjustment results, and the prompt words are updated, thereby generating an image that is more in line with the user's expectations.
The process of user adjustment of prompt words is simplified, and the efficiency and accuracy of generating images that meet user needs is improved.
Smart Images

Figure CN2024123080_26062025_PF_FP_ABST
Abstract
Description
Method, apparatus, device and medium for generating an image
[0001] This application claims priority to the Chinese invention patent application entitled “Methods, devices, apparatus and media for generating images” and application number 2023117448315, filed on December 18, 2023, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] Exemplary implementations of the present disclosure generally relate to image generation, and more particularly to methods, devices, apparatuses, and computer-readable storage media for adjusting prompt words to generate images that better meet user requirements. Background Art
[0003] Machine learning techniques have been widely used to solve visual tasks. Currently, some technical solutions have been proposed for generating images based on text descriptions (e.g., prompt words). However, the generated images may not meet user expectations, requiring users to constantly adjust the text descriptions. Therefore, it is desirable to provide a simple and effective method to guide users to continuously refine the content of the prompt words, thereby generating images that better meet their needs.
[0004] Summary of the Invention
[0005] In a first aspect of the present disclosure, a method for generating an image is provided. In this method, a first conversation message is provided, prompting a user to specify an element to be adjusted in a first image, the first image being generated based on a first prompt word. A second conversation message is received, specifying an adjustment result for the element to be adjusted. A second image generated based on a second prompt word is provided, the second prompt word being obtained by updating the first prompt word using the element to be adjusted and the adjustment result.
[0006] In a second aspect of the present disclosure, a device for generating an image is provided. The device includes: a message providing module configured to provide a first conversation message, the first conversation message prompting a user to specify an element to be adjusted of a first image, the first image being generated based on a first prompt; a receiving module configured to receive a second conversation message, the second conversation message specifying an adjustment result of the element to be adjusted; and an image providing module configured to provide a second image generated based on a second prompt, the second prompt being obtained by updating the first prompt using the element to be adjusted and the adjustment result.
[0007] In a third aspect of the present disclosure, an electronic device is provided. The electronic device includes: at least one processing unit; and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to the first aspect of the present disclosure.
[0008] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the processor implements the method according to the first aspect of the present disclosure.
[0009] It should be understood that the content described in this summary section is not intended to limit the key features or important features of the implementation of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other features, advantages and aspects of various implementations of the present disclosure will become more apparent hereinafter with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein:
[0011] FIG1 shows a block diagram of an application environment according to an exemplary implementation of the present disclosure;
[0012] FIG2 illustrates a block diagram for generating an image according to some implementations of the present disclosure;
[0013] FIG3 illustrates a block diagram of interactions in a generation process according to some implementations of the present disclosure;
[0014] FIG4 shows a block diagram of an image generated based on an adjusted prompt word according to some implementations of the present disclosure;
[0015] FIG5 is a block diagram illustrating a process of adjusting prompt words according to some implementations of the present disclosure;
[0016] FIG6 illustrates a block diagram of a process for specifying adjustment results according to some implementations of the present disclosure;
[0017] FIG7 illustrates a block diagram of a process for adjusting prompt words according to some implementations of the present disclosure;
[0018] FIG8 shows a block diagram of a process for determining an original prompt word according to some implementations of the present disclosure;
[0019] FIG9 shows a flowchart of a method for generating an image according to some implementations of the present disclosure;
[0020] FIG10 shows a block diagram of an apparatus for generating an image according to some implementations of the present disclosure; and
[0021] FIG11 shows a block diagram of a device capable of implementing various implementations of the present disclosure. DETAILED DESCRIPTION
[0022] The following describes implementations of the present disclosure in more detail with reference to the accompanying drawings. Although certain implementations of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the implementations described herein. Rather, these implementations are provided to provide a more thorough and complete understanding of the present disclosure. It should be understood that the drawings and implementations of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0023] In the description of the implementation of the present disclosure, the term "including" and similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "based at least in part on". The term "an implementation" or "the implementation" should be understood as "at least one implementation". The term "some implementations" should be understood as "at least some implementations". The following may also include other explicit and implicit definitions. As used herein, the term "model" can represent the association relationship between various data. For example, the above-mentioned association relationship can be obtained based on a variety of technical solutions currently known and / or to be developed in the future.
[0024] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.
[0025] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0026] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0027] As an optional but non-limiting implementation, in response to receiving a user's active request, a prompt message may be sent to the user, for example, in the form of a pop-up window, in which the prompt message may be presented in text form. Furthermore, the pop-up window may also include a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0028] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0029] As used herein, the term "in response to" refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of executing a subsequent action executed in response to the event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is satisfied. For example, in some cases, a subsequent action may be executed immediately upon the occurrence of the event or the satisfaction of the condition; in other cases, the subsequent action may be executed some time after the occurrence of the event or the satisfaction of the condition.
[0030] Sample Environment
[0031] Machine learning technology has been widely used in visual tasks. For example, a variety of machine learning models have been proposed for generating images based on text. A user can use text to describe the image they wish to generate. FIG1 shows a block diagram 100 of an application environment according to an exemplary implementation of the present disclosure. As shown in FIG1 , in an interface 130 of an image generation tool, a user can enter a prompt word 110, and the image generation tool will then return an image 120 generated based on the prompt word 110.
[0032] Specifically, the prompt 110 may include "dining table, coffee cup." The generated image 120 will then include the content specified by the prompt, for example, a cup of coffee placed on the dining table. However, the initial prompt may not fully describe the user's needs, which can lead to the image generated by the machine learning model sometimes not meeting the user's requirements. For example, the user may not be satisfied with the background of image 120 or may wish to generate an image in a different style. In this case, the user has to manually adjust the text of the prompt.
[0033] However, during the adjustment process, the user may not have thought about which elements in the image need to be adjusted, which may cause the user to repeatedly rewrite the prompt words but still fail to obtain the desired image. In this case, it is hoped that the user can be guided to find the content they want to adjust, and then gradually guide the user to obtain a satisfactory image.
[0034] Overview of Image Generation
[0035] In order to at least partially address the deficiencies in the prior art, according to an exemplary implementation of the present disclosure, a method for generating an image is proposed. Referring to FIG. 2 , an overview of an exemplary implementation of the present disclosure is described, which shows a block diagram 200 for generating an image according to some implementations of the present disclosure. As shown in FIG. 2 , in the interface 240 of the image generation tool, a prompt word 110 (e.g., referred to as a first prompt word) for specifying the image to be generated can be received, and a first image (e.g., image 120) can be generated based on the prompt word 110. Further, the image 120 can be provided, and a first conversation message 210 can be provided near the image 120 (e.g., immediately below the image 120). For example, the user can be asked, "Which image elements need to be adjusted?", thereby prompting the user to specify the elements to be adjusted of the first image.
[0036] After receiving the first conversation message 210, the user may enter a second conversation message 220, which may specify the adjustment result of the element to be adjusted. For example, in the example of FIG2 , the user may specify to change the style of the image to a line style. The image generation tool may then receive the second conversation message 220 and provide an image 230 (e.g., referred to as a second image).
[0037] It should be understood that the second image here is generated based on the second prompt word, and the second prompt word is obtained by updating the first prompt word using the element to be adjusted and the adjustment result. Specifically, the image generation tool can extract the adjustment result of the element to be adjusted from the received second conversation message 220. In the example of Figure 2, the element to be adjusted is the "style of the image", and the "adjustment result" is the "line style". The extracted information can be used to modify the first prompt word to obtain the second prompt word "dining table, coffee cup, line style". Then, a new image generated using the second prompt word can be provided. At this time, the user does not have to rewrite the prompt word, but can obtain the adjusted prompt word and the image drawn in the line style through a simple dialogue with the image generation tool.
[0038] Using the exemplary implementation of this disclosure, the conversational message can gradually guide the user to find the image element they wish to adjust, then specify the adjustment result for the element to be adjusted, thereby modifying the prompt word. In this way, the user can be gradually guided to refine the prompt word, thereby generating an image that better meets the user's expectations.
[0039] Detailed process of image generation
[0040] An overview of the image generation process has been described with reference to FIG2 . Further details will be described below with reference to other figures. According to an exemplary implementation of the present disclosure, the first conversation message can be provided in a variety of ways. For example, the first conversation message can be provided immediately after the first image is generated, inquiring about which elements of the generated image the user is dissatisfied with. Alternatively and / or additionally, the first conversation message can be provided upon receiving negative feedback regarding the first image. In this way, excessive conversation messages can be prevented from disrupting the user's normal use.
[0041] For further details, see FIG3 , which illustrates a block diagram 300 of interactions during a generation process according to some implementations of the present disclosure. As shown in interface 350 of FIG3 , the image generation tool can receive prompts 110 from a user and generate an image 120 according to the prompts 110. The user may be dissatisfied with the image 120 and enter negative feedback 310, indicating that the image 120 is not the style they desired. Upon receiving negative feedback 310 regarding the generated image 120, the image generation tool may provide a first conversation message 320 to prompt the user to enter further requirements.
[0042] It should be understood that the first conversation message 320 in this example can be generated based on a semantic analysis of the negative feedback 310. For example, keywords in the negative feedback can be extracted to identify potential image elements to be adjusted. If the negative feedback indicates that the element to be adjusted is "style," a subsequent conversation message can ask the user to further specify the adjustment results for "style," for example, "Please specify the style of the image." At this point, the user can enter a second conversation message 330 and specify that an image with "line style" be output. At this point, the image generation tool can provide an image 340 generated using the modified prompt word (i.e., "dining table, coffee cup, line style").
[0043] According to an example implementation of the present disclosure, if the negative feedback is a general statement such as "not good" but does not specify the specific content of the element to be adjusted, the subsequent conversation message can require the user to further specify the element to be adjusted and the expected adjustment result. It should be understood that Figure 3 only schematically illustrates an example of obtaining the element to be adjusted and the expected adjustment result in a round of conversation. Alternatively and / or additionally, the user's needs can be gradually refined in one or more rounds of conversation to obtain more accurate adjustment requirements.
[0044] It should be understood that FIG3 merely illustrates the specific process of modifying the prompt word, using adjustment of image style as an example of the element to be adjusted. Alternatively and / or additionally, the element to be adjusted may include at least one of the following: the style of the first image, the background, and at least one attribute of an object in the foreground. Specifically, image analysis may be performed on the generated image 120 to extract various aspects of the image, thereby providing one or more candidates for the element to be adjusted. For example, image recognition processing may be performed on image 120 to identify the image style, background, and various objects in the foreground, etc.
[0045] For example, the image generation tool may provide the following conversation message to ask the user which elements he wants to adjust: Please select the elements you want to adjust (multiple selections are allowed): style, background, coffee cup, plate, table. The user can select one or more elements, at which point a subsequent conversation message specifying the adjustment results of the selected elements may be presented. Assuming the user selects "style", the image generation tool may provide a further conversation message: Please select the desired style: line style, oil painting style, black and white style, and so on. Assuming the user selects "oil painting style", an image in the oil painting style may be generated. Utilizing the example implementation of the present disclosure, the user's needs may be gradually refined based on a multi-faceted analysis of the image 120, thereby generating appropriate prompt words, thereby generating an image that better meets the user's expectations in a simpler and more effective manner.
[0046] FIG4 shows a block diagram 400 of an image generated based on adjusted prompt words according to some implementations of the present disclosure. Assuming a user specifies to modify the image background to the ocean, image 410 may be provided. Image 410 is generated based on the adjusted prompt words "dining table, coffee cup, ocean background." Assuming a user specifies to modify the image background to the starry sky, image 420 may be provided. Image 420 is generated based on the adjusted prompt words "dining table, coffee cup, starry sky background." In this way, the user does not have to manually re-enter the prompt words, but can modify the prompt words through a simple dialogue with the image generation tool, thereby obtaining an image that better meets the user's needs.
[0047] According to an example implementation of the present disclosure, a page for editing a prompt word can be provided to the user. For example, a second prompt word after adjustment can be presented, and the difference between the second prompt word and the first prompt word can be shown. In this way, it is convenient for the user to adjust the prompt word in a visual manner. For more details, see Figure 5, which shows a block diagram 500 of the process of adjusting the prompt word according to some implementations of the present disclosure.
[0048] As shown in FIG5 , the image generation tool may provide an interface 540 , which may include a dialog message 510 , inquiring the user whether further editing of the prompt word is desired. As shown in block 516 , the difference between the prompt word before and after adjustment may be highlighted using different fonts, colors, font sizes, etc., so that the user can confirm whether the adjusted prompt word meets their needs. Furthermore, a control 512 may be provided to initiate further editing; alternatively and / or additionally, a control 514 may be provided to directly submit the adjusted prompt word.
[0049] According to an exemplary implementation of the present disclosure, upon receiving a confirmation operation for the adjusted second prompt word (e.g., upon detecting that the user has pressed control 514), a second image may be provided. Alternatively and / or additionally, upon detecting that the user has pressed control 512, the current prompt word may be directly copied to conversation message 520 and used as the second prompt word.
[0050] The user can modify the conversation message 520 to generate a new prompt word (e.g., a third prompt word). For example, as shown in box 522, the user can add "background changed to the sea." In this case, the third prompt word may include: "dining table, coffee cup, line style, background changed to the sea." The image generation tool can generate and provide an image 530 based on the third prompt word. Using the example implementation of the present disclosure, the user can continuously refine and adjust the prompt word according to their needs, thereby obtaining an image that better meets their needs.
[0051] According to an exemplary implementation of the present disclosure, after providing a new image to a user, the user may be further asked which image elements need to be modified. Specifically, after providing a second image, a third conversation message may be provided, prompting the user to specify another element of the second image to be adjusted. A fourth conversation message may be received, specifying another adjustment result for the other element to be adjusted. A fourth image generated based on a fourth prompt word may be provided, the fourth prompt word being obtained by updating the second prompt word using the other element to be adjusted and the other adjustment result.
[0052] Continuing with the example above, the user can specify to remove the tray underneath the coffee cup. A new prompt will be generated: "Dining table, coffee cup, line style, no tray," and the image generation tool will generate a new image. This allows the user to further adjust unsatisfactory aspects of the image and create a final image that meets their needs.
[0053] It should be understood that the above only describes the process for specifying the adjustment result of a certain element to be adjusted using text as an example. Alternatively and / or additionally, the adjustment result can be specified in the form of text and / or images. See Figure 6 for more details, which shows a block diagram 600 of the process for specifying the adjustment result according to some implementations of the present disclosure. As shown in Figure 6, assuming that the user wants to adjust the background of the image, a variety of prompts can be provided to the user to obtain the user-specified adjustment result 610. For example, the control 620 can allow the user to specify the adjustment result in text. For example, the user can enter keywords such as "change the background to the sea" in the text input control 622 to specify the specific content of the background image.
[0054] Alternatively and / or additionally, control 630 may allow the user to specify the adjustment result in an image format. For example, the user may click on a predefined image 632, etc., to specify the specific content of the background image. Alternatively and / or additionally, a control for specifying a background image may be provided to the user, allowing the user to select a background image from a photo album on the client device or from another remote location on the network.
[0055] It should be understood that the various elements to be adjusted described above are merely illustrative; alternatively and / or additionally, the user may specify other elements to be adjusted. For example, assuming the generated image is of a person, the user may specify at least one of the following: face replacement, hairstyle replacement, clothing replacement, jewelry replacement, etc. Assuming the user desires a face replacement, a designated control may be provided to allow the user to specify an image to replace the current person's face, and so on. Using the example implementations of this disclosure, the user may be allowed to specify adjustment results in a variety of ways, and the specified adjustment results may be used to modify the prompt word, thereby outputting an image that better meets the user's expectations.
[0056] According to an example implementation of the present disclosure, semantic analysis can be performed on prompt words, and the prompt words can be presented in a structured format. See FIG7 for more details, which shows a block diagram 700 of a process for adjusting prompt words according to some implementations of the present disclosure. As shown in FIG7 , interface 730 may include an editing control 710, in which the prompt word is split into two parts: a description word, which is used to specify an element in the image to be generated; and the value of the description word. A prompt word can include one or more description words. For example, the prompt word in FIG7 includes three description words: foreground, style, and background.
[0057] Controls 712, 714, and 716 may be provided to specify the values of the respective descriptive words. For example, the text "dining table, coffee cup" may be entered in control 712 to specify the foreground of the image, the text "realistic style" may be entered in control 714 to specify the style of the image, and an image may be entered in control 716 to specify the background of the image, and so on.
[0058] It should be understood that FIG7 merely schematically illustrates some examples of descriptive words. Alternatively and / or additionally, the descriptive words may further include, but are not limited to, the style, foreground, background, hue, time, location, character, etc. of the image. For example, a warm-toned image can be generated by setting "hue = warm color," and an image of the night time period can be generated by setting "time = night," and so on. Furthermore, if it is desired to generate a character image, the descriptive words may further specify specific information such as the character's age, hair color, clothing, etc. By utilizing the example implementation of the present disclosure, various aspects of the generated image can be explicitly specified, thereby facilitating the refinement of the prompt words and enabling the image generation tool to more accurately understand the user's refined needs.
[0059] The above describes the process of adjusting the first prompt word input by the user to generate a second prompt word that better meets the user's needs. In the context of this disclosure, the method for obtaining the original first prompt word is not limited. For example, the user can directly input the first prompt word. Alternatively and / or additionally, the user can select an image template to use the prompt word used to generate the image template as the first prompt word.
[0060] See FIG8 for more details on obtaining prompt words, which shows a block diagram 800 of the process for determining the original prompt word according to some implementations of the present disclosure. As shown in FIG8 , an interface 810 of the image generation tool can provide at least one image template indicating the image to be generated. Each image template can have a predetermined prompt word, and the user can select a template to generate a similar image. Specifically, template 820 can be used to generate an image related to a coffee theme. The prompt word of template 820 can be, for example, "On the dining table, a cup of coffee..." The user can press control 822 to generate a similar image. Specifically, when the user presses control 822, the user can directly enter the prompt word "On the dining table, a cup of coffee..." According to an example implementation of the present disclosure, if the user interactively confirms the prompt word template, the prompt word template can be determined as the first prompt word. In other words, if the user confirms, the first prompt word can be submitted to the image generation tool, and the corresponding image is generated.
[0061] As shown in FIG8 , template 830 can be used to generate an image related to a comic character. The prompt words of template 830 can be, for example, “sweet girl, comic style, wireframe…” The user can press control 832 to generate a similar image. Specifically, when the user presses control 832, the prompt words “sweet girl, comic style, wireframe…” can be directly input. Alternatively and / or additionally, the user can perform further editing to use the edited prompt word template as the first prompt word. For example, the user can modify the prompt words in the template to “sweet girl, comic style, gray background” to generate the corresponding image.
[0062] Using the example implementations of this disclosure, users do not need to enter prompt words themselves. Instead, they can obtain prompt words from a predefined template, thereby generating images more quickly and efficiently. After the user enters a first prompt word using the template, a corresponding first image can be generated, and a first conversation message can be provided after the first image. This first conversation message can prompt the user to specify the elements to be adjusted in the first image, and then the prompt word can be gradually refined in the manner described above.
[0063] Using the exemplary implementation of this disclosure, the conversational message can gradually guide the user to find the image element they wish to adjust, then specify the adjustment result for the element to be adjusted, thereby modifying the prompt word. In this way, the user can be gradually guided to refine the prompt word, thereby generating an image that better meets the user's expectations.
[0064] Example Process
[0065] FIG9 illustrates a flow chart of a method 900 for generating an image, according to some implementations of the present disclosure. At block 910, a first session message is provided, prompting a user to specify an element to be adjusted in a first image generated based on a first prompt. At block 920, a second session message is received, specifying an adjustment result for the element to be adjusted. At block 930, a second image generated based on a second prompt is provided, the second prompt being obtained by updating the first prompt using the element to be adjusted and the adjustment result.
[0066] According to an example implementation of the present disclosure, providing the first conversation message includes: providing the first conversation message in response to receiving negative feedback on the first image.
[0067] According to an exemplary implementation of the present disclosure, the element to be adjusted includes at least any one of the following: the style of the first image, the background, and at least one attribute of an object in the foreground.
[0068] According to an example implementation of the present disclosure, at least one property of the object is obtained based on recognizing the first image.
[0069] According to an exemplary implementation of the present disclosure, the method further includes: providing a second prompt word, wherein the difference between the second prompt word and the first prompt word is highlighted.
[0070] According to an exemplary implementation of the present disclosure, providing the second image includes: providing the second image in response to receiving a confirmation operation for the second prompt word.
[0071] According to an example implementation of the present disclosure, the method further includes: in response to receiving an update operation for the second prompt word, updating the second prompt word to generate a third prompt word; and providing a third image generated based on the third prompt word.
[0072] According to an example implementation of the present disclosure, the method further includes at least any one of the following: acquiring a first prompt word in response to user input; and using the prompt word used to generate the image template as the first prompt word in response to the user selecting the image template.
[0073] According to an example implementation of the present disclosure, the method further includes: providing a third conversation message, the third conversation message being used to prompt the user to specify another element to be adjusted of the second image; receiving a fourth conversation message, the fourth conversation message being used to specify another adjustment result of the other element to be adjusted; and providing a fourth image generated based on a fourth prompt word, the fourth prompt word being obtained by updating the second prompt word using the other element to be adjusted and the other adjustment result.
[0074] Example devices and equipment
[0075] Figure 10 shows a block diagram of an apparatus 1000 for generating an image according to some implementations of the present disclosure. The apparatus 1000 includes: a message providing module 1010 configured to provide a first conversation message, the first conversation message prompting a user to specify an element to be adjusted of a first image generated based on a first prompt word; a receiving module 1020 configured to receive a second conversation message, the second conversation message specifying an adjustment result of the element to be adjusted; and an image providing module 1030 configured to provide a second image generated based on a second prompt word, the second prompt word being obtained by updating the first prompt word using the element to be adjusted and the adjustment result.
[0076] According to an exemplary implementation of the present disclosure, the message providing module includes: a feedback-based providing module configured to provide a first conversation message in response to receiving negative feedback on a first image.
[0077] According to an exemplary implementation of the present disclosure, the element to be adjusted includes at least any one of the following: the style of the first image, the background, and at least one attribute of an object in the foreground.
[0078] According to an example implementation of the present disclosure, at least one property of the object is obtained based on recognizing the first image.
[0079] According to an exemplary implementation of the present disclosure, the apparatus further includes: a prompt word providing module configured to provide a second prompt word, wherein the difference between the second prompt word and the first prompt word is highlighted.
[0080] According to an exemplary implementation of the present disclosure, the image providing module includes: a confirmation-based providing module configured to provide a second image in response to receiving a confirmation operation for a second prompt word.
[0081] According to an example implementation of the present disclosure, the apparatus further includes: an updating module configured to, in response to receiving an update operation for the second prompt word, update the second prompt word to generate a third prompt word; and the image providing module is further configured to: provide a third image generated based on the third prompt word.
[0082] According to an example implementation of the present disclosure, the device further includes at least any one of the following: a first acquisition module, configured to acquire a first prompt word in response to a user's input; and a second acquisition module, configured to use the prompt word used to generate the image template as the first prompt word in response to the user selecting the image template.
[0083] According to an example implementation of the present disclosure, the message providing module is further configured to provide a third conversation message, which is used to prompt the user to specify another element to be adjusted of the second image; the receiving module is further configured to receive a fourth conversation message, which specifies another adjustment result of the other element to be adjusted; and the image providing module is further configured to provide a fourth image generated based on a fourth prompt word, which is obtained by updating the second prompt word using another element to be adjusted and another adjustment result.
[0084] FIG11 shows a block diagram of a device 1100 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 1100 shown in FIG11 is merely exemplary and should not be construed as limiting the functionality and scope of the implementations described herein. The computing device 1100 shown in FIG11 can be used to implement the methods described above.
[0085] As shown in FIG11 , computing device 1100 is in the form of a general-purpose computing device. Components of computing device 1100 may include, but are not limited to, one or more processors or processing units 1110, memory 1120, storage device 1130, one or more communication units 1140, one or more input devices 1150, and one or more output devices 1160. Processing unit 1110 may be a real or virtual processor and is capable of performing various processes according to a program stored in memory 1120. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to increase the parallel processing capabilities of computing device 1100.
[0086] The computing device 1100 typically includes a plurality of computer storage media. Such media can be any available media accessible to the computing device 1100, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 1120 can be a volatile memory (e.g., registers, cache, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 1130 can be a removable or non-removable medium and can include a machine-readable medium, such as a flash drive, a disk, or any other medium that can be used to store information and / or data (e.g., training data for training) and can be accessed within the computing device 1100.
[0087] The computing device 1100 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 11 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a "floppy disk") and an optical drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 1120 may include a computer program product 1125 having one or more program modules configured to perform various methods or actions of various implementations of the present disclosure.
[0088] The communication unit 1140 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 1100 can be implemented as a single computing cluster or multiple computing machines that can communicate via a communication connection. Thus, the computing device 1100 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or other network nodes.
[0089] Input device 1150 may be one or more input devices, such as a mouse, keyboard, or trackball. Output device 1160 may be one or more output devices, such as a display, a speaker, or a printer. Computing device 1100 may also communicate with one or more external devices (not shown) via communication unit 1140, as needed, such as storage devices, display devices, or the like, with one or more devices that allow a user to interact with computing device 1100, or with any device that allows computing device 1100 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0090] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, wherein the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.
[0091] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products implemented according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0092] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, such that when these instructions are executed by the processing unit of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0093] Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0094] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple implementations of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part for a module, program segment or instruction, and a part for a module, program segment or instruction comprises one or more executable instructions for realizing the logical function of the specification. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be realized by a special hardware-based system that performs the function or action of the specification, or can be realized by a combination of special hardware and computer instructions.
[0095] While various implementations of the present disclosure have been described above, the foregoing description is intended to be illustrative, not exhaustive, and not limited to the disclosed implementations. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described implementations. The terminology used herein is selected to best explain the principles of the implementations, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the various implementations disclosed herein.
Claims
1. A method for generating an image, comprising: Providing a first conversation message, wherein the first conversation message is used to prompt a user to specify an element to be adjusted of a first image, wherein the first image is generated based on a first prompt word; receiving a second session message, wherein the second session message specifies an adjustment result of the element to be adjusted; as well as A second image generated based on a second prompt word is provided, where the second prompt word is obtained by updating the first prompt word by using the element to be adjusted and the adjustment result.
2. The method of claim 1 , wherein providing the first session message comprises: In response to receiving negative feedback for the first image, the first conversation message is provided. 3 . The method according to claim 1 , wherein the element to be adjusted comprises at least any one of the following: style of the first image, background, and at least one attribute of an object in the foreground. The method of claim 3 , wherein the at least one property of the object is obtained based on recognizing the first image.
5. The method according to claim 4, further comprising: The second prompt word is provided, and the difference between the second prompt word and the first prompt word is highlighted.
6. The method of claim 5, wherein providing the second image comprises: In response to receiving a confirmation operation for the second prompt word, providing the second image.
7. The method according to claim 4, further comprising: In response to receiving an update operation for the second prompt word, updating the second prompt word to generate a third prompt word; as well as A third image generated based on the third prompt word is provided.
8. The method according to claim 1, further comprising at least one of the following: Responding to the user's input, acquiring the first prompt word; and In response to the user selecting an image template, a prompt word used to generate the image template is used as the first prompt word.
9. The method according to claim 1, further comprising: providing a third session message, wherein the third session message is used to prompt the user to specify another element of the second image to be adjusted; receiving a fourth session message, wherein the fourth session message specifies another adjustment result of the another element to be adjusted; as well as A fourth image generated based on a fourth prompt word is provided, where the fourth prompt word is obtained by updating the second prompt word by using the another element to be adjusted and the another adjustment result.
10. A device for generating an image, comprising: A message providing module is configured to provide a first conversation message, wherein the first conversation message is used to prompt a user to specify an element to be adjusted of a first image, wherein the first image is generated based on a first prompt word; A receiving module, configured to receive a second session message, wherein the second session message specifies an adjustment result of the element to be adjusted; as well as The image providing module is configured to provide a second image generated based on a second prompt word, where the second prompt word is obtained by updating the first prompt word by using the element to be adjusted and the adjustment result.
11. An electronic device, comprising: at least one processing unit; as well as At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to any one of claims 1 to 9 when executed by the at least one processing unit. 12 . A computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the processor is caused to implement the method according to claim 1 .
Citation Information
Patent Citations
Image generation method and device, equipment, medium and program
CN116824020A
Image generation method and device, electronic equipment and storage medium
CN116843795A
Method and device for generating image, equipment and medium
CN117671067A
Generating ground truth annotations corresponding to digital image editing dialogues for training state tracking models
US20200312298A1