Electronic device, method, and non-transitory computer-readable storage medium for generating image by using trained model

By enabling user interaction to modify visual objects generated from handwritten input, the electronic device addresses the mismatch between user intention and display, improving user satisfaction.

WO2026155315A1PCT designated stage Publication Date: 2026-07-23SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2025-09-26
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing electronic devices using trained models to generate images from handwritten input often display visual objects differently from the user's intention, causing user discomfort.

Method used

The electronic device allows users to input handwritten strokes, processes them through a trained model to generate a second image, and enables user interaction to modify parts of the visual object displayed, providing options for changing these parts based on user input.

Benefits of technology

This approach addresses user discomfort by allowing users to adjust the displayed visual objects to align with their intentions, enhancing user satisfaction and control over the generated image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025015142_23072026_PF_FP_ABST
    Figure KR2025015142_23072026_PF_FP_ABST
Patent Text Reader

Abstract

This electronic device may comprise: at least one processor including a processing circuit; a display; and a memory for storing one or more programs configured to be individually or collectively executed by the at least one processor, and including one or more storage media, wherein the one or more programs cause the electronic device to: identify an input for obtaining a second image by using a first image; obtain information on the first image by providing the first image to a trained model on the basis of the input; obtain a keyword included in the information and options related to the keyword; obtain the second image by providing a prompt including the information to the trained model; display, through the display, UI objects indicating the second image and the options, respectively; and display, through the display, a third image on the basis of a user input for at least one UI object among the UI objects.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device, method, and non-transient computer-readable storage medium for generating images using a trained model

[0001] The present disclosure relates to an electronic device, a method, and a non-transient computer-readable storage medium for generating an image using a trained model.

[0002] Artificial intelligence can simulate neural activities of a person (or organism), such as perception and / or inference, and can be implemented by hardware, software, or a combination of said hardware and said software designed to perform calculations for simulating neural activities.

[0003] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.

[0004] An electronic device is described. The electronic device may include at least one processor comprising a processing circuit, a display, and a memory comprising one or more storage media configured to store one or more programs configured to be executed individually or collectively by the at least one processor. The one or more programs may include instructions that cause the electronic device to identify an input for acquiring a second image using a first image. The one or more programs may include instructions that cause the electronic device to provide the first image to a trained model based on the input. The one or more programs may include instructions that cause the electronic device to acquire information about the first image generated by the trained model using the first image. The one or more programs may include instructions that cause the electronic device to acquire a keyword included in the information and options regarding the keyword. The above one or more programs may include instructions that cause the electronic device to provide the trained model with a prompt containing the information. The above one or more programs may include instructions that cause the electronic device to acquire the second image, which is generated by the trained model using the prompt and includes a visual object representing the keyword. The above one or more programs may include instructions that cause the electronic device to display user interface (UI) objects indicating the second image and the options, respectively, through the display.The above one or more programs may include instructions that cause the electronic device to receive user input for at least one UI object among the UI objects. The above one or more programs may include instructions that cause the electronic device to display, based on the user input, a third image including another visual object expressing a keyword having an option indicated by the at least one UI object through the display.

[0005] A method is described. The method may be performed within an electronic device including a display. The method may include an operation of identifying an input for acquiring a second image using a first image. The method may include an operation of providing the first image to a trained model based on the input. The method may include an operation of acquiring information about the first image generated by the trained model using the first image. The method may include an operation of acquiring a keyword and options regarding the keyword included in the information. The method may include an operation of providing a prompt containing the information to the trained model. The method may include an operation of acquiring the second image, which is generated by the trained model using the prompt and includes a visual object representing the keyword. The method may include an operation of displaying user interface (UI) objects that indicate the second image and the options, respectively, through the display. The method may include an operation of receiving user input for at least one of the UI objects. The above method may include the operation of displaying a third image, which includes another visual object expressing a keyword having an option indicated by at least one UI object, through the display based on the user input.

[0006] A non-transient computer-readable storage medium is described. The non-transient computer-readable storage medium may store one or more programs. The one or more programs may include instructions that cause the electronic device to identify an input for acquiring a second image using a first image. The one or more programs may include instructions that cause the electronic device to provide the first image to a trained model based on the input when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to acquire information about the first image generated by the trained model using the first image when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to acquire a keyword and options regarding the keyword when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to provide the trained model with a prompt containing the information when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to acquire the second image, which is generated by the trained model using the prompt and includes a visual object representing the keyword, when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to display user interface (UI) objects indicating the second image and the options, respectively, through the display when executed by the electronic device.The above one or more programs may include instructions that cause the electronic device to receive user input for at least one UI object among the UI objects when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to display, through the display, a third image including another visual object expressing a keyword having an option indicated by the at least one UI object, based on the user input when executed by the electronic device.

[0007] The above and other aspects, features, and advantages of specific embodiments of the present disclosure will become more apparent from the following description in conjunction with the accompanying drawings. In the drawings:

[0008] Figure 1 illustrates an example of an image displayed based on handwriting input.

[0009] Figure 2 is a block diagram of an exemplary electronic device.

[0010] FIG. 3a is a flowchart illustrating exemplary operations of an electronic device for acquiring information about a first image.

[0011] FIG. 3b illustrates an example of a user interface (UI) that receives input for acquiring a second image using a first image.

[0012] FIG. 4 is a flowchart illustrating exemplary operations of an electronic device for displaying a second image and UI (user interface) objects.

[0013] Figure 5 illustrates an example of user input for the style of the second image.

[0014] FIGS. 6a and FIGS. 6b illustrate examples of user input for displaying UI objects that indicate options, respectively.

[0015] FIG. 6c illustrates an example of text input for changing parts of an object.

[0016] FIGS. 7a and 7b illustrate examples of user input for displaying a UI object that indicates additional options.

[0017] FIG. 8 is a flowchart illustrating exemplary operations of an electronic device for displaying a third image.

[0018] FIG. 9a illustrates an example of displaying a third image.

[0019] FIG. 9b illustrates an example of a third image obtained based on text input.

[0020] FIGS. 10a, FIGS. 10b, and FIGS. 10c illustrate examples of text input and image input received along with handwritten input.

[0021] FIG. 10d illustrates an example of displaying a third image within another application.

[0022] FIG. 11 is a block diagram of an electronic device in a network environment according to various embodiments.

[0023] FIG. 12 illustrates an example of a generative artificial intelligence system according to one embodiment.

[0024] Hereinafter, embodiments of the present disclosure are described in detail with reference to the drawings so that those skilled in the art can easily practice them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and brevity.

[0025] Figure 1 illustrates an example of an image displayed based on handwriting input.

[0026] Referring to FIG. 1, the electronic device (100) may be a device available for receiving handwriting input (e.g., strokes) or may be capable of corresponding. For example, the electronic device (100) may be one of various types of mobile devices such as smartphones having various form factors (e.g., bar-type smartphones, foldable-type smartphones, or rollable-type smartphones), tablets, wearable devices, cellular phones, personal computers (PCs) (e.g., laptops or desktops), and / or other similar computing devices that include circuits (or circuitry) for providing the operation of receiving handwriting input.

[0027] For example, the electronic device (100) may include a display (110) (e.g., the display (230) of FIG. 2). In a first state (105), the electronic device (100) may receive handwritten input through the display (110). For example, the electronic device (100) may receive handwritten input based on a pointing device such as a fingertip, stylus, digitizer, and / or a mouse for controlling the position of a cursor that is in contact with the display (110). The electronic device (100) may identify at least one stroke (115) based on the handwritten input. At least one processor (210) may display a first image (120) containing at least one stroke (115) through the display (230).

[0028] The electronic device (100) can obtain information about the first image (120) by providing the first image (120) to a trained model. For example, the trained model may include the generative AI (artificial intelligence) model (1230) of FIG. 12. For example, at least one processor (210) may generate a prompt containing information about the first image (120). For example, at least one processor (210) may obtain the second image (130) by providing the training model with a prompt containing information about the first image (120). The prompt may include text requesting the generation of the second image (130).

[0029] Based on acquiring a second image (130), the electronic device (100) may transition from a first state (105) to a second state (125). In the second state (125), at least one processor (210) may display the second image (130) through a display (230). The second image (130) may include a visual object (135) that represents or corresponds to at least one stroke (115) contained in the first image (120). The visual object (135) may include parts (140) of the visual object (135). For example, parts (140) of the visual object (135) may be features of or correspond to the visual object (135). In the second image (130), parts (140) of the visual object (135) may be represented differently from the intention of the user who provided the handwriting input. As parts (140) of the visual object (135) within the second image (130) are displayed differently from the user's intention, the user may feel discomfort. A solution may be required to resolve the user's discomfort caused by parts (140) of the visual object (135) that are displayed differently from the user's intention.

[0030] To address or resolve this inconvenience, the electronic device (100) may receive user input to change at least one part (140) of the parts (135) of the visual object (135). Based on the user input, the electronic device (100) may display a third image through the display (230) in which at least one part (140) of the visual object (135) has been changed. Information regarding the first image (120) may be used to determine the parts (140) of the visual object (135) and options regarding the parts (140) of the visual object (135). The electronic device (100) may perform operations exemplified within the description of FIGS. 3a through 10b to display the third image in which at least one part (140) of the visual object (135) has been changed. The electronic device (100) may include components for performing said operations. said components may be exemplified in FIG. 2.

[0031] Figure 2 is a block diagram of an exemplary electronic device.

[0032] Referring to FIG. 2, the electronic device (200) may be one of various types of mobile devices, such as smartphones having various form factors (e.g., bar-type smartphones, foldable-type smartphones, or rollable-type smartphones), tablets, wearable devices, cellular phones, personal computers (PCs) (e.g., laptops and / or desktops), and / or other similar computing devices. For example, the electronic device (200) may include or correspond to the electronic device (100) of FIG. 1. For example, the electronic device (200) may include at least a part of the electronic device (1101) of FIG. 11 or correspond to at least a part of the electronic device (1101) of FIG. 11. For example, the electronic device (200) may include at least one processor (210) (e.g., processor (1120) of FIG. 11), memory (220) (e.g., memory (1130) of FIG. 11), and display (230) (e.g., display module (1160) of FIG. 11).

[0033] According to one embodiment, at least one processor (210) may include a processing circuit. For example, at least one processor (210) may include a central processing unit (e.g., including a processing circuit). For example, at least one processor (210) may include a graphic processing unit (e.g., including a processing circuit) and / or a neural processing unit (e.g., including a processing circuit). For example, at least one processor (210) may be an application processor or may correspond. For example, at least one processor (210) may be configured to control a memory (220) and a display (230). At least one processor (210) may be configured to execute instructions stored in the memory (220) individually or collectively to cause an electronic device (200) (or electronic device (100)) to perform at least some of the operations illustrated in the description of FIG. 1. At least one processor (210) may be configured to execute instructions stored in memory (220) to cause the electronic device (200) to perform at least some of the operations illustrated in the description of FIGS. 3a through 10d.

[0034] According to one embodiment, the term “processor” as used herein, including in the claims, may include various processing circuits comprising at least one processor, and one or more of the at least one processor may be configured to perform the various functions described below in a distributed manner, individually and / or collectively. As used below, where “processor,” “at least one processor,” and “one or more processors” are described as being configured to perform various functions, these terms encompass, for example, but not limited to, situations where one processor performs some of the cited functions and other processor(s) perform other parts of the cited functions, and also situations where one processor can perform all of the cited functions. Additionally, the at least one processor may include a combination of processors that perform the enumerated / disclosed various functions, for example, in a distributed manner. The at least one processor may execute program instructions to achieve or perform the various functions.

[0035] According to one embodiment, the memory (220) may include one or more storage media. For example, the memory (220) may store various data used by at least one component of the electronic device (200) (e.g., at least one processor (210) and / or display (230)). For example, the data may include input data or output data for software and related commands. The memory (220) may include volatile memory or non-volatile memory.

[0036] According to one embodiment, the display (230) may output visualized information under the control of at least one processor (210). For example, the display (230) may include a flat panel display (FPD) and / or electronic paper. The FPD may include a liquid crystal display (LCD), a plasma display panel (PDP), and / or one or more light emitting diodes (LEDs). For example, the LEDs may include organic LEDs (OLEDs). The display (230) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by the touch. For example, the display (230) may be configured to receive handwritten input and user input. For example, the display (230) may be configured to display an image. For example, a display (230) that supports touch functions may be referred to as a touchscreen. The display (230) may further include a structure capable of detecting input using a stylus pen, such as EMR (electro-magnetic resonance) or AES (active electrostatic solution). For example, handwriting input can be performed using a stylus pen.

[0037] The electronic device (200) illustrated in FIG. 2 can perform at least some of the operations shown in FIG. 3a through FIG. 10d. For example, the operations shown in FIG. 3a through FIG. 10d can be caused by (or in) the electronic device (200) under the control of at least one processor (210).

[0038] FIG. 3a is a flowchart illustrating exemplary operations of an electronic device for acquiring information about a first image.

[0039] Referring to FIG. 3a, in operation 300, at least one processor (210) can identify an input for obtaining a second image (e.g., a second image (603) in FIG. 6a) using a first image (e.g., a first image (505) in FIG. 5) through a display (230). For example, the input may include a handwriting input provided by a user. For example, the handwriting input may be referred to as a drawing input. For example, the handwriting input may be received based on a pointing device such as a fingertip in contact with the display (230), a stylus, a digitizer, and / or a mouse for controlling the position of a cursor. For example, the handwriting input may include a user gesture of drawing a stroke using a fingertip or a pointing device. For example, handwritten input can be received through a UI (user interface) displayed through the display (230) or through an image displayed through the display (230).

[0040] At least one processor (210) can identify at least one stroke based on hand input. For example, at least one stroke may correspond to a trajectory and / or path dragged by a fingertip, stylus, and / or digitizer while the fingertip, stylus, and / or digitizer is in contact with the display (230). For example, at least one stroke may correspond to a trajectory and / or path of a cursor and / or mouse pointer moved within the display (230) by a pointing device, such as a mouse, for controlling the position of the cursor.

[0041] At least one processor (210) can acquire a first image including at least one stroke. For example, at least one stroke included in the first image may include a rendered stroke. For example, at least one processor (210) can render at least one stroke so that at least one stroke appears to be drawn using a pen in the first image.

[0042] For example, an input for obtaining a second image using a first image may include an input for selecting the first image as the image to be used to obtain the second image from among the images stored in the electronic device (200). For example, at least one processor (210) may display images stored in the electronic device (200) through a display (230). At least one processor (210) may receive an input for selecting the first image from among the displayed images. For example, the input may include a touch input made on the first image. For example, the input may be received through a display (230) (e.g., a touchscreen). However, the present disclosure is not limited to the exemplary embodiments above.

[0043] FIG. 3b illustrates an example of a user interface (UI) that receives input for acquiring a second image using a first image.

[0044] Referring to FIG. 3b, the state (335) may be a state in which a UI (340) for generating an image using a trained model is displayed or corresponds to a state in which a UI (340) for generating an image using a trained model is displayed. In the state (335), at least one processor (210) may display the UI (340) through a display (230). For example, the UI (340) may be displayed as application software for generating an image using a trained model is executed (in the foreground). For example, the UI (340) may include an indicator (345) indicating that the trained model is being used to generate an image.

[0045] For example, the UI (340) may include UI objects (350). For example, the UI objects (350) may be used to determine the input to be received through an input field (355) within the UI (340). For example, a UI object (350-1) may indicate handwritten input, a UI object (350-2) may indicate image input, and a UI object (350-3) may indicate text input.

[0046] For example, at least one processor (210) may receive handwritten input through an input field (355) based on user input for a UI object (350-1). For example, at least one processor (210) may acquire a first image including at least one stroke based on at least one stroke identified through the handwritten input. For example, at least one processor (210) may receive image input through an input field (355) based on user input for a UI object (350-2). For example, the image input may include an input that determines the first image among the images stored in the electronic device (200) as the image to be used to acquire the second image. For example, at least one processor (210) may receive text input through an input field (355) based on user input for a UI object (350-3). For example, text input may be received via a virtual keyboard (or soft keyboard) or via a microphone of an electronic device (200). However, it is not limited thereto.

[0047] Referring again to FIG. 3a, in operation 310, at least one processor (210) can display a first image through a display (230). For example, the first image may include at least one stroke identified according to handwriting input. For example, by displaying the first image, at least one processor (210) may enable the user to view at least one stroke identified according to handwriting input.

[0048] As another example, the first image may include an image determined by the input among images stored in the electronic device (200). For example, at least one processor (210) may enable the user to view the first image determined by the input by displaying the first image.

[0049] In operation 320, at least one processor (210) may provide a first image containing at least one stroke to a trained model. For example, the trained model may include a generative AI (artificial intelligence) model, a machine learning model, and / or a deep learning model. For example, the trained model may include a large language model (LLM). For example, the trained model may refer to the generative AI model (1230) of FIG. 12. For example, the trained model may be contained within an electronic device (200) or within a server (e.g., the server (1108) of FIG. 11). For example, at least one processor (210) may transmit the first image to the server using a communication circuit to provide the first image to the trained model contained within the server. For example, the server may receive the first image from the electronic device (200). The server can provide (or input) the first image received from the electronic device (200) to the trained model.

[0050] At least one processor (210) may provide a first prompt along with a first image to a trained model. For example, the first prompt may be stored in memory (220) or generated by a prompt generator. For example, the prompt generator may include the prompt design component (1221) of FIG. 12. For example, the first prompt may be a prompt requesting a description of the first image or may correspond to a prompt. For example, the first prompt may include text requesting a description of the first image (e.g., “Describe the object”). For example, the first prompt may further include text requesting keywords included in the description of the first image (e.g., “Extract keywords from the description of the object”). For example, the keywords may be text corresponding to or may correspond to features of an object composed of at least one stroke in the first image. For example, the first prompt may further include text requesting variations (or options) regarding a keyword included in the description of the first image (e.g., “Tell me several changeable options for the keyword”). For example, variations regarding the keyword may include, but are not limited to, color, texture, proportion, shape, size, and / or appearance regarding the feature.

[0051] For example, the first prompt may further include exemplary text of a keyword included in the description of the first image (e.g., “Features of a flower include petals, stems, and leaves.”). For example, the first prompt may further include exemplary text of variations (or options) regarding the keyword (e.g., “The petals may be red, yellow, or blue.”). At least one processor (210) may cause the trained model to generate the keyword and / or variations regarding the keyword included in the description of the first image according to the exemplary texts by providing the first prompt, which includes exemplary text of the keyword included in the description of the first image and / or exemplary text of variations regarding the keyword.

[0052] In operation 330, at least one processor (210) can obtain information about the first image generated using the first image by a trained model. For example, if the trained model is included in a server, the server can transmit information about the first image generated using the first image by the trained model to the electronic device (200). At least one processor (210) can obtain information about the first image by receiving information about the first image from the server using a communication circuit. For example, information about the first image may include text describing the first image (e.g., “A simple line drawing a tulip. The tulip has one stem with two leaves. It is connected to the base of the stem.”).

[0053] For example, at least one processor (210) may obtain keywords included in the information (e.g., “tulip,” “stem,” and / or “leaf”). For example, the keywords included in the information may refer to features of the object described by the information. For example, the features of the object described by the information may be or correspond to changeable parts of the object described by the information. For example, at least one processor (210) may obtain options (or variations) regarding the keywords (e.g., “tulip: red, yellow, purple”, “stem: green, brown, black”, and / or “leaf: green, brown, yellow”). For example, options (or variations) regarding the keywords may include texts that are changeable (or addable) regarding the keywords. For example, options regarding the keywords may be or correspond to options (or variations) of the changeable parts of the object. For example, a trained model may be used to obtain keywords and options (or variations) regarding the keywords included in the information. For example, at least one processor (210) may provide information about the first image back to the trained model. For example, at least one processor (210) may obtain keywords and options regarding the keywords included in the information generated by the trained model using information about the first image.

[0054] According to another embodiment, if the first prompt provided to the trained model includes text requesting a description of the first image, text requesting a keyword included in the description of the first image, and text requesting variations regarding the keyword, at least one processor (210) can obtain the keyword included in the information and options regarding the keyword, along with information about the first image generated by the trained model.

[0055] For example, operations 320 and 330 may be performed after operation 310 is performed. For example, operations 320 and 330 may be performed while the first image is displayed. For example, operations 320 and 330 may be performed while operation 310 is performed, or before operation 310 is performed. However, they are not limited thereto. For example, at least one processor (210) may acquire a second image using information about the first image. Acquiring a second image using information about the first image is exemplified in the description of FIG. 4.

[0056] FIG. 4 is a flowchart illustrating exemplary operations of an electronic device for displaying a second image and UI (user interface) objects.

[0057] Referring to FIG. 4, in operation 400, at least one processor (210) may provide a second prompt containing information about a first image to a trained model. For example, the trained model may include a model for image generation (e.g., a diffusion model). For example, the trained model may correspond to a trained model used to generate information about the first image, or may be different from a trained model used to generate information about the first image. For example, the trained model may be contained within an electronic device (200) or contained within a server. For example, if the trained model is contained within a server, at least one processor (210) may provide the second prompt to the trained model by transmitting the second prompt to the server using a communication circuit.

[0058] For example, the second prompt may be stored in memory (220) or generated by a prompt generator. The prompt generator may include the prompt design component (1221) of FIG. 12. For example, the second prompt may include text requesting the creation of an image based on information about the first image (e.g., “A simple line drawing a tulip. The tulip has one stem with two leaves. The stem is all connected to the ground. Draw a picture of this”).

[0059] In operation 410, at least one processor (210) may obtain a second image generated by a trained model using a second prompt. For example, the second image may be derived from a second prompt containing information about a first image. For example, the second image may include an object representing information about the first image. For example, the object representing information about the first image may include parts of an object representing each of the keywords included in the information about the first image. For example, the parts of an object representing each of the keywords included in the information about the first image may be referred to as features of the object.

[0060] For example, the style of the second image may be determined based on user input for the style of the second image. At least one processor (210) may receive user input for the style of the second image before acquiring the second image. User input for the style of the second image is exemplified in the description of FIG. 5.

[0061] Figure 5 illustrates an example of user input for the style of the second image.

[0062] Referring to FIG. 5, the state (500) may be a state in which a UI for handwriting input is displayed or may correspond to a state in which a UI for handwriting input is displayed. In the state (500), at least one processor (210) may display UI objects (515) for determining (or identifying) the style of a second image to be generated using a trained model through a display (230) within the UI for handwriting input. For example, the UI objects (515) may include a UI object for a style corresponding to watercolor, a UI object for a style corresponding to illustration, a UI object for a style corresponding to pop art, a UI object for a style corresponding to sketch, a UI object for a style corresponding to a three-dimensional cartoon, and a UI object for a style corresponding to oil painting. For example, when an image input is received, at least one processor (210) may identify whether the image identified according to the image input includes a human face. At least one processor (210) may further display UI objects for a style corresponding to comics, UI objects for a style corresponding to 3D characters, UI objects for a style corresponding to sketches, and UI objects for a style corresponding to watercolors, based on identifying that the image identified according to the image input includes a human face. However, it is not limited thereto. At least one processor (210) may receive user input (525) for one of the UI objects (520). For example, the user input (525) may be a user input for a style of the second image, or may correspond to one. For example, the user input (525) may include a touch input having a contact point on the UI object (520). The user input (525) may be received through a display (230) (e.g., a touchscreen).For example, user input (525) may be received through an external electronic device (e.g., a mouse) connected to the electronic device (200). For example, user input (525) may include voice (or speech) input received through a microphone of the electronic device (200) (e.g., the audio module (1170) of FIG. 11). However, it is not limited thereto.

[0063] For example, user input (525) may be received before handwritten input is received, or after handwritten input is received. At least one processor (210) may receive handwritten input after receiving user input (525), or receive user input (525) after receiving handwritten input. At least one processor (210) may display a first image (505) containing at least one stroke (510) identified according to handwritten input through a display (230) based on handwritten input. At least one processor (210) may change the appearance (e.g., color) of a UI object (520) for which user input (525) is received based on user input (525). At least one processor (210) may indicate that input for a UI object (520) has been received by displaying the UI object (520) with the changed appearance (e.g., color).

[0064] For example, at least one processor (210) may display a UI object (530) for generating an image within a UI for manual input via a display (230). At least one processor (210) may receive user input for the UI object (530). User input for the UI object (530) may be received after manual input and user input (525) have been received. At least one processor (210) may obtain information about the first image, keywords included in the information, and options regarding the keywords by providing the first image (505) and the first prompt to a trained model based on user input for the UI object (530). For example, obtaining information about the first image, keywords included in the information, and options regarding the keywords may refer to operations 320 to 330 of FIG. 3A. At least one processor (210) can acquire a second image by providing a second prompt containing information about the first image to a trained model. For example, the second prompt may further include other information about the style indicated by the UI object (530) that received the user input (525) (e.g., a style corresponding to watercolor). For example, the second image may be derived from the second prompt. For example, the second image may have the style indicated by the UI object (530) (e.g., a style corresponding to watercolor).

[0065] Referring again to FIG. 4, in operation 420, at least one processor (210) can display a second image through a display (230). For example, parts of an object included in the second image may be displayed differently from the user's intention. For example, it may be required to change parts of an object included in the second image that are displayed differently from the user's intention based on user input.

[0066] At least one processor (210) can display information about a keyword included in information about the first image through a display (230). For example, since the keyword is represented by a part of an object in the second image, at least one processor (210) can indicate a changeable part of an object in the second image by displaying the keyword. At least one processor (210) can display UI objects through the display (230) that indicate options regarding the keyword in conjunction with the keyword. For example, the UI objects may indicate options of the part of the object in the second image that represents the keyword. For example, the UI objects may be displayed based on user input. User input for displaying the UI objects is exemplified in the description of FIGS. 6a and 6b.

[0067] FIGS. 6a and FIGS. 6b illustrate examples of user input for displaying UI objects that indicate options, respectively.

[0068] Referring to FIG. 6a, the state (600) may be a state in which the second image (603) is displayed or may correspond to a state in which the second image (603) is displayed. In the state (600), at least one processor (210) may display UI objects (606) together with the second image (603) through a display (230). At least one processor (210) may receive user input for the UI objects (606). User input for the UI objects (606) may include touch input having a contact point on the UI objects (606). For example, user input for the UI objects (606) may be received through a display (230) (e.g., a touchscreen).

[0069] For example, at least one processor (210) may save (or clip) the second image (603) to the clipboard based on user input for the UI object (606-1). For example, the user input for the UI object (606-1) may be a user input for copying the second image (603) or may correspond to it. For example, at least one processor (210) may share (or transmit) the second image (603) based on user input for the UI object (606-2). At least one processor (210) may receive user input for a UI object (606-2) and then receive user input determining an external electronic device (or user, or user ID) to share (or transmit) the second image (603). For example, at least one processor (210) may share (or transmit) the second image (603) to the determined external electronic device (or user, or user ID). At least one processor (210) may store the second image (603) in the electronic device (200) (or in memory (220)) based on user input for a UI object (606-3). As another example, at least one processor (210) may store the object (605) in the electronic device (200) by cropping the object (605) within the second image (603) based on user input for a UI object (606-3). For example, The object (605) can be stored in the electronic device (200) as a sticker, and the object (605) stored as a sticker can be added to another image or transmitted to an external electronic device through a messenger application.

[0070] For example, at least one processor (210) may further display an indicator (607) along with the second image (603) through the display (230). For example, the indicator (607) may indicate that a scroll input (e.g., a left-right scroll input) may be received through the second image (603). At least one processor (210) may display the first image (e.g., the first image (505) of FIG. 5) or other images generated using the first image based on the scroll input received on the second image (603). For example, the indicator (607) may represent the image being displayed through the display (230) among the second image (603), the first image, and other images.

[0071] For example, at least one processor (210) may further display UI objects (608) and UI objects (609) along with a second image (603) through a display (230). For example, at least one processor (210) may receive user input for UI objects (608) and UI objects (609). For example, UI objects (608) may represent the first image. For example, at least one processor (210) may modify (or change) the first image based on user input for UI objects (608). For example, at least one processor (210) may receive text input based on user input for UI objects (609). At least one processor (210) may generate a prompt using text identified through text input. At least one processor (210) can obtain another image generated from a trained model using the second image (603) and the prompt by providing the second image (603) and the prompt to the trained model.

[0072] At least one processor (210) may display a UI object (610) for changing (or modifying) a second image (603) via a display (230). At least one processor (210) may receive user input (615) for the UI object (610). For example, the user input (615) may include touch input having a contact point on the UI object (610). The user input (615) may be received via the display (230) (e.g., a touchscreen). For example, the user input (615) may be received via an external electronic device (e.g., a mouse) connected to the electronic device (200). For example, the user input (615) may include voice (or verbal) input received via a microphone of the electronic device (200). However, it is not limited thereto.

[0073] Based on user input (615), the electronic device (200) can transition from state (600) to state (620). In state (620), at least one processor (210) can display a window (625) based on user input (615). For example, the window (625) may include information (630) about keywords included in information about the first image. For example, the information about keywords (630) may include text indicating the keywords (e.g., “tulip”, “stem”, and / or “leaves”). The keywords may each represent parts of an object (605) in the second image (603). At least one processor (210) may indicate the changeable parts of the object (605) in the second image (603) by displaying the information about keywords (630).

[0074] For example, the window (625) may further include UI objects (635) that indicate options regarding the keywords. At least one processor (210) may display UI objects (635) that indicate options in conjunction with information (630) regarding the keywords within the window (625). For example, the UI objects (635) that indicate options may include text indicating options (e.g., “red tulip”, “yellow tulip”, and / or “purple tulip”). For example, at least one processor (210) may indicate options for a part of an object (605) within a second image (603) representing the keyword by displaying the UI objects (635-1) in conjunction with information (630-1) regarding the keyword.

[0075] Referring to FIG. 6b, the state (640) may be a state in which the second image (603) is displayed or may correspond to a state in which the second image (603) is displayed. In the state (640), at least one processor (210) may display UI objects (645) that indicate keywords on the second image (603) via the display (230). For example, the UI objects (645) may include text indicating the keywords (e.g., “tulip”, “stem”, and / or “leaves”). The UI objects (645) may be displayed in conjunction with parts (650) of the object (605) in the second image (603) that represent the keywords. For example, the UI object (645-1) may be displayed in conjunction with a part (650-1) of the object (605) that represents the keyword indicated by the UI object (645-1). For example, a UI object (645-2) may be indicated in association with a part (650-2) of an object (605) that represents a keyword indicated by the UI object (645-2). For example, a UI object (645-3) may be indicated in association with a part (650-3) of an object (605) that represents a keyword indicated by the UI object (645-3). At least one processor (210) can intuitively indicate the changeable parts of an object (605) by indicating the UI objects (645) in association with parts (650) of an object (605) that represent keywords.

[0076] At least one processor (210) may receive user input (655) for UI objects (645). For example, user input (655) for UI objects (645) may be user input for changing parts of an object (605) that represent keywords indicated by the UI objects (645), or may correspond to such input. For example, user input (655) may include touch input having a contact point on the UI objects (645). User input (655) may be received via a display (230) (e.g., a touchscreen). For example, user input (655) may be received via an external electronic device (e.g., a mouse) connected to the electronic device (200). For example, user input (655) may include voice (or verbal) input received via a microphone of the electronic device (200), but is not limited thereto.

[0077] Based on user input (655) for a UI object (645-1), the electronic device (200) may transition from state (640) to state (660) (shown in FIG. 6b). In state (660), at least one processor (210) may display UI objects (635-1) that indicate options in association with the UI object (645-1) to which the user input (655) was received, based on the user input (655). The options indicated by each UI object (635-1) may be options regarding the keyword indicated by the UI object (645-1) or may correspond to it. For example, the UI objects (635-1) indicating each option may include text indicating each option (e.g., “red tulip”, “yellow tulip”, and / or “purple tulip”). As UI objects (635-1) are displayed in association with an object (645-1) indicating a keyword, options for the part of the object (605) in the second image (603) representing the keyword can be indicated.

[0078] For example, options directed by executable objects (635-1) may not include options intended by the user. For example, at least one processor (210) may receive text input directing options intended by the user to modify parts (650) of object (605). Text input directing options intended by the user is described with reference to FIG. 6c.

[0079] FIG. 6c illustrates an example of text input for changing parts of an object.

[0080] Referring to FIG. 6c, the state (663) may be a state in which the second image (603) is displayed or may correspond to a state in which the second image (603) is displayed. In the state (663), at least one processor (210) may display information (665) indicating keywords on the second image (603) via the display (230). For example, the information (665) may include text indicating keywords (e.g., “tulip”, “stem”, and / or “leaves”). For example, the information (665) may be displayed in conjunction with parts (650) of an object (605) within the second image (603) that represent the keywords. For example, the first information (665-1) may be displayed in conjunction with parts (650-1) of an object (605) that represent the keywords indicated by the first information (665-1). For example, the second information (665-2) may be indicated in association with a part (650-2) of an object (605) representing a keyword indicated by the second information (665-2). For example, the third information (665-3) may be indicated in association with a part (650-3) of an object (605) representing a keyword indicated by the third information (665-3). At least one processor (210) can intuitively indicate the changeable parts of the object (605) by indicating the information (665) in association with parts (650) of the object (605) representing the keywords.

[0081] For example, at least one processor (210) may display a text input field (670) through a display (230). For example, the text input field (670) may be displayed simultaneously with the second image (603). For example, at least one processor (210) may receive text input through the text input field (670). For example, the text input may be an input for changing parts (650) of an object (605) that represent keywords indicated by information (665), or may correspond to.

[0082] For example, text input may be received via a virtual keyboard. For example, the virtual keyboard may be displayed via a display (230) based on touch input having a contact point on a text input field (670). For example, the text input may be received via an external electronic device (e.g., a keyboard) connected to the electronic device (200). For example, the text input may include voice input (or verbal input) received via a microphone of the electronic device (200). However, it is not limited thereto.

[0083] For example, text corresponding to the text input may include keywords indicated by the information (665) (e.g., “tulip”, “stem”, and “leaves”). For example, text corresponding to the text input may include text indicating changes to parts (650) of an object (605) in the second image (603). For example, text corresponding to the text input may include additional keywords not included in the information (665). For example, text corresponding to the text input may include text indicating changes to parts of an object (605) in the second image (603) expressed by additional keywords.

[0084] For example, the text corresponding to the text input may include text instructing the deletion (or removal) of parts (650) of an object (605) within the second image (603). For example, the text corresponding to the text input may include text instructing the addition of another object within the second image (603). For example, at least one processor (210) may generate a prompt for modifying the second image (603) using the text received through the text input field (670).

[0085] For example, at least one processor (210) can modify parts (650) of an object (605) in the second image (603) in detail using text received through a text input field (670). For example, at least one processor (210) can modify (or delete, or add) parts (650) of an object (605) in the second image (603) and other parts of an object (605) using text received through a text input field (670).

[0086] For example, the options indicated by the UI objects (635-1) in FIGS. 6a and 6b may not include options intended by the user. For example, it may be required to display a UI object indicating additional options. Displaying a UI object indicating additional options is exemplified in the description of FIGS. 7a and 7b.

[0087] FIGS. 7a and 7b illustrate examples of user input for displaying a UI object that indicates additional options.

[0088] Referring to FIG. 7a, the state (700) may be a state in which UI objects (635-1) indicating options that are fewer than the acquired options are displayed, or may correspond to such a state. In the state (700), at least one processor (210) may display a reference number of UI objects (635-1) through a display (230). For example, the reference number may be the number of UI objects (635-1) that can be displayed in conjunction with information (630-1) about the keyword, or may correspond to such a state. The reference number may be predetermined or set (or changed) by the user. For example, if UI objects (635-1) exceeding the reference number are displayed, the area occupied by the UI objects (635-1) within the UI may be excessively large, or there may be insufficient area to display UI objects indicating options related to other keywords. Even if the number of options regarding keywords included in the information about the first image exceeds the standard number, at least one processor (210) can display the standard number of UI objects (635-1) in conjunction with the information about the keywords (630-1).

[0089] For example, the options indicated by the reference number of UI objects (635-1) may not include the option intended by the user. At least one processor (210) may display a UI object (705) for additional options in conjunction with the UI objects (635-1) (or information about keywords (630-1)) via the display (230). At least one processor (210) may receive user input (710) for the UI object (705). For example, the user input (710) may include a touch input having a contact point on the UI object (705). The user input (710) may be received via the display (230) (e.g., a touchscreen). For example, the user input (710) may be received via an external electronic device (e.g., a mouse) connected to the electronic device (200). For example, user input (710) may include voice (or verbal) input received through the microphone of the electronic device (200). However, it is not limited thereto.

[0090] At least one processor (210) may, based on user input (710) for a UI object (705), display a UI object indicating additional options via a display (230), or display a UI object indicating additional options in place of the UI objects (635-1). However, it is not limited thereto. At least one processor (210) may provide the user with a changeable additional option for the part of the object expressing the keyword by displaying a UI object indicating additional options.

[0091] Referring to FIG. 7b, in state (715), at least one processor (210) may display UI objects (635-1) indicating information (630-1) about a keyword and options, respectively, through a display (230). The UI objects (635-1) may be displayed in conjunction with the information (630-1) about the keyword. For example, the options indicated by the UI objects (635-1) may not include options intended by the user. At least one processor (210) may further display UI objects (720) for additional options in conjunction with the information (630-1) about the keyword (or UI objects (635-1)) through the display (230).

[0092] At least one processor (210) may receive user input (725) for a UI object (720). For example, user input (725) may include touch input having a contact point on the UI object (720). User input (725) may be received via a display (230) (e.g., a touchscreen). For example, user input (725) may be received via an external electronic device (e.g., a mouse) connected to the electronic device (200). For example, user input (725) may include voice (or verbal) input received via a microphone of the electronic device (200). However, it is not limited thereto.

[0093] At least one processor (210) may display an input field for additional options via a display (230) based on user input (725) for a UI object (720). At least one processor (210) may receive user input for additional options via the input field. For example, user input for additional options may include text input and / or voice (or verbal) input. For example, user input for additional options may be received via a virtual keyboard or via a microphone of an electronic device (200). However, it is not limited thereto.

[0094] At least one processor (210) may, based on user input for additional options, display additional UI objects indicating additional options via the display (230), or replace UI objects indicating additional options with UI objects (635-1). However, it is not limited thereto. For example, additional options may be determined by text (or voice) identified according to user input. At least one processor (210) may provide the user with changeable additional options for parts of objects expressing keywords by displaying UI objects indicating additional options.

[0095] At least one processor (210) can receive user input for one of the UI objects (635-1) that indicate options. At least one processor (210) can display a third image based on user input for the UI object. Displaying a third image based on user input for the UI object is exemplified in the description of FIG. 8.

[0096] FIG. 8 is a flowchart illustrating exemplary operations of an electronic device for displaying a third image.

[0097] Referring to FIG. 8, in operation 800, at least one processor (210) may receive user input for at least one UI object among UI objects that indicate options. For example, user input for at least one UI object may include touch input having a contact point on at least one UI object. User input for at least one UI object may be received through a display (230) (e.g., a touchscreen). For example, user input for at least one UI object may be received through an external electronic device (e.g., a mouse) connected to the electronic device (200). For example, user input for at least one UI object may include voice (or verbal) input received through a microphone of the electronic device (200). However, it is not limited thereto.

[0098] According to another embodiment, at least one processor (210) may bypass (or refrain from, or skip, or interrupt, or not generate) the generation of a second image and display UI objects through a display (230). For example, at least one processor (210) may bypass (or refrain from, or skip, or interrupt, or not display) the display of the second image and display a third image based on user input for one of the UI objects.

[0099] In operation 810, at least one processor (210) may provide a third prompt to a trained model, which includes information about a first image and information about an option directed by at least one UI object, based on user input for at least one UI object. For example, the trained model may include a model for image generation (e.g., a diffusion model). For example, the trained model may be different from the trained model used to generate information about the first image or the trained model used to generate information about the first image. For example, the trained model may be contained within an electronic device (200) or contained within a server. For example, at least one processor (210) may provide the third prompt to the trained model by transmitting the third prompt to a server using a communication circuit.

[0100] For example, the third prompt may be generated by a prompt generator including the prompt design component (1221) of FIG. 12. For example, the third prompt may include an option indicated by at least one UI object and text requesting the creation of an image according to information about the first image (e.g., “A simple line to draw a tulip. A purple tulip has one stem with two leaves. The stem is connected to the ground. Draw a picture of this”). For example, the third prompt may be a prompt that further includes information about the option indicated by at least one UI object in the second prompt, or may correspond to it. For example, the third prompt may further include text corresponding to text input through the text input field (670) of FIG. 6c.

[0101] According to another embodiment, at least one processor (210) may provide a first image to a trained model along with a third prompt. For example, a third image generated by the trained model using the third prompt may include a visual object having a shape different from the visual object contained in the second image. For example, it may be required to obtain a third image containing an object having a shape corresponding to the shape of the visual object contained in the second image. The shape may represent a keyword having an option indicated by at least one UI object. For example, at least one processor (210) may perform a control scale by further providing the first image to the trained model. At least one processor (210) may apply weights to the first image by performing a control scale. At least one processor (210) may cause the trained model to generate a third image containing a visual object having a shape corresponding to the shape of the visual object contained in the second image according to said weights by applying weights to the first image.

[0102] In operation 820, at least one processor (210) may obtain a third image generated by a trained model using a third prompt. For example, the third image may be derived from a third prompt that includes information about a first image and information about an option indicated by at least one UI object. For example, the third image may include an object that represents information about the first image. For example, the object that represents information about the first image may include parts of an object that represent each of the keywords included in the information about the first image. For example, the part of the object that represents information about the first image may represent a keyword having an option indicated by at least one UI object.

[0103] In operation 830, at least one processor (210) may display a third image through a display (230). For example, an object contained within the third image may include a portion of an object expressing a keyword having an option indicated by at least one UI object. For example, an object contained within the third image may include an expression intended by the user. At least one processor (210) may provide an object containing the expression intended by the user by displaying the third image. Displaying the third image is exemplified in the description of FIG. 9a.

[0104] FIG. 9a illustrates an example of displaying a third image.

[0105] Referring to FIG. 9a, the state (900) may be a state in which UI objects (635) indicating options are displayed or correspond to each other. In the state (900), at least one processor (210) may display the UI objects (635) in association with information (630) regarding keywords through a display (230). At least one processor (210) may receive user input (910) for one of the UI objects (635-1) displayed in association with information (630-1) regarding keywords. For example, the user input (910) may include a touch input having a contact point on the UI object (905-1). The user input (910) may be received through a display (230) (e.g., a touchscreen). For example, user input (910) may be received through an external electronic device (e.g., a mouse) connected to the electronic device (200). For example, user input (910) may include voice (or verbal) input received through a microphone of the electronic device (200). However, it is not limited thereto.

[0106] For example, at least one processor (210) may further receive user input for one UI object (905-2) among UI objects (635-2) displayed in association with information (630-2) about keywords and / or user input for one UI object (905-3) among UI objects (635-3) displayed in association with information (630-3) about keywords. For example, at least one processor (210) may change the appearance (e.g., color) of the UI object (905-1) (or UI object (905-2), or UI object (905-3)) for which user input (910) was received, based on user input (910) (or user input for UI object (905-2), or user input for UI object (905-3)). At least one processor (210) can indicate that input for a UI object (905-1) (or UI object (905-2), or UI object (905-3)) is received by displaying a UI object (905-1) (or UI object (905-2), or UI object (905-3)) with a changed appearance (e.g., color).

[0107] For example, at least one processor (210) may display a UI object (915) for creating (or regenerating) an image through a display (230). For example, at least one processor (210) may receive user input (920) for the UI object (915). For example, user input (920) for the UI object (915) may be received after user input (910), user input for the UI object (905-2), and / or user input for the UI object (905-3) have been received. For example, user input (920) may include touch input having a contact point on the UI object (915). User input (920) may be received through a display (230) (e.g., a touchscreen). For example, user input (920) may be received through an external electronic device (e.g., a mouse) connected to the electronic device (200). For example, user input (920) may include voice (or verbal) input received through the microphone of the electronic device (200). However, it is not limited thereto.

[0108] Based on user input (920), the electronic device (200) may transition from state (900) to state (925). In state (925), at least one processor (210) may display information (930) indicating that a third image (945) is being generated via a display (230) based on user input (920). For example, the information (930) may include text indicating that the third image (945) is being generated (e.g., “generating”). For example, the information (930) may include text for an option indicated by the UI object (905-1) to which user input (910) is received (e.g., “purple tulip”). However, it is not limited thereto.

[0109] At least one processor (210) may provide a third prompt to a trained model based on user input (920), the third prompt including information about a first image and information about an option indicated by a UI object. At least one processor (210) may acquire a third image generated by the trained model using the third prompt. For example, acquiring the third image (945) may be described in the description of operations 810 and 820 of FIG. 8.

[0110] For example, at least one processor (210) may display a UI object (935) for stopping the generation of a third image through a display (230). At least one processor (210) may receive user input regarding the UI object (935). At least one processor (210) may stop (or refrain from, or skip, or bypass, or not generate) the generation of the third image based on the user input regarding the UI object (935).

[0111] For example, information (930) indicating that a third image (945) is being generated may be displayed until the third image (945) is acquired. Based on acquiring the third image (945), the electronic device (200) may transition from state (925) to state (940). In state (940), at least one processor (210) may display the third image (945) through a display (230) based on acquiring the third image (945). An object (950) within the third image (945) may represent keywords having options indicated by UI objects (905) to which user inputs are received. For example, the object (950) may include a part (955-1) of the object (950) that represents a keyword having an option indicated by the UI object (905-1) (e.g., “purple tulip”). For example, the object (950) may include a part (955-2) of the object (950) that represents a keyword having an option (e.g., “brown stem”) indicated by the UI object (905-2). For example, the object (950) may include a part (955-3) of the object (950) that represents a keyword having an option (e.g., “green leaves”) indicated by the UI object (905-3). For example, parts (955) of the object (950) may correspond to expressions intended by the user. For example, at least one processor (210) may provide an object (950) that includes expressions intended by the user by displaying a third image (945).

[0112] For example, at least one processor (210) may display a UI object (960) for changing (or modifying) the third image (945) through a display (230). For example, the UI object (960) may reference the UI object (610) of FIG. 6a. At least one processor (210) may receive user input for the UI object (960). At least one processor (210) may display a window (e.g., the window (625) of FIG. 6a) based on the user input (615). For example, at least one processor (210) may further change (or modify) the third image (945) through the window.

[0113] FIG. 9b illustrates an example of a third image obtained based on text input.

[0114] Referring to FIG. 9b, the state (965) may be a state in which the second image (603) is displayed or may correspond to a state in which the second image (603) is displayed. For example, at least one processor (210) may display a text input field (970) through a display (230). For example, the text input field (970) may be displayed simultaneously with the second image (603). For example, at least one processor (210) may receive text input through the text input field (970). For example, the text input may be an input for changing an object (605) within the second image (603) or may correspond to a state in which the second image (603) is displayed.

[0115] For example, text input may be received via a virtual keyboard. For example, the virtual keyboard may be displayed via a display (230) based on touch input having a contact point on a text input field (970). For example, the text input may be received via an external electronic device (e.g., a keyboard) connected to the electronic device (200). For example, the text input may include voice input (or verbal input) received via a microphone of the electronic device (200). However, it is not limited thereto.

[0116] For example, at least one processor (210) can generate a prompt corresponding to the text input based on the text input. For example, at least one processor (210) can provide the prompt and the second image (603) to a trained model. For example, at least one processor (210) can obtain a third image (945) generated using the prompt and the second image (603) from the trained model.

[0117] The electronic device (200) can transition from state (965) to state (975) based on acquiring a third image (945). In state (975), at least one processor (210) can display the third image (945) through a display (230). For example, the third image (945) may be an image modified from the second image (603) based on text input received through a text input field (970), or may correspond to it. For example, by displaying the third image (945), at least one processor (210) can provide an object (950) containing representations intended by the user.

[0118] FIGS. 10a, FIGS. 10b, and FIGS. 10c illustrate examples of text input and image input received along with handwritten input.

[0119] Referring to FIG. 10a, the state (1000) may be a state in which text input is received or may correspond. In the state (1000), at least one processor (210) may display a UI object (1005) for text input through a display (230). At least one processor (210) may receive user input (1010) for the UI object (1005). For example, the user input (1010) may include a touch input having a contact point on the UI object (1005). The user input (1010) may be received through the display (230) (e.g., a touchscreen). For example, the user input (1010) may be received through an external electronic device (e.g., a mouse) connected to the electronic device (200). For example, the user input (1010) may include voice (or verbal) input received through a microphone of the electronic device (200). However, it is not limited to this.

[0120] At least one processor (210) can display an input field (1015) through a display (230) based on user input (1010). At least one processor (210) can receive text input (1025) through the input field (1015). For example, text input (1025) may be received through a virtual keyboard (1020) (or a soft keyboard). For example, text input (1025) may be received through an external electronic device (e.g., a keyboard) connected to the electronic device (200). For example, text input (1025) may be received through a microphone of the electronic device (200). However, it is not limited thereto. For example, text input (1025) may be received along with handwritten input. For example, at least one processor (210) may receive a text input (1025) and then receive a handwritten input, or receive a text input (1025) and then receive a handwritten input.

[0121] At least one processor (210) may provide a second prompt to a trained model based on a text input (1025), the second prompt including information including text identified according to the text input (1025) (e.g., “good morning”) and information about a first image (e.g., the first image (505) of FIG. 5). At least one processor (210) may obtain a second image (1022) generated by the trained model using the second prompt.

[0122] The electronic device (200) can transition from state (1000) to state (1021) based on acquiring a second image (1022). In state (1021), at least one processor (210) can display the second image (1022) and UI objects (e.g., UI objects (635) of FIG. 6a) through a display (230). For example, at least one processor (210) can further display an indicator (1024) along with the second image (1022) through the display (230). For example, the indicator (1024) can indicate that a scroll input (e.g., a left-right scroll input) can be received through the second image (1022). At least one processor (210) may display text identified through text input (1025) or other images generated using said text based on scroll input received on the second image (1022). For example, an indicator (1024) may display an image being displayed through the display (230) among the second image (1022), said text, and other images.

[0123] For example, at least one processor (210) may further display UI objects (1026) and UI objects (1027) through a display (230). For example, at least one processor (210) may receive user input for UI objects (1026) and UI objects (1027). For example, UI objects (1026) may include text identified through text input (1025). At least one processor (210) may display text identified through text input (1025) through a display (230) based on user input for UI objects (1026). For example, at least one processor (210) may further receive image input based on user input for UI objects (1027). For example, image input may include user input for selecting one image among images stored in the electronic device (200). For example, at least one processor (210) can obtain a third image generated from a trained model using an image identified through an image input and a second prompt by providing the image identified through the image input and a second prompt to a trained model based on an image input.

[0124] Referring to FIG. 10b, the state (1030) may be a state in which a plurality of images (1035) stored in the electronic device (200) (or stored in association with a user account) are displayed or may correspond to. In the state (1030), at least one processor (210) may receive user input (1045) for at least one image (1040) among the plurality of images (1035) stored in the electronic device (200) (or stored in association with a user account). For example, the user input (1045) may include a touch input having a contact point on the image (1040). The user input (1045) may be received through a display (230) (e.g., a touchscreen). For example, the user input (1045) may be received through an external electronic device (e.g., a mouse) connected to the electronic device (200). For example, user input (1045) may include voice (or verbal) input received through the microphone of the electronic device (200). However, it is not limited thereto.

[0125] For example, at least one processor (210) may display at least one image (1040) indicated by the user input (1045) through a display (230) based on receiving the user input (1045). At least one processor (210) may receive handwriting input while at least one image (1040) is displayed. At least one processor (210) may acquire a first image including at least one stroke identified according to the handwriting input and at least one image (1040) based on the handwriting input received while at least one image (1040) is displayed. At least one processor (210) may acquire a second image by providing a second prompt containing information about the first image to a trained model. At least one processor (210) may display the second image and UI objects. For example, displaying the second image and UI objects may refer to operations 400 to 420 of FIG. 4.

[0126] According to another embodiment, at least one processor (210) may receive a text input (1025) and a user input (1045) for at least one image (1040). At least one processor (210) may receive the user input (1045) after receiving the text input (1025), or receive the text input (1025) after receiving the user input (1045). At least one processor (210) may provide a second prompt to a trained model based on the text input (1025), the second prompt including information containing text identified according to the text input (1025) and information about at least one image (1040). At least one processor (210) may obtain a second image (1055) generated by the trained model using the second prompt.

[0127] The electronic device (200) can transition from state (1030) to state (1050) based on acquiring a second image (1055). In state (1050), at least one processor (210) can display the second image (1055) and UI objects through a display (230). For example, displaying the second image and UI objects may refer to operations 400 to 420 of FIG. 4.

[0128] For example, at least one processor (210) may further display an indicator (1060) along with the second image (1055) through the display (230). For example, the indicator (1060) may indicate that a scroll input (e.g., left-right scroll input) may be received through the second image (1055). At least one processor (210) may display text identified through text input (e.g., “starry night”) or other images generated using at least one image (1040) based on the scroll input received on the second image (1055). For example, the indicator (1060) may represent the image being displayed through the display (230) among the second image (1055), the text, and other images.

[0129] For example, at least one processor (210) may further display UI objects (1065) and UI objects (1070) through a display (230). For example, at least one processor (210) may receive user input for UI objects (1065) and UI objects (1070). For example, UI objects (1065) may include text identified through text input (e.g., “starry night”). At least one processor (210) may display text identified through text input (e.g., “starry night”) through a display (230) based on user input for UI objects (1065). For example, UI objects (1070) may include at least one image (1040) for which user input (1045) has been received. At least one processor (210) can display at least one image (1040) through a display (230) based on user input for a UI object (1070).

[0130] Referring to FIG. 10c, the state (1075) may be a state in which a web application is running or may correspond to one. In the state (1075), at least one processor (210) may display a web page (1076) through a display (230). For example, the web page (1076) may include a first image (1077). For example, the first image (1077) may be displayed based on the markup language of the web page (1076). For example, the markup language may include HTML (hypertext markup language) and / or XML (extensible markup language).

[0131] For example, at least one processor (210) may receive (or identify) an input (1078) for a first image (1077) in a webpage (1076). For example, the input (1078) for the first image (1077) may include a touch input having a contact point on the first image (1077). For example, the input (1078) for the first image (1077) may be referred to as a long press input for the first image (1077). For example, the long press input for the first image (1077) may be a touch input having a contact point on the first image (1077) that is maintained for a threshold time, or may correspond to one. For example, the input (1078) for the first image (1077) may include input received regarding the first image (1077) through an external electronic device (e.g., a mouse or keyboard) connected to the electronic device (200). For example, the input (1078) for the first image (1077) may include voice input (or verbal input) received through a microphone of the electronic device (200). However, it is not limited thereto.

[0132] For example, at least one processor (210) may display a pop-up window (or floating window) (1079) through a display (230) based on an input (1078) for a first image (1077). For example, the pop-up window (1079) may be displayed in conjunction with the first image (1077). For example, the pop-up window (1079) may include a UI object (1080) that directs a change (or edit) of the first image (1077). For example, at least one processor (210) may receive (or identify) an input (1081) for the UI object (1080). For example, the input (1081) for the UI object (1080) may be referenced as an input for creating (or obtaining) a second image (1084) using the first image (1077). For example, input (1081) for a UI object (1080) may include a touch input having a contact point on the UI object (1080). For example, input (1081) for a UI object (1080) may be received through a display (230) (e.g., a touchscreen). For example, input (1081) for a UI object (1080) may include input received regarding a first image (1077) through an external electronic device (e.g., a mouse or keyboard) connected to the electronic device (200). For example, input (1078) for a first image (1077) may include voice input (or verbal input) received through a microphone of the electronic device (200). However, it is not limited thereto.

[0133] The electronic device (100) can transition from state (1075) to state (1082) based on input (1081). In state (1082), at least one processor (210) can execute an application for generating an image (e.g., a second image (1084)) using a trained model based on input (1081). For example, at least one processor (210) can display a screen (1083) of the application for generating an image using the trained model through a display (230). For example, the screen (1083) may be displayed in window mode (or picture-in-picture (PIP)). For example, the screen (1083) may be displayed superimposed on a web page (1076).

[0134] For example, at least one processor (210) may provide a first image (1077) to a trained model through a screen (1083). For example, at least one processor (210) may obtain a second image (1084) generated by the trained model using the first image (1077). For example, at least one processor (210) may display the second image (1084) through a display (230). For example, the second image (1084) may be displayed on the screen (1083) of an application for generating an image using the trained model.

[0135] For example, the screen (1083) may include a UI object (1085). For example, the UI object (1085) may display a first image (1077). At least one processor (210) may display images stored in the electronic device (200) through a display (230) based on user input regarding the UI object (1085). For example, at least one processor (210) may further provide at least one image among the images stored in the electronic device (200) to a trained model. For example, at least one processor (210) may obtain a third image generated by the trained model using at least one image among the images stored in the electronic device (200).

[0136] For example, the screen (1083) may include a text input field (1093). For example, at least one processor (210) may generate a prompt corresponding to the text input based on text input through the text input field (1093). For example, at least one processor (210) may further provide the prompt to a trained model. For example, at least one processor (210) may further use the prompt to obtain a third image generated.

[0137] FIG. 10d illustrates an example of displaying a third image within another application.

[0138] Referring to FIG. 10d, the state (1086) may be a state in which the first screen (1087) and the second screen (1088) are displayed simultaneously, or may correspond to such a state. For example, the first screen (1087) and the second screen (1088) may be displayed as a split view. For example, the first screen (1087) and the second screen (1088) may be displayed simultaneously based on picture-in-picture (PIP). For example, the first screen (1087) may be displayed superimposed on the second screen (1088). For example, the second screen (1088) may be displayed superimposed on the first screen (1087). For example, the first screen (1087) may include the screen of an application for using images, such as a note application, a messenger application, an electronic document application, or an image editing application. For example, the second screen (1088) may include a screen of an application for generating an image using a trained model.

[0139] In state (1086), at least one processor (210) can simultaneously display a first screen (1087) and a second screen (1088) through a display (230). For example, at least one processor (210) can display the first screen (1087) and the second screen (1088) as a split view. For example, at least one processor (210) can display an image (1089) generated from a trained model in the second screen (1088). For example, the screen (1088) may include a UI object (1090) for storing the image (1089). For example, at least one processor (210) can store the image (1089) in an electronic device (200) based on user input to the UI object (1090). For example, the image (1089) may be stored as a sticker object. For example, at least one processor (210) can obtain (or store) a sticker object for an object by cropping the object along the boundary of the object in the image (1089).

[0140] For example, at least one processor (210) may receive (or identify) an input (1091) for displaying an image (1089) within a first screen (1087). For example, the input (1091) may be referred to as a drag input for the image (1089). For example, the input (1091) may be a sequence or correspondence comprising a touch input having a contact point on the image (1089) (or an object within the image (1089)), a drag input having a contact point moving from the image (1089) (or an object within the image (1089)) to the first screen (1087), and a touch input releasing the contact point on the first screen (1087). For example, the input (1091) may be received via a display (230) (e.g., a touchscreen).

[0141] Based on input (1091), the electronic device (200) can switch from state (1086) to state (1092). In state (1092), at least one processor (210) can display an image (1089) through a first screen (1087) based on input (1091). For example, the image (1089) can be displayed in an input field of the first screen (1087). For example, the image (1089) can be displayed in the first screen (1087) as a sticker object representing an object within the image (1089).

[0142] For example, at least one processor (210) can bypass a sequence including the operation of storing the image (1089) in the second screen (1088) and the operation of loading the image (1089) in the first screen (1087) by displaying the image (1089) in the first screen (1087) based on the input (1091). For example, at least one processor (210) can enhance the user experience (UX) for the image (1089) by displaying the image (1089) in the first screen (1087) based on the input (1091).

[0143] FIG. 11 is a block diagram of an electronic device in a network environment according to various embodiments.

[0144] Referring to FIG. 11, in a network environment (1100), an electronic device (1101) may communicate with an electronic device (1102) through a first network (1198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (1104) or a server (1108) through a second network (1199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (1101) may communicate with the electronic device (1104) through a server (1108). According to one embodiment, the electronic device (1101) may include a processor (1120), memory (1130), input module (1150), sound output module (1155), display module (1160), audio module (1170), sensor module (1176), interface (1177), connection terminal (1178), haptic module (1179), camera module (1180), power management module (1188), battery (1189), communication module (1190), subscriber identification module (1196), or antenna module (1197). In some embodiments, at least one of these components (e.g., connection terminal (1178)) may be omitted from the electronic device (1101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (1176), camera module (1180), or antenna module (1197)) may be integrated into a single component (e.g., display module (1160)).

[0145] The processor (1120) can, for example, execute software (e.g., program (1140)) to control at least one other component (e.g., hardware or software component) of the electronic device (1101) connected to the processor (1120) and perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (1120) can store commands or data received from other components (e.g., sensor module (1176) or communication module (1190)) in volatile memory (1132), process the commands or data stored in volatile memory (1132), and store the resulting data in non-volatile memory (1134). According to one embodiment, the processor (1120) may include a main processor (1121) (e.g., a central processing unit or an application processor) or an auxiliary processor (1123) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (1101) includes a main processor (1121) and an auxiliary processor (1123), the auxiliary processor (1123) may be configured to use less power than the main processor (1121) or to be specialized for a specified function. The auxiliary processor (1123) may be implemented separately from the main processor (1121) or as part thereof.

[0146] The auxiliary processor (1123) may control at least some of the functions or states associated with at least one component of the electronic device (1101) (e.g., display module (1160), sensor module (1176), or communication module (1190)) on behalf of the main processor (1121) while the main processor (1121) is in an inactive (e.g., sleep) state, or together with the main processor (1121) while the main processor (1121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (1123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (1180) or communication module (1190)). According to one embodiment, the auxiliary processor (1123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (1101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (1108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.

[0147] The memory (1130) can store various data used by at least one component of the electronic device (1101) (e.g., processor (1120) or sensor module (1176)). The data may include, for example, software (e.g., program (1140)) and input or output data for related commands. The memory (1130) may include volatile memory (1132) or non-volatile memory (1134).

[0148] The program (1140) may be stored as software in memory (1130) and may include, for example, an operating system (1142), middleware (1144), or an application (1146).

[0149] The input module (1150) can receive commands or data to be used for a component of the electronic device (1101) (e.g., processor (1120)) from outside the electronic device (1101) (e.g., user). The input module (1150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0150] The sound output module (1155) can output a sound signal to the outside of the electronic device (1101). The sound output module (1155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.

[0151] The display module (1160) can visually provide information to an external (e.g., user) of the electronic device (1101). The display module (1160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (1160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.

[0152] The audio module (1170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (1170) can acquire sound through the input module (1150) or output sound through the sound output module (1155) or an external electronic device (e.g., electronic device (1102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (1101).

[0153] The sensor module (1176) can detect the operating state of the electronic device (1101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (1176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0154] The interface (1177) may support one or more specified protocols that can be used for the electronic device (1101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (1102)). According to one embodiment, the interface (1177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0155] The connection terminal (1178) may include a connector through which the electronic device (1101) can be physically connected to an external electronic device (e.g., electronic device (1102)). According to one embodiment, the connection terminal (1178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0156] The haptic module (1179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (1179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.

[0157] The camera module (1180) can capture still images and video. According to one embodiment, the camera module (1180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0158] The power management module (1188) can manage the power supplied to the electronic device (1101). According to one embodiment, the power management module (1188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).

[0159] The battery (1189) can supply power to at least one component of the electronic device (1101). According to one embodiment, the battery (1189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0160] The communication module (1190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (1101) and an external electronic device (e.g., electronic device (1102), electronic device (1104), or server (1108)), and the performance of communication through the established communication channel. The communication module (1190) may include one or more communication processors that operate independently of the processor (1120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1190) may include a wireless communication module (1192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (1194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (1104) via a first network (1198) (e.g., a short-range communication network such as Bluetooth, Wi-Fi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (1199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1192) can identify or authenticate the electronic device (1101) within a communication network such as the first network (1198) or the second network (1199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (1196).

[0161] The wireless communication module (1192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (1192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (1192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beamforming, or large-scale antenna. The wireless communication module (1192) can support various requirements specified in the electronic device (1101), external electronic device (e.g., electronic device (1104)), or network system (e.g., second network (1199)). According to one embodiment, the wireless communication module (1192) can support a Peak data rate (e.g., 20 Gbps or more) for realizing eMBB, loss coverage (e.g., 164 dB or less) for realizing mMTC, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for realizing URLLC.

[0162] An antenna module (1197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (1197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (1197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (1198) or a second network (1199), may be selected from the plurality of antennas, for example, by a communication module (1190). A signal or power may be transmitted or received between the communication module (1190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (1197).

[0163] According to various embodiments, the antenna module (1197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.

[0164] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.

[0165] According to one embodiment, commands or data may be transmitted or received between the electronic device (1101) and an external electronic device (1104) through a server (1108) connected to a second network (1199). Each of the external electronic devices (1102, or 1104) may be the same or a different type of device as the electronic device (1101). According to one embodiment, all or part of the operations performed on the electronic device (1101) may be performed on one or more of the external electronic devices (1102, 1104, or 1108). For example, if the electronic device (1101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (1101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (1101). The electronic device (1101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (1101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (1104) may include an Internet of Things (IoT) device. The server (1108) may be an intelligent server using machine learning and / or neural networks.According to one embodiment, an external electronic device (1104) or server (1108) may be included within the second network (1199). The electronic device (1101) may be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0166] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.

[0167] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as "coupled" or "connected" to another (e.g., 2nd) component, with or without the terms "functionally" or "communicationly," it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0168] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0169] Various embodiments of the present document may be implemented as software (e.g., program (2040)) comprising one or more instructions stored in a storage medium (e.g., internal memory (2036) or external memory (2038)) readable by a machine (e.g., electronic device (2001)). For example, a processor (e.g., processor (2020)) of the machine (e.g., electronic device (2001)) may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-transient' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0170] According to one embodiment, the method according to the various embodiments disclosed herein may be provided as included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0171] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

[0172] FIG. 12 illustrates an example of a generative artificial intelligence system according to one embodiment.

[0173] Referring to FIG. 12, the AI ​​system (1200) may include an input / output interface (1210), an AI framework (1220), a generative AI model (1230), and / or a knowledge repository (1290).

[0174] The input / output interface (1210) can receive input. The input may include user input and / or data acquired or generated by an electronic device (e.g., the electronic device (200) or electronic device (1101) described above). The data may include images, videos, and / or sensor data generated by at least one processor of the electronic device (e.g., at least one processor (210) or processor (1120)), such as illuminance data around the electronic device acquired from a sensor or sensor hub (e.g., auxiliary processor (1123), attitude data (or orientation data) of the electronic device, temperature inside the electronic device (e.g., display (230)), or temperature of at least one processor (210), size information of the display area of ​​the display (230), and / or images acquired through an image sensor of the electronic device (e.g., included in a camera module (1180)). The user input may include natural language, touch data obtained through a touch circuit included within the display panel (e.g., used to identify input from a finger and / or stylus), an image displayed (and / or to be displayed) on the display panel, and / or video. By example, without limitation, the user input may be received by an input / output interface (1210) along with context information. The context information may be additional information obtained in relation to the user input or may correspond to it. The context information may be related to the state at the time the user input is received (e.g., the state of the electronic device and / or the state of the surroundings of the electronic device (e.g., user state)). For example, the context information may include information about one or more software applications executed within the electronic device at the time the user input is received.For example, the above situation information may include information about the location of the electronic device (or the location of the user of the electronic device) when the user input is received. For example, the user input may be integrated with the situation information. For example, the user input with the situation information integrated as the input may be received by the input / output interface (1210).

[0175] The input / output interface (1210) may transmit (or provide) an output. The output may include a result (or result information) generated or obtained by the AI ​​system (1200) based on at least part of the input. The format of the output may vary. For example, the output may include natural language. For example, the output may include content (e.g., media content and / or multimedia content). For example, the output may include actions related to the user of the electronic device. For example, the output may have a format according to the user settings of the electronic device.

[0176] The AI ​​framework (1220) can be used to obtain information (or data) about the input from the input / output interface (1210) and to control one or more components related to the AI ​​system (1200) using the obtained information.

[0177] For example, a prompt design component (1221) within an AI framework (1220) can generate or obtain prompts for a generative AI model (1230) (e.g., including a large language model (LLM) or a large multimodal model (LMM)) using the acquired information. For example, the prompt design component (1221) may be an AI component that uses a learning algorithm and / or a neural network to provide prompts that are enhanced over time, or may correspond to one. For example, the prompt design component (1221) can generate or obtain prompts by accessing a knowledge component (e.g., a knowledge repository (1290)) containing user preference data, a prompt library, and / or prompt examples using the acquired information. The generated prompts may be provided to the generative AI model (1230) (e.g., including an LLM or LMM).

[0178] For example, an API / plugin management component (1222) within the AI ​​framework (1220) may be used to support communication for additional information requested (or induced) in relation to the prompt provided (or to be provided) to the generative AI model (1230). For example, the API / plugin management component (1222) may be used to create or establish a channel for communication with various data sources (e.g., knowledge repository (1290)). For example, the API / plugin management component (1222) may support access to at least some of the data sources. For example, the API / plugin management component (1222) may be used to request another component (e.g., application / service component (1280)) that performs feedback (or response) according to the prompt. As a non-limiting example, information obtained (or generated) through the API / plugin management component (1222) may be provided to the prompt design component (1221) for generating a prompt. As a non-limiting example, information obtained (or generated) through the API / plugin management component (1222) may be provided to the generative AI model (1230).

[0179] For example, an improvement component (1223) within the AI ​​framework (1220) can at least partially tune (or adjust) (or change) the result (e.g., content) obtained (or output) from the generative AI model (1230). For example, the improvement component (1223) can determine or verify whether the content obtained from the generative AI model (1230) is related to the input. For example, the improvement component (1223) can determine or verify whether the content obtained from the generative AI model (1230) contains biased content. For example, the improvement component (1223) can determine or verify whether the content obtained from the generative AI model (1230) contains harmful content. For example, the improvement component (1223) can support or assist in performing additional processing to improve the content obtained from the generative AI model (1230). For example, the improvement component (1223) may support providing a hint to the user to improve the content.

[0180] The generative AI model (1230) may be an artificial intelligence neural network that generates feedback in response to a prompt, or may respond. For example, the feedback may include additional data and / or information relative to the prompt, but may also include new content relative to the prompt. For example, the generative AI model (1230) may include a model that generates images and / or a model that generates language. For example, the model that generates images may include a generative adversarial network (GAN) and / or a variational autoencoder (VAE). For example, the model that generates images may include a diffusion-based generative model (e.g., a transformer VAE). For example, the model that generates language may include CHAT-GPT 3 and / or CHAT-GPT 4. For example, the generative AI model (1230) may include an LMM that generates the feedback by recognizing characters, images, and / or voice.

[0181] As an example without limitation, the AI ​​framework (1220) and / or generative AI model (1230) may be included within an AI module (e.g., including a processing circuit) within the electronic device (200). For example, the AI ​​module may be operatively coupled with at least one processor of the electronic device (200) (e.g., at least one processor (210) or processor (1120)). For example, the AI ​​module may be operatively coupled with a display driving circuit of the electronic device. For example, the AI ​​module may be operatively coupled with a sensor hub of the electronic device for one or more sensors within the electronic device.

[0182] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure pertains.

[0183] An electronic device as described above (e.g., the electronic device (200) of FIG. 2) may include at least one processor (e.g., at least one processor (210) of FIG. 2) comprising a processing circuit, a display (e.g., the display (230) of FIG. 2), and a memory (e.g., the memory (220) of FIG. 2) comprising one or more storage media configured to store one or more programs configured to be executed individually or collectively by said at least one processor. The one or more programs may include instructions that cause said electronic device to identify an input for acquiring a second image (e.g., the second image (603) of FIG. 6a) using a first image (e.g., the first image (505) of FIG. 5). The one or more programs may include instructions that cause said electronic device to provide said first image to a trained model based on said input. The above one or more programs may include instructions that cause the electronic device to obtain information about the first image generated by the trained model using the first image. The above one or more programs may include instructions that cause the electronic device to obtain a keyword included in the information and options regarding the keyword. The above one or more programs may include instructions that cause the electronic device to provide the trained model with a prompt containing the information. The above one or more programs may include instructions that cause the electronic device to obtain the second image generated by the trained model using the prompt and including a visual object representing the keyword.The above one or more programs may include instructions that cause the electronic device to display, through the display, user interface (UI) objects (e.g., UI objects (635) of FIG. 6a) indicating the second image and the options, respectively. The above one or more programs may include instructions that cause the electronic device to receive user input (e.g., user input (910) of FIG. 9a) for at least one UI object among the UI objects (e.g., at least one UI object (905-1) of FIG. 9a). The above one or more programs may include instructions that cause the electronic device to display, through the display, a third image (e.g., third image (945) of FIG. 9a) containing another visual object (e.g., visual object (950) of FIG. 9a) representing a keyword having an option indicated by the at least one UI object, based on the user input.

[0184] For example, the input may include handwritten input received through the display. The first image may include at least one stroke identified according to the handwritten input.

[0185] For example, the above input may include an input that determines the first image among the images stored in the electronic device as the image to be used to acquire the second image.

[0186] For example, the one or more programs may include instructions that cause the electronic device to provide the trained model with the first image another prompt requesting a description of the first image. The one or more programs may include instructions that cause the electronic device to obtain the information generated by the trained model using the first image and the other prompt.

[0187] For example, the one or more programs may include instructions that cause the electronic device to provide the trained model with the first image another prompt requesting a description of the first image and a keyword included in the description. The one or more programs may include instructions that cause the electronic device to obtain the information generated by the trained model using the first image and the other prompt, the keyword included in the information, and the options.

[0188] For example, the one or more programs may include instructions that cause the electronic device to generate another prompt including the information and other information regarding the option indicated by the at least one UI object based on the user input. The one or more programs may include instructions that cause the electronic device to provide the other prompt to the trained model. The one or more programs may include instructions that cause the electronic device to acquire the third image generated by the trained model using the other prompt. The one or more programs may include instructions that cause the electronic device to display the third image through the display.

[0189] For example, the above one or more programs may include instructions that cause the electronic device to provide the first image together with the other prompt to the trained model. The above one or more programs may include instructions that cause the electronic device to obtain the third image, which is generated by the trained model using the first image and the other prompt and includes the other visual object having a shape corresponding to the visual object included in the second image and expressing the keyword having the option. The above one or more programs may include instructions that cause the electronic device to display the third image through the display.

[0190] For example, the one or more programs may include instructions that cause the electronic device to receive other user input for the style of the second image to be generated. The one or more programs may include instructions that cause the electronic device to generate the prompt, which further includes other information about the style, based on the other user input. The one or more programs may include instructions that cause the electronic device to provide the prompt to the trained model. The one or more programs may include instructions that cause the electronic device to acquire the second image, which is generated by the trained model using the prompt, includes the visual object, and has the style.

[0191] For example, the one or more programs may include instructions that cause the electronic device to receive text input. The one or more programs may include instructions that cause the electronic device to provide the trained model with the prompt, which further includes other information about the text identified according to the text input, based on the text input. The one or more programs may include instructions that cause the electronic device to acquire the second image generated by the trained model using the prompt.

[0192] For example, the one or more programs may include instructions that cause the electronic device to receive the handwriting input while the fourth image is displayed through the display. The one or more programs may include instructions that cause the electronic device to acquire the first image, which includes the at least one stroke identified according to the handwriting input and the fourth image, based on the handwriting input received while the fourth image is displayed. The one or more programs may include instructions that cause the electronic device to provide the first image to the trained model.

[0193] For example, the trained model may include a large language model (LLM) and a model for image generation. The one or more programs may include instructions that cause the electronic device to provide the first image to the LLM. The one or more programs may include instructions that cause the electronic device to obtain the information generated by the LLM using the first image. The one or more programs may include instructions that cause the electronic device to obtain the keyword included in the information and the options regarding the keyword. The one or more programs may include instructions that cause the electronic device to provide the prompt to the model for image generation. The one or more programs may include instructions that cause the electronic device to obtain the second image generated by the model for image generation using the prompt.

[0194] For example, the one or more programs may include instructions that cause the electronic device to display the first image through the display. The one or more programs may include instructions that cause the electronic device to receive other user input for generating the second image while displaying the first image. The one or more programs may include instructions that cause the electronic device to provide the first image to the trained model based on the other user input.

[0195] For example, the one or more programs may include instructions that cause the electronic device to receive other user input for changing the appearance of the visual object included in the second image while displaying the second image. The one or more programs may include instructions that cause the electronic device to display the UI objects based on the other user input.

[0196] For example, the one or more programs may include instructions that cause the electronic device to display the second image, other information indicating the keyword, and the UI objects through the display. The UI objects may be displayed in conjunction with the other information.

[0197] For example, the other information mentioned above may be displayed in conjunction with a part of the visual object corresponding to the keyword in the second image.

[0198] For example, the one or more programs may include instructions that cause the electronic device to receive other user input regarding the other information. The one or more programs may include instructions that cause the electronic device to display the UI objects in conjunction with the other information through the display based on the other user input.

[0199] For example, the above one or more programs may include instructions that cause the electronic device to display the second image, the UI objects, and a text input field through the display. The above one or more programs may include instructions that cause the electronic device to receive text input through the text input field. The above one or more programs may include instructions that cause the electronic device to receive user input for at least one UI object among the UI objects. The above one or more programs may include instructions that cause the electronic device to display the third image through the display, based on the text input and the user input, the third image including the other visual object expressing the keyword having the option indicated by the at least one UI object and text corresponding to the text input.

[0200] For example, the one or more programs may include instructions that cause the electronic device to generate another prompt including the text, the information, and other information regarding the option indicated by the at least one UI object, based on the text input and the user input. The one or more programs may include instructions that cause the electronic device to provide the other prompt to the trained model. The one or more programs may include instructions that cause the electronic device to acquire the third image generated by the trained model using the other prompt. The one or more programs may include instructions that cause the electronic device to display the third image through the display.

[0201] An electronic device as described above may include at least one processor comprising a processing circuit, a display, and a memory comprising one or more storage media configured to store one or more programs configured to be executed individually or collectively by said at least one processor. The one or more programs may include instructions that cause said electronic device to receive a text input for acquiring a first image. The one or more programs may include instructions that cause said electronic device to acquire a keyword contained in text identified according to said text input and options regarding said keyword. The one or more programs may include instructions that cause said electronic device to provide a prompt containing said text to a trained model. The one or more programs may include instructions that cause said electronic device to acquire a first image including a visual object representing said keyword, which is generated by said trained model using said prompt. The above one or more programs may include instructions that cause the electronic device to display UI objects indicating the first image and the options, respectively, through the display. The above one or more programs may include instructions that cause the electronic device to receive user input for at least one UI object among the UI objects. The above one or more programs may include instructions that cause the electronic device to display, through the display, a second image including another visual object expressing a keyword having an option indicated by the at least one UI object, based on the user input.

[0202] For example, the one or more programs may include instructions that cause the electronic device to receive other user input for changing the appearance of the visual object included in the first image while displaying the first image. The one or more programs may include instructions that cause the electronic device to display the UI objects based on the other user input.

[0203] For example, the one or more programs may include instructions that cause the electronic device to receive other user input for the style of the first image to be generated. The one or more programs may include instructions that cause the electronic device to generate the prompt, which further includes information about the style, based on the other user input. The one or more programs may include instructions that cause the electronic device to provide the prompt to the trained model. The one or more programs may include instructions that cause the electronic device to acquire the first image, which is generated by the trained model using the prompt, includes the visual object, and has the style.

[0204] For example, the one or more programs may include instructions that cause the electronic device to generate another prompt containing information about the option indicated by the text and the at least one UI object based on the user input. The one or more programs may include instructions that cause the electronic device to provide the other prompt to the trained model. The one or more programs may include instructions that cause the electronic device to acquire the second image generated by the trained model using the other prompt. The one or more programs may include instructions that cause the electronic device to display the second image through the display.

[0205] The method described above may be performed within an electronic device including a display. The method may include an operation of identifying an input for acquiring a second image using a first image. The method may include an operation of providing the first image to a trained model based on the input. The method may include an operation of acquiring information about the first image generated by the trained model using the first image. The method may include an operation of acquiring a keyword and options regarding the keyword included in the information. The method may include an operation of providing a prompt containing the information to the trained model. The method may include an operation of acquiring the second image, which is generated by the trained model using the prompt and includes a visual object representing the keyword. The method may include an operation of displaying user interface (UI) objects that indicate the second image and the options, respectively, through the display. The method may include an operation of receiving user input for at least one of the UI objects. The above method may include the operation of displaying a third image, which includes another visual object expressing a keyword having an option indicated by at least one UI object, through the display based on the user input.

[0206] For example, the input may include handwritten input received through the display. The first image may include at least one stroke identified according to the handwritten input.

[0207] For example, the above input may include an input that determines the first image among the images stored in the electronic device as the image to be used to acquire the second image.

[0208] For example, the above method may include an action of providing another prompt requesting a description of the first image to the trained model along with the first image. The above method may include an action of obtaining the information generated by the trained model using the first image and the other prompt.

[0209] For example, the above method may include an operation of providing the trained model, together with the first image, another prompt requesting a description of the first image and a keyword included in the description. The above method may include an operation of obtaining the information generated by the trained model using the first image and the other prompt, the keyword included in the information, and the options.

[0210] For example, the above method may include an action of generating another prompt based on the user input, the other prompt including the information and other information regarding the option indicated by the at least one UI object. The above method may include an action of providing the other prompt to the trained model. The above method may include an action of obtaining the third image generated by the trained model using the other prompt. The above method may include an action of displaying the third image through the display.

[0211] For example, the above method may include the action of providing the first image together with the other prompt to the trained model. The above method may include the action of obtaining the third image, which is generated by the trained model using the first image and the other prompt, and includes the other visual object that has a shape corresponding to the visual object included in the second image and expresses the keyword having the option. The above method may include the action of displaying the third image through the display.

[0212] For example, the method may include receiving other user input for the style of the second image to be generated. The method may include generating the prompt, which further includes other information regarding the style, based on the other user input. The method may include providing the prompt to the trained model. The method may include acquiring the second image, which is generated by the trained model using the prompt, includes the visual object, and has the style.

[0213] For example, the above method may include an operation of receiving text input. The above method may include an operation of providing the trained model with the prompt, which further includes other information regarding the text identified according to the text input, based on the text input. The above method may include an operation of acquiring the second image generated by the trained model using the prompt.

[0214] For example, the above method may include an operation of receiving the handwriting input while the fourth image is displayed through the display. The above method may include an operation of acquiring the first image including the at least one stroke identified according to the handwriting input and the fourth image, based on the handwriting input received while the fourth image is displayed. The above method may include an operation of providing the first image to the trained model.

[0215] For example, the trained model may include a large language model (LLM) and a model for image generation. The method may include an operation of providing the first image to the LLM. The method may include an operation of obtaining the information generated by the LLM using the first image. The method may include an operation of obtaining the keyword included in the information and the options regarding the keyword. The method may include an operation of providing the prompt to the model for image generation. The method may include an operation of obtaining the second image generated by the model for image generation using the prompt.

[0216] For example, the above method may include an operation of displaying the first image through the display. The above method may include an operation of receiving another user input for generating the second image while displaying the first image. The above method may include an operation of providing the first image to the trained model based on the other user input.

[0217] For example, the method may include receiving other user input to change the appearance of the visual object included in the second image while displaying the second image. The method may include displaying the UI objects based on the other user input.

[0218] For example, the above method may include an operation of displaying the second image, other information indicating the keyword, and the UI objects through the display. The UI objects may be displayed in conjunction with the other information.

[0219] For example, the other information mentioned above may be displayed in conjunction with a part of the visual object corresponding to the keyword in the second image.

[0220] For example, the above method may include an operation of receiving other user input regarding the other information. The above method may include an operation of displaying the UI objects in conjunction with the other information through the display based on the other user input.

[0221] For example, the above method may include an operation of displaying the second image, the UI objects, and a text input field through the display. The above method may include an operation of receiving text input through the text input field. The above method may include an operation of receiving user input for at least one UI object among the UI objects. The above method may include an operation of displaying the third image through the display, based on the text input and the user input, the third image including the other visual object expressing the keyword having the option indicated by the at least one UI object and text corresponding to the text input.

[0222] For example, the above method may include an action of generating another prompt based on the text input and the user input, the other prompt including the text, the information, and other information regarding the option indicated by the at least one UI object. The above method may include an action of providing the other prompt to the trained model. The above method may include an action of obtaining the third image generated by the trained model using the other prompt. The above method may include an action of displaying the third image through the display.

[0223] The method described above may be performed within an electronic device including a display. The method may include an operation of receiving text input for acquiring a first image. The method may include an operation of acquiring a keyword included in text identified according to the text input and options regarding said keyword. The method may include an operation of providing a prompt containing said text to the trained model. The method may include an operation of acquiring a first image including a visual object representing said keyword, which is generated by the trained model using said prompt. The method may include an operation of displaying UI objects indicating said first image and said options, respectively, through the display. The method may include an operation of receiving user input for at least one UI object among said UI objects. The method may include an operation of displaying a second image through the display, based on said user input, which includes another visual object representing a keyword having an option indicated by said at least one UI object.

[0224] For example, the above method may include an operation of receiving another user input to change the appearance of the visual object included in the first image while displaying the first image. The above method may include an operation of displaying the UI objects based on the other user input.

[0225] For example, the method may include an action of receiving another user input for the style of the first image to be generated. The method may include an action of generating the prompt, which further includes information about the style, based on the other user input. The method may include an action of providing the prompt to the trained model. The method may include an action of acquiring the first image, which is generated by the trained model using the prompt, includes the visual object, and has the style.

[0226] For example, the method may include an action of generating another prompt containing information about the option indicated by the text and at least one UI object based on the user input. The method may include an action of providing the other prompt to the trained model. The method may include an action of obtaining the second image generated by the trained model using the other prompt. The method may include an action of displaying the second image (945) through the display (230).

[0227] A non-transient computer-readable storage medium as described above may store one or more programs. The one or more programs may include instructions that cause the electronic device to identify an input for acquiring a second image using a first image when executed by the electronic device having a display. The one or more programs may include instructions that cause the electronic device to provide the first image to a trained model based on the input when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to acquire information about the first image generated by the trained model using the first image when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to acquire a keyword and options regarding the keyword when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to provide the trained model with a prompt containing the information when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to acquire the second image, which is generated by the trained model using the prompt and includes a visual object representing the keyword, when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to display user interface (UI) objects indicating the second image and the options, respectively, through the display when executed by the electronic device.The above one or more programs may include instructions that cause the electronic device to receive user input for at least one UI object among the UI objects when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to display, through the display, a third image including another visual object expressing a keyword having an option indicated by the at least one UI object, based on the user input when executed by the electronic device.

[0228] For example, the input may include handwritten input received through the display. The first image may include at least one stroke identified according to the handwritten input.

[0229] For example, the above input may include an input that determines the first image among the images stored in the electronic device as the image to be used to acquire the second image.

[0230] For example, the one or more programs may include instructions that cause the electronic device to provide the trained model, along with the first image, another prompt requesting a description of the first image when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain the information generated by the trained model using the first image and the other prompt when executed by the electronic device.

[0231] For example, the one or more programs may include instructions that cause the electronic device to provide the trained model, together with the first image, another prompt requesting a description of the first image and a keyword included in the description, when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain the information generated by the trained model using the first image and the other prompt, the keyword included in the information, and the options when executed by the electronic device.

[0232] For example, the one or more programs may include instructions that cause the electronic device to generate another prompt, which includes the information and other information regarding the option indicated by the at least one UI object, based on the user input when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to provide the other prompt to the trained model when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to acquire the third image generated by the trained model using the other prompt when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to display the third image through the display when executed by the electronic device.

[0233] For example, the above one or more programs may include instructions that cause the electronic device to provide the first image along with the other prompt to the trained model when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to obtain the third image, which is generated by the trained model using the first image and the other prompt when executed by the electronic device and includes the other visual object having a shape corresponding to the visual object included in the second image and expressing the keyword having the option. The above one or more programs may include instructions that cause the electronic device to display the third image through the display when executed by the electronic device.

[0234] For example, the one or more programs may include instructions that cause the electronic device to receive other user input for the style of the second image to be generated when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to generate the prompt, which further includes other information about the style, based on the other user input, when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to provide the prompt to the trained model when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain the second image, which is generated by the trained model using the prompt, includes the visual object, and has the style, when executed by the electronic device.

[0235] For example, the one or more programs may include instructions that cause the electronic device to receive text input when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to provide the trained model with the prompt, which further includes other information regarding the text identified according to the text input, based on the text input when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to acquire the second image generated by the trained model using the prompt when executed by the electronic device.

[0236] For example, the one or more programs may include instructions that cause the electronic device to receive the hand input while the fourth image is displayed through the display when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to acquire the first image, which includes the at least one stroke identified according to the hand input and the fourth image, based on the hand input received while the fourth image is displayed when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to provide the first image to the trained model when executed by the electronic device.

[0237] For example, the trained model may include a large language model (LM) and a model for image generation. The one or more programs may include instructions that cause the electronic device to provide the first image to the LLM when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain the information generated by the LLM using the first image when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain the keyword included in the information and the options regarding the keyword when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to provide the prompt to the model for image generation when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain the second image generated by the model for image generation using the prompt when executed by the electronic device.

[0238] For example, the one or more programs may include instructions that cause the electronic device to display the first image through the display when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to receive other user input to generate the second image while displaying the first image when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to provide the first image to the trained model based on the other user input when executed by the electronic device.

[0239] For example, the one or more programs may include instructions that cause the electronic device to receive other user input to change the appearance of the visual object included in the second image while displaying the second image when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to display the UI objects based on the other user input when executed by the electronic device.

[0240] For example, the one or more programs may include instructions that cause the electronic device to display the second image, other information indicating the keyword, and the UI objects through the display when executed by the electronic device. The UI objects may be displayed in conjunction with the other information.

[0241] For example, the other information mentioned above may be displayed in conjunction with a part of the visual object corresponding to the keyword in the second image.

[0242] For example, the one or more programs may include instructions that cause the electronic device to receive other user input regarding the other information when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to display the UI objects in conjunction with the other information through the display based on the other user input when executed by the electronic device.

[0243] For example, the above one or more programs may include instructions that cause the electronic device to display the second image, the UI objects, and a text input field through the display when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to receive text input through the text input field when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to receive user input for at least one UI object among the UI objects when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to display the third image through the display, based on the text input and the user input, the third image including the other visual object expressing the keyword having the option indicated by the at least one UI object and text corresponding to the text input when executed by the electronic device.

[0244] For example, the one or more programs may include instructions that cause the electronic device to generate another prompt including the text, the information, and other information regarding the option indicated by the at least one UI object, based on the text input and the user input. The one or more programs may include instructions that cause the electronic device to provide the other prompt to the trained model. The one or more programs may include instructions that cause the electronic device to acquire the third image generated by the trained model using the other prompt. The one or more programs may include instructions that cause the electronic device to display the third image through the display.

[0245] A non-transient computer-readable storage medium as described above may store one or more programs. The one or more programs may include instructions that cause the electronic device to receive text input for acquiring a first image when executed by the electronic device having a display. The one or more programs may include instructions that cause the electronic device to acquire a keyword contained within text identified according to the text input and options regarding said keyword when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to provide a prompt containing said text to the trained model when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to acquire a first image that includes a visual object representing said keyword, which is generated by the trained model using said prompt when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to display UI objects indicating the first image and the options, respectively, through the display when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to receive user input for at least one UI object among the UI objects when executed by the electronic device. The above one or more programs may include instructions that cause the electronic device to display a second image, including another visual object expressing an option indicated by the at least one UI object, through the display based on the user input when executed by the electronic device.

[0246] For example, the one or more programs may include instructions that cause the electronic device to receive other user input to change the appearance of the visual object included in the first image while displaying the first image when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to display the UI objects based on the other user input when executed by the electronic device.

[0247] For example, the one or more programs may include instructions that cause the electronic device to receive other user input for the style of the first image to be generated when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to generate the prompt, which further includes information about the style, based on the other user input when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to provide the prompt to the trained model when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to obtain the first image, which is generated by the trained model using the prompt, includes the visual object, and has the style, when executed by the electronic device.

[0248] For example, the one or more programs may include instructions that cause the electronic device to generate another prompt, which includes information about the option indicated by the text and the at least one UI object, based on the user input when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to provide the other prompt to the trained model when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to acquire the second image generated by the trained model using the other prompt when executed by the electronic device. The one or more programs may include instructions that cause the electronic device to display the second image through the display when executed by the electronic device.

[0249] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs.

Claims

1. In an electronic device, At least one processor including a processing circuit; Display; and Memory comprising one or more programs configured to be executed individually or collectively by at least one processor, and including one or more storage media, The above one or more programs are: Identifying an input for acquiring a second image using a first image; Based on the above input, the first image is provided to the trained model; Acquire information about the first image generated using the first image by the above-mentioned trained model; Obtain keywords included in the above information and options regarding said keywords; Provide a prompt containing the above information to the trained model; Acquiring the second image, which includes a first visual object expressing the keyword, generated by the above-trained model using the above-trained prompt; Through the above display, UI (user interface) objects indicating the second image and the options, respectively are displayed; Receiving user input for at least one UI object among the above UI objects; and Based on the above user input, to display a third image including a second visual object expressing a keyword having an option indicated by the at least one UI object through the display, Instructions including those that cause the above electronic device Electronic device.

2. In Claim 1, The above input is, It includes handwriting input received through the above display, and The first image above is, including at least one stroke identified according to the above handwritten input, Electronic device.

3. In Claim 1, The above input is, Includes an input for determining, among the images stored in the electronic device, the first image as the image to be used to acquire the second image. Electronic device.

4. In Claim 1, The above one or more programs are: Providing another prompt requesting a description of the first image to the trained model along with the first image; and To obtain the information generated by the above-mentioned trained model using the above-mentioned first image and the above-mentioned other prompt, Instructions including those that cause the above electronic device Electronic device.

5. In Claim 1, The above one or more programs are: Providing the trained model with the first image another prompt requesting a description of the first image and keywords included in the description; and To obtain the information generated by the above-trained model using the above-trained image and the above-trained other prompt, the above-mentioned keywords included in the information, and the above-mentioned options, Instructions including those that cause the above electronic device Electronic device.

6. In Claim 1, The above one or more programs are: Based on the above user input, generate another prompt including the above information and other information regarding the above option indicated by the above at least one UI object; Providing the above other prompts to the above-mentioned trained model; Acquire the third image generated using the other prompt by the above-mentioned trained model; and To display the third image through the above display, Instructions including those that cause the above electronic device Electronic device.

7. In Claim 6, The above one or more programs are: Providing the above first image to the trained model along with the above other prompt; A third image is obtained that includes a second visual object expressing the keyword having the option, which is generated by the above-trained model using the first image and the other prompt, and the shape of the second visual object corresponds to the shape of the first visual object included in the second image; and To display the third image through the above display, Instructions including those that cause the above electronic device Electronic device.

8. In Claim 1, The above one or more programs are: Receiving other user input for the style of the second image to be generated above; Based on the above other user input, generate the above prompt including additional information about the above style; Providing the above prompt to the above-mentioned trained model; and To obtain the second image generated by the above-trained model using the above-trained prompt, including the first visual object, and having the above-trained style, Instructions including those that cause the above electronic device Electronic device.

9. In Claim 1, The above one or more programs are: Receive text input; Based on the above text input, the training model is provided with the above prompt, which further includes other information regarding the text identified according to the above text input; and To acquire the second image generated by the above-mentioned trained model using the above-mentioned prompt; Instructions including those that cause the above electronic device Electronic device.

10. In Claim 1, The above one or more programs are: While the fourth image is displayed through the above display, the above handwritten input is received; Based on the handwriting input received while the fourth image is displayed, the first image including the at least one stroke identified according to the handwriting input and the fourth image is obtained; and To provide the above first image to the above-trained model, Instructions including those that cause the above electronic device Electronic device.

11. In Claim 1, The above-mentioned trained model is, Includes a large language model (LM) and a model for image generation; The above one or more programs are: Providing the above first image to the LLM; Acquire the information generated using the first image by the above LLM; Obtaining the above keywords included in the above information and the above options regarding the above keywords; Providing the above prompt to the model for generating the above image; and To obtain the second image generated using the prompt by the model for generating the image above, Instructions including those that cause the above electronic device Electronic device.

12. In Claim 1, The above one or more programs are: Displaying the first image through the above display; Receiving other user input to generate the second image while displaying the first image; and Based on the other user input above, to provide the first image to the trained model, Instructions including those that cause the above electronic device Electronic device.

13. In Claim 1, The above one or more programs are: Receiving other user input to change the appearance of the first visual object included in the second image while displaying the second image; and To display the UI objects based on the other user input mentioned above, Instructions including those that cause the above electronic device Electronic device.

14. In Claim 1, The above one or more programs are: To display the second image, other information indicating the keyword, and the UI objects through the above display, Includes instructions that cause the above electronic device, The above UI objects are displayed in conjunction with the above other information, Electronic device.

15. In Claim 1, The above one or more programs are: Displaying the second image, the UI objects, and a text input field through the above display; Receive text input through the above text input field; Receiving user input for at least one UI object among the above UI objects; and Based on the text input and the user input, to display the third image including the second visual object expressing the keyword having the option indicated by the text corresponding to the text input and the option indicated by the at least one UI object, through the display. Instructions including those that cause the above electronic device Electronic device.