Electronic device, and method for generating image corresponding to drag input of user of electronic device

The electronic device facilitates image generation by allowing users to drag words between text and image windows, addressing the challenge of conveying complex descriptions and ensuring accurate image representation.

WO2026084386A1PCT designated stage Publication Date: 2026-04-23SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2025-10-13
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Users face difficulties in accurately conveying complex image descriptions using text prompts due to challenges in selecting appropriate vocabulary and expressing object relationships, leading to discrepancies between generated images and desired outcomes.

Method used

An electronic device with a display and processor that allows users to drag words between text and image windows, generating prompts based on the distance and position of word inputs to create images reflecting the intended linguistic context.

Benefits of technology

Enables users to generate images that accurately reflect their intended context without complex textual descriptions, improving the alignment between user intent and generated output.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025016028_23042026_PF_FP_ABST
    Figure KR2025016028_23042026_PF_FP_ABST
Patent Text Reader

Abstract

In an electronic device and an operating method of the electronic device, according to an embodiment, the electronic device may display, on a display, a text window that receives a plurality of word inputs including a first word and a second word, and an image display window displayed separately from the text window. The electronic device may receive a first input of dragging the first word to the image display window. The electronic device may generate, in response to the first input, a first prompt instructing to generate a first image including a first object corresponding to the first word. The electronic device may generate the first image on the basis of the first prompt. The electronic device may receive a second input of dragging the second word to the image display window. The electronic device may generate, in response to the second input, a second prompt instructing to generate a second image including a second object and the first object on the basis of the first word, the second word, and a distance between a first position to which the first word has been dragged according to the first input and a second position to which the second word has been dragged according to the second input. The electronic device may generate the second image on the basis of the second prompt.
Need to check novelty before this filing date? Find Prior Art

Description

electronic device and a method for generating an image corresponding to the user's drag input of the electronic device

[0001] The present disclosure relates to an electronic device and a method of operating the electronic device, and more specifically to a technique for generating an image corresponding to a user's drag input.

[0002] With the development of digital technology, various types of electronic devices such as smartphones, digital cameras, and / or wearable devices are widely used. To support and enhance the functionality of these electronic devices, the hardware and / or software parts of the devices are continuously being developed.

[0003] For example, portable electronic devices (hereinafter referred to as "electronic devices"), such as smartphones, have become capable of incorporating various functions. Electronic devices include touchscreen-based displays to allow users easy access to various functions, and can provide screens for various applications through the display.

[0004] Recently, with the rapid advancement of big data and deep learning technologies, artificial intelligence (AI) technology has been implemented in electronic devices and is also being applied to intelligent personalized services that analyze specific data and integrate and utilize information from various fields tailored to the user. For example, users can control electronic devices through voice conversation and perform searches, queries, and responses regarding specific information using a knowledge base powered by deep learning. Furthermore, with the evolution of AI technology, generative AI is being implemented. Generative AI refers to AI technology that creates new, similar content using existing content such as text, audio, and / or images. For instance, generative AI can represent AI technology capable of generating content (e.g., text, audio, images, and / or videos) corresponding to a given input.

[0005] The information described above may be provided as related art for the purpose of aiding understanding of this document. None of the above is to be claimed as prior art related to this document, nor can it be used to determine prior art.

[0006] Users of electronic devices utilizing image generation AI may need to clearly input detailed information about the image they wish to generate via text prompts. However, users may experience difficulty in selecting vocabulary or expressions that accurately describe the desired image. Users may also struggle to describe the arrangement or relationships between various objects in text. For instance, while a user may input multiple words, it may be difficult to clearly convey the specific linguistic context between them. Furthermore, if a user inputs unclear prompts, a discrepancy may arise between the image generated by the electronic device and the image desired by the user. For instance, when attempting to implement complex scenes or interactions between objects, the electronic device may fail to generate the desired image because it cannot produce prompts that reflect the user's intent based solely on text input.

[0007] The technical problems to be solved in this document are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this invention belongs from the description below.

[0008] According to one embodiment, the electronic device may include a display. The electronic device may include a memory that stores at least one computer program containing instructions. The electronic device may include at least one processor. When the instructions are executed individually or collectively by the at least one processor, the electronic device may display on the display a text window that receives a plurality of word inputs including a first word and a second word, and an image display window that is displayed separately from the text window. The instructions may receive a first input that drags the first word to the image display window. The electronic device may generate a first prompt that instructs to create a first image containing a first object corresponding to the first word in response to the first input. The instructions may create the first image based on the first prompt. The electronic device may receive a second input that drags the second word to the image display window. The electronic device may generate a second prompt that instructs the generation of a second image including the second object and the first object, based on the distance between the first word, the second word, and the first position where the first word is dragged according to the first input and the second position where the second word is dragged according to the second input, in response to the second input. The electronic device may generate the second image based on the second prompt.

[0009] A method of operation of an electronic device according to one embodiment may include an operation of displaying on a display a text window that receives a plurality of word inputs including a first word and a second word, and an image display window that is displayed separately from the text window. A method of operation of an electronic device may include an operation of receiving a first input that drags the first word to the image display window. A method of operation of an electronic device may include an operation of generating a first prompt that instructs to generate a first image including a first object corresponding to the first word in response to the first input. A method of operation of an electronic device may include an operation of generating a first image based on the first prompt. A method of operation of an electronic device may include an operation of receiving a second input that drags the second word to the image display window. A method of operation of an electronic device may include an operation of generating a second prompt that instructs to generate a second image including the second object and the first object in response to the second input, based on the distance between the first word, the second word, a first position where the first word was dragged according to the first input, and a second position where the second word was dragged according to the second input. The method of operation of the electronic device may include the operation of generating a second image based on the second prompt.

[0010] An electronic device can receive user input in which the user enters a word and drags the entered word. The electronic device can determine the distance between the locations where multiple words are dragged. Based on the determined distance, the electronic device can determine the linguistic context between the words and generate a prompt instructing the user to create (or modify) an image based on the determined linguistic context. The user can generate an image that reflects the linguistic context between the words without having to create a complex prompt.

[0011] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.

[0012] FIG. 1 is a block diagram of an exemplary electronic device capable of performing the operations described in this document.

[0013] Figure 2 is a schematic diagram of an exemplary artificial intelligence system.

[0014] FIG. 3 is a block diagram of an electronic device according to one embodiment.

[0015] FIG. 4 is a drawing illustrating a display of an electronic device according to one embodiment.

[0016] FIG. 5 is a drawing illustrating an electronic device that generates an image in response to user input according to one embodiment.

[0017] FIG. 6 is a drawing illustrating a display of an electronic device according to one embodiment.

[0018] FIGS. 7A, FIGS. 7B, FIGS. 7C, FIGS. 7D, and FIGS. 7E are drawings illustrating a display of an electronic device that displays an image on a display in response to user input according to one embodiment.

[0019] FIG. 8 is a drawing illustrating a display of an electronic device that receives user input for dragging an object into a text window or user input for modifying an object, according to one embodiment.

[0020] FIGS. 9A, FIGS. 9B, FIGS. 9C, FIGS. 9D, and FIGS. 9E are drawings illustrating a display of an electronic device that displays an image on a display in response to user input according to one embodiment.

[0021] FIGS. 10a, FIGS. 10b, FIGS. 10c, and FIGS. 10d are drawings illustrating an electronic device that generates different images depending on the position where a word is dragged according to one embodiment.

[0022] FIG. 11 is a method flowchart of an electronic device according to one embodiment.

[0023] FIG. 1 is a block diagram of an exemplary electronic device (100) capable of performing the operations described in this document.

[0024] Referring to FIG. 1, the electronic device (100) may be one of various forms of electronic devices, such as a notebook (190), smartphones (191) having various form factors (e.g., a bar-type smartphone (191-1), a foldable-type smartphone (191-2), or a sliderable (or rollable)-type smartphone (191-3)), a tablet (192), a cellular phone (not shown), and other similar computing devices (not shown). The components, their relationships, and their functions illustrated in FIG. 1 are illustrative only and are not intended to limit the implementations described or claimed herein. The electronic device (100) may be referred to as a mobile device, a user device, a multifunction device, a portable device, or a server.

[0025] The electronic device (100) may include components comprising at least one processor (110) (hereinafter referred to as processor (110)), at least one memory (120) (hereinafter referred to as memory (120)), at least one display (140) (hereinafter referred to as display (140)), at least one image sensor (150) (hereinafter referred to as image sensor (150)), at least one communication circuit (160) (hereinafter referred to as communication circuit (160)), and / or at least one sensor (170) (hereinafter referred to as sensor (170)). The components are merely exemplary. For example, the electronic device (100) may include other components (e.g., power management integrated circuitry (PMIC), audio processing circuit, antenna, rechargeable battery, or input / output interface). For example, some components may be omitted from the electronic device (100). For example, some components may be integrated into a single component.

[0026] The processor (110) may be implemented as one or more IC (integrated circuit (or circuitry)) chips and may perform various data processing operations. The processor (110) may include at least one electrical circuit and may process instructions (or programs, data, etc.) stored in memory (120) individually or collectively in a distributed manner. The processor (110) may include a processor assembly comprising one or more processing circuits. The processor (110) may include any processing circuit that is operative to control the performance and operations of one or more components of the electronic device (100) (e.g., memory (120), display (140), image sensor (150), communication circuit (160), and / or sensor (170)). For example, the processor (110) (e.g., application processor (AP)) may be implemented as a system on chip (SoC) (e.g., a single chip or chipset). For example, the processor (110) may be implemented with a plurality of cores (or at least one core circuit), a plurality of chips, or a plurality of chipsets. For example, the processor (110) may include one or more processing circuits. For example, the processor (110) may include one or more processing circuits configured to perform the various functions of the present disclosure individually and / or collectively. As an example without limitation, at least a portion of the processor (110) may be included in a first chip of the electronic device (100), and at least another portion of the processor (110) may be included in a second chip of the electronic device (100) different from the first chip of the electronic device (100).

[0027] For example, the processor (110) may include a central processing unit (111), a graphics processing unit (112), a neural processing unit (113), an image signal processor (114), a display controller (115), a memory controller (116), a storage controller (117), a communication processor (118), and / or a sensor interface (119). These components of the processor (110) are merely exemplary. For example, the processor (110) may include other components. For example, some components of the processor (110) may be omitted from the processor (110). For example, some components of the processor (110) may be included as separate components of the electronic device (100) outside of the processor (110). For example, some components of the processor (110) (e.g., memory controller (116)) may be included in other components (e.g., at least part of memory (120), an interface (e.g. available for connection to at least one component of the electronic device (100)), a display (140) and / or an image sensor (150)).

[0028] The processor (110) may cause other components of the electronic device (100) to perform various operations by executing instructions stored in memory (120). The CPU (111) (or central processing circuit) may be configured to control the components of the processor (110) based on the execution of instructions stored in memory (120) (e.g., volatile memory (121) and / or non-volatile memory (122)). The GPU (112) (or graphics processing circuit) may be configured to execute parallel operations (e.g., rendering). The NPU (113) (or neural processing circuit, or AI (artificial intelligence) chip) may be configured to execute operations for an artificial intelligence model (e.g., convolution computation). An ISP (114) (or image signal processing circuit) may be configured to process a raw image acquired through an image sensor (150) into a format suitable for a component within the electronic device (100) or a component of the processor (110). A display controller (115) (or display control circuit, or DPU (display processing unit)) may be configured to process an image acquired from a CPU (111), GPU (112), ISP (114), or memory (120) (e.g., volatile memory (121)) into a format suitable for a display (140). A memory controller (116) (or memory control circuit) may be configured to control reading data from the volatile memory (121) and writing data to the volatile memory (121). A storage controller (117) (or storage control circuit) may be configured to control reading data from the non-volatile memory (122) and writing data to the non-volatile memory (122).The CP (118) (communication processing circuit) may be configured to process data obtained from a component of the processor (110) into a format suitable for transmitting to another electronic device via the communication circuit (160), or to process data obtained from another electronic device via the communication circuit (160) into a format suitable for processing by the component of the processor (110). For example, the communication circuit (160) may include one or more communication circuits. The sensor interface (119) (or sensing data processing circuit, sensor hub) may be configured to process data regarding the state of the electronic device (100) and / or the state around the electronic device (100), obtained through the sensor (170), into a format suitable for the component of the processor (110).

[0029] Memory (120) may include one or more storage media (or one or more storage devices). For example, memory (120) may include a memory assembly comprising one or more storage media. For example, the one or more storage media may include a hard drive, a permanent memory such as flash memory, read-only memory (ROM) (e.g., non-volatile memory (122)), a semi-permanent memory such as random access memory (RAM) (e.g., volatile memory (121)), any other suitable type of storage (or storage assembly), or any combination thereof. Memory (120) may include a cache memory, which is one or more different types of memory used to temporarily store data for a function or feature of the electronic device (100). As an example not limited to, the cache memory may be included within the processor (110). The memory (120) may be fixedly embedded within the electronic device (100) or incorporated into one or more suitable types of components (e.g., a SIM (subscriber identity module) card and / or an SD (secure digital) card) that can be repeatedly inserted into and removed from the electronic device (100).

[0030] For example, memory (120) may store one or more software applications, such as operating system (or system) software applications, firmware software applications, driver software applications, plugin (e.g., add-in, add-on, and / or applet) software applications, and / or any other suitable software applications. For example, the one or more software applications may include instructions executable by the processor (110). For example, memory (120) may store instructions that can be called by an application programming interface (API). For example, memory (120) may store instructions within a library.

[0031] FIG. 2 is a schematic diagram of an exemplary artificial intelligence system. Referring to FIG. 2, an artificial intelligence system according to one embodiment may include a user interface (210), a database (220), an application and service component (230), an AI framework (240), and a generative AI model (250).

[0032] A user interface (210) may receive a user query. The input may include user input and / or data obtained or generated by an electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (300) of FIG. 3). The data may include images, videos, and / or sensor data generated by at least one processor of the electronic device (e.g., at least one processor (110)), such as light intensity data around the electronic device obtained from a sensor (170) or sensor hub, attitude data (or orientation data) of the electronic device, temperature inside the electronic device (e.g., the temperature of the display (140) or the temperature of at least one processor (110)), size information of the display area of ​​the display (140), and / or images obtained through an image sensor (150) of the electronic device. For example, a user query may be in the form of natural language, touch data obtained through a touch circuit included in the display (140) (e.g., used to identify input from a finger and / or stylus), an image, audio, and / or video. Additionally, context information may be transmitted along with the user query. Context information may include various side information related to the time when the user query is input into the artificial intelligence system. Examples include application information currently being used by the user or location information of the user. As another example, the user query may be a non-natural language input that does not generate natural language, such as a design request or modification. Additionally, a mixed form of the natural language, image, sound, and context information described above is also possible.Additionally, the user interface (210) may output results of the artificial intelligence system to the user. The output may include results (or result information) generated or obtained by the artificial intelligence system based on at least part of the input. The output may be in the form of natural language or specific content, and may also be provided in the form of an action requested by the user. For example, the output may have a format according to the user settings of the electronic device.

[0033] The AI ​​framework (240) can receive a user query and coordinate and control each component necessary to perform the user's intent. The AI ​​framework (240) may include a prompt design component (242), an APIs / Plugins Management component (244), and an output modification component (246).

[0034] User queries or actions entered in the user interface (210) can be transmitted to a prompt design component (242). The prompt design component (242) can be used to generate prompts suitable for input into a large language model (LLM), a large vision model (LVM), or large multimodal models. The prompt design component (242) may be an AI component that uses machine learning algorithms or neural networks to develop better prompts over time. The prompt design component (242) can generate prompts by accessing a database (220) (e.g., a knowledge component) containing user preference data, a prompt library, and prompt examples, and can transmit them to the large language model (LLM), large vision model (LVM), and / or large multimodal model (LMM).

[0035] The application and plugin management component (244) can perform the role of communicating with external information when there is a request for additional information when user input is transmitted as input to the generative AI model (250). The application and plugin management component (244) establishes a channel to communicate with the outside of the artificial intelligence system through an application programming interface (API), thereby enabling access to various data sources. For example, the application and plugin management component (244) can be used to request other components (e.g., application and service components (230)) that perform feedback (or response) according to the prompt. The acquired information can be used to generate a prompt by the prompt design component (242) together with the user input, or can be used as input to the generative AI model (250). Additionally, the application and plugin management component (244) can request the action through the API if the application or service needs to perform an action that ultimately executes the user query rather than an intermediate result. Information obtained from an external source can be transmitted as input to a generative AI model (250) along with user input.

[0036] The output modification component (246) can finely tune (or adjust) (or change) the output output from the generative AI model (250). For example, the output modification component (246) can determine the relevance (e.g., score) between the output (e.g., content) of the generative AI model and the user input. For example, the output modification component (246) can verify whether the content generated through the Large Language Model (LLM), Large Vision Model (LVM), or Large Multimodal Model (LMM) contains the aforementioned relevance, biased information (e.g., selective information), or harmful information (e.g., violent content or profanity). Additionally, the output modification component (246) can determine the extent to which the output matches the desired result and, if additional processing is required, proceed with that process. Additionally, the output modification component (246) can configure and provide a hint to the user to avoid unwanted output.

[0037] A generative AI model (250) generally refers to an artificial intelligence neural network that generates new forms of data based on user input information. A generative AI model (250) may include an image generation model and / or a language generation model. An image generation model may include a generative adversarial network (GAN) and / or a variational autoencoder (VAE). An example of an image generation model is a diffusion-based generative model that uses the structures of a VAE and a transformer. Additionally, a language generation model is a model trained to output the statistically most appropriate output value based on input values, and representative examples include models such as CHAT-GPT 3 and CHAT-GPT 4. Additionally, it may include large multimodal models (LMMs) capable of recognizing various forms of data input, such as text, images, and voice, and generating new data corresponding to them.

[0038] In one embodiment, the AI ​​framework (240) and / or generative AI model (250) may be included within an AI module (e.g., including a processing circuit) within the electronic device. For example, the AI ​​module may be operatively coupled with at least one processor of the electronic device (e.g., at least one processor (110) or processor (1310)). For example, the AI ​​module may be operatively coupled with a sensor hub of the electronic device for one or more sensors within the electronic device.

[0039] FIG. 3 is a block diagram of an electronic device according to one embodiment.

[0040] The electronic device (300) may include a processor (310), memory (320), and / or a display (330). Various embodiments of this document may be implemented even if some of the illustrated configurations are omitted or substituted with other configurations. The electronic device (300) may further include at least some of the configurations and / or functions of the electronic device (100) of FIG. 1 in addition to the illustrated configurations. At least some of each configuration of the electronic device (300) illustrated (or not illustrated) may be operatively, functionally, and / or electrically connected.

[0041] The display (330) may include a configuration identical or similar to at least one display (140) of FIG. 1. According to one embodiment, the display (330) may display various images provided by the processor (310). According to one embodiment, the display (330) may visually provide various screens related to the application being executed and its use (e.g., contents screen, application execution screen, menu screen, and / or function execution screen) under the control of the processor (310).

[0042] The display (330) may be combined with a touch sensor, a pressure sensor capable of measuring the intensity of the touch, and / or a touch panel (e.g., a digitizer) that detects a magnetic field-based stylus pen. According to one embodiment, the display (330) may detect touch input, air gesture input, and / or hovering input (or proximity input) by measuring a change in a signal (e.g., voltage, light intensity, resistance, electromagnetic signal, and / or charge quantity) at a specific location on the display (330) based on the touch sensor, pressure sensor, and / or touch panel. For example, the display (330) may include a touchscreen that detects touch and / or proximity touch (or hovering) input using a part of the user's body (e.g., a finger) or an input device (e.g., a stylus pen).

[0043] The display (330) may include, but is not limited to, a liquid crystal display (LCD), a light-emitting diode (LED), an organic light-emitting diode (OLED) display (330), and / or an active matrix OLED (AMOLED) display (330), a micro electro mechanical systems (MEMS) display (330), or an electronic paper display (330). According to one embodiment, the display (330) may include a flexible display (330).

[0044] The processor (310) may include at least one processing circuitry, and the processor (310) may include at least one processor (310) (e.g., an application processor (310) or a communication processor (310)). The operations described in FIGS. 1 to 10 may be performed individually or collectively by at least one processor included in the processor (310).

[0045] The memory (320) can store at least one computer program, and at least one computer program may include instructions that can be executed by the processor (310). The operations described in FIGS. 1 to 10 can be performed according to the execution of instructions contained in the memory (320).

[0046] According to one embodiment, instructions may be stored as software on memory (320) and executed by a processor (310). For example, instructions may include control commands such as arithmetic and logical operations, data movement, and / or input / output that can be recognized by the processor (310). According to one embodiment, the software may include various applications that can provide various functions (or services) (e.g., routine functions, call functions, message functions, messenger functions, email functions, social networking service (SNS) functions, search functions, media (e.g., video and / or music) playback functions, game functions, and / or wireless communication functions) in an electronic device (300).

[0047] According to one embodiment, the memory (320) can store software. The memory (320) can store program modules (e.g., client modules) that support various applications and intelligent services.

[0048] According to one embodiment, the memory (320) can store various data used by at least one component (e.g., processor (310)) of the electronic device (300). In one embodiment, the data may include, for example, input data or output data for software and commands related to the software.

[0049] In one embodiment, the data may include various data (e.g., training data, prompt data, context, and / or a learning model) to support the electronic device (300) in editing and generating (e.g., regenerating or reconstructing) artificial intelligence-based data (e.g., an image). In one embodiment, the data may include information regarding various settings to support the electronic device (300) in controlling the operation of editing and / or generating artificial intelligence-based data (e.g., an image).

[0050] In one embodiment, the data may include various learning data and / or parameters obtained based on the user's learning through interaction with the user. In one embodiment, the data may include various schemas (or algorithms, models, networks, or functions) to support AI-based image editing and / or generation operations.

[0051] For example, a schema for supporting artificial intelligence-based image editing and / or generation operations in an electronic device (300) may include a neural network. In one embodiment, the neural network may include a neural network model based on at least one of an artificial neural network (ANN), a convolutional neural network (CNN), a region with convolutional neural network (R-CNN), a region proposal network (RPN), a recurrent neural network (RNN), a stacking-based deep neural network (S-DNN), a state-space dynamic neural network (S-SDNN), a Deconvolution Network, a deep belief network (DBN), a restricted Boltzmann machine (RBM), a long short-term memory (LSTM) network, a classification network, a plain residual network, a dense network, a hierarchical pyramid network, and / or a fully convolutional network. According to one embodiment, the type of neural network model is not limited to the examples described above.

[0052] The processor (310) can display an application screen on the display (330). For example, the processor (310) can activate the display (330) to display an application screen in response to a touch input or a hard key input on the display (330). The application may refer to an application configured to display an image in at least a portion of the area on the display (330).

[0053] The processor (310) may display a text window on the display (330) that displays at least one word. The text window may refer to an area on the display (330) configured to receive user input for entering a word or to display at least one word entered by the user.

[0054] The processor (310) can receive user input that inputs a word on the display (330). The processor (310) can detect user input that inputs a word on the display (330) using at least one sensor (e.g., a touch sensor).

[0055] The processor (310) can display at least one word entered by the user on the text window in response to receiving user input entering a word into the text window.

[0056] The processor (310) can display an image display window on the display (330) that is displayed separately from the text window. The image display window may refer to an area on the display (330) set to display an image.

[0057] The processor (310) can receive user input that drags a word to an image display window. Depending on the user input, an area where an object corresponding to the word is displayed can be determined.

[0058] The processor (310) can generate a prompt that instructs the creation of an image containing an object corresponding to the word in response to user input dragging the word to an image display window. The processor (310) can determine the location where the word was dragged according to the user input. For example, the processor (310) can determine that the input for dragging the word has ended when the user's touch input on the display (330) ends, and can determine the location where the word was dragged. The prompt generated by the processor (310) can instruct the creation of an image containing an object corresponding to the word at the location where the word was dragged.

[0059] According to one embodiment, the processor (310) may generate a prompt to perform image generation including an object based on a character (e.g., a word) entered by a user, information related to an object corresponding to the character (e.g., a word), and parameter information related to an application. The character (e.g., a word) entered by the user may be used as a prompt source for image generation including an object. The character may include character-based content. The processor (310) may generate a prompt to perform image generation including an object based on a word entered by the user and / or information related to an object corresponding to the word.

[0060] According to one embodiment, the processor (310) may generate a prompt that instructs to generate an image with an object corresponding to the word as the background when the word is a word indicating a background.

[0061] The processor (310) may display a preview area of ​​an object corresponding to the prompt until user input dragging a word ends. For example, the processor (310) may display the border of an object instructed to be created according to the prompt. The preview area of ​​the object may be displayed at the location where the word was dragged according to user input.

[0062] The processor (310) can generate an image containing an object corresponding to a word based on a prompt instructing the generation of an image containing an object corresponding to a word in response to the termination of user input in which the word is dragged to an image display window. The processor (310) can generate an image containing an object corresponding to a word at the location where the word was dragged. The processor (310) can display the generated image on the display (330).

[0063] For example, the processor (310) may generate (or acquire) an image containing an object in relation to a generated prompt. The processor (310) may receive a prompt source (e.g., a first word, a second word) related to generating an image containing an object based on interaction with the user, and may generate (e.g., regenerate or reconstruct) an image containing an object on a server or on a device based on the prompt source. The processor (310) may provide the prompt to a generative AI on a device and / or server to execute an image generation process based on the generated prompt. The processor (310) may provide a prompt requesting image generation (e.g., a question or instruction to be entered into the generative AI) to the generative AI. The processor (310) may generate (or acquire) an image according to the image generation process executed in relation to the prompt by the on-device AI (e.g., the generative AI model of FIG. 2). The processor (310) can receive (or obtain) an image from the server containing an object according to an image generation process executed in relation to a prompt in the server artificial intelligence.

[0064] For example, the processor (310) may generate (or acquire) an image based on the generated prompt and the additional image. The additional image may refer to either a previously generated image or an image in which some areas of the previously generated image are masked. The processor (310) may generate an image in which at least one object mark included in the additional image is retained. If the additional image is an image in which some areas are masked, the processor (310) may generate an image in which the masked areas are inpainted.

[0065] It should be understood that the operation of the electronic device (300) (or processor (310)) of the present disclosure generating (or modifying) a prompt and the operation of generating (or modifying) an image based on the prompt may be applied to the operation of the processor (310) of FIG. 3 generating a prompt and the operation of generating an image based on the prompt.

[0066] When the processor (310) receives user input to drag a second word, which is distinct from the first word that has been dragged, to an image display window, it can determine whether the generated image contains an object. According to one embodiment, if the first word that has been dragged to the image display window is present, the processor (310) can determine that the generated image contains an object.

[0067] The processor (310) can determine whether to display the first object corresponding to the first word and the second object corresponding to the second word in a mutually related form based on the distance between the first position where the first word is dragged and the second position where the second word is dragged according to user input dragging the word to the image display window, when the object is included in the generated image.

[0068] Displaying the first object and the second object in a mutually related form may refer to inferring (or determining) a linguistic context that linguistically connects the first word and the second word, and displaying an image containing the first object and the second object corresponding to the inferred (or determined) linguistic context that linguistically connects the first word and the second word. Displaying the first object and the second object in a mutually unrelated form may refer to displaying an image containing the first object included in an image generated in response to receiving a first input when no previously generated image exists, and the second object included in an image generated in response to receiving a second input when no previously generated image exists.

[0069] According to one example, the processor (310) may determine that the first word and the second word cannot be linguistically connected when checking (or determining) the linguistic context of the first word and the second word. If the first word and the second word cannot be linguistically connected, the processor (310) may determine to display the first object and the second object in a non-related form. According to one example, if the first word and the second word can be linguistically connected, the processor (310) may determine to display the first object and the second object in a related form.

[0070] The processor (310) can determine the linguistic context of the first word and the second word based on the distance between the first position and the second position. The processor (310) can determine the linguistic context that is most likely to exist in a situation separated by the distance between the first position and the second position among the linguistic contexts of the first word and the second word as the linguistic context that linguistically connects the first word and the second word.

[0071] The processor (310) can determine the linguistic context of the first word and the second word based on the distance between the first location and the second location on the server (301) or on the device. The processor (310) can use an artificial intelligence model on the device and / or server (301) (e.g., an artificial intelligence model trained to infer the linguistic context between words) to determine the linguistic context of the first word and the second word based on the distance between the first location and the second location determined by the artificial intelligence model of the server (301). The processor (310) can receive (or obtain) from the server (301) the linguistic context of the first word and the second word according to the distance between the first location and the second location determined by the artificial intelligence model of the server (301). The server (301) that performs the action of determining the linguistic context of the first word and the second word may be a server distinct from the server that generates an image related to the prompt.

[0072] According to one embodiment, the processor (310) can determine the linguistic context of the first word and the second word based on at least some features (e.g., size, pose) of the first object or the second object included in the generated image when the first object or the second object is included in the generated image. The processor (310) can determine the linguistic context with the highest probability of existence among the linguistic contexts of the first object and the second object corresponding to at least some features (e.g., size, pose) of the first object or the second object as well as the distance between the first position and the second position. For example, even if the distance between the first position and the second position is the same, the processor (310) may determine the linguistic context of the first word and the second word differently depending on whether the size of the first object or the second object displayed in the generated image is larger or smaller.

[0073] According to one embodiment, the processor (310) can determine the linguistic context of the first word and the second word based on a second location where the word (e.g., the second word) is dragged according to user input that drags the word (e.g., the second word) to an image display window. The processor (310) can determine the linguistic context of the first word and the second word based on the coordinates of the second location on the image display window. For example, the processor (310) can determine the linguistic context of the first word and the second word differently depending on whether the second location is less than a specified distance from the boundary of the image display window. The linguistic context of the first word and the second word may reflect that if the second location is less than a specified distance from the boundary of the image display window, at least a part of the second object corresponding to the second word is not included in the image.

[0074] According to one embodiment, the linguistic context among a plurality of words dragged onto the display (330) may be determined based on the order in which each of the plurality of words is dragged onto the image display window. When the processor (310) receives user input to drag a specific word (e.g., a second word) onto the image display window, it may generate a prompt that reflects the previously dragged word (or object) and the location where the previously dragged word (or object) was dragged. Since the processor (310) generates a prompt that reflects the previously dragged word (or object) and the location where the previously dragged word (or object) was dragged, the prompt generated by the processor (310) may vary depending on the previously dragged word (or object) and the location of the previously dragged word (or object).

[0075] If the processor (310) determines to display the second object and the first object in a mutually related form, it may generate a prompt instructing to create an image containing the second object and the first object in a mutually related form. For example, the processor (310) may generate a prompt including a first word, a second word, and the linguistic context of the first word and the second word. The processor (310) may generate a prompt instructing to generate an image including a first object and a second object corresponding to the linguistic context of the first word and the second word. If the processor (310) determines to display the second object and the first object in a non-related form, it may generate a prompt instructing to generate an image including the second object and the first object in a non-related form. According to one embodiment, the processor (310) may generate a prompt instructing to modify a previously generated image. The operation of the processor (310) (or electronic device (300)) of the present disclosure generating a prompt instructing to generate an image may be replaced with the operation of generating a prompt instructing to modify a previously generated image.

[0076] The processor (310) can generate an image based on the generated prompt and display it on the display (330).

[0077] The processor (310) can receive user input to drag one of the multiple objects (e.g., a second object) included in a previously generated image, and can determine a third location where one of the multiple objects (e.g., a second object) was dragged according to the user input.

[0078] The processor (310) may determine whether to display the second object and the first object in a related form based on the distance between the third position and the first position (or the second position). If the processor (310) determines to display the second object and the first object in a related form, it may generate a prompt instructing to create an image containing the second object and the first object in a related form. For example, the processor (310) may determine the linguistic context of the first word and the second word based on the distance between the third position and the first position (or the second position) and generate a prompt containing the linguistic context of the first word and the second word.

[0079] The processor (310) may display a preview area of ​​an object corresponding to a prompt until user input to drag one of a plurality of objects (e.g., a second object) ends. For example, the processor (310) may display the border of the second object that is instructed to be created according to the prompt. The preview area may be displayed at the location (e.g., a third location) where the object (e.g., the second object) was dragged according to user input.

[0080] The processor (310) may generate an image including the second object and the first object and display it on the display (330) in response to the termination of user input dragging one of the multiple objects (e.g., the second object). The second object and the first object may be objects generated to reflect the linguistic context of the first word and the second word based on the distance between the third position and the first position (or the second position).

[0081] The processor (310) can display at least a portion of a modified prompt in a text window to reflect the linguistic context of the first word and the second word based on the distance between the third position and the first position (or second position).

[0082] According to one embodiment, the processor (310) may generate a prompt instructing the creation of an image containing objects included in the created image, excluding the dragged objects, in response to receiving user input to drag objects included in the created image into a text window. For example, the processor (310) may mask the area corresponding to the dragged objects in the created image. The processor (310) may generate a prompt instructing the creation of an image in which the masked area is inpainted.

[0083] According to one embodiment, the processor (310) can generate an image containing a plurality of objects included in a previously generated image excluding the dragged object, based on an additional image in which the area where the dragged object is displayed is masked and a generated prompt. The processor (310) can display the generated image on a display (330).

[0084] The processor (310) may display on the display (330) an area capable of receiving a user input that instructs to modify an object in response to receiving a user input that selects an object included in the generated image. The user input that selects an object included in the generated image may be a user input that is distinct from a user input that drags an object included in the generated image. For example, the user input that selects an object included in the generated image may be a user input that touches an object for a specified amount of time.

[0085] The processor (310) can transform an object in response to receiving user input that instructs the transformation of an object included in a generated image. The user input that instructs the transformation of an object may be an input that instructs the area (width, height, and / or depth) where the object is displayed and the location where the object is displayed. The processor (310) can transform the object in response to the input that instructs the area (width, height, and / or depth) where the object is displayed and the location where the object is displayed.

[0086] FIG. 4 is a drawing illustrating a display of an electronic device according to one embodiment.

[0087] The electronic device (300) can display an application screen on the display (330). For example, the electronic device (300) can activate the display (330) to display an application screen in response to a touch input on the display (330) or an input to a hard key. The application may refer to an application configured to display an image in at least a portion of the area on the display (330).

[0088] The electronic device (300) may display a text window (420) that displays at least one word on the display (330). The text window (420) may refer to an area on the display (330) configured to receive user input for entering a word or to display at least one word entered by the user.

[0089] The electronic device (300) can receive user input that inputs a word (421) on the display (330). The electronic device (300) can detect user input that inputs a word (421) on the display (330) using at least one sensor (e.g., a touch sensor).

[0090] In response to receiving user input (421) of a word being entered into a text window (420), the electronic device (300) can display at least one word (422) entered by the user on the text window (420).

[0091] The electronic device (300) can display an image display window (410) that is displayed separately from the text window (420) on the display (330). The image display window (410) may refer to an area on the display (330) set to display an image.

[0092] FIG. 5 is a drawing illustrating an electronic device that generates an image in response to user input according to one embodiment.

[0093] An electronic device (e.g., the electronic device (300) of FIG. 3) can receive user input to input a word into a text window (e.g., the text window (420) of FIG. 4) and display at least one word (e.g., a first word (501), a second word (502)) entered by the user on the text window.

[0094] The electronic device (300) can receive user input (e.g., first input (515), second input (525)) that drags a word (e.g., first word (501), second word (502)) to an image display window (e.g., image display window (410) of FIG. 4). Depending on the user input (e.g., first input (515), second input (525)), the location and / or size of the area where an object (e.g., first object (511), second object (522)) corresponding to the word (e.g., first word (501), second word (502)) is displayed may be determined.

[0095] The electronic device (300) may generate a prompt instructing the creation of an image containing an object corresponding to a word (e.g., a first object (511)) in response to user input (515) dragging a word (e.g., a first word (501)) to an image display window (410). The electronic device (300) may determine the location where the word (e.g., a first word (501)) was dragged according to the user input (e.g., a first location where the first word (501) was dragged). For example, the electronic device (300) may determine that the first input (515) dragging the first word (515) has ended when the user's touch input on the display (330) ends, and may determine the first location. The prompt generated by the electronic device (300) may instruct the creation of an image containing the first object (511) at the first location. The first location may be a location identified by coordinates (e.g., xy coordinates) on the image display window (410), a distance from the edge of the display (330), and / or a [relative] distance from an object displayed on the image display window (410). According to one example, the electronic device (300) may generate a prompt reflecting that at least part of the first object corresponding to the first word (501) is not included in the image if the first location is less than a specified distance from the boundary of the image display window.

[0096] The electronic device (300) may generate a prompt instructing to create a first object based on coordinates on the image display window (410) of a first position. The prompt instructing to create a first object may vary depending on the coordinates on the image display window (410) where the first word (501) is dragged. If the first position is a position adjacent to a boundary (e.g., left boundary) on the image display window (410), the electronic device (300) may generate a prompt instructing to create a first object formed in a direction not adjacent to the boundary (e.g., right). If the first position is a position adjacent to a substantial central part on the image display window (410), the electronic device (300) may generate a prompt instructing to create a first object formed in any direction. The electronic device (300) may generate a prompt instructing the creation of a first object based on a location within an image display area that is distinct from the location adjacent to the boundary and the location adjacent to the substantial central part, as well as a location adjacent to the boundary and the location adjacent to the substantial central part. The electronic device (300) may create an object according to the generated prompt. The electronic device (300) may create objects having different shapes and / or arrangements corresponding to words representing the same object according to coordinates on the image display window (410).

[0097] The electronic device (300) can generate an image (510) containing an object corresponding to a word (e.g., a first object (511)) based on a prompt instructing the generation of an image containing an object corresponding to a word (e.g., a first object (511)) in response to the termination of user input in which a word (e.g., a first word (501)) is dragged to an image display window (410). The electronic device (300) can generate an image (510) containing an object corresponding to a word (e.g., a first object (511)) at a first location. The electronic device (300) can display the generated image (510) on a display (330). According to one example, the electronic device (300) may generate an image containing a first object (511) in a form in which at least a part of the first object corresponding to the first word (501) is not included in the image based on a prompt reflecting that at least a part of the first object is not included in the image.

[0098] When the electronic device (300) receives user input (e.g., second input (525)) that drags a word (e.g., second word (502)) to the image display window (410), it can determine whether an object (e.g., first object (511)) is included in the previously generated image. According to one embodiment, the electronic device (300) can determine that an object (e.g., first object (511)) is included in the previously generated image if there is a word (e.g., first word (501)) that has been dragged to the image display window (410).

[0099] The electronic device (300) can determine whether to display the second object (522) and the first object (511) in a mutually related form based on the distance between the first position where the object (e.g., the first object (511)) of the generated image is dragged and the second position where the word (e.g., the second word (502)) is dragged, according to user input in which the word (e.g., the second word (502)) is dragged to the image display window (410), when the generated image contains an object (e.g., the first object (511)).

[0100] Displaying the first object (511) and the second object (522) in a mutually related form may refer to inferring (or determining) a linguistic context that linguistically connects the first word (501) and the second word (502), and displaying an image containing the first object and the second object corresponding to the linguistic context that linguistically connects the inferred (or determined) first word (501) and the second word (502). Displaying the first object and the second object in a mutually unrelated form may refer to including and displaying the first object created by dragging the first word to the image display area (410) when there is no previously created image (e.g., there is no object and / or image displayed in the image display area (410)) and the second object created by dragging the first word to the image display area (410) when there is no previously created image (e.g., there is no object and / or image displayed in the image display area (410)) in a single image.

[0101] The electronic device (300) can determine the linguistic context of the first word (501) and the second word (502) based on the distance between the first location and the second location. The electronic device (300) can use an artificial intelligence model (e.g., an artificial intelligence model trained to infer the linguistic context between words) to determine the linguistic context of the first word (501) and the second word (502) based on the distance between the first location and the second location. For example, the electronic device can use on-device artificial intelligence or use server artificial intelligence and receive results from a server. The electronic device (300) can determine the linguistic context of the first word (501) and the second word (502) based on the distance between the first location and the second location. The electronic device (300) can determine the linguistic context that is most likely to exist among the linguistic contexts of the first word (501) and the second word (502) when separated by the distance between the first location and the second location. Words (e.g., the first word (501) and the second word (502)) can be connected by various linguistic contexts that linguistically link the words (the first word (501) and the second word (502)). Each of these various linguistic contexts may have a different probability of existence depending on the distance between the words (the first word (501) and the second word (502)). For example, a specific linguistic context (e.g., the linguistic context where the first word (501) is 'person' and the second word (502) is 'clothes', 'person' wearing 'clothes') may have a lower probability of existence as the distance between the words (the first word (501) and the second word (502)) increases. As another example, a specific linguistic context (e.g., a linguistic context where the first word (501) is 'person' and the second word (502) is 'clothes', a 'person' holding 'clothes') may have a higher probability of existence as the distance between the words (first word (501) and second word (502)) approaches a specific value.

[0102] The electronic device (300) can determine the distance between words (first word (501) and second word (502)) and the distance between words (first word (501) and second word (502)), and among the linguistic contexts between words (first word (501) and second word (502)), the linguistic context between words that is most likely to exist when the words (first word (501) and second word (502)) are separated by the distance between words (first word (501) and second word (502)).

[0103] According to one embodiment, the electronic device (300) can determine the linguistic context differently as the distance between the first position and the second position changes. The linguistic context of the first word (e.g., person) and the second word (e.g., cup) when the distance between the first position and the second position is the first distance (e.g., holding a cup close to the body) and the linguistic context of the first word (e.g., person) and the second word (e.g., reaching out to hold a cup) when the distance between the first position and the second position is the second distance (e.g., reaching out to hold a cup) may be different.

[0104] According to one embodiment, the electronic device (300) may determine the linguistic context of the first word (501) and the second word (502) differently depending on whether the second position is less than the boundary of the image display window (410) and a specified distance. For example, the linguistic context of the first word (501) and the second word (502) may reflect that if the second position is less than the boundary of the image display window and a specified distance, at least a part of the second object corresponding to the second word (502) is not included in the image.

[0105] According to one embodiment, the linguistic context between the first word (501) and the second word (502) dragged to the display (330) may be determined based on the order in which the first word (501) and the second word (502) are dragged onto the image display window. When the electronic device (300) receives user input to drag a specific word (e.g., the second word (502)) onto the image display window, it may determine the linguistic context based on the previously dragged word (e.g., the first word (501)) and the location (e.g., the first location) where the previously dragged word (e.g., the first word (501)) was dragged.

[0106] According to one example, the electronic device (300) may determine that the first word (501) and the second word (502) cannot be linguistically connected when determining the linguistic context of the first word (501) and the second word (502). The electronic device (300) may determine that the first object and the second object are not linguistically connected when the first word (501) and the second word (502) cannot be linguistically connected. According to one example, the electronic device (300) may determine that the first object (511) and the second object (522) are linguistically connected when the first word (501) and the second word (502) can be linguistically connected. When the first object (511) and the second object (522) are displayed in a mutually related form, the prompt may be a prompt that reflects the linguistic context of the first word and the second word (e.g., draw a hand holding a pencil).

[0107] If the electronic device (300) determines to display the second object and the first object in a non-related form, it may generate a prompt instructing to generate an image (520) including the second object (522) and the first object (511) in a non-related form.

[0108] The electronic device (300) can generate an image (520) based on the generated prompt and display it on the display.

[0109] In response to receiving a user input (e.g., a third input (535)) that drags one of a plurality of objects (e.g., a second object (5220)) included in a previously generated image (530), the electronic device (300) can determine a third location where one of the plurality of objects (e.g., a second object (5220)) was dragged according to the user input.

[0110] The electronic device (300) may determine whether to display the second object (5220) and the first object (511) in a mutually related form based on the distance between the third position and the first position (or the second position). If the electronic device (300) determines to display the second object (5220) and the first object (511) in a mutually related form, it may generate a prompt instructing to create an image containing the second object (5220) and the first object (511) in a mutually related form. For example, the electronic device (300) may generate a prompt containing the first word (501), the second word (502), and the linguistic context of the first word (501) and the second word (502). The electronic device (300) may generate a prompt instructing to create an image containing the first object and the second object corresponding to the linguistic context of the first word (501) and the second word (502). A prompt instructing to create an image including a second object (5220) and a first object (511) of a mutually related form may be a prompt instructing to modify a previously created image (530).

[0111] According to one example, the electronic device (300) may determine a linguistic context based on the distance between a specific area (e.g., mouth) of a first object (e.g., person) (511) included in a pre-generated image and a third location where a second object (5220) (e.g., cup) is dragged. Among the linguistic contexts between words (first word (501) and second word (502)), the linguistic context related to a specific area (e.g., mouth) of the first object (e.g., person) may vary depending on the distance between the specific area (e.g., mouth) of the first object (e.g., person) (511) and the third location. For example, the probability of the existence of a linguistic context (e.g., 'holding a cup to the mouth') reflecting information corresponding to a specific area may increase as the distance between the specific area (e.g., mouth) of the first object (e.g., person) (511) and the third location decreases. The electronic device (300) can determine a linguistic context (e.g., 'holding a cup to the mouth') that reflects information corresponding to a specific area (e.g., mouth) of the first object (e.g., person) (511) when the distance between that specific area (e.g., mouth) and the third location is close. The electronic device (300) can determine a linguistic context (e.g., 'holding a cup in the hand') that reflects information corresponding to another area (e.g., hand) of the first object (e.g., person) (511) when the distance between that specific area (e.g., hand) and the third location is close. Although the above examples are described for a specific area of ​​the first object (511), it should be understood that they may also apply to a specific area of ​​the second object (5220).

[0112] According to one example, a linguistic context can be determined based on the relative arrangement of a first object (e.g., person) (511) and a second object (e.g., dog) (5220) included in a previously generated image. Among the linguistic contexts between words (first word (501) and second word (502)), a specific linguistic context (e.g., 'dog dragging person') may vary depending on the relative arrangement of the first object (e.g., person) (511) and the second object (e.g., dog) (5220). For example, the probability of the existence of a specific linguistic context (e.g., 'a dog dragging a person') may be higher when the second object (e.g., a dog) (5220) is placed in the second direction (e.g., front) of the first object (e.g., person) (511) than when the second object (e.g., a dog) (5220) is placed in the first direction (e.g., back) of the first object (e.g., person) (511). The electronic device (300) can determine the specific linguistic context (e.g., 'a dog dragging a person') as the linguistic context of the first word (501) and the second word (502) when the second object (e.g., a dog) (5220) is dragged in the second direction (e.g., front) of the first object (e.g., person) (511). For another example, the probability of the existence of a different linguistic context (e.g., 'a person dragging a dog') may be higher when the second object (e.g., a dog) (5220) is placed in the first direction (e.g., back) of the first object (e.g., a person) (511) than when the second object (e.g., a dog) (5220) is placed in the second direction (e.g., front) of the first object (e.g., a person) (511). The electronic device (300) can determine the different linguistic context (e.g., 'a person dragging a dog') as the linguistic context of the first word (501) and the second word (502) when the second object (e.g., a dog) (5220) is dragged in the first direction (e.g., back) of the first object (e.g., a person dragging a dog).It should be understood that the above example may also apply to cases where the relative placement is in a direction other than the second direction (e.g., front) and the first direction (e.g., back) (e.g., up, down, left, right). It should also be understood that it may apply to a specific area (e.g., legs) of the first object (e.g., person) (511) and the second object (e.g., box) (5220) (or a specific area of ​​the second object). For example, the electronic device (300) may determine a specific linguistic context (e.g., 'a person sitting in a box') from the linguistic context of the first word (501) and the second word (502) when the second object (e.g., box) (5220) is dragged in a third direction (e.g., down) of a specific area (e.g., legs) of the first object (e.g., person) (511).

[0113] The electronic device (300) may display a preview area (5320) of an object corresponding to the prompt (e.g., a second modified object (542)) until user input for dragging one of a plurality of objects (e.g., a second object (5220)) ends. For example, the electronic device (300) may display the border of the second modified object (542) that is instructed to be created according to the prompt. The preview area (5320) may be displayed at the location (e.g., a third location) where the object (e.g., the second object (5220)) was dragged according to user input.

[0114] The electronic device (300) may generate an image (540) including a second modified object (542) and a first modified object (541) in response to the termination of user input for dragging one of a plurality of objects (e.g., a second object (5220)), and display it on the display (330). The second modified object (542) and the first modified object (541) may be objects generated to reflect the linguistic context of the first word (501) and the second word (502) based on the distance between the third position and the first position (or the second position).

[0115] The electronic device (300) can display at least a portion (503) of a modified prompt instructing to create an image (540) including a second modified object (542) and a first modified object (541) in a text window.

[0116] FIG. 6 is a drawing illustrating a display (330) of an electronic device (300) according to one embodiment.

[0117] The electronic device (300) can receive a first input (615) that drags a first word (e.g., hand) (602) to an image display window (e.g., image display window (410) of FIG. 4). The area where a first object (611) corresponding to the first word (e.g., hand) (602) is displayed can be moved according to the first input (615).

[0118] The electronic device (300) may generate a prompt instructing the creation of an image containing an object (e.g., a first object (611)) corresponding to the first word in response to a first input (615) that drags the first word (e.g., hand) (602) to an image display window (410). The electronic device (300) may determine a first location where the first word (e.g., hand) (602) was dragged according to the first input (615). For example, the electronic device (300) may determine that the first input (615) that drags the first word (e.g., hand) (602) has ended when the user's touch input on the display (330) ends, and may determine the first location. The prompt generated by the electronic device (300) may instruct the creation of an image containing the first object at the first location.

[0119] The electronic device (300) can generate an image (610) containing a first object (611) corresponding to a first word (e.g., hand) (602) based on a prompt instructing to generate an image containing a first object (611) corresponding to a first word (e.g., hand) (602) in response to the termination of a first input (615) of dragging a first word (e.g., hand) (602) to an image display window (410). The electronic device (300) can generate an image (610) containing a first object (611) corresponding to a first word (e.g., hand) (602) at a first location.

[0120] The electronic device (300) can receive a second input (625) that drags a second word (e.g., pencil) (601) to an image display window (410). The area where a second object (622) corresponding to the second word (e.g., pencil) (601) is displayed can be moved according to the second input (625).

[0121] When the electronic device (300) receives a second input (625) that drags a second word (e.g., pencil) (601) to an image display window (410), it can check whether an object (e.g., a first object (611)) is included in a previously generated image (610).

[0122] The electronic device (300) can determine whether to display the second object (622) and the first object (611) in a mutually related manner based on the distance between the second position where the second word (e.g., pencil) (601) is dragged and the second position where the second word (e.g., pencil) (601) is dragged, according to a second input (625) that drags the first word (e.g., hand) (602) to the image display window (410), when the first object (e.g., first object (611)) is included in the previously generated image (610).

[0123] The electronic device (300) can determine the linguistic context with the highest probability of existence among the linguistic contexts corresponding to the case where the first word (e.g., hand) (602) and the second word (e.g., pencil) (601) are separated by a distance between the first position and the second position. When determining the linguistic context of the first word (e.g., hand) (602) and the second word (e.g., pencil) (601), the electronic device (300) may determine that the first word (e.g., hand) (602) and the second word (e.g., pencil) (601) cannot be linguistically connected. If the first word (e.g., hand) (602) and the second word (e.g., pencil) (601) cannot be linguistically connected, the electronic device (300) may determine to display the first object (611) and the second object (622) in a non-related form.

[0124] If the electronic device (300) determines to display the second object (622) and the first object (611) in a non-related form, it may generate a prompt instructing to generate an image including the second object (622) and the first object (611) in a non-related form. Based on the generated prompt, the electronic device (300) may generate an image (620) and display it on the display (330).

[0125] In response to receiving a third input (635) that drags a second object (6220) included in a previously generated image, the electronic device (300) can determine a third location where one of a plurality of objects (e.g., the second object (6220)) is dragged according to the third input (635).

[0126] The electronic device (300) may determine whether to display the second object (622) and the first object (611) in a related form based on the distance between the third position and the first position. If the electronic device (300) determines to display the second object (622) and the first object (611) in a related form, it may generate a prompt instructing to create an image containing the second object and the first object in a related form. For example, the electronic device (300) may generate a prompt instructing to modify a previously created image (620).

[0127] The electronic device (300) can generate (or modify) an image (630) containing a second object (632) and a first object (631) of a mutually related form in response to the termination of a third input dragging a second object (6220), and display it on a display (330).

[0128] The electronic device (300) may display at least a portion (603) of a modified prompt instructing to create an image (630) including a second object (632) and a first object (631) of a mutually related form in a text window (e.g., text window (420) of FIG. 4).

[0129] FIGS. 7A, FIGS. 7B, FIGS. 7C, FIGS. 7D, and FIGS. 7E are drawings illustrating a display of an electronic device that displays an image on a display in response to user input according to one embodiment.

[0130] The electronic device (300) can display a text window (e.g., the text window (420) of FIG. 4) on the display (330). The text window (420) may refer to an area on the display (330) set to receive user input for entering a word or to display at least one word entered by the user.

[0131] In response to receiving user input to input a word into a text window (420), the electronic device (300) can display at least one word entered by the user (e.g., a first word (701), a second word (702), a third word (703), or a fourth word (704)) on the text window (420).

[0132] The electronic device (300) can display an image display window (410) that is displayed separately from the text window (420) on the display (330). The image display window (410) may refer to an area on the display (330) set to display an image.

[0133] The electronic device (300) can receive user input (715) that drags the first word (701) to the image display window (410). The area where the first object corresponding to the first word (701) is displayed can be moved according to the user input (715).

[0134] The electronic device (300) may generate a prompt that instructs the creation of an image containing a first object in response to user input (715) that drags a first word (701) to an image display window (410). The electronic device (300) may check the location where the first word (701) was dragged according to the user input (715). For example, when the user's touch input on the display (330) ends, the electronic device (300) may confirm that the user input (715) that drags the first word (701) has ended and check the location where the first word (701) was dragged. The prompt generated by the electronic device (300) may instruct the creation of an image (720) containing a first object (721) at the location where the first word (701) was dragged.

[0135] According to one embodiment, the electronic device (300) may generate a prompt instructing to create an image with an object corresponding to the first word (701) as the background when the first word (701) is a word indicating a background (e.g., sky, living room).

[0136] The electronic device (300) may display a preview area (7110) of a first object corresponding to a prompt until user input (715) of dragging a first word (701) is terminated. For example, the electronic device (300) may display a border of the first object instructed to be created according to the prompt. The preview area (7110) of the first object may be displayed at the location where the first word (701) was dragged according to the user input (715).

[0137] The electronic device (300) can generate an image (720) containing a first object corresponding to the first word (701) based on a prompt instructing to generate an image containing a first object corresponding to the first word (701) in response to the termination of user input (715) of dragging the first word (701) to an image display window (410). The electronic device (300) can generate an image (720) containing a first object (721) corresponding to the first word (701) at the location where the first word (701) was dragged. The electronic device (300) can display the generated image (720) on a display (330).

[0138] Referring to FIG. 7b, when the electronic device (300) receives user input (735) of dragging the fourth word (704) to the image display window (410), it can check whether the previously generated image (720) contains an object (e.g., the first object (721)).

[0139] The electronic device (300) can determine the linguistic context with the highest probability of existence among the linguistic contexts of the fourth word (704) and the first word (701) based on the distance between the location where the first word (701) is dragged and the location where the fourth word (704) is dragged, when an object (e.g., the first object (721)) is included in the previously generated image (720). For example, among the linguistic contexts of the first object (e.g., sky) and the fourth object (e.g., climber), the linguistic context in which a climber is walking against the sky background may have the highest probability of existence.

[0140] The electronic device (300) may generate a prompt that instructs the generation of an image including a first object and a fourth object. For example, the electronic device (300) may generate a prompt that includes a first word (701), a fourth word (704), and the linguistic context of the first word (701) and the fourth word (704).

[0141] The electronic device (300) may display a preview area (7340) of a fourth object corresponding to a prompt until user input dragging the fourth word (704) ends. For example, the electronic device (300) may display the border of the fourth object instructed to be created according to the prompt. The preview area (7340) of the fourth object may be displayed at the location where the fourth word (704) was dragged according to user input (735).

[0142] The electronic device (300) can generate an image (740) based on a prompt instructing to generate an image including a fourth object and a first object in response to the termination of user input (735) of dragging a fourth word (704) to an image display window (410). The electronic device (300) can generate an image (740) including a fourth object (744) corresponding to the fourth word (704) at the location where the fourth word (704) was dragged. The electronic device (300) can display the generated image (740) on a display (330).

[0143] Referring to FIG. 7c, when the electronic device (300) receives user input (755) that drags the third word (703) to the image display window (410), it can check whether the previously generated image (740) contains an object (e.g., the first object (721) or the fourth object (744)).

[0144] The electronic device (300) can determine the linguistic context of the previously dragged words (e.g., first word (701), fourth word (704)) and the third word (703) based on the position where the third word (703) is dragged according to user input (755) of dragging the third word (703) to the image display window (410) when the previously generated image (740) contains an object (e.g., first word (701), fourth word (704)). The electronic device can determine the linguistic context with the highest probability of existence among the linguistic contexts of the first word (701), fourth word (704), and third word (703) based on the distance between the position where the previously dragged words (e.g., first word (701), fourth word (704)) are dragged and the position where the third word (703) is dragged. For example, among the linguistic contexts of the first object (e.g., sky), the fourth object (e.g., climber), and the third object (e.g., rock wall), the probability of existence of a context in which a climber is walking against the backdrop of the sky and the rock wall may be the highest.

[0145] The electronic device (300) may generate a prompt that instructs the generation of an image including a first object, a fourth object, and a third object. For example, the electronic device (300) may generate a prompt that includes a first word (701), a fourth word (704), a third word (703), and a linguistic context that linguistically links the first word (701), the fourth word (704), and the third word (703).

[0146] The electronic device (300) may display a preview area (7530) of a third object corresponding to a prompt until user input dragging the third word (703) ends. For example, the electronic device (300) may display the border of the third object instructed to be created according to the prompt. The preview area (7530) of the third object may be displayed at the location where the third word (703) was dragged according to user input (755).

[0147] The electronic device (300) can generate an image (760) including a first object (721), a fourth object (744), and a third object (763) based on a prompt instructing to generate an image including a first object, a fourth object, and a third object in response to the termination of user input (755) of dragging a third word (703) to an image display window (410). The electronic device (300) can generate an image (760) including a third object (763) corresponding to the third word (703) at the location where the third word (703) was dragged. The electronic device (300) can display the generated image (760) on a display (330).

[0148] Referring to FIG. 7d, the electronic device (300) receives a user input (775) for dragging a third object (7630) (e.g., the third object (763) in FIG. 7c) among a plurality of objects included in a previously generated image, and can determine the location where the third object was dragged according to the user input (775).

[0149] The electronic device (300) can determine the linguistic context of the previously dragged words (e.g., the first word (701), the fourth word (704)) and the third word (703) based on the position where the third object (7630) is dragged according to user input (775) of dragging the third object (7630).

[0150] The electronic device can determine the linguistic context with the highest probability of existence among the linguistic contexts of the first word (701), the fourth word (704), and the third word (703) based on the distance between the location where the previously dragged words (e.g., the first word (701), the fourth word (704)) were dragged and the location where the third object (7630) was dragged. For example, among the linguistic contexts of the first object (e.g., sky), the fourth object (e.g., climber), and the third object (e.g., rock wall), the context in which a climber is walking against the background of the sky and the right rock wall may have the highest probability of existence.

[0151] The electronic device (300) may generate a prompt that instructs the generation of an image including a first object, a fourth object, and a third object. For example, the electronic device (300) may generate a prompt that includes a first word (701), a fourth word (704), a third word (703), and a linguistic context that linguistically links the first word (701), the fourth word (704), and the third word (703).

[0152] The electronic device (300) may display a preview area (7730) of the third object corresponding to the prompt until user input for dragging the third object ends. For example, the electronic device (300) may display the border of the third object instructed to be created according to the prompt. The preview area (7730) of the third object may be displayed at the location where the third object (7630) was dragged according to user input.

[0153] The electronic device (300) can generate an image (780) including the first object (721), the fourth object (744), and the third object (783) based on a prompt instructing to generate an image including the first object, the fourth object, and the third object in response to the termination of user input (775) for dragging the third object (7630). The electronic device (300) can generate an image (780) including the third object (783) corresponding to the third word (703) at the location where the third object was dragged. The electronic device (300) can display the generated image (780) on the display (330).

[0154] Referring to FIG. 7e, the electronic device (300) receives a user input (795) to drag a third object among a plurality of objects included in a previously generated image, and can determine the location where the third object (7830) (e.g., the third object (783) in FIG. 7d) was dragged according to the user input (795).

[0155] The electronic device can reconfirm the linguistic context of previously dragged words (e.g., first word (701), fourth word (704)) and third word (703) based on the position where the third object (7830) was dragged according to user input (795) of dragging the third object (7830) again.

[0156] The electronic device (300) can reconfirm the linguistic context with the highest probability of existence among the linguistic contexts of the first word (701), the fourth word (704), and the third word (703) based on the distance between the location where the previously dragged words (e.g., the first word (701), the fourth word (704)) were dragged and the location where the third object (7830) was dragged again. For example, among the linguistic contexts of the first object (e.g., sky), the fourth object (e.g., climber), and the third object (e.g., rock wall), the context in which a climber is walking against the background of the sky and the left rock wall may have the highest probability of existence.

[0157] The electronic device (300) may generate a prompt that instructs the generation of an image including a first object, a fourth object, and a third object. For example, the electronic device (300) may generate a prompt that includes a first word (701), a fourth word (704), a third word (703), and a linguistic context that linguistically links the first word (701), the fourth word (704), and the third word (703).

[0158] The electronic device (300) may display a preview area (7930) of the third object corresponding to the prompt until user input to drag the third object again ends. For example, the electronic device (300) may display the border of the third object instructed to be created according to the prompt. The preview area (7930) of the third object may be displayed at the location where the third object (7830) was dragged according to user input (795).

[0159] The electronic device (300) can generate an image (7100) including the first object (721), the fourth object (744), and the third object (7103) based on a prompt instructing to generate an image including the first object, the fourth object, and the third object in response to the termination of user input (795) for dragging the third object (7830) again. The electronic device (300) can generate an image (7100) including the third object (7103) corresponding to the third word (703) at the location where the third object was dragged. The electronic device (300) can display the generated image (7100) on the display (330).

[0160] FIG. 8 is a drawing illustrating a display of an electronic device that receives user input for dragging an object into a text window or user input for modifying an object, according to one embodiment.

[0161] In response to receiving user input (815) of dragging one of a plurality of objects (e.g., a fourth object (7440)) into a text window (420), the electronic device (300) may generate a prompt instructing to create an image containing a plurality of objects (e.g., a first object (721) and a third object (7103) of FIG. 7e) included in a previously created image (e.g., an image (7100) of FIG. 7e) excluding the dragged object (7440). For example, the electronic device (300) may mask an area corresponding to the dragged object (7440) in the previously created image (e.g., an image (7100) of FIG. 7e). The electronic device (300) may generate a prompt instructing to create an image in which the masked area is inpainted.

[0162] According to one embodiment, the electronic device (300) can generate an image (810) comprising a plurality of objects (e.g., the first object (721) and the third object (7103) of FIG. 7e) included in a previously generated image (e.g., the image (7100) of FIG. 7e) excluding the dragged object (7440), based on an additional image in which the area where the dragged object (7440) is displayed is masked and a generated prompt. The electronic device (300) can display the generated image (810) on a display (330).

[0163] The electronic device (300) may display on the display (330) an area (820) for receiving user input that instructs to change an object in response to receiving user input that selects an object (813) included in the generated image (810). User input that selects an object (813) included in the generated image (810) may be a user input that is distinct from user input that drags an object included in the generated image (810). For example, user input that selects an object included in the generated image may be a user input that touches an object for a specified amount of time.

[0164] The electronic device (300) can change the object (813) in response to receiving user input that instructs the change of the object included in the generated image. The user input that instructs the change of the object may be an input that indicates the area (width, height, and / or depth) where the object (813) is displayed and the location where the object (813) is displayed. The electronic device (300) can change the object (813) in response to the input that indicates the area (width, height, and / or depth) where the object (813) is displayed and the location where the object (813) is displayed.

[0165] The electronic device (300) may generate a prompt instructing the creation of an image that includes an area (width, height, and / or depth) where an object is displayed according to user input and an object (e.g., a third object) corresponding to the location where the object is displayed. For example, the electronic device (300) may mask the area where the object (813) is displayed on the created image. The electronic device (300) may inpaint the masked area and generate a prompt instructing the creation of an image that includes an area (width, height, and / or depth) and a location corresponding to the area (width, height, and / or depth) and the location indicated by user input.

[0166] The electronic device (300) can generate an image (815) containing an object (823) corresponding to an area (width, height, and / or depth) and location indicated by user input, based on an additional image in which the area where the object is displayed is masked and a generated prompt. The electronic device (300) can display the generated image (815) on a display (330).

[0167] FIGS. 9A, FIGS. 9B, FIGS. 9C, FIGS. 9D, and FIGS. 9E are drawings illustrating a display of an electronic device that displays an image on a display in response to user input according to one embodiment.

[0168] In FIGS. 9a, 9b, 9c, 9d, and 9e, content that overlaps with what is described in FIGS. 7a, 7b, 7c, 7d, or 7e will be omitted.

[0169] Referring to FIG. 9a, when the electronic device (300) receives user input dragging the second word (902) to an image display window, it can check whether an object is included in a previously generated image (e.g., image (815) of FIG. 8).

[0170] The electronic device (300) can determine the linguistic context of the previously dragged words (e.g., first word (901), third word (903)) and the second word (902) based on the position where the second word (902) is dragged according to user input (915) that drags the second word (902) to the image display window (410) when the object is included in the previously generated image.

[0171] The electronic device (300) can determine the linguistic context with the highest probability of existence among the linguistic contexts of the first word (901), the third word (903), and the second word (902) based on the distance between the location where the previously dragged words (e.g., the first word (901), the third word (903)) were dragged and the location where the second word (902) was dragged. For example, among the linguistic contexts of the first object (e.g., sky), the third object (e.g., rock wall), and the second object (e.g., tree), the linguistic context in which a rock wall with the sky as a background and a tree located higher than the rock wall may have the highest probability of existence.

[0172] According to one embodiment, the linguistic context between a plurality of words (e.g., a second word (902), a third word (903)) may be determined based on the order in which each of the plurality of words is dragged onto an image display window. When the electronic device (300) receives user input to drag a specific word (e.g., a second word (902)) onto an image display window, it may generate a prompt that reflects the previously dragged word (or object) and the location where the previously dragged word (or object) was dragged.

[0173] For example, according to FIG. 9a, when the electronic device (300) receives user input to drag a second word (902) onto an image display window (410) that displays a previously generated image containing a third object, it may generate a prompt to generate an image containing the second object. The electronic device (300) may determine the linguistic context of the third word (903) and the second word (902) corresponding to the second word (902) and the location where the second word (902) was dragged, as well as the location where the third word (903) and the third word (903) were dragged, and may generate a prompt corresponding to the determined linguistic context. For example, the electronic device (300) can determine the linguistic context of the second word (902) and the third word (903) that determines that the location where the second word (902) is dragged is located higher on the image display window (410) than the location where the third word (903) is dragged, and that the second object is located higher on the image display window (410) than the third object. The electronic device (300) can generate a prompt corresponding to the determined linguistic context and generate an image (920) based on the generated prompt.

[0174] Referring to FIG. 9b, the electronic device (300) can determine the linguistic context of the previously dragged words (e.g., first word (901), second word (902), third word (903)) and the fourth word (904) based on the position where the fourth word (904) is dragged according to user input dragging the fourth word (904) to an image display window (410).

[0175] According to one embodiment, the electronic device (300) can determine the linguistic context of the previously dragged words (e.g., first word (901), second word (902), third word (903)) and the fourth word (904) based on the coordinates on the image display window (410) where the fourth word (904) is dragged. For example, the electronic device (300) can determine the linguistic context of the previously dragged words (e.g., first word (901), second word (902), third word (903)) and the fourth word (904) differently depending on whether the coordinates on the image display window (410) where the fourth word (904) is dragged are less than a specified distance from the boundary of the image display window (410).

[0176] For example, the linguistic context connecting the first word (901) and the fourth word (904) may reflect that if the position where the fourth word (904) is dragged is less than the boundary of the image display window (410) and a specified distance, at least a portion of the fourth object corresponding to the fourth word (904) is not included in the image. The electronic device (300) may generate a prompt corresponding to the determined linguistic context and generate an image (940) containing the fourth object (944) based on the generated prompt.

[0177] Referring to FIG. 9c, the electronic device (300) receives a user input (955) that drags a fourth object (9440) (e.g., the fourth object (944) in FIG. 9B) included in a previously generated image (940), and can determine the location where the fourth object (9440) was dragged according to the user input (955).

[0178] The electronic device (300) can determine the linguistic context of the previously dragged words (e.g., first word (901), second word (902), third word (903)) and fourth word (904) based on the position where the fourth object (9440) is dragged according to user input (955) of dragging the fourth object (9440).

[0179] For example, the electronic device (300) can determine a linguistic context in which a climber climbs the side of a rock wall with the sky and trees as a background. The electronic device (300) can generate a prompt corresponding to the determined linguistic context and generate an image (960) including a first object (911), a second object (922), a third object (913), and a fourth object (964) based on the generated prompt.

[0180] Referring to FIG. 9d, the electronic device (300) receives a user input (975) that drags the fourth object (9640) (e.g., the fourth object (964) in FIG. 9c) included in the previously generated image (960), and can determine the location where the fourth object (9640) was dragged according to the user input (975).

[0181] The electronic device (300) can reconfirm the linguistic context of the previously dragged words (e.g., first word (901), second word (902), third word (903)) and fourth word (904) based on the position where the fourth object (9640) was dragged according to user input (975) of dragging the fourth object (9640) again.

[0182] The electronic device (300) can reconfirm the linguistic context with the highest probability of existence among the linguistic contexts of the first word (901), the second word (902), the third word (903), and the fourth word (904) based on the distance between the location where the previously dragged words (e.g., the first word (901), the second word (902), the third word (903)) were dragged and the location where the fourth object was dragged again. For example, the electronic device (300) can determine a linguistic context in which a climber climbs the central part of a rock wall with the sky and trees as a background. The electronic device (300) can generate a prompt corresponding to the determined linguistic context and generate an image based on the generated prompt.

[0183] Referring to FIG. 9e, the electronic device (300) receives a user input (995) to drag a fourth object included in a previously generated image again, and can determine the location where the fourth object (9840) was dragged according to the user input (995).

[0184] The electronic device (300) can reconfirm the linguistic context of the previously dragged words (e.g., first word (901), second word (902), third word (903)) and fourth word (904) based on the position where the fourth object (9840) was dragged according to user input (995) of dragging the fourth object (9840) again.

[0185] The electronic device (300) can reconfirm the linguistic context with the highest probability of existence among the linguistic contexts of the first word (901), the second word (902), the third word (903), and the fourth word (904) based on the distance between the location where the previously dragged words (e.g., the first word (901), the second word (902), the third word (903)) were dragged and the location where the fourth object was dragged again. For example, the electronic device (300) can determine a linguistic context in which a climber is standing on a rock wall with the sky and trees as a background. The electronic device (300) can generate a prompt corresponding to the determined linguistic context and generate an image including the first object (911), the second object (922), the third object (913), and the fourth object (9104) based on the generated prompt.

[0186] The electronic device (300) can display at least a portion (905) of the generated prompt on a text window (420).

[0187] FIGS. 10a, FIGS. 10b, FIGS. 10c, and FIGS. 10d are drawings illustrating an electronic device that generates different images depending on the position where a word is dragged according to one embodiment.

[0188] Referring to FIG. 10a, the electronic device (300) can generate an image (1010) containing a word (e.g., a first word (1001), a second word (1002), or a third word (1003)) and an object (e.g., a first object (1011), a second object (1012), or a third object (1013)) in response to dragging the word onto an image display window (410), and display it on a display (330).

[0189] Referring to FIG. 10b, the electronic device (300) can display an image (1020) containing a second object (1012) with a third object (1013) as the background on a display (330).

[0190] The electronic device (300) can receive user input to drag a first word (1001) to an image display window (410). The location where the first word (1001) is dragged may be a first location (10211), a second location (10212), a third location (10213), or a fourth location (10214), but is not limited to the above examples.

[0191] Referring to FIG. 10c, the electronic device (300) can receive user input to drag a first word (1001) to a first position (10211).

[0192] The electronic device (300) may display a notification window (1005) indicating that the distance between the location where the first word (1001) was dragged and the location where the second word (1002) was dragged is greater than or equal to a specified value in response to receiving user input to drag the first word (1001) to the first location (10211). The notification window (1005) may be displayed in a portion of the display (330) and may be displayed on a different layer from the image display window (410) or the text window (420). The notification window (1005) may include text (1006) for checking whether to display the first object (1011) and an area (1007, 1008) for receiving user input.

[0193] When the electronic device (300) receives user input selecting an area (1007) that instructs to display a first object (1011) included in a notification window, it may generate a prompt that instructs to generate an image containing the first object (1011). For example, the first object (1011) may be an object that is not related to an object (e.g., a second object (1012)) included in a previously generated image (1020). Based on the generated prompt, the electronic device (300) may generate an image (1030) containing the first object (1011) and the second object (1012) that are not related to each other.

[0194] When the electronic device (300) receives user input selecting an area (1008) that instructs not to display the first object (1011) included in the notification window, it can display a previously generated image (1020) on the display (330).

[0195] According to FIG. 10d, the electronic device (300) can receive user input to drag the first word (1001) to a second position (10212), a third position (10213), or a fourth position (10214).

[0196] The electronic device (300) can determine the linguistic context of the first word and the second word based on the distance between the location where the first word is dragged (e.g., the second location (10212), the third location (10213), or the fourth location (10214)) and the location where the second word is dragged. For example, the distance between the second location (10212) and the location where the second word (1002) is dragged may correspond to a linguistic context in which the second object (1012) is holding the first object (1011) with one arm extended. The distance between the third location (10213) and the location where the second word (1002) is dragged may correspond to a linguistic context in which the second object (1012) is holding the first object (1011) with both hands. The distance between the fourth position and the position where the second word (1002) is dragged can correspond to a linguistic context in which the second object (1012) is holding the first object (1011) around the mouth with both hands.

[0197] The electronic device (300) can generate a prompt corresponding to a determined linguistic context and display the generated image (e.g., image (1040), image (1050), or image (1060)) on the display (330).

[0198] FIG. 11 is a method flowchart of an electronic device according to one embodiment.

[0199] An electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (300) of FIG. 3) may, in operation 1110, display on the display a text window that receives a plurality of word inputs including a first word and a second word, and an image display window that is displayed separately from the text window.

[0200] The electronic device (300) can display an application screen on the display (330). For example, the electronic device (300) can activate the display (330) to display an application screen in response to a touch input on the display (330) or an input to a hard key. The application may refer to an application configured to display an image in at least a portion of the display (330).

[0201] The electronic device (300) may display a text window on the display (330) that displays at least one word. The text window may refer to an area on the display (330) set to receive user input for entering a word or to display at least one word entered by the user.

[0202] The electronic device (300) can receive user input that inputs a word on the display (330). The electronic device (300) can detect user input that inputs a word on the display (330) using at least one sensor (e.g., a touch sensor).

[0203] The electronic device (300) can display at least one word entered by the user on the text window in response to receiving user input entering a word into the text window.

[0204] The electronic device (300) can display an image display window on the display (330) that is displayed separately from the text window. The image display window may refer to an area on the display (330) set to display an image.

[0205] The electronic device (300) can receive user input by dragging a word to an image display window. Depending on the user input, an area where an object corresponding to the word is displayed can be determined.

[0206] In operation 1120, the electronic device (300) receives a first input that drags a first word to an image display window and can generate a first prompt that instructs to generate a first image containing a first object corresponding to the first word in response to the first input.

[0207] The electronic device (300) can generate a prompt that instructs the creation of an image containing an object corresponding to a word in response to user input dragging a word to an image display window. The electronic device (300) can determine the location where the word was dragged according to user input. For example, the electronic device (300) can determine that the input for dragging the word has ended when the user's touch input on the display (330) ends, and can determine the location where the word was dragged. The prompt generated by the electronic device (300) can instruct the creation of an image containing an object corresponding to the word at the location where the word was dragged.

[0208] According to one embodiment, the electronic device (300) may generate a prompt to perform image generation including an object based on a character (e.g., a word) entered by a user, information related to an object corresponding to the character (e.g., a word), and parameter information related to an application. The character (e.g., a word) entered by the user may be used as a prompt source for image generation including an object. The character may include character-based content. The electronic device (300) may generate a prompt to perform image generation including an object based on a word entered by the user and / or information related to an object corresponding to the word.

[0209] According to one embodiment, the electronic device (300) can generate a prompt that instructs to generate an image with an object corresponding to the word as the background when the word is a word indicating a background.

[0210] The electronic device (300) can display a preview area of ​​an object corresponding to a prompt until user input dragging a word ends. For example, the electronic device (300) can display the border of an object instructed to be created according to the prompt. The preview area of ​​the object can be displayed at the location where the word was dragged according to user input.

[0211] The electronic device (300) can generate a first image based on a first prompt in operation 1130.

[0212] The electronic device (300) can generate an image containing an object corresponding to a word based on a prompt instructing the generation of an image containing an object corresponding to a word in response to the termination of user input in which the word is dragged to an image display window. The electronic device (300) can generate an image containing an object corresponding to a word at the location where the word was dragged. The electronic device (300) can display the generated image on the display (330).

[0213] For example, the electronic device (300) can generate (or acquire) an image containing an object in relation to a generated prompt. The electronic device (300) receives a prompt source (e.g., a first word, a second word) related to generating an image containing an object based on interaction with a user, and can generate (e.g., regenerate or reconstruct) an image containing an object on a server or on a device based on the prompt source. The electronic device (300) can provide the prompt to a generative artificial intelligence on a device and / or server to execute an image generation process based on the generated prompt. The electronic device (300) can provide a prompt requesting image generation (e.g., a question or instruction to be entered into the generative artificial intelligence) to the generative artificial intelligence. The electronic device (300) can generate (or acquire) an image according to the image generation process executed in relation to the prompt by the on-device artificial intelligence (e.g., the generative AI model of FIG. 2). The electronic device (300) can receive (or obtain) an image from the server containing an object according to an image generation process executed in relation to a prompt in the server artificial intelligence.

[0214] For example, the electronic device (300) may generate (or acquire) an image based on a generated prompt and an additional image. The additional image may refer to either a generated image or an image in which some areas of the generated image are masked. The electronic device (300) may generate an image in which at least one object mark included in the additional image is retained. If the additional image is an image in which some areas are masked, the electronic device (300) may generate an image in which the masked areas are inpainted.

[0215] It should be understood that the operation of the electronic device (300) of the present disclosure generating (or modifying) a prompt and the operation of generating (or modifying) an image based on the prompt may be subject to the description of the operation of the electronic device (300) of FIG. 3 generating a prompt and the operation of generating an image based on the prompt.

[0216] In operation 1140, the electronic device receives a second input for dragging a second word to the image display window, and in response to the second input, may generate a second prompt instructing to create a second image including a second object and a first object based on the distance between a first position where the first word was dragged according to the first input and a second position where the second word was dragged according to the second input.

[0217] The electronic device (300) can check whether an object is included in a previously generated image when it receives user input to drag a second word, which is distinct from a previously dragged first word, to an image display window. According to one embodiment, the electronic device (300) can check whether an object is included in a previously generated image if there is a first word that has been dragged to an image display window.

[0218] The electronic device (300) can determine whether to display a first object corresponding to a first word and a second object corresponding to a second word in a mutually related form based on the distance between a first position where a first word is dragged and a second position where a second word is dragged according to user input dragging the word to an image display window, when an object is included in a previously generated image.

[0219] Displaying the first object and the second object in a mutually related form may refer to inferring (or determining) a linguistic context that linguistically connects the first word and the second word, and displaying an image containing the first object and the second object corresponding to the inferred (or determined) linguistic context that linguistically connects the first word and the second word. Displaying the first object and the second object in a mutually unrelated form may refer to displaying an image containing the first object included in an image generated in response to receiving a first input when no previously generated image exists, and the second object included in an image generated in response to receiving a second input when no previously generated image exists.

[0220] According to one example, the electronic device (300) may determine that the first word and the second word cannot be linguistically connected when checking (or determining) the linguistic context of the first word and the second word. If the first word and the second word cannot be linguistically connected, the electronic device (300) may determine to display the first object and the second object in a non-related form. According to one example, if the first word and the second word can be linguistically connected, the electronic device (300) may determine to display the first object and the second object in a related form.

[0221] The electronic device (300) can determine the linguistic context of the first word and the second word based on the distance between the first position and the second position. The electronic device (300) can determine the context with the highest probability of existence among the linguistic contexts corresponding to the case where the first word and the second word are separated by the distance between the first position and the second position as the linguistic context that linguistically connects the first word and the second word.

[0222] The electronic device (300) can determine the linguistic context of the first word and the second word based on the distance between the first location and the second location on the server (301) or on the device. The electronic device (300) may use an artificial intelligence model on the device and / or server (301) (e.g., an artificial intelligence model trained to infer the linguistic context between words) to determine the linguistic context of the first word and the second word based on the distance between the first location and the second location determined by the artificial intelligence model of the server (301). The electronic device (300) may receive (or obtain) the linguistic context of the first word and the second word from the server (301) based on the distance between the first location and the second location determined by the artificial intelligence model of the server (301). The server (301) that performs the action of determining the linguistic context of the first word and the second word may be a server distinct from the server that generates an image related to the prompt.

[0223] According to one embodiment, the electronic device (300) can determine the linguistic context of the first word and the second word based on at least some features (e.g., size, pose) of the first object or the second object included in the generated image when the first object or the second object is included in the generated image. The electronic device (300) can determine the linguistic context with the highest probability of existence among the linguistic contexts of the first object and the second object corresponding to at least some features (e.g., size, pose) of the first object or the second object as well as the distance between the first position and the second position. For example, even if the distance between the first position and the second position is the same, the electronic device (300) may determine the linguistic context of the first word and the second word differently depending on whether the size of the first object or the second object displayed in the generated image is larger or smaller.

[0224] According to one embodiment, the electronic device (300) can determine the linguistic context of the first word and the second word based on a second location where the word (e.g., the second word) is dragged according to user input that drags the word (e.g., the second word) to an image display window. The electronic device (300) can determine the linguistic context of the first word and the second word based on the coordinates of the second location on the image display window. For example, the electronic device (300) can determine the linguistic context of the first word and the second word differently depending on whether the second location is less than a specified distance from the boundary of the image display window. The linguistic context of the first word and the second word may reflect that, when the second location is less than a specified distance from the boundary of the image display window, at least a part of the second object corresponding to the second word is not included in the image.

[0225] According to one embodiment, the linguistic context between a plurality of words dragged to the display (330) may be determined based on the order in which each of the plurality of words is dragged onto the image display window. When the electronic device (300) receives user input to drag a specific word (e.g., a second word) onto the image display window, it may generate a prompt that reflects the previously dragged word (or object) and the location where the previously dragged word (or object) was dragged. Since the electronic device (300) generates a prompt that reflects the previously dragged word (or object) and the location where the previously dragged word (or object) was dragged, the prompt generated by the electronic device (300) may vary depending on the previously dragged word (or object) and the location of the previously dragged word (or object).

[0226] If the electronic device (300) determines to display the second object and the first object in a mutually related form, it may generate a prompt instructing to create an image containing the second object and the first object in a mutually related form. For example, the electronic device (300) may generate a prompt containing a first word, a second word, and the linguistic context of the first word and the second word. The electronic device (300) may generate a prompt instructing to create an image containing the first object and the second object corresponding to the linguistic context of the first word and the second word. If the electronic device (300) determines to display the second object and the first object in a mutually unrelated form, it may generate a prompt instructing to create an image containing the second object and the first object in a mutually unrelated form. According to one embodiment, the electronic device (300) may generate a prompt instructing to modify a previously created image. The operation of generating a prompt that instructs the electronic device (300) of the present disclosure to generate an image can be replaced with the operation of generating a prompt that instructs the generated image to be modified.

[0227] The electronic device (300) can generate a second image based on a second prompt in operation 1150.

[0228] The electronic device (300) can generate an image based on the generated prompt and display it on the display (330).

[0229] The electronic device (300) can receive user input to drag one of the multiple objects (e.g., a second object) included in a previously generated image, and can determine a third location where one of the multiple objects (e.g., a second object) was dragged according to the user input.

[0230] The electronic device (300) may determine whether to display the second object and the first object in a related form based on the distance between the third position and the first position (or the second position). If the electronic device (300) determines to display the second object and the first object in a related form, it may generate a prompt instructing to create an image containing the second object and the first object in a related form. For example, the electronic device (300) may determine the linguistic context of the first word and the second word corresponding to the distance between the third position and the first position (or the second position), and generate a prompt containing the linguistic context of the first word and the second word.

[0231] The electronic device (300) may display a preview area of ​​an object corresponding to a prompt until user input for dragging one of a plurality of objects (e.g., a second object) ends. For example, the electronic device (300) may display the border of the second object that is instructed to be created according to the prompt. The preview area may be displayed at the location (e.g., a third location) where the object (e.g., the second object) was dragged according to user input.

[0232] The electronic device (300) may generate an image including the second object and the first object and display it on the display (330) in response to the termination of user input in which one of the plurality of objects (e.g., the second object) is dragged. The second object and the first object may be objects generated to reflect the linguistic context of the first word and the second word based on the distance between the third position and the first position (or the second position).

[0233] The electronic device (300) can display at least a portion of a modified prompt in a text window to reflect the linguistic context of the first word and the second word based on the distance between the third position and the first position (or second position).

[0234] According to one embodiment, the electronic device (300) may generate a prompt instructing the creation of an image containing objects included in the created image, excluding the dragged objects, in response to receiving user input to drag objects included in the created image into a text window. For example, the electronic device (300) may mask the area corresponding to the dragged objects in the previously created image. The electronic device (300) may generate a prompt instructing the creation of an image in which the masked area is inpainted.

[0235] According to one embodiment, the electronic device (300) can generate an image including a plurality of objects included in a previously generated image excluding the dragged object, based on an additional image in which the area where the dragged object is displayed is masked and a generated prompt. The electronic device (300) can display the generated image on a display (330).

[0236] The electronic device (300) may display on the display (330) an area capable of receiving a user input that instructs to modify an object in response to receiving a user input that selects an object included in the generated image. The user input that selects an object included in the generated image may be a user input that is distinct from a user input that drags an object included in the generated image. For example, the user input that selects an object included in the generated image may be a user input that touches an object for a specified amount of time.

[0237] The electronic device (300) can deform an object in response to receiving user input that instructs the deformation of an object included in a generated image. The user input that instructs the deformation of an object may be an input that instructs the area (width, height, and / or depth) where the object is displayed and the location where the object is displayed. The electronic device (300) can deform the object in response to the input that instructs the area (width, height, and / or depth) where the object is displayed and the location where the object is displayed.

[0238] According to one embodiment, the electronic device may include a display. The electronic device may include a memory that stores at least one computer program containing instructions. The electronic device may include at least one processor. When the instructions are executed individually or collectively by the at least one processor, the electronic device may display on the display a text window that receives a plurality of word inputs including a first word and a second word, and an image display window that is displayed separately from the text window. The instructions may receive a first input that drags the first word to the image display window. The electronic device may generate a first prompt that instructs to create a first image containing a first object corresponding to the first word in response to the first input. The instructions may create the first image based on the first prompt. The electronic device may receive a second input that drags the second word to the image display window. The electronic device may generate a second prompt that instructs the generation of a second image including the second object and the first object, based on the distance between the first word, the second word, and the first position where the first word is dragged according to the first input and the second position where the second word is dragged according to the second input, in response to the second input. The electronic device may generate the second image based on the second prompt.

[0239] In an electronic device according to one embodiment, instructions may determine whether to display the second object and the first object in a mutually related form based on the distance between the first position and the second position. If it is determined to display the second object and the first object in a mutually related form, the second prompt generated may be a prompt instructing to generate the second image to include the second object and the first object in a mutually related form.

[0240] In an electronic device according to one embodiment, when it is determined to display the second object and the first object in a mutually unrelated form, the generated second prompt may be a prompt that instructs to generate a second image including the first object included in the first image and a second object in a form not mutually related to the first object.

[0241] In an electronic device according to one embodiment, instructions may determine a linguistic context that linguistically connects the first word and the second word, corresponding to the distance between the first position and the second position. The generated second prompt may be a prompt that instructs to generate a second image by reflecting the linguistic context that linguistically connects the first word and the second word.

[0242] In an electronic device according to one embodiment, the instructions may determine a linguistic context that linguistically connects the first word and the second word based on the order of the first input and the second input.

[0243] In an electronic device according to one embodiment, the instructions may determine a third location where the second object is dragged according to the third input in response to receiving a third input that drags the second object included in the second image. The instructions may determine whether to display the second object and the first object in a related form based on the distance between the third location and the first location. If the instructions determine to display the second object and the first object in a related form, the instructions may generate a third prompt instructing to modify the second image to include the second object and the first object in a related form. The instructions may modify the second image based on the third prompt.

[0244] In an electronic device according to one embodiment, the third prompt may include information instructing to inpaint an area corresponding to the second object on the second image.

[0245] In an electronic device according to one embodiment, the instructions may include an area in the image display window for receiving user input that instructs to change at least one object. The instructions may change at least one of the size, position, or angle at which the at least one object is displayed on the display based on user input entered into the area for receiving user input that instructs to change at least one object.

[0246] In an electronic device according to one embodiment, the second prompt is,

[0247] If the coordinates of the second position on the image display window are less than the boundary of the image display window and a specified distance, it may be a prompt that instructs to generate the second image in a form that does not include at least a part of the second object.

[0248] In an electronic device according to one embodiment, the instructions may display at least a portion of a second prompt instructing to generate the second image on a text display window on the display.

[0249] In an electronic device according to one embodiment, the instructions may include the second image on the image display window and display it on the display.

[0250] A method of operation of an electronic device according to one embodiment may include an operation of displaying on a display a text window that receives a plurality of word inputs including a first word and a second word, and an image display window that is displayed separately from the text window. A method of operation of an electronic device may include an operation of receiving a first input that drags the first word to the image display window. A method of operation of an electronic device may include an operation of generating a first prompt that instructs to generate a first image including a first object corresponding to the first word in response to the first input. A method of operation of an electronic device may include an operation of generating a first image based on the first prompt. A method of operation of an electronic device may include an operation of receiving a second input that drags the second word to the image display window. A method of operation of an electronic device may include an operation of generating a second prompt that instructs to generate a second image including the second object and the first object in response to the second input, based on the distance between the first word, the second word, a first position where the first word was dragged according to the first input, and a second position where the second word was dragged according to the second input. The method of operation of the electronic device may include the operation of generating a second image based on the second prompt.

[0251] A method of operation of an electronic device according to one embodiment may include an operation of determining whether to display the second object and the first object in a mutually related form based on the distance between the first position and the second position. If it is determined to display the second object and the first object in a mutually related form, the second prompt generated may be a prompt instructing to generate the second image to include the second object and the first object in a mutually related form.

[0252] In a method of operation of an electronic device according to one embodiment, when it is determined to display the second object and the first object in a form that is not mutually related, the generated second prompt may be a prompt that instructs to generate a second image including the first object included in the first image and a second object in a form that is not mutually related to the first object.

[0253] A method of operating an electronic device according to one embodiment may include an operation of determining a linguistic context that linguistically connects the first word and the second word, corresponding to the distance between the first position and the second position. The generated second prompt may be a prompt that instructs to generate a second image by reflecting the linguistic context that linguistically connects the first word and the second word.

[0254] A method of operation of an electronic device according to one embodiment may include an operation of determining a linguistic context that linguistically connects the first word and the second word based on the order of the first input and the second input.

[0255] A method of operation of an electronic device according to one embodiment may include, in response to receiving a third input that drags the second object included in the second image, an operation of checking a third position where the second object is dragged according to the third input. A method of operation of the electronic device may include an operation of determining whether to display the second object and the first object in a mutually related form based on the distance between the third position and the first position. A method of operation of the electronic device may include, if it is determined to display the second object and the first object in a mutually related form, an operation of generating a third prompt that instructs to modify the second image to include the second object and the first object in a mutually related form. A method of operation of the electronic device may include an operation of modifying the second image based on the third prompt.

[0256] In a method of operating an electronic device according to one embodiment, the third prompt may include information instructing to inpaint an area corresponding to the second object on the second image.

[0257] In a method of operation of an electronic device according to one embodiment, the method may include an operation of including and displaying an area receiving user input that instructs to change at least one object in the image display window. The method of operation of the electronic device may include an operation of changing at least one of the size, position, or angle at which the at least one object is displayed on the display based on user input entered into the area receiving user input that instructs to change the at least one object.

[0258] In a method of operating an electronic device according to one embodiment, the second prompt may be a prompt that instructs the second image to be generated in a form that does not include at least a part of the second object when the coordinates of the second position on the image display window are less than the boundary of the image display window and a specified distance.

[0259] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.

[0260] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any possible combination of items listed together in the corresponding phrase. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0261] As used in this document, the term "module" may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be a component formed as a whole, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0262] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (100)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (100)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' merely means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0263] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0264] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the components of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to the integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device, display; Memory for storing at least one computer program including instructions; and It includes at least one processor, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, A text window receiving multiple word inputs including a first word and a second word, and an image display window displayed separately from the text window are displayed on the display, and Receiving a first input that drags the first word to the image display window, and generating a first prompt that instructs to generate a first image including a first object corresponding to the first word in response to the first input, A first image is generated based on the first prompt above, and Receives a second input for dragging the second word to the image display window, and in response to the second input, generates a second prompt instructing to generate a second image including the second object and the first object based on the distance between the first word, the second word, the first position where the first word was dragged according to the first input, and the second position where the second word was dragged according to the second input. An electronic device that generates a second image based on the above second prompt.

2. In claim 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, Based on the distance between the first position and the second position, determine whether to display the second object and the first object in a mutually related form, and An electronic device in which, when it is determined to display the second object and the first object in a mutually related form, the second prompt generated is a prompt that instructs to generate the second image to include the second object and the first object in a mutually related form.

3. In Paragraph 2, An electronic device in which, when it is decided to display the second object and the first object in a mutually unrelated form, the generated second prompt is a prompt that instructs to generate a second image including the first object included in the first image and the second object in a form not mutually related to the first object.

4. In claim 2, when the instructions are executed individually or collectively by the at least one processor, the electronic device, Based on the distance between the first position and the second position, a linguistic context is determined that linguistically connects the first word and the second word, and An electronic device in which the second prompt generated above is a prompt that instructs to generate a second image by reflecting a linguistic context that linguistically connects the first word and the second word.

5. In claim 4, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that determines a linguistic context for linguistically connecting the first word and the second word based on the order of the first input and the second input.

6. In paragraph 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, In response to receiving a third input for dragging the second object included in the second image, a third position to which the second object was dragged according to the third input is checked, and It determines whether to display the second object and the first object in a mutually related form based on the distance between the third position and the first position, and If it is decided to display the second object and the first object in a mutually related form, a third prompt is generated to instruct the second image to be modified to include the second object and the first object in a mutually related form, and An electronic device that modifies a second image based on the third prompt above.

7. In Clause 6, the third prompt is, An electronic device comprising information instructing to inpaint an area corresponding to the second object on the second image.

8. In claim 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An area receiving user input that instructs to change at least one object is included in the image display window and displayed, and An electronic device that changes at least one of the size, position, or angle at which the at least one object is displayed on a display, based on user input entered in an area receiving user input that instructs to change the at least one object.

9. In claim 1, the second prompt is, An electronic device, which is a prompt instructing to generate the second image in a form that does not include at least a part of the second object when the coordinates of the second position on the image display window are less than a specified distance from the boundary of the image display window.

10. In claim 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that displays at least a portion of a second prompt instructing to generate the second image on a text display window on the display.

11. In claim 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that includes the second image above on the image display window and displays it on the display.

12. In a method of operating an electronic device, An operation to display on a display a text window that receives multiple word inputs including a first word and a second word, and an image display window that is displayed separately from the text window; An action of receiving a first input that drags the first word to the image display window; an action of generating a first prompt that instructs to generate a first image including a first object corresponding to the first word in response to the first input; The operation of generating a first image based on the first prompt above; An action of receiving a second input for dragging the second word to the image display window; an action of generating a second prompt corresponding to the second input, which instructs to generate a second image including the second object and the first object based on the distance between the first word, the second word, the first position where the first word was dragged according to the first input, and the second position where the second word was dragged according to the second input; and A method of operation of an electronic device comprising the operation of generating a second image based on the second prompt above.

13. In claim 12, the method of operating the electronic device is, The operation includes determining whether to display the second object and the first object in a mutually related form based on the distance between the first position and the second position, and A method of operation of an electronic device in which, when it is decided to display the second object and the first object in a mutually related form, the second prompt generated is a prompt that instructs to generate the second image to include the second object and the first object in a mutually related form.

14. A method of operation of an electronic device according to claim 13, wherein, when it is determined to display the second object and the first object in a mutually unrelated form, the generated second prompt is a prompt that instructs to generate a second image including the first object included in the first image and the second object in a form not mutually related to the first object.

15. In claim 13, the method of operating the electronic device is, Based on the distance between the first position and the second position, the operation includes determining a linguistic context that linguistically connects the first word and the second word, and A method of operation of an electronic device, wherein the second prompt generated above is a prompt that instructs to generate a second image by reflecting a linguistic context that linguistically connects the first word and the second word.

Citation Information

Patent Citations

  • Display device

    KR1020250016921A

  • Nitrogen oxide reduction system and reduction method using the same

    KR1020260020004A

  • Semiconductor package

    KR1020260037488A

  • buoy and ship route guide system using the same

    KR102490585B1

  • Merging multiple images as input to an ai image generation algorithm

    US20240193821A1