Image-generating method, electronic apparatus, and program and storage medium

An AI model generates dynamic graphic elements from text input, addressing the limitations of existing communication technologies by creating personalized and engaging stickers with visual, audio, and haptic effects, enhancing user expression and interaction.

WO2026155551A1PCT designated stage Publication Date: 2026-07-23SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2026-01-14
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing communication technologies lack the ability to generate dynamic and personalized graphic elements, such as stickers, that effectively convey user emotions and personality, limiting the richness and engagement of online interactions.

Method used

An AI model is utilized to generate images based on text input, allowing for the creation of dynamic graphic elements like stickers that incorporate visual, audio, and haptic effects, leveraging multimodal models and deep learning to personalize user experiences.

Benefits of technology

The AI model enables the generation of engaging and personalized graphic elements that enhance user expression and communication, providing a richer and more interactive experience in messaging and social media platforms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2026000848_23072026_PF_FP_ABST
    Figure KR2026000848_23072026_PF_FP_ABST
Patent Text Reader

Abstract

According to an embodiment, provided is a method comprising an operation of receiving a text input of a user. The method may include an operation of acquiring description information for generating an image on the basis of the text input. The method may include an operation of acquiring a visual object on the basis of the description information. The method may include an operation of requesting image generation based on a prompt including the description information. The method may include an operation of displaying the visual object while the image is being generated on the basis of the prompt. The method may include an operation of acquiring an image generated on the basis of the prompt. The method may include an operation of displaying the generated image.
Need to check novelty before this filing date? Find Prior Art

Description

Method, electronic device, program, and storage medium for generating an image

[0001] The present disclosure relates to a method, electronic device, program, and storage medium for generating an image, and more specifically, to generating an image using an artificial intelligence (AI) model.

[0002] The development of information and communication technology has transformed the ways in which users form social relationships and communicate with one another in online environments. For example, people communicate with others using messaging services or social network services (SNS), and these digital means of communication are actively used not only for everyday conversations but also in various fields such as work, education, and leisure.

[0003] People express their opinions by uploading text, audio, images, and videos to messaging services or social media. Moving beyond simple text-based communication, graphic elements such as stickers, emoticons, and emojis are being used to more effectively express user emotions and enrich conversations. Graphic elements (e.g., stickers) can be static images, but are not limited to this; they can be dynamic images composed of multiple images, such as videos or animated images. Graphic elements can be combined with audio or haptic effects and output together.

[0004] Graphic elements (e.g., stickers) can not only visually convey the user's emotions or the context of a conversation, but also personalize the user experience by revealing the user's personality and make communication more interesting. Users can create graphic elements (e.g., stickers) to add their own personality.

[0005] Recently, artificial intelligence systems capable of achieving human-level intelligence are being utilized in various fields. Unlike conventional rule-based smart systems, artificial intelligence systems are systems in which machines learn, make judgments, and become smarter on their own. As artificial intelligence systems improve in recognition accuracy and gain a more accurate understanding of user preferences with continued use, existing rule-based smart systems are gradually being replaced by deep learning-based artificial intelligence systems.

[0006] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art in relation to the present disclosure.

[0007] According to one embodiment, a method may be provided that includes an operation of receiving text input from a user. The method may include an operation of obtaining descriptive information for generating an image based on the text input. The method may include an operation of obtaining a visual object based on the descriptive information. The method may include an operation of requesting the generation of an image based on a prompt including the descriptive information. The method may include an operation of displaying the visual object while the image is being generated based on the prompt. The method may include an operation of obtaining the generated image based on the prompt. The method may include an operation of displaying the generated image.

[0008] According to one embodiment, an electronic device may be provided comprising one or more processors including a memory for storing instructions and processing circuitry. When the instructions are executed individually or jointly by the one or more processors, the electronic device may receive text input from a user. When the instructions are executed individually or jointly by the one or more processors, the electronic device may obtain descriptive information for generating an image based on the text input. When the instructions are executed individually or jointly by the one or more processors, the electronic device may obtain a visual object based on the descriptive information. When the instructions are executed individually or jointly by the one or more processors, the electronic device may request the generation of an image based on a prompt containing the descriptive information. When the instructions are executed individually or jointly by the one or more processors, the electronic device may display the visual object while the image is being generated based on the prompt. When the instructions are executed individually or jointly by the one or more processors, the electronic device may obtain an image generated based on the prompt. The above instructions may display the generated image when executed individually or jointly by the one or more processors.

[0009] A computer-readable non-transitory recording medium according to one embodiment of the present invention may store at least one instruction and / or instruction that causes an electronic device to perform the method or operation of the electronic device described above when executed.

[0010] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.

[0011] FIG. 1 is a block diagram of an electronic device according to one embodiment.

[0012] FIG. 2 is a diagram illustrating the process of acquiring and displaying visual objects and images based on text input in an electronic device according to one embodiment.

[0013] FIG. 3 is a flowchart of a method according to one embodiment.

[0014] FIG. 4 is a diagram illustrating the process of displaying a visual object using an artificial intelligence model according to one embodiment.

[0015] FIG. 5 is a diagram illustrating the process of generating an image based on descriptive information according to one embodiment.

[0016] FIG. 6 is a diagram illustrating the process of displaying a visual object using keywords according to one embodiment.

[0017] FIG. 7 is a diagram illustrating the process of displaying a visual object using effect information according to one embodiment.

[0018] FIG. 8 is a diagram illustrating the process of generating a second image associated with a first image according to one embodiment.

[0019] FIG. 9 is a diagram illustrating the process of selecting a second image to be generated using second description information.

[0020] FIG. 10 is a diagram illustrating the process of receiving text input to generate a sticker image to be used in a messenger application according to one embodiment.

[0021] FIG. 11 is a diagram illustrating the process of receiving text input and drawing input for generating an image according to one embodiment.

[0022] FIG. 12 is a diagram illustrating the process of generating a sticker image by referring to conversation history in a messenger application according to one embodiment.

[0023] FIG. 13 is a diagram illustrating the process of modifying an image generated based on text input according to one embodiment.

[0024] FIG. 14 is a diagram showing an example of a system including a generative artificial intelligence model according to one embodiment.

[0025] Embodiments of the present disclosure are described below in detail with reference to the attached drawings so that those skilled in the art can easily implement them. However, the present disclosure may be embodied in various different forms and is not limited to the embodiments described herein. Furthermore, in order to clearly explain the present disclosure in the drawings, parts unrelated to the explanation have been omitted, and similar parts throughout the specification are denoted by similar reference numerals.

[0026] The terms used in this disclosure are described in their current, general form considering the functions mentioned herein; however, they may refer to various other terms depending on the intent of those skilled in the art, case law, or the emergence of new technologies. Accordingly, the terms used in this disclosure should not be interpreted solely by their names, but should be interpreted based on the meaning of the terms and the overall content of this disclosure.

[0027] Additionally, terms such as the first, second, third, ..., Nth may be used to describe various components, but the components should not be limited by these terms. These terms are used for the purpose of distinguishing one component from another.

[0028] Throughout the specification, when a part is described as being "connected" to another part, this includes not only cases where they are "directly connected," but also cases where they are "electrically connected" with other components interposed between them. Furthermore, when a part is described as "including" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.

[0029] Phrases such as "in one embodiment" appearing in various places in this disclosure do not necessarily refer to the same embodiment.

[0030] One embodiment of the present disclosure may be represented by functional block configurations and various processing steps. Some or all of these functional blocks may be implemented by various numbers of hardware and / or software configurations that execute specific functions. For example, the functional blocks of the present disclosure may be implemented by one or more microprocessors or by circuit configurations for a specific function. Additionally, for example, the functional blocks of the present disclosure may be implemented in various programming or scripting languages. The functional blocks may be implemented as algorithms executed on one or more processors. Furthermore, the present disclosure may employ prior art for electronic configuration, signal processing, and / or data processing. Terms such as "mechanism," "element," "means," and "configuration" may be used broadly and are not limited to mechanical and physical configurations.

[0031] Furthermore, the connecting lines or connecting members between the components depicted in the drawings are merely illustrative of functional connections and / or physical or circuit connections. In the actual device, connections between components may be represented by various alternative or added functional connections, physical connections, or circuit connections.

[0032] Artificial intelligence technology consists of machine learning (e.g., deep learning) and elemental technologies utilizing machine learning.

[0033] Machine learning is an algorithmic technology that classifies and learns the features of input data on its own, and the elemental technology is a technology that mimics functions such as cognition and judgment of the human brain by utilizing deep learning machine learning algorithms, and consists of the fields of linguistic understanding, visual understanding, reasoning / prediction, knowledge representation, and motion control.

[0034] The various fields where artificial intelligence technology is applied are as follows. Linguistic understanding is a technology that recognizes, applies, and processes human language and text, and includes natural language processing, machine translation, dialogue systems, question answering, and speech recognition / synthesis. Visual understanding is a technology that perceives and processes objects like human vision, and includes object recognition, object tracking, image search, person recognition, scene understanding, spatial understanding, and image enhancement. Inference and prediction is a technology that judges information to logically infer and predict, and includes knowledge / probability-based inference, optimization prediction, preference-based planning, and recommendation. Knowledge representation is a technology that automatically processes human experiential information into knowledge data, and includes knowledge construction (data generation / classification) and knowledge management (data utilization). Motion control is a technology that controls the autonomous driving of vehicles and the movement of robots, and includes motion control (navigation, collision, driving) and manipulation control (behavior control).

[0035] The first artificial intelligence model may be a multimodal model that learns and processes relationships between data of various types or various modalities, such as text and images. The first artificial intelligence model may be a large multimodal model (LMM) trained using text data and image data, and the first artificial intelligence model may generate an image or generate text associated with an image based on a text prompt, an image prompt, or a prompt consisting of text and an image input to the first artificial intelligence model. The first artificial intelligence model is not limited to an LMM, and the first artificial intelligence model may be a generative artificial intelligence model trained to generate descriptive information that elaborates on the text input based on the user's text input.

[0036] The second AI model may be an AI model trained to generate an image based on an input text prompt. The second AI model may be a text-to-image model and may be a combination of a transformer-based language model and a diffusion-based image generation model, but is not limited thereto.

[0037] The first artificial intelligence model may be any artificial intelligence model capable of performing bidirectional interaction between text and image, and the second artificial intelligence model may be any artificial intelligence model capable of performing unidirectional interaction that generates an image based on text.

[0038] The descriptive information generated by the first artificial intelligence model refers to text information describing an image to be generated based on the user's text input. The descriptive information can be generated by the first artificial intelligence model by applying a prompt containing the user's text input to the first artificial intelligence model. A prompt containing the descriptive information can be input to the second artificial intelligence model, and the second artificial intelligence model can generate an image described by the descriptive information included in the prompt based on the input prompt.

[0039] The image generated by the second artificial intelligence model may be a static image, but is not limited thereto, and may be a dynamic image made up of multiple images, for example, a video or an animated image. The image may be a graphic element, such as a sticker, for use in messaging services or social media.

[0040] A visual object is a graphic element used to describe the process of creating an image and may contain information about what the image to be created is. A visual object may consist of text, but is not limited thereto. For example, a visual object may be implemented in a form that combines text with visual effects, audio effects, or haptic effects.

[0041] The present disclosure will be described in detail below with reference to the attached drawings.

[0042] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to various embodiments.

[0043] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).

[0044] The processor (120) can control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., a program (140)), and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.

[0045] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.

[0046] The number of processors (120) may be one or more. For example, the processor (120) may have the structure of a multi-core processor such as a dual core, a quad core, or a hexa core.

[0047] The processor (120) can control the operations of the electronic device (101) by executing instructions stored in the memory (130). For example, the processor (120) may correspond to a plurality of processors that divide and collectively perform a plurality of operations among the processors.

[0048] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).

[0049] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0050] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0051] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.

[0052] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.

[0053] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).

[0054] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0055] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0056] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0057] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.

[0058] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0059] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).

[0060] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0061] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).

[0062] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) may support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.

[0063] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).

[0064] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.

[0065] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.

[0066] According to one embodiment, commands or data may be transmitted or received between an electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In one embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0067] An electronic device (101) according to one embodiment may correspond to the electronic devices (201, 401, 601, 701, 901) of FIGS. 1 to 9, and the electronic device (101) may perform the operations of the electronic devices (201, 401, 601, 701, 901) of FIGS. 1 to 9.

[0068] According to one embodiment, an electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may be provided, comprising a memory (130) for storing instructions; and one or more processors (120) including processing circuitry. When the instructions are executed individually or jointly by the one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may receive text input (2011, 510) from a user. When the above instructions are executed individually or jointly by the one or more processors (120), they may obtain descriptive information (520, 620, 720, 820) for generating an image (2060, 560, 860, 9060) based on the text input (2011, 510). When the above instructions are executed individually or jointly by the one or more processors (120), they may obtain a visual object (2040, 6040, 7040, 9040) based on the descriptive information (520, 620, 720, 820). When the above commands are executed individually or jointly by the one or more processors (120), they may request the creation of an image (2060, 560, 860, 9060) based on a prompt (522, 822) containing the description information (520, 620, 720, 820). When the above commands are executed individually or jointly by the one or more processors (120), they may display the visual object (2040, 6040, 7040, 9040) while the image (2060, 560, 860, 9060) is being created based on the prompt (522, 822).When the above commands are executed individually or jointly by the one or more processors (120), they may enable the acquisition of an image (2060, 560, 860, 9060) generated based on the prompt (522, 822). When the above commands are executed individually or jointly by the one or more processors (120), they may enable the display of the generated image (2060, 560, 860, 9060).

[0069] According to one embodiment, the visual object (2040, 6040, 7040, 9040) may include text (7040). When the instructions are executed individually or jointly by the one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may receive a first user input corresponding to the visual object (2040, 6040, 7040, 9040). When the above commands are executed individually or jointly by one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may display a plurality of user interface (UI) items for changing the text (7040) of the visual object (2040, 6040, 7040, 9040). When the above commands are executed individually or jointly by the one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may change the text (7040) of the visual object (2040, 6040, 7040, 9040) by receiving a second user input through the plurality of UI items. When the above commands are executed individually or jointly by one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may modify the description information (520, 620, 720, 820) based on the modified text (7040) of the visual object (2040, 6040, 7040, 9040).When the above instructions are executed individually or jointly by the one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may obtain the image (2060, 560, 860, 9060) generated based on the prompt (522, 822) containing the modified description information (520, 620, 720, 820).

[0070] According to one embodiment, the visual object (2040, 6040, 7040, 9040) may include text (7040). When the instructions are executed individually or jointly by the one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may receive a first user input corresponding to the visual object (2040, 6040, 7040, 9040). When the above commands are executed individually or jointly by one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may display a plurality of user interface (UI) items for changing the text (7040) of the visual object (2040, 6040, 7040, 9040). When the above commands are executed individually or jointly by the one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may change the text (7040) of the visual object (2040, 6040, 7040, 9040) by receiving a second user input through the plurality of UI items. When the above commands are executed individually or jointly by the one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may obtain other descriptive information based on the modified text (7040) of the visual object (2040, 6040, 7040, 9040).When the above instructions are executed individually or jointly by the one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may obtain another image generated based on another prompt containing the other description information. When the above instructions are executed individually or jointly by the one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may display the other image obtained.

[0071] According to one embodiment, the description information (520, 620, 720, 820) may include text describing an image to be generated based on the prompt. When the commands are executed individually or jointly by one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may perform sentence analysis on the text to identify the main word or subject of the text. The displayed visual object (2040, 6040, 7040, 9040) may include the identified subject or the acquired keyword.

[0072] According to one embodiment, the description information (520, 620, 720, 820) may include text describing an image to be generated based on the prompt. When the commands are executed individually or jointly by one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may obtain keywords of the text of the description information (520, 620, 720, 820). The displayed visual object may include the obtained keywords.

[0073] According to one embodiment, the keyword may include the subject of the text of the description information (520, 620, 720, 820) and a predicate or modifier representing the subject.

[0074] According to one embodiment, when the instructions are executed individually or jointly by one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may obtain effect information indicating at least one of a visual effect, an audio effect, or a haptic effect to be applied to the visual object (2040, 6040, 7040, 9040). When the above commands are executed individually or jointly by the one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may output at least one of the visual effect, the audio effect, or the haptic effect while displaying the visual object (2040, 6040, 7040, 9040).

[0075] According to one embodiment, the prompt (522, 822) may be a second prompt (522, 822). When the instructions are executed individually or jointly by one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may apply the first prompt (512) containing the text input (2011, 510) to the first artificial intelligence (AI) model (402, 502, 602, 702, 802) to obtain the description information (520, 620, 720, 820) from the first AI model (402, 502, 602, 702, 802). When the above commands are executed individually or jointly by one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may apply the second prompt (522, 822) to the second artificial intelligence model (404, 504, 804, 1450) to obtain the image (2060, 560, 860, 9060) generated by the second AI model (404, 504, 804, 1450).

[0076] According to one embodiment, when the instructions are executed individually or jointly by one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may generate the first prompt (512) including the text input (2011, 510) based on a template. When the above commands are executed individually or jointly by one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may input the first prompt (512) into the first artificial intelligence model (402, 502, 602, 702, 802) to obtain the description information (520, 620, 720, 820) for generating the image (2060, 560, 860, 9060). The template may include information regarding a constraint for unifying the subject in the text of the description information (520, 620, 720, 820).

[0077] According to one embodiment, when the instructions are executed individually or jointly by one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may obtain effect information indicating at least one of a visual effect, an audio effect, or a haptic effect to be applied to the image (2060, 560, 860, 9060) by applying a third prompt (632, 732) containing the description information (520, 620, 720, 820) to the first artificial intelligence model (402, 502, 602, 702, 802) or the third artificial intelligence model (602, 702). When the above commands are executed individually or jointly by the one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may output at least one of the visual effect, the audio effect, or the haptic effect while displaying the generated image (2060, 560, 860, 9060).

[0078] According to one embodiment, the description information (520, 620, 720, 820) may be the first description information (820). When the above commands are executed individually or jointly by one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may obtain second description information (880) that describes at least one of a previous image, a next image, or an alternative image of the image (2060, 560, 860, 9060) described by the first description information (820) by applying a fourth prompt (842) containing the first description information (820) to the first artificial intelligence model (402, 502, 602, 702, 802) or the fourth artificial intelligence model (802). When the above commands are executed individually or jointly by the one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may obtain at least one of the previous image, subsequent image, or alternative image generated by applying the fifth prompt (852) containing the second description information (880) to the second artificial intelligence model (404, 504, 804, 1450) or the fifth artificial intelligence model (804). When the above instructions are executed individually or jointly by the one or more processors (120), the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301) may display at least one of the generated previous image, the subsequent image, or the alternative image.

[0079] According to one embodiment, the first artificial intelligence model (402, 502, 602, 702, 802) and the second artificial intelligence model (404, 504, 804, 1450) may be placed on a server that communicates with the electronic device (101, 201, 401, 601, 701, 901, 1101, 1201, 1301).

[0080] FIG. 2 is a diagram illustrating the process of acquiring and displaying visual objects and images based on text input in an electronic device according to one embodiment.

[0081] Graphic elements such as stickers, emoticons, and emojis can not only visually convey a user's emotions or the context of a conversation, but also personalize the user experience by revealing the user's personality and make communication more interesting. Users can create graphic elements (e.g., stickers) to add their own personality. Graphic elements (e.g., stickers) may be static images, but are not limited to this; they may be dynamic images made up of multiple images, such as videos or animated images.

[0082] Artificial intelligence models can be utilized to generate images such as stickers. Users can explain the image they wish to create to the AI ​​model, and the model can generate the image by understanding and translating the user's description. The time required for a generative AI model to generate an image is longer than the time required to generate text. This is because, even if the file sizes of the text and image to be generated are the same, image generation requires higher computational complexity and computational power. Therefore, while the AI ​​model is generating an image, a notification message such as "Generating image" may be displayed.

[0083] However, from the user's perspective, it is difficult to predict what the final generated image will look like. If the user's description is insufficient, the time required for the AI ​​model to interpret it is short; however, because the AI ​​generates an image with greater freedom of interpretation, it is more difficult for the user to predict the final result, and due to the lack of specific guidelines, an image different from the user's intent may be output. Providing rich descriptions to the AI ​​model entails the user's time and effort.

[0084] The notification message "Generating image" displayed while an image is being generated by an AI model does not provide any hints about the image to be created, making it difficult for the user to predict the generated image.

[0085] According to one embodiment, while an image is being generated by an artificial intelligence model, a visual object corresponding to the image generation process may be displayed, and through the displayed visual object, a hint regarding the image to be generated may be provided to the user. For example, the visual object may include text describing the image generation process.

[0086] Refer to Fig. 3 for convenience of explanation.

[0087] FIG. 3 is a flowchart of a method according to one embodiment.

[0088] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.

[0089] According to one embodiment, operations 310, 320, 330, 340, 350, and 360 can be understood as being performed in a processor (not shown) of an electronic device (e.g., electronic device (201) of FIG. 2).

[0090] According to one embodiment, in operation 310, the electronic device (201) can receive text input (2011) for generating an image.

[0091] Referring to FIG. 2, an electronic device (201) may display an image generation interface (2010) for generating images such as stickers used in messaging services or social media. The image generation interface (2010) may include an input interface (2012) for receiving text input (2011). The text input (2011) may consist of words that outline the image (e.g., sticker) to be generated. The text input (2011) may be received through the input interface (2012) of the electronic device (201). The text input (2011) may be entered into the input interface (2012) by a user interacting with the keypad (2030) of the electronic device (201), but is not limited thereto. For example, the text input (2011) may be received by the electronic device (201) by recording the user's voice and converting the recorded voice into text. Text input (2011) can be received from the electronic device (201) by copying at least a portion of the text displayed on the electronic device (201).

[0092] According to one embodiment, an image can be generated based on text input (2011) by using an artificial intelligence model. The artificial intelligence model may be included within an electronic device (201) or within a server that communicates with the electronic device (201).

[0093] According to one embodiment, the image generation interface (2010) may include buttons (2022 and 2024) for selecting the style of the image to be generated. The style of the image to be generated may be distinguished according to the method of representation of the image or design elements constituting the image, and the style of the image may include, for example, a doodle style, an illustration style, a pixel art style, and a 3D style, but is not limited thereto. FIG. 2 illustrates two style buttons, but is not limited thereto, and various number of style buttons may be included in the image generation interface (2010), and style buttons may be omitted and an appropriate style may be automatically determined according to the content of the text input (2011) entered into the input interface (2012).

[0094] According to one embodiment, the image generation interface (2010) may include a generation button (2020) for requesting an artificial intelligence model to generate an image. By pressing the generation button (2020), an image may be generated based on a text input (2011) entered into an input interface (2012). When the text input (2011) entered into the input interface (2012) is applied to the artificial intelligence model as a prompt, even if the artificial intelligence model is well trained to generate an image based on the given text, an image deviating from the user's intent may be generated if the text input (2011) is not sufficiently detailed. To generate an image that matches the user's intent, descriptive information may be obtained based on the text input (2011), and a prompt containing the obtained descriptive information may be applied to the artificial intelligence model.

[0095] According to one embodiment, in operation 320, the electronic device (201) can obtain descriptive information based on the text input (2011) received in operation 310. According to one embodiment, by applying a prompt including the text input (2011) to an artificial intelligence model, descriptive information for generating an image can be obtained, which will be described later with reference to FIGS. 4 and FIGS. 5.

[0096] According to one embodiment, after description information is acquired, the electronic device (201) may request the generation of an image based on a prompt containing the description information. The electronic device (201) may request the generation of an image from an artificial intelligence model, and the generation of an image may be requested by inputting a prompt containing the description information to the artificial intelligence model. The artificial intelligence model may be placed in the electronic device (201) or in a server communicating with the electronic device (201). The request for the generation of an image may be performed before acquiring a visual object (2040) based on the description information, but is not limited thereto, and the two operations may be performed in parallel or the order thereof may be reversed. For example, while a visual object (2040) is acquired through sentence analysis of the text of the description information, the generation of an image based on the description information may be requested. While the image is being generated by this request, the visual object (2040) acquired through sentence analysis of the text of the description information may be displayed. While the image is being generated, keyword and / or effect information is generated based on the description information, and a visual object (2040) based on the generated keyword and / or effect information can be obtained and displayed. After the generation of the image is completed, the display of the visual object is terminated and the generated image (2060) can be displayed.

[0097] According to one embodiment, in operation 330, the electronic device (201) can acquire a visual object (2040) based on the description information acquired in operation 320. According to one embodiment, the visual object (2040) can be acquired from the description information, which will be described later with reference to FIG. 4 and FIG. 5. According to one embodiment, the visual object (2040) can be acquired from keywords generated by applying the description information to an artificial intelligence model, which will be described later with reference to FIG. 6.

[0098] According to one embodiment, in operation 340, the electronic device (201) can display the visual object (2040) obtained in operation 330. According to one embodiment, effect information generated by applying description information to an artificial intelligence model can be applied to the visual object (2040), which will be described later with reference to FIG. 7.

[0099] According to one embodiment, a visual object (2040) corresponding to the image generation process may be displayed until the image generation by the artificial intelligence model is completed. As illustrated in FIG. 2, when the text input (2011) is "animal drinking bubble tea," the visual object (2040) may include text expressing that an image is being generated, such as "Drawing Panda...". A loading interface (2014) for displaying the visual object (2040) may be included in the image generation interface (2010). The loading interface (2014) may be referred to as a loading screen (2014).

[0100] Through the visual object (2040), the user can see that the animal drinking bubble tea in the image to be generated by the artificial intelligence model is a panda. According to one embodiment, through the visual object (2040), the user can predict the image to be generated more specifically than the text input (2011) they entered.

[0101] According to one embodiment, while a visual object (2040) is displayed, the generate button (2020) may be disabled, but is not limited thereto. For example, when text of the displayed visual object (2040) is selected and modified by a user, the generate button (2020) may be enabled so that the modified content is reflected in the description information and an image can be generated based on the modified description information.

[0102] According to one embodiment, in operation 350, the electronic device (201) can obtain an image (2060) generated based on the description information obtained in operation 320. According to one embodiment, by applying a prompt containing the description information to an artificial intelligence model, the artificial intelligence model can be made to generate an image (2060), which will be described later with reference to FIGS. 4 and FIGS. 5.

[0103] According to one embodiment, in operation 360, the electronic device (201) can display the image (2060) obtained in operation 350. According to one embodiment, when the creation of the image (2060) is completed or when the created image (2060) is obtained, the electronic device (201) can stop displaying the visual object (2040) and display the image (2060) obtained in operation 350. An image interface (2016) for displaying the obtained image (2060) may be included in the image creation interface (2010). According to one embodiment, the image interface (2016) may include a text input field (2018) representing text input entered by a user for the creation of the image (2060). Through the text input field (2018), the user can compare the text input entered by the user with the image (2060) generated by it at a glance. According to one embodiment, the text input field (2018) may represent text of a visual object obtained based on description information. According to one embodiment, the user can easily modify the image (2060) through the text input field (2018), which will be described later with reference to FIG. 13.

[0104] According to one embodiment, while the generated image (2060) is being displayed, the generate button (2020) may be disabled, but is not limited thereto. For example, when a text input field (2018) is selected by a user and the text input is modified, the generate button (2020) may be enabled so that the modified content is reflected in the description information and an image can be generated based on the modified description information.

[0105] According to one embodiment, the image interface (2016) may include a button (2019) indicating the style of the generated image (2060). When the button (2019) is selected by the user, options for changing the current style of the generated image (2060) to another style may be provided to the user, thereby allowing the user to easily change the style of the generated image (2060).

[0106] According to one embodiment, the image interface (2016) may include a phrase indicating that the image (2060) was generated by an artificial intelligence model, but the phrase may be omitted.

[0107] According to one embodiment, a method may be provided that includes the operation of receiving a user's text input (2011, 510). The method may include the operation of obtaining descriptive information (520, 620, 720, 820) for generating an image based on the text input (2011, 510). The method may include the operation of obtaining a visual object (2040, 6040, 7040, 9040) based on the descriptive information (520, 620, 720, 820). The method may include the operation of requesting the generation of an image based on a prompt (522, 822) containing the descriptive information (520, 620, 720, 820). The above method may include an operation of displaying the visual object (2040, 6040, 7040, 9040) while the image (2060, 560, 860, 9060) is being generated based on the prompt (522, 822). The above method may include an operation of acquiring the generated image (2060, 560, 860, 9060) based on the prompt (522, 822). The above method may include an operation of displaying the generated image (2060, 560, 860, 9060).

[0108] According to one embodiment, the visual object (2040, 6040, 7040, 9040) may include text (7040). The method may include an operation of receiving a first user input corresponding to the visual object (2040, 6040, 7040, 9040). The method may include an operation of displaying a plurality of user interface (UI) items for changing the text (7040) of the visual object (2040, 6040, 7040, 9040). The method may include an operation of changing the text (7040) of the visual object (2040, 6040, 7040, 9040) by receiving a second user input through the plurality of UI items. The above method may include an operation of modifying the description information (520, 620, 720, 820) based on the modified text (7040) of the visual object (2040, 6040, 7040, 9040). The operation of obtaining the generated image (2060, 560, 860, 9060) may include an operation of obtaining the generated image (2060, 560, 860, 9060) based on the prompt (522, 822) containing the modified description information.

[0109] According to one embodiment, the visual object (2040, 6040, 7040, 9040) may include text (7040). The method may include an operation of receiving a first user input corresponding to the visual object (2040, 6040, 7040, 9040). The method may include an operation of displaying a plurality of user interface (UI) items for changing the text (7040) of the visual object (2040, 6040, 7040, 9040). The method may include an operation of changing the text (7040) of the visual object (2040, 6040, 7040, 9040) by receiving a second user input through the plurality of UI items. The above method may include an operation of obtaining other descriptive information based on the modified text (7040) of the visual object (2040, 6040, 7040, 9040). The above method may include an operation of obtaining another image generated based on another prompt containing the other descriptive information. The above method may include an operation of displaying the other image obtained.

[0110] According to one embodiment, the description information (520, 620, 720, 820) may include text describing an image to be generated based on the prompt (522, 822). The method may include an operation of performing sentence analysis on the text to identify the main word or subject of the text. The displayed visual object (2040, 6040, 7040, 9040) may include the identified subject.

[0111] According to one embodiment, the description information (520, 620, 720, 820) may include text describing an image to be generated based on the prompt (522, 822). The method may include an operation of obtaining keywords of the text of the description information (520, 620, 720, 820). The displayed visual object (2040, 6040, 7040, 9040) may include the obtained keywords.

[0112] According to one embodiment, the keyword may include the subject of the text of the description information (520, 620, 720, 820) and a predicate or modifier representing the subject.

[0113] According to one embodiment, the method may include an operation of obtaining effect information indicating at least one of a visual effect, an audio effect, or a haptic effect to be applied to the visual object (2040, 6040, 7040, 9040). An operation of displaying the visual object (2040, 6040, 7040, 9040) may include an operation of outputting at least one of the visual effect, the audio effect, or the haptic effect while displaying the visual object (2040, 6040, 7040, 9040).

[0114] According to one embodiment, the prompt (522, 822) may be a second prompt (522, 822). The operation of obtaining the description information (520, 620, 720, 820) may include applying a first prompt (512) containing the text input (2011, 510) to a first artificial intelligence (AI) model (402, 502, 602, 702, 802) to obtain the description information (520, 620, 720, 820) from the first AI model (402, 502, 602, 702, 802). The operation of acquiring the generated image may include applying the second prompt (522, 822) to the second artificial intelligence model (404, 504, 804, 1450) to acquire the image generated by the second artificial intelligence model (404, 504, 804, 1450).

[0115] According to one embodiment, the operation of obtaining the description information (520, 620, 720, 820) may include the operation of generating the first prompt (512) including the text input (2011, 510) based on a template. The operation of obtaining the description information (520, 620, 720, 820) may include the operation of inputting the first prompt (512) into the first artificial intelligence model (402, 502, 602, 702, 802) to obtain the description information (520, 620, 720, 820) for generating the image (2060, 560, 860, 9060). The above template may include information regarding a constraint for unifying the subject in the text of the above description information (520, 620, 720, 820).

[0116] According to one embodiment, the method may include an operation of obtaining effect information indicating at least one of a visual effect, an audio effect, or a haptic effect to be applied to the image (2060, 560, 860, 9060) by applying a third prompt (632, 732) containing the description information (520, 620, 720, 820) to the first artificial intelligence model (402, 502, 602, 702, 802) or the third artificial intelligence model (602, 702). The operation of displaying the generated image (2060, 560, 860, 9060) may include an operation of outputting at least one of the visual effect, the audio effect, or the haptic effect while displaying the generated image (2060, 560, 860, 9060).

[0117] According to one embodiment, the description information (520, 620, 720, 820) may be the first description information (520, 620, 720, 820). The method may include the operation of obtaining second description information (880) that describes at least one of the previous image, next image, or alternative image of the image (2060, 560, 860, 9060) described by the first description information (520, 620, 720, 820) by applying a fourth prompt (842) containing the first description information (820) to the first artificial intelligence model (402, 502, 602, 702, 802) or the fourth artificial intelligence model (802). The above method may include an operation of obtaining at least one of the previous image, subsequent image, or alternative image generated by applying a fifth prompt (852) including the second description information (880) to the second artificial intelligence model (404, 504, 804, 1450) or the fifth artificial intelligence model (804). The above method may include an operation of displaying at least one of the generated previous image, subsequent image, or alternative image.

[0118] According to one embodiment, the method may be performed by an electronic device. The first artificial intelligence model (402, 502, 602, 702, 802) and the second artificial intelligence model (404, 504, 804, 1450) may be placed on a server that communicates with the electronic device.

[0119] FIG. 4 is a diagram illustrating the process of displaying a visual object using an artificial intelligence model according to one embodiment.

[0120] In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.

[0121] According to one embodiment, operations 410, 412, 420, 440, 422, and 450 can be understood as being performed in a processor (not shown) of an electronic device (401).

[0122] According to one embodiment, in operation 410, the electronic device (401) can receive text input from a user. Since operation 410 is substantially the same as operation 310 of FIG. 3, a redundant description is omitted.

[0123] According to one embodiment, in operation 412, the electronic device (401) may apply a first prompt to the first artificial intelligence model (402). According to one embodiment, the first prompt may include text input received in operation 410, which will be described later with reference to FIG. 5.

[0124] According to one embodiment, the first artificial intelligence model (402) may be a multimodal model that learns and processes relationships between data of various types or various modalities, such as text and images. The first artificial intelligence model (402) may be a large multimodal model (LMM) learned using text data and image data, and the first artificial intelligence model (402) may generate an image or generate text associated with an image based on a text prompt, an image prompt, or a prompt composed of text and an image input to the first artificial intelligence model (402). For example, the first artificial intelligence model (402) may generate descriptive information composed of text based on a first prompt containing text input from a user. The first prompt containing text input from a user and the descriptive information generated by the first artificial intelligence model (402) based on the first prompt will be described later with reference to FIG. 5.

[0125] The first artificial intelligence model (402) is not limited to an LMM, and the first artificial intelligence model (402) may be a generative artificial intelligence model trained to generate descriptive information that elaborates on the text input based on the user's text input.

[0126] According to one embodiment, in operation 420, the electronic device (401) can obtain description information generated by the first artificial intelligence model (402) based on the first prompt applied in operation 412. Since operation 420 is substantially the same as operation 320 of FIG. 3, a redundant description is omitted.

[0127] According to one embodiment, in operation 422, the electronic device (401) may apply a second prompt to the second artificial intelligence model (404). According to one embodiment, the second prompt may include descriptive information obtained in operation 420, which will be described later with reference to FIG. 5.

[0128] The second artificial intelligence model (404) may be an artificial intelligence model trained to generate an image based on an input text prompt. The second artificial intelligence model (404) may be a text-to-image model and may be a combined form of a transformer-based language model and a diffusion-based image generation model, but is not limited thereto. Based on a second prompt containing descriptive information, the second artificial intelligence model (404) may interpret the descriptive information and generate an image that meets the conditions included in the descriptive information.

[0129] The first artificial intelligence model (402) can be any artificial intelligence model capable of performing bidirectional interaction between text and image, and the second artificial intelligence model (404) can be any artificial intelligence model capable of performing unidirectional interaction that generates an image based on text.

[0130] According to one embodiment, in operation 440, the electronic device (401) can display a visual object based on the description information obtained in operation 420. Since operation 440 is substantially the same as operation 340 of FIG. 3, a redundant description is omitted. According to one embodiment, the visual object (2040) can be obtained from the description information, which will be described later with reference to FIG. 5. According to one embodiment, the visual object (2040) can be obtained from keywords generated by applying the description information to an artificial intelligence model, which will be described later with reference to FIG. 6. According to one embodiment, effect information generated by applying the description information to an artificial intelligence model can be applied to the visual object (2040), which will be described later with reference to FIG. 7.

[0131] According to one embodiment, in operation 450, the electronic device (401) can acquire an image generated by the second artificial intelligence model (404) based on the second prompt applied in operation 422. Since operation 450 is substantially the same as operation 350 of FIG. 3, a redundant description is omitted.

[0132] FIG. 4 illustrates that description information is generated by a first artificial intelligence model (402) and an image is generated by a second artificial intelligence model (404), but the first artificial intelligence model (402) and the second artificial intelligence model (404) may be operated in an integrated manner. For example, in operation 412, the electronic device (401) may apply a first prompt to the integrated artificial intelligence model. In operation 420, the electronic device (401) may obtain description information generated by the integrated artificial intelligence model based on a prompt, and in operation 422, apply a second prompt containing the description information to the integrated artificial intelligence model, but is not limited thereto. For example, the integrated artificial intelligence model may generate description information based on a prompt received from the electronic device (401) (e.g., a first prompt), and then, without transmitting the description information to the electronic device (401), generate an image based on another prompt containing the description information (e.g., a second prompt) and transmit it to the electronic device (401). According to one embodiment, in order for a visual object to be displayed on an electronic device (401) while an image is being generated, an integrated artificial intelligence model may perform sentence analysis on the text of the descriptive information to identify the subject of the text and transmit the identified subject to the electronic device (401). Thus, a visual object containing the identified subject may be displayed on the electronic device (401).

[0133] The integrated artificial intelligence model may be implemented to receive the first and second prompts described above, as well as the third, fourth, and fifth prompts described below, and to output results processed according to the received prompts. For example, the integrated artificial intelligence model may extract keywords of the description information or generate effect information based on the third prompt containing description information. For example, the integrated artificial intelligence model may generate other description information based on the fourth prompt containing description information, or generate other images based on the fifth prompt containing other description information. Among the results output by the integrated artificial intelligence model, the generated images and keywords of the description information may be transmitted to the electronic device (401).

[0134] In order for a visual object to be displayed on an electronic device (401) while an image is being generated, the integrated artificial intelligence model may acquire keywords of descriptive information, for example, the subject and predicates or modifiers representing the subject in the descriptive information, and transmit the acquired keywords to the electronic device (401). Thus, a visual object containing the keywords can be displayed on the electronic device (401).

[0135] The integrated artificial intelligence model may be a multimodal model, for example, an LMM, that learns and processes relationships between data of various types or various modalities, such as text and images. The integrated artificial intelligence model is not limited to an LMM, and the integrated artificial intelligence model may be pre-trained to generate descriptive information that elaborates on text input based on a user's text input, and may be pre-trained to generate an image based on the descriptive information. FIG. 5 is a diagram illustrating the process of generating an image based on descriptive information according to one embodiment.

[0136] Referring to FIG. 5, the first prompt (512) input to the first artificial intelligence model (502) may include text input (510), and the first prompt (512) may be generated by filling the input field of the template (511) with text input (510) or by replacing at least a part of the template (511) with text input (510). According to one embodiment, the template (511) may be stored in an electronic device, and the electronic device may store a plurality of templates. The plurality of templates may each correspond to buttons for selecting a style (e.g., 2022 and 2024 in FIG. 2), and the template (511) corresponding to the button pressed by the user may be used in the first prompt (512) and input to the first artificial intelligence model (502). Refer to Table 1 further to describe the template (511).

[0137]

[0138] Table 1 shows an example of a template (511). The template (511) may include a guideline or goal to define the overall direction of the work to be processed according to a prompt containing the template (511). According to one embodiment, the role of the first artificial intelligence model (502) (e.g., creation of a prompt) may be defined by the guideline. Referring to Table 1, the guideline of the template (511) may include creating a prompt for creating a sticker image, i.e., descriptive information (520), and the use of the sticker image may be defined by the guideline of the template (511). Referring to Table 1, the goal of the template (511) may include creating descriptive information based on a text input (510) provided in the template (511). The text input (510) provided in Table 1 is the text input (510) "animal drinking bubble tea" exemplified in FIG. 2. The first artificial intelligence model (502) can interpret a prompt containing a template (511) and perform a given task (e.g., generation of descriptive information (520)) according to the interpreted instructions.

[0139] According to one embodiment, the template (511) may include a constraint applied in the process of the first artificial intelligence model (502) generating description information (520) based on a first prompt (512) containing the template (511), and may include a constraint applied in the process of generating an image (560) by the description information (520). For example, the constraint may include a condition for limiting the length of the prompt, such as the prompt being composed of a maximum of n sentences, and thus the time for the first artificial intelligence model (502) to generate the description information (520) may be shortened. According to one embodiment, the content that must be included in the prompt and the content that must not be included may be defined by the constraint. For example, constraints may include conditions such as k details to be merged into the image (560) to be generated being included in the prompt, words that may violate the content policy not being included in the prompt, / 'character / ', / 'speech bubble / ', / 'script / ', / 'word / ' not being included in the prompt, and content associated with the text not being depicted by the prompt.

[0140] According to one embodiment, the constraint may restrict the background of the prompt to be generated, that is, the image (560) to be generated by the description information (520), and by such constraint, a background other than a sticker may not be allowed or may be restricted to a white background.

[0141] According to one embodiment, the constraint may define the style of the image (560) to be generated by the prompt (e.g., a doodle style, an illustration style, a pixel art style, or a 3D style). According to one embodiment, the template (511) may be stored in an electronic device, and the electronic device may store a plurality of templates that each define different styles. The plurality of templates may each correspond to buttons for selecting a style (e.g., 2022 and 2024 in FIG. 2), and the template (511) corresponding to the button pressed by the user may be used in the first prompt (512) and input into the first artificial intelligence model (502). The constraint may include conditions regarding quantitative features of the prompt or image (560) as well as conditions regarding qualitative features.

[0142] The template (511) is not limited to Table 1, and other templates (511) of the prompt for generating sticker images may be used.

[0143] According to one embodiment, the constraint may include a constraint for clearly expressing the subject of the sentence of the description information (520), that is, the prompt to be generated. If the subject of the sentence included in the description information (520) is clear, the word corresponding to the subject of the sentence can be identified, and the identified word can be included in the visual object and displayed. Therefore, even before the image (560) is generated by the description information (520), the user can know in advance the content of the image (560) to be generated through the word included in the visual object. The process of determining the visual object from the description information (520) is explained with further reference to Table 2.

[0144]

[0145] Table 2 shows an example of descriptive information (520). The descriptive information (520) of Table 2 is generated by applying a first prompt (512) in which the text input (510) of FIG. 2 “animal drinking bubble tea” is entered into the template (511) of Table 1 to the first artificial intelligence model (502). The descriptive information (520) consists of text that satisfies the instructions and constraints of the first prompt (512) defined by the template (511), and the text describes an image (560) to be generated. For example, the descriptive information (520) consists of three sentences, and the three sentences describe a panda wearing a straw hat and loudly drinking bubble tea from a boba cup, and represent an image (560) drawn in a doodle style with a white background.

[0146] According to one embodiment, the electronic device can perform sentence analysis on sentences included in the description information (520), and the words constituting the sentences can be classified according to grammar. For example, the subject of the sentences, modifiers that limit the subject, and predicates indicating the action or appearance of the subject can be identified. "Panda" is identified as the subject of the first sentence in Table 2.

[0147] According to one embodiment, the electronic device can perform sentence analysis on the first sentence included in the description information (520), identify the word corresponding to the subject of the sentence, and display the identified word as a visual object. According to one embodiment, the visual object may include a fixed message and an identification message, and the subject of the first sentence included in the description information (520) may be included in the visual object as an identification message. The format of the identification message (e.g., the first letter of the word is capitalized) may be predefined. Referring to FIG. 2, in the visual object of "Drawing Panda ...", "Panda" corresponds to the identification message, and the remaining parts excluding the identification message, "Drawing" and " ...", correspond to the fixed message. According to one embodiment, the electronic device can perform sentence analysis on the first sentence included in the description information (520), identify the words corresponding to the subject of the sentence and the modifier limiting the subject, and display the identified words as a visual object. For example, "cute panda" is identified as the subject of the first sentence in Table 2 and the modifier limiting the subject.

[0148] Since sentence analysis can be performed without the help of an artificial intelligence model outside the electronic device, sentence analysis can be performed and completed immediately after the description information (520) is acquired.

[0149] According to one embodiment, the constraint may include a constraint for unifying the subjects of the sentences of the prompt to be generated, i.e., the description information (520). The description information (520) generated according to this constraint may include sentences in which the subjects are all identical, and it can be expected that the word corresponding to the subject of the sentence will be the same regardless of which sentence in the description information (520) is subjected to sentence analysis. The word identified as the subject may be included in a visual object and displayed. Since the result output by the artificial intelligence model is generated according to a probability distribution, a result that does not meet the constraint may be generated. Accordingly, the electronic device may be implemented to identify the word with the most overlap among the words identified as the subject and display that word as a visual object.

[0150] According to one embodiment, a visual object may be obtained from a keyword generated by applying description information (520) to a first artificial intelligence model (502) or a separate artificial intelligence model, which will be described later with reference to FIG. 6. According to one embodiment, effect information generated by applying description information (520) to a first artificial intelligence model (502) or a separate artificial intelligence model may be applied to a visual object, which will be described later with reference to FIG. 7.

[0151] In the process of generating description information (520), the template (511) used for the first prompt (512) input into the first artificial intelligence model (502) may be referred to as a description generation template (511). That is, the description generation template (511) refers to the template (511) of the first prompt (512) input into the first artificial intelligence model (502), and the first prompt (512), which is completed by filling the description generation template (511) with user input text, can be input into the first artificial intelligence model (502).

[0152] Referring again to FIG. 5, the electronic device can acquire description information (520) generated by the first artificial intelligence model (502), and the electronic device can input a second prompt (522) containing the acquired description information (520) to the second artificial intelligence model (504). The second artificial intelligence model (504) generates an image (560) according to the description information (520) included in the second prompt (522), and the electronic device can acquire and display the generated image (560). The electronic device can display a visual object until the generation of the image (560) is completed. The electronic device can display a visual object while the image (560) is being generated.

[0153] According to one embodiment, the first artificial intelligence model (502) can process a task in the corresponding language according to the received language, or process a task by converting the received language into a main language, and can output the result of the task in the received language or in the main language.

[0154] In the present disclosure, the electronic device may be a smartphone or tablet including a flat display, but is not limited thereto. For example, the electronic device (701) may be a foldable device, a rollable device, or a sliderable device including a flexible display, and may be an extended reality (XR) device including a head-mounted display (HMD) or an optical waveguide display. The XR device includes a virtual reality (VR) device, an augmented reality (AR) device, and a mixed reality (MR) device.

[0155] According to one embodiment, an image (e.g., image (560) of FIG. 5) generated based on depiction information (e.g., depiction information (520) of FIG. 5) may be a 2D image, but is not limited thereto, and may be, for example, a 3D image.

[0156] According to one embodiment, the XR device may receive text input from a user (e.g., text input (510) of FIG. 5), and the received text input may be input into an image generation interface provided by the XR device (e.g., image generation interface (2010) of FIG. 2). The XR device may obtain descriptive information based on the user's text input. According to one embodiment, the XR device may input a prompt containing the user's text input (e.g., first prompt (512) of FIG. 5) to an artificial intelligence model (e.g., first artificial intelligence model (502) of FIG. 5)) and obtain descriptive information generated by the artificial intelligence model (e.g., descriptive information (520) of FIG. 5). If the artificial intelligence model is an integrated artificial intelligence model implemented to perform both descriptive information and image generation, the artificial intelligence model may generate an image based on the generated descriptive information and transmit it to the XR device. When the AI ​​model is implemented to generate description information and another AI model (e.g., the second AI model (504) of FIG. 5) generates an image, the other AI model can receive the description information, generate an image, and transmit the generated image to an XR device. FIG. 6 is a diagram illustrating the process of displaying a visual object using keywords according to one embodiment.

[0157] According to one embodiment, the electronic device (601) may apply a third prompt (632) containing descriptive information (620) to a first artificial intelligence model (602) or a separate artificial intelligence model to extract keywords (670) of the descriptive information (620). The separate artificial intelligence model may be an artificial intelligence model trained to extract keywords (670) from a given text. The separate artificial intelligence model may be included in the electronic device (601) or in a server communicating with the electronic device (601).

[0158] Referring to FIG. 6, an electronic device (601) displays an image generation interface (6010), and the image generation interface (6010) includes a loading interface (6014) in which a visual object (6040) is displayed. To display the visual object (6040), the electronic device (601) applies a third prompt (632) containing acquired description information (620) to a first artificial intelligence model (602), and the first artificial intelligence model (602) can analyze the given third prompt (632), i.e., the description information (620), to extract or generate a keyword (670). The electronic device (601) can display the visual object (6040) using the keyword (670) identified by the first artificial intelligence model (602). According to one embodiment, a constraint limiting the number of words constituting the keyword (670) may be included in the third prompt (632) so that the user can recognize the visual object (6040) displayed on the loading interface (6014) at a glance. For example, a constraint limiting the keyword (670) to being extracted or generated from description information (620) and the keyword (670) to being described with four or fewer words may be included in the third prompt (632). According to one embodiment, a constraint unifying the subject of the words constituting the keyword (670) may be included in the third prompt (632) so that the user can specifically imagine the image to be generated.

[0159] Referring to Table 2, "cute panda", "panda wearing hat", "panda slurping bubble tea", "panda with rosy cheeks", "panda on white background", "panda in doodle style", and "panda with playful expression" can be extracted or generated as keywords (670) of the description information (620).

[0160] According to one embodiment, a visual object (6040) may include a fixed message and an identification message, and an identified keyword (670) may be included in the visual object (6040) as an identification message. The format of the identification message (e.g., the first letter of a word is capitalized) may be predefined. Referring to FIG. 6, in the visual object (6040) of "Drawing Cute Panda ...", "Cute Panda" corresponds to the identification message, and the remaining parts excluding the identification message, "Drawing" and " ...", correspond to the fixed message. According to one embodiment, an electronic device (601) may alternately display the identified keywords (670). For example, visual objects (6040) each containing the identified keywords (670) may be displayed sequentially or randomly. Thus, the user can know in advance the content of the image to be generated by the second artificial intelligence model through the alternately displayed keywords (670).

[0161] FIG. 7 is a diagram illustrating the process of displaying a visual object using effect information according to one embodiment.

[0162] According to one embodiment, the electronic device (701) may apply a third prompt (732) containing description information (720) to a first artificial intelligence model (702) or a separate artificial intelligence model to generate effect information (770) from the description information (720). The separate artificial intelligence model may be an artificial intelligence model trained to generate effect information (770) from a given text. The separate artificial intelligence model may be included in the electronic device (701) or in a server communicating with the electronic device (701).

[0163] Referring to FIG. 7, an electronic device (701) displays an image generation interface (7010), and the image generation interface (7010) includes a loading interface (7014) in which a visual object (7040) is displayed. To display the visual object (7040), the electronic device (701) applies a third prompt (732) containing acquired description information (720) to a first artificial intelligence model (702), and the first artificial intelligence model (702) can generate effect information (770) by analyzing the given third prompt (732), i.e., the description information (720). The electronic device (701) can display the visual object (740) using the effect information (770) generated by the first artificial intelligence model (702).

[0164] A third prompt (732) input to the first artificial intelligence model (702) so that the first artificial intelligence model (702) can generate effect information (770) may include descriptive information (720), and the third prompt (732) may be generated by filling the input field of the template (731) with descriptive information (720) or by replacing at least a part of the template (731) with descriptive information (720). The template (731) may be stored in an electronic device (701). Refer further to Table 3 to describe the template (731).

[0165]

[0166] Table 3 shows an example of a template (731). The template (731) may include instructions or goals to define the overall direction of the work to be processed according to the prompt containing the template (731). The template (731) may be referred to as an effect generation template. That is, the effect generation template (731) refers to the template (731) of the third prompt (732) that is input into the first artificial intelligence model (702), and the third prompt (732), which is completed by filling the effect generation template (731) with the description information (720), may be input into the first artificial intelligence model (702).

[0167] The template (731) may include instructions for creating effects to be output on an electronic device along with a loading screen while an image is being generated. For example, the template (731) may include examples corresponding to visual effects, audio effects, and haptic effects, and the first artificial intelligence model (702) may refer to the examples to generate effect information (770) representing visual effects, audio effects, and haptic effects to be output on an electronic device along with a loading screen while an image is being generated. The examples included in the template (731) may function as constraints to be applied in the process of generating the effect information (770).

[0168] Visual effects, audio effects, and haptic effects may be associated with the image to be generated. According to one embodiment, the visual effect may include setting a background of a color and / or pattern corresponding to the theme of the image to be generated as the background of the loading screen. For example, if the image to be generated is a "lion," the color and pattern corresponding to the lion may be set to a background color of "0xFFCC33" and a wide horizontal line pattern of "0xB8860B". According to one embodiment, the audio effect may include setting a song corresponding to the theme of the image to be generated as background music for the loading screen. For example, if the image to be generated is a "lion," "The lion sleeps tonight," one of the famous songs about lions, may be set as background music. According to one embodiment, the audio effect may include outputting a sound effect associated with the image to be generated upon completion of loading. For example, if the image to be generated is a "lion," the roar of a lion may be output as an audio effect. According to one embodiment, the haptic effect may include outputting a vibration associated with the image to be generated upon completion of loading. Haptic effects can be associated with audio effects. For example, if the image to be generated is a "lion," vibrations may be output according to the intensity of the roar while the sound of the lion's roar is being output.

[0169] Outputting the effect expressed by the effect information (770) from the electronic device is explained with further reference to Table 4. Table 4 shows the effect information (770) generated by the first artificial intelligence model (702).

[0170]

[0171] Effect information (770) generated by the first artificial intelligence model (702) is composed of text, and the electronic device (701) acquires the effect information (770) composed of text. According to one embodiment, the electronic device (701) may be equipped with a module, an artificial intelligence model, and / or an application programming interface (API) for changing the settings of the electronic device or controlling the electronic device based on natural language commands so that the electronic device (701) can output visual effects, audio effects, and haptic effects based on the acquired effect information (770). Accordingly, the electronic device can perform operations based on natural language commands included in the acquired effect information (770). For example, the visual effects included in the effect information (770) may be output from the loading interface (7014) on the condition that the generation of an image is completed in the image generation interface (7010). For example, the electronic device can interpret the natural language command corresponding to a visual effect, "display 'Drawing Cute Panda ..." on the loading screen," and accordingly display the corresponding text (7040) within the loading interface (7014). For example, the electronic device can interpret the natural language commands corresponding to a visual effect, "set the background of the loading screen to '0xFFFFFF'" and "set a dotted pattern of color '0x000000' on the background of the loading screen," and accordingly set the background (7042) of the loading interface (7014). For example, the electronic device can interpret the natural language commands corresponding to an audio effect, "play the music 'Kung fu fighting' on the loading screen" and "play the sound of drinking bubble tea for 4 seconds when loading is complete," and accordingly play the music on the loading interface (7014) and output a sound effect (7044) when loading is complete.For example, the electronic device can interpret a natural language command corresponding to a haptic effect, "output vibration according to the sound of drinking bubble tea played when loading is complete," and accordingly output vibration according to the sound effect (7044) when loading is complete. As the sound effect (7044) and vibration are output, the acquired image may be displayed on the loading interface (7014), or the device may switch from the loading interface (7014) to an image interface where the acquired image is displayed.

[0172] The effect information (770) of Table 4 is provided as an example generated based on the template (731) of Table 3, and the effect information (770) may represent more specific effects as shown in Table 5. The electronic device can perform an action corresponding to the interpreted natural language command among the natural language commands included in the effect information (770).

[0173]

[0174] According to one embodiment, the electronic device may display buttons for selecting effects to be applied to a visual object, and these buttons may be displayed together with an input interface for receiving text input (e.g., 2012 in FIG. 2) or with a loading interface (7014). For example, buttons for selecting effects may be displayed below the loading interface (7014).

[0175] According to one embodiment, the first artificial intelligence model (702) can process a task in the corresponding language according to the received language, or process a task by converting the received language into a main language, and output the result of the task in the received language or output it in the main language. The electronic device can interpret the language of the effect information (770) as is or interpret it by converting it into another language, and perform an operation according to the interpretation result.

[0176] FIG. 7 illustrates that the effects of the effect information (770) are applied to a visual object (7040), but is not limited thereto, and these effects may be applied to a generated image or output together with the generated image. For example, animation effects or 3D effects, such as the panda's facial expression change or bubble tea movement effect in Table 5, may be applied to the generated image, and audio effects may be output together with the image. Therefore, even if the generated image is not an animated image, animation effects may be applied to provide a lively image to the user. According to one embodiment, animation effects or audio effects may be output when the generated image is selected on an electronic device (701). When a message containing the generated image is transmitted to another electronic device, animation effects or audio effects may be output when the message is viewed by a user on the other electronic device or when the image is selected on the other electronic device.

[0177] According to one embodiment, effect information (e.g., effect information (770) of FIG. 7) generated based on description information (e.g., description information (720) of FIG. 7) may include 3D effects as exemplified in Table 5, and an image with 3D effects applied may be displayed. FIG. 8 is a diagram illustrating the process of generating a second image associated with a first image according to one embodiment.

[0178] The first image (860) refers to an image (e.g., 2060 in FIG. 2, 560 in FIG. 5) generated through the processes described in FIG. 1 through 7, and the descriptive information used to generate the first image (860) is referred to as the first descriptive information (820), the visual object to be displayed on the electronic device while the first image (860) is generated is referred to as the first visual object, and the effect information applied to the first visual object or the first image (860) may be referred to as the first effect information. A second prompt (822) containing the first descriptive information (820) is applied to the second artificial intelligence model (804), and the second artificial intelligence model (804) can generate the first image (860) based on the first descriptive information (820).

[0179] According to one embodiment, the electronic device may generate a fourth prompt (842) containing first description information (820) and obtain second description information (880) by applying the generated fourth prompt (842) to a first artificial intelligence model (802) or a separate artificial intelligence model. The separate artificial intelligence model may be an artificial intelligence model trained to output second description information (880) that describes a situation associated with the first description information (820) based on the given first description information (820). The separate artificial intelligence model may be included in the electronic device or in a server communicating with the electronic device.

[0180] According to one embodiment, the second description information (880) may describe a previous image, a subsequent image, or an alternative image of the first image (860). The second description information (880) may describe a past situation before reaching the situation expressed by the first description information (820), describe a future situation after the situation expressed by the first description information (820), or describe an alternative situation that replaces the situation expressed by the first description information (820).

[0181] According to one embodiment, the electronic device can generate a fifth prompt (852) including acquired second description information (880) and apply the generated fifth prompt (852) to a second artificial intelligence model (804) to obtain a second image (862) generated by the second artificial intelligence model (804). The description information used to generate the second image (862) is referred to as the second description information (880), the visual object to be displayed on the electronic device while the second image (862) is generated is referred to as the second visual object, and the effect information applied to the second visual object or the second image (862) may be referred to as the second effect information.

[0182] According to one embodiment, a first effect information may be generated by applying a prompt (e.g., the third prompt (732) of FIG. 7) containing first description information (820) to a first artificial intelligence model (802), and the generated first effect information may be applied to a first visual object or output from an electronic device together with the first visual object. In the first effect information, animation effects or 3D effects may be applied to the generated first image (860), and audio effects may be output together with the first image (860).

[0183] According to one embodiment, second effect information may be generated by applying a prompt (e.g., the third prompt (732) of FIG. 7) containing second description information (880) to the first artificial intelligence model (802), and the generated second effect information may be applied to a second visual object or output from an electronic device together with the second visual object. In the second effect information, animation effects or 3D effects may be applied to the generated second image (862), and audio effects may be output together with the second image (862).

[0184] Table 6 shows exemplary first description information (820).

[0185]

[0186] Table 7 shows an exemplary template including the first descriptive information (820) of Table 6. The template of Table 7 may be included in the fourth prompt (842), and the template may include instructions and constraints for describing the next image of the image represented by the first descriptive information (820).

[0187]

[0188] Table 8 shows exemplary second descriptive information (880) generated by a fourth prompt (842) containing the template of Table 7. The second descriptive information (880) consists of text that satisfies the instructions and constraints listed in the template.

[0189]

[0190] Table 9 shows an exemplary template including the first descriptive information (820) of Table 6. The template of Table 9 may be included in the fourth prompt (842), and the template may include instructions and constraints for describing a previous image of the image represented by the first descriptive information (820).

[0191]

[0192] Table 10 shows exemplary second descriptive information (880) generated by a fourth prompt (842) containing the template of Table 9. The second descriptive information (880) consists of text that satisfies the instructions and constraints listed in the template.

[0193]

[0194] Table 11 shows an exemplary template including the first descriptive information (820) of Table 6. The template of Table 11 may be included in the fourth prompt (842), and the template may include instructions and constraints for describing an alternative image of the image represented by the first descriptive information (820).

[0195]

[0196] Table 12 shows exemplary second descriptive information (880) generated by a fourth prompt (842) containing the template of Table 11. The second descriptive information (880) consists of text that satisfies the instructions and constraints listed in the template.

[0197]

[0198] FIG. 9 is a diagram illustrating the process of selecting a second image to be generated using second description information.

[0199] Referring to Tables 8, 10 and 12, the electronic device (901) may acquire second depiction information describing a second image representing a previous situation, a next situation, or an alternative situation of the first image (9060), and may display buttons (9080 or 9082) together with the loading interface (9014) or with the image interface (9016) based on the second depiction information. A visual object (9040) may be displayed in the loading interface (9014), and the generated first image (9060) or second image may be displayed in the image interface (9016).

[0200] According to one embodiment, the buttons (9080 or 9082) may include a keyword for generating a second image associated with a first image (9060), and by selecting the button the user wants, the user can transmit the second description information to a second artificial intelligence model to generate a second image associated with the first image (9060), and the generated second image can be used as a sticker image in an electronic device (901).

[0201] Buttons (9080 or 9082) displayed together with the loading interface (9014) or the image interface (9016) may be generated based on second descriptive information representing a previous situation of the first image (9060) to be generated, for example, the buttons (9080) may include keywords representing a previous situation of the first image (9060) to be generated, such as "a panda arriving at a bubble tea shop," "a panda ordering bubble tea," "a panda watching the process of making bubble tea," "a panda promising to go to a bubble tea shop with friends," or "the moment a panda tired from the hot weather discovers a bubble tea shop." The buttons (9080) generated based on the second descriptive information may be displayed before the generation of the first image (9060) is completed.

[0202] Buttons (9080 or 9082) displayed together with the loading interface (9014) or the image interface (9016) may be generated based on second descriptive information indicating the next situation of the first image (9060) to be generated or generated, for example, the buttons (9080 or 9082) may include keywords indicating the next situation of the first image (9060) to be generated or generated, such as "with friends in front of a bubble tea shop," "panda bathing in bubble tea pearls," "panda working at a bubble tea farm," or "panda holding an empty bubble tea cup and feeling regretful." Buttons (9080) generated based on the second descriptive information may be shown before the generation of the first image (9060) is completed, but are not limited thereto, and buttons (9082) generated based on the second descriptive information may be displayed together with the generated first image (9060).

[0203] Buttons (9080 or 9082) displayed together with the loading interface (9014) or the image interface (9016) may be created based on second descriptive information representing an alternative situation of the first image (9060) to be created or created, and the buttons (9080 or 9082) may include other details that replace the details of the first image (9060), for example, keywords representing an alternative situation of the first image (9060) to be created or created, such as "polar bear," "ice cream," "bucket hat." Buttons (9080) created based on the second descriptive information may be shown before the creation of the first image (9060) is completed, but are not limited thereto, and buttons (9082) created based on the second descriptive information may be displayed together with the created first image (9060).

[0204] The image interface (9016) may include a phrase (9062) indicating that the image (9060) was generated by an artificial intelligence model, but the phrase (9062) may be omitted.

[0205] FIG. 10 is a diagram illustrating the process of receiving text input to generate a sticker image to be used in a messenger application according to one embodiment.

[0206] Referring to FIG. 10, a sticker image can be generated through a messenger application (1001) running on an electronic device. The messenger application (1001) may include a top tool bar (1002), a chat screen (1004), a keypad (1030A), and a keypad toolbar (1030B).

[0207] The keypad (1030A) may include an on-screen keyboard for receiving text input from a user. The keypad (1030A) may change according to the form of input required by the messenger application (1001), and the keypad toolbar (1030B) may also change according to the keypad (1030A).

[0208] The top toolbar (1002) may include, but is not limited to, a back button (10022), chat partner profile information (10024), chat partner contact information (10026), and option button (10028) as illustrated in FIG. 10.

[0209] The keypad toolbar (1030B) may include, but is not limited to, buttons for performing actions associated with the keypad as illustrated in FIG. 10, such as an AI button (1031), a sticker button (1032), a copy button (1033), a paste button (1034), a settings button (1035), and a more button (1036). When the AI ​​button (1031) is selected, AI options supported by the messenger application (1001) may be provided through the chat screen (1004) or the keypad (1030A).

[0210] Referring to FIG. 10, when the sticker button (1032) is selected on the keypad toolbar (1030B), the keypad toolbar (1030B) can be switched to a keypad (1030C) where various stickers are listed so that the user can browse and select various stickers. On the keypad toolbar (1030D) associated with the keypad (1030C), buttons for performing sticker-related actions can be displayed. The keypad toolbar (1030D) may include, but is not limited to, buttons for performing sticker-related actions, such as a button (10321) for returning to the previous keypad (1030A), a button (10322) for generating stickers based on artificial intelligence, buttons (10323, 10324, 10325, 10326, and 10327) for selecting various types of stickers, and a button (10328) for adding new stickers.

[0211] According to one embodiment, when a sticker button (1032) is selected and a button (10322) for generating a sticker based on artificial intelligence is selected, the aforementioned image generation interface (1010) may be displayed, but is not limited thereto, and the image generation interface (1010) may be displayed through other paths. For example, when an artificial intelligence button (1031) is selected, artificial intelligence options supported by the messenger application (1001) are provided, and among the provided options, an option for generating a sticker based on artificial intelligence may be included.

[0212] Referring to FIG. 10, when the image generation interface (1010) is displayed, it can be switched to a keypad (1030A) including an on-screen keyboard for receiving text input from a user.

[0213] The image generation interface (1010) may include a back button (1021) and a completion button (1023) for inputting the generated image into a message.

[0214] According to one embodiment, the number of characters of the user's text input used to generate an image may be limited, and the number of characters entered and the maximum number of characters (1015) may be indicated in the image generation interface (1010).

[0215] The image generation interface (1010) may include an input interface (1012), a generation button (1020) for requesting an artificial intelligence model to generate an image, and buttons (1022 and 1024) for selecting the style of the image to be generated. Since this is identical to the input interface (2012), generation button (2020), and style buttons (2022 and 2024) of FIG. 2, a redundant description is omitted. As shown in FIG. 10, the generation button (1020) may be disabled when no text is entered by a user into the input interface (1012), and the generation button (1020) may be enabled when text is entered by a user.

[0216] When a user inputs text input (e.g., text input (2011) of FIG. 2) to create a sticker image through the keypad (1030A) and selects the create button (1020), description information is created based on the text input, and while the sticker image is created based on the description information, a visual object can be acquired and displayed based on the description information.

[0217] FIG. 11 is a diagram illustrating the process of receiving text input and drawing input for generating an image according to one embodiment.

[0218] According to one embodiment, the user's text input may include not only text input directly entered by the user, but also text input converted from drawing input directly drawn by the user.

[0219] Referring to FIG. 11, the image generation interface (1110) may include a text input interface (1112A) for receiving text input from a user and a drawing input interface (1112B) for receiving drawing input from a user. A phrase (1113A) that prompts text input from a user may be displayed in the text input interface (1112A). A phrase (1113B) that prompts drawing input from a user may be displayed in the drawing input interface (1112B).

[0220] Referring to FIG. 11, the image generation interface (1110) may include, but is not limited to, a pen option button (1121) for changing the characteristics of the drawing input (e.g., line color, thickness, or transparency), an edit button (1123) for editing the drawing input such as copy, paste, or cut, and a more button (1125) for providing other options related to the drawing input.

[0221] According to one embodiment, the electronic device (1101) may generate a prompt including a drawing input drawn by a user. The prompt may be generated from a template that includes instructions or purposes for defining the overall direction of the task to be processed. The purpose of the template may include generating descriptive information based on the user's text input and / or drawing input entered through the text input interface (1112A) and / or the drawing input interface (1112B).

[0222] According to one embodiment, a prompt including a user's text input and drawing input may be input to an artificial intelligence model (e.g., the first artificial intelligence model (502) of FIG. 5 or an integrated artificial intelligence model). The artificial intelligence model may be a multimodal model that learns and processes relationships between data of various types or various modalities, such as text and images, and may be an LMM trained using text data and image data. Based on the input prompt, the artificial intelligence model may generate descriptive information describing the user's text input and drawing input, and may extract keywords of the descriptive information or generate effect information based on the generated descriptive information. As the method for generating descriptive information is described with reference to FIG. 4 and FIG. 5, the method for extracting keywords is described with reference to FIG. 6, and the method for generating effect information is described with reference to FIG. 7, redundant descriptions are omitted.

[0223] According to one embodiment, the generated descriptive information may be input to an artificial intelligence model (e.g., the second artificial intelligence model (504) of FIG. 5 or an integrated artificial intelligence model). The artificial intelligence model may be pre-trained to generate an image based on the input text prompt. The artificial intelligence model may generate an image (e.g., a sticker image) based on the input descriptive information and transmit the generated image to an electronic device (1101). While the image is being generated by the artificial intelligence model, the electronic device (1101) may acquire and display a visual object based on the descriptive information. The method of generating an image based on the descriptive information and displaying a visual object acquired based on the descriptive information while the image is being generated has been described above with reference to FIG. 2, FIG. 3, FIG. 4, FIG. 5, FIG. 6, and FIG. 7, so a redundant description is omitted.

[0224] According to one embodiment, the artificial intelligence model can immediately generate an image based on the user's drawing input without outputting descriptive information, and can modify or improve the generated image by adding details based on the user's text input to the generated image and transmit it to the electronic device (1101). In order for a visual object to be displayed on the electronic device (1101) while the image is being generated, the artificial intelligence model extracts keywords based on the user's text input and drawing input and transmits them to the electronic device (1101), and the electronic device (1101) can display a visual object based on the keywords. The method of displaying a visual object based on keywords has been described above with reference to FIG. 7, so a redundant description is omitted.

[0225] The image generation interface (1110) may include a generation button (1120) for requesting an artificial intelligence model to generate an image and a dropdown button (1126) for selecting the style of the image to be generated. When the dropdown button (1126) is selected, various style options that can be applied to the image to be generated may be displayed.

[0226] FIG. 12 is a diagram illustrating the process of generating a sticker image by referring to conversation history in a messenger application according to one embodiment.

[0227] According to one embodiment, a sticker image can be generated through a messenger application executed on an electronic device (1201), and the sticker image can be generated based on the conversation history in which the user participated in the messenger application.

[0228] For example, when a messenger application is executed on an electronic device (1201) and an image generation interface (1210) is called on a chat screen for conversing with one or more chat partners, the electronic device (1201) can recommend an appropriate sticker image to the user based on the conversation history with the chat partners. Referring to FIG. 12, if a message or context celebrating someone's birthday is found in the conversation history, the electronic device (1201) can recommend an image (1213B) indicating that the birthday is being celebrated to the user. The recommended image (1213B) can be displayed on an image input interface (1212B). The electronic device (1201) can recommend multiple images to the user, and a dot indicator (12107) indicating the position of the current image among the multiple images can be displayed on the image generation interface (1210). According to one embodiment, the recommended image (1213B) can be edited by the user through the image input interface (1212B).

[0229] A prompt including text and images entered through a text input interface (1212A) and an image input interface (1212B) may be input to an artificial intelligence model (e.g., the first artificial intelligence model (502) of FIG. 5 or an integrated artificial intelligence model), and the artificial intelligence model may generate descriptive information describing the text input and image input based on the input prompt, and may extract keywords of the descriptive information or generate effect information based on the generated descriptive information. The method of generating descriptive information is described with reference to FIG. 4 and FIG. 5, the method of extracting keywords is described with reference to FIG. 6, and the method of generating effect information is described with reference to FIG. 7, so redundant descriptions are omitted. The generated descriptive information may be input to an artificial intelligence model (e.g., the second artificial intelligence model (504) of FIG. 5 or an integrated artificial intelligence model), and an image may be generated. While the image is being generated, the electronic device (1201) may acquire and display a visual object based on the descriptive information. A method for generating an image based on descriptive information and displaying a visual object obtained based on descriptive information while the image is being generated has been described above with reference to FIGS. 2, 3, 4, 5, 6, and 7, so a redundant description is omitted.

[0230] According to one embodiment, the artificial intelligence model can modify or improve the recommended image (1213B) and transmit it to the electronic device (1201) by adding details to the recommended image (1213B) based on the user's text input without outputting descriptive information. In order for a visual object to be displayed on the electronic device (1201) while the image is being completed, the artificial intelligence model extracts keywords based on text input and image input and transmits them to the electronic device (1201), and the electronic device (1201) can display a visual object based on the keywords. The method of displaying a visual object based on keywords has been described above with reference to FIG. 7, so a redundant description is omitted.

[0231] Referring to FIG. 12, the image generation interface (1210) may include a button (1221) for calling a keypad (1230), a completion button (1223) for inputting the generated image into a message, and a button (1225) for shrinking or hiding the image generation interface (1210) via swipe input.

[0232] Referring to FIG. 12, the image generation interface (1210) may include a generate button (1220) for requesting an artificial intelligence model to generate an image and a dropdown button (1226) for selecting the style of the image to be generated. When the dropdown button (1226) is selected, various style options that can be applied to the image to be generated may be displayed. As illustrated in FIG. 12, the generate button (1220) may be disabled when no text is entered by a user in the text input interface (1212A), and the generate button (1220) may be enabled when text is entered by a user.

[0233] FIG. 13 is a diagram illustrating the process of modifying an image generated based on text input according to one embodiment.

[0234] FIG. 13 assumes that the image generated in FIG. 2 is modified. As shown in FIG. 2, a sticker image of a panda (1360A) can be generated based on the text input "animal drinking bubble tea".

[0235] Referring to FIG. 13, an electronic device (1301) can display an image (1360A) generated through an image generation interface (1310). The generated image (1360A) can be displayed in an image interface (1316A) of the image generation interface (1310). The image interface (1316A) may include a text input field (1318A) and a button (1319A) indicating a style applied to the image, and FIG. 13 illustrates that the existing text input "animal drinking bubble tea" in the text input field (1318A) has been modified to "cute tiger drinking bubble tea". When the text input displayed in the text input field (1318A) is modified, or when the button (1319) indicating a style applied to the image (1360A) is selected and the style is modified, the image generation button (1320) can be reactivated and selected.

[0236] When the activated image generation button (1320) is selected, the electronic device (1301) can acquire descriptive information for generating an image based on modified text input, and while the image is being generated based on the acquired descriptive information, it can display the visual object acquired based on the descriptive information. An image (1360B) can be generated based on the modified text input and displayed in the image interface (1316B). The image interface (1316B) may include a button (1319B) and a text input field (1318B) indicating a style applied to the image (1360B), and the text input field (1318B) may contain the modified text input. In FIG. 13, the image generation button (1320) displayed together with the image interface (1316B) is shown as disabled, but when the text input "cute tiger drinking bubble tea" displayed in the text input field (1318B) is modified, the disabled image generation button (1320) can be reactivated.

[0237] An image interface (1316B) that displays an image (1360B) generated based on modified text input can be separated from an image interface (1360A) that displays an image (1360A) generated based on existing text input, and the two interfaces (1316A and 1316B) can be scrolled or displayed alternately via swipe input.

[0238] FIG. 14 is a diagram showing an example of a system including a generative artificial intelligence model according to one embodiment.

[0239] The artificial intelligence system (1400) may include a user query / response interface (1410), an AI framework (1420), an application / service component (1430), knowledge repositories (1440), and / or a generative AI model (1450).

[0240] Referring to FIG. 14, a user query / response interface (1410) may receive input. The input may include user input and / or data obtained or generated by an electronic device (e.g., electronic device (101, 201, 401, 601, 701, 901)). The data may include images, videos, and / or sensor data generated by at least one processor of the electronic device (e.g., at least one processor (120)) (e.g., illuminance data around the electronic device obtained from a sensor or sensor hub (e.g., auxiliary processor (123)), attitude data (or orientation data) of the electronic device, temperature inside the electronic device (e.g., display module (160)) or temperature of at least one processor (120)), size information of the display area of ​​the display module (160), and / or images obtained through an image sensor of the electronic device (e.g., included in a camera module (180)). For example, user input may take the form of biosignals, natural language, touch data acquired through touch circuits included within the display module (e.g., used to identify input from a finger and / or stylus), images, and / or videos. Additionally, context information may be transmitted along with the user input. Context information may include various additional information at the time of user input. For example, this could include information about the application currently being used by the user or the user's location information. Furthermore, user input may take the form of a mixture of the aforementioned biosignals, natural language, images, sounds, and context information. Additionally, user input may take the form of non-natural language input, such as input for selecting a menu.

[0241] The user query / response interface (1410) can output results from a generative artificial intelligence system to the user. The output may include results (or result information) generated or obtained by the artificial intelligence system (1400) based on at least part of the input. The output may be in the form of natural language or specific content, and may also be provided in a form such as an action requested by the user. For example, the output may have a format according to the user settings of the electronic device. For example, the output may include information regarding an emergency situation in which the user is facing. The output may include detailed information about the emergency situation, for example, information indicating the type of emergency situation, information regarding the user's surrounding circumstances, personal status (health status) information, location information, information regarding the user's request, and information indicating the truthfulness of the user's request.

[0242] The AI ​​framework (1420) can receive input from the user and coordinate and control each component necessary to perform the user's intent based on the user's query.

[0243] User input received from the user query / response interface (1410) can be transmitted to a prompt design component (1421). The prompt design component (1421) can be used to generate a prompt suitable for inputting the user input into a generative AI model (1450) (e.g., a large language model (LLM), a large vision model (LVM), and / or large multimodal models (LMM)).

[0244] The prompt design component (1421) may be an AI component that uses machine learning algorithms or neural networks to develop better prompts over time. The prompt design component (1421) may generate prompts by accessing a knowledge component (e.g., knowledge repository (1440)) containing user preference data, a prompt library, and prompt examples based on user input, and may pass the generated prompts to a generative AI model (1450) (e.g., LLM, LVM, and / or LMM).

[0245] The APIs / Plugins management component (1423) can perform the role of communicating with external information when there is a request for additional information when user input is passed as input to the generative AI model (1450). The APIs / Plugins management component (1423) establishes a channel to communicate with the outside of the AI ​​interface via APIs, and can enable access to various data sources (e.g., knowledge repository (1440)) through the established channel. For example, the APIs / Plugins management component (1423) can be used to request another component (e.g., application / service component (1430)) that performs feedback (or response) according to the prompt. If the application or service needs to perform an action that ultimately executes the user's input rather than an intermediate result, the APIs / Plugins management component (1423) can request that action from the application / service component (1430) via APIs. Information obtained from the outside may be used to generate a prompt in the prompt design component (1421) along with user input, or it may be passed as input to the generative model.

[0246] A refiner component (e.g., output modification component (1425)) can fine-tune (or adjust) (or modify) the output produced by a generative AI model (1450) (e.g., LLM, LVM, and / or LMM). For example, the refiner component can verify whether the content generated by the generative AI model (1450) (e.g., LLM, LVM, and / or LMM) is irrelevant, contains biased content, or contains harmful content. Additionally, the refiner component can determine the extent to which the output matches the desired result and, if additional processing is required, proceed with that process. Furthermore, the refiner component can configure and provide hints to the user to help avoid unwanted outputs.

[0247] A generative AI model (1450) generally refers to an artificial intelligence neural network that generates new forms of data based on user input information. A generative AI model (1450) may include a model that generates images and / or a model that generates language. Models that generate images include, but are not limited to, GANs (generative adversarial networks) and VAEs (variational autoencoders), and examples include diffusion-based generative models that use VAEs and Transformer structures. Models that generate language are models trained to output the most statistically appropriate output value based on input values, and examples include models such as CHAT-GPT 3 and CHAT-GPT 4. There are also LMMs that can recognize various forms of data input, such as sensing data, biosignals, text, images, and voice, and generate new data corresponding to them.

[0248] In one embodiment, the AI ​​framework (1420) and / or generative AI model (1450) may be included in a server communicating with the electronic device, but is not limited thereto, and may be included in an AI module (e.g., including a processing circuit) within the electronic device. For example, the AI ​​module may be operatively coupled with at least one processor of the electronic device (e.g., at least one processor (120)). For example, the AI ​​module may be operatively coupled with a sensor hub of the electronic device for one or more sensors within the electronic device.

[0249] According to one embodiment, the electronic device may be configured to include at least some of the user query / response interface (1410) AI framework (1420), application / service component (1430), knowledge repository (1440), or generative AI model (1450) of FIG. 14. According to one embodiment, at least some of the user query / response interface (1410) AI framework (1420), application / service component (1430), knowledge repository (1440), or generative AI model (1450) of FIG. 14 may be included in another electronic device (e.g., a server, a wearable electronic device, or another user's electronic device) that communicates with the electronic device.

[0250] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.

[0251] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0252] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0253] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0254] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM It can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0255] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In electronic devices: Memory for storing instructions; and It includes one or more processors including processing circuitry, and When the above instructions are executed individually or collectively by the one or more processors, the electronic device: Receive user text input, Based on the above text input, descriptive information for generating an image is obtained, and Based on the above description information, a visual object is obtained, and Requesting the generation of an image based on a prompt containing the above description information, and Display the visual object while the image is being generated based on the above prompt, and Acquire the image generated based on the above prompt, and An electronic device that displays the above-mentioned generated image.

2. In Paragraph 1, The above visual object includes text, and When the above instructions are executed individually or jointly by the one or more processors, the electronic device: Receiving a first user input corresponding to the above visual object, and Displaying a plurality of user interface (UI) items for changing the text of the above visual object, and By receiving a second user input through the above plurality of UI items, the text of the visual object is changed, and Modify the description information based on the modified text of the visual object, and An electronic device for obtaining the image generated based on the prompt containing the modified description information.

3. In Paragraph 1, The above visual object includes text, and When the above instructions are executed individually or jointly by the one or more processors, the electronic device: Receiving a first user input corresponding to the above visual object, and Displaying a plurality of user interface (UI) items for changing the text of the above visual object, and By receiving a second user input through the above plurality of UI items, the text of the visual object is changed, and Obtaining other descriptive information based on the modified text of the above visual object, Obtaining another image generated based on another prompt containing the above other descriptive information, and An electronic device capable of displaying another image obtained above.

4. In Paragraph 1, The above description information includes text describing an image to be generated based on the above prompt, and When the above instructions are executed individually or jointly by the one or more processors, the electronic device: Perform sentence analysis on the above text to identify the main word or subject of the above text; The above-mentioned displayed visual object is an electronic device comprising the identified subject or the acquired keyword.

5. In Paragraph 1, The above description information includes text describing an image to be generated based on the above prompt, and When the above instructions are executed individually or jointly by the one or more processors, the electronic device: To obtain keywords of the text of the above description information, and The above-mentioned displayed visual object is an electronic device comprising the above-mentioned acquired keyword.

6. In Paragraph 5, The above keyword is an electronic device comprising the subject of the text of the above descriptive information and a predicate or modifier representing the subject.

7. In Paragraph 1, When the above instructions are executed individually or jointly by the one or more processors, the electronic device: Acquire effect information representing at least one of a visual effect, an audio effect, or a haptic effect to be applied to the above visual object, and An electronic device that outputs at least one of the visual effect, the audio effect, or the haptic effect while displaying the visual object.

8. In Paragraph 1, The above prompt is the second prompt, and When the above instructions are executed individually or jointly by the one or more processors, the electronic device: A first prompt including the above text input is applied to a first artificial intelligence (AI) model to obtain the description information from the first artificial intelligence model, and A method of applying the above-mentioned second prompt to a second artificial intelligence model to obtain the above-mentioned image generated by the above-mentioned second artificial intelligence model.

9. In Paragraph 8, When the above instructions are executed individually or jointly by the one or more processors, the electronic device: Generate the first prompt including the text input based on the template, and The above first prompt is input into the above first artificial intelligence model to obtain the above description information for generating the above image, and The above template is an electronic device that includes information regarding a constraint for unifying the subject in the text of the above description information.

10. In Paragraph 8, When the above instructions are executed individually or jointly by the one or more processors, the electronic device: By applying a third prompt including the above description information to the first artificial intelligence model or the third artificial intelligence model, effect information indicating at least one of a visual effect, an audio effect, or a haptic effect to be applied to the image is obtained, and An electronic device that outputs at least one of the visual effect, the audio effect, or the haptic effect while displaying the generated image.

11. In Paragraph 8, The above description information is the first description information, and When the above instructions are executed individually or jointly by the one or more processors, the electronic device: By applying a fourth prompt including the first description information to the first artificial intelligence model or the fourth artificial intelligence model, second description information is obtained that describes at least one of a previous image, a next image, or an alternative image of the image described by the first description information. At least one of the previous image, subsequent image, or alternative image generated by applying the fifth prompt including the second description information to the second artificial intelligence model or the fifth artificial intelligence model is obtained, and An electronic device capable of displaying at least one of the previously generated image, the subsequent image, or the alternative image.

12. In Paragraph 8, The first artificial intelligence model and the second artificial intelligence model are an electronic device deployed on a server that communicates with the electronic device.

13. Action of receiving user text input; An operation to obtain descriptive information for generating an image based on the above text input; An operation to acquire a visual object based on the above description information; An action requesting the generation of an image based on a prompt containing the above description information; An action of displaying the visual object while an image is being generated based on the above prompt; The operation of acquiring an image generated based on the above prompt; and A method comprising the operation of displaying the generated image above.

14. In Paragraph 13, The above visual object includes text, and The above method is: An operation to receive a first user input corresponding to the above visual object; An operation to display a plurality of user interface (UI) items for changing the text of the above visual object; An operation to change the text of the visual object by receiving a second user input through the plurality of UI items above; It includes an operation to modify the description information based on the modified text of the visual object, and The operation of acquiring the above-mentioned generated image is: A method comprising the operation of obtaining the image generated based on the prompt including the modified description information.

15. A computer-readable, non-transient recording medium storing a computer program comprising instructions, wherein the instructions, when executed by an electronic device, cause the electronic device: Action of receiving user text input; An operation to obtain descriptive information for generating an image based on the above text input; An operation to acquire a visual object based on the above description information; An action requesting the generation of an image based on a prompt containing the above description information; An action of displaying the visual object while an image is being generated based on the above prompt; The operation of acquiring an image generated based on the above prompt; and A recording medium that enables a method including the operation of displaying the generated image.