Method for providing images and electronic device performing method
The electronic device uses a generative AI model to process drawing and voice inputs with contextual data, addressing the lack of dynamic image generation by creating contextually rich AI images.
Patent Information
- Application Number
- PCT/KR2024/018955
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-21
- Filing Date
- 2024-11-27
- Publication Date
- 2025-07-24
AI Technical Summary
Existing technologies lack efficient methods for generating images using generative artificial intelligence models that incorporate both drawing and voice inputs, along with contextual information such as time and location, to create dynamic and contextually rich images.
An electronic device employs a generative AI model to process a combination of drawing and voice inputs, incorporating additional information like time and location, to generate AI images that reflect user interactions and contextual data.
The solution enables the creation of dynamic AI images that accurately represent user inputs and contextual information, enhancing the richness and relevance of generated content.
Smart Images

Figure KR2024018955_24072025_PF_FP_ABST
Abstract
Description
Image providing method and electronic device for doing so
[0001] Various embodiments of the present invention disclose a method for providing an image generated based on a generative artificial intelligence (AI) model and an electronic device for performing the same.
[0002] Generative models that can generate media such as text and images in response to prompts are being used in a variety of applications, including art and development.
[0003] According to one embodiment, an electronic device includes a display, at least one processor, and a memory storing instructions executable by the at least one processor, wherein when the instructions are executed by the at least one processor, the electronic device causes at least: receiving a first user input including at least one of a first drawing input or a first voice input, determining first data based on the first user input and first additional information about the first user input, wherein the first additional information includes at least one of first visual information about a time at which the first user input was received or first location information about a location at which the first drawing input was received; receiving a second user input including at least one of a second drawing input or a second voice input, determining second data based on the second user input and second additional information about the second user input, wherein the second additional information includes at least one of second visual information about a time at which the second user input was received or second location information about a location at which the second drawing input was received; and generating a generative artificial intelligence (AI) model. By using the method, an AI image can be generated based on at least a portion of the first data and the second data, and the AI image can be displayed through the display.
[0004] In one embodiment, a method performed by an electronic device including a display; at least one processor; and a memory storing instructions executable by the at least one processor, the method comprising: receiving a first user input including at least one of a first drawing input or a first voice input; determining first data based on the first user input and first additional information about the first user input, wherein the first additional information includes at least one of first visual information about a time at which the first user input was received or first location information about a location at which the first drawing input was received; receiving a second user input including at least one of a second drawing input or a second voice input; determining second data based on the second user input and second additional information about the second user input, wherein the second additional information includes at least one of second visual information about a time at which the second user input was received or second location information about a location at which the second drawing input was received; The method may include an operation of generating an AI image based on at least a portion of the first data and the second data using a generative AI (artificial intelligence) model; and an operation of displaying the AI image through the display.
[0005] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.
[0006] FIG. 2 is a diagram illustrating data for image generation according to one embodiment.
[0007] Figure 3 is a flowchart of an AI image providing method according to one embodiment.
[0008] FIG. 4 is a diagram for explaining an operation of displaying an AI image according to one embodiment.
[0009] FIGS. 5A and 5B are flowcharts of a method for providing objects corresponding to each drawing, according to one embodiment.
[0010] FIG. 6 is a drawing for explaining an operation of displaying a graphic object according to one embodiment.
[0011] FIG. 7 is a diagram for explaining an operation of displaying an AI object according to one embodiment.
[0012] Figure 8 is a flowchart of a video providing method according to one embodiment.
[0013] FIG. 9 is a drawing for explaining an operation of providing a video according to one embodiment.
[0014] FIG. 10 is a flowchart of a method for providing an edited AI image according to one embodiment.
[0015] FIG. 11 is a flowchart of a method for providing an edited AI image for a stored image according to one embodiment.
[0016] FIG. 12 is a diagram illustrating an operation of providing an edited AI image according to one embodiment.
[0017] Hereinafter, an embodiment of the present document may be described with reference to the attached drawings.
[0018] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.
[0019] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0020] The processor (120) may control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing, for example, software (e.g., a program (140)), and may perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculation, the processor (120) may store a command or data received from another component (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the command or data stored in the volatile memory (132), and store the resulting data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0021] The auxiliary processor (123) may control at least a part of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, in the electronic device (101) itself where artificial intelligence is performed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0022] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0023] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0024] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0025] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. According to one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0026] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0027] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0028] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0029] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0030] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0031] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0032] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0033] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0034] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0035] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0036] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0037] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0038] In one embodiment, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0039] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0040] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service by itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0041] FIG. 2 is a diagram illustrating data for image generation according to one embodiment.
[0042] According to one embodiment, the image providing system (hereinafter, system) (2) may include an electronic device (e.g., the electronic device (101) of FIG. 1). For example, the electronic device may include at least one processor (e.g., the processor (120) of FIG. 1), a communication module (e.g., the communication module (190) of FIG. 1), a memory (e.g., the memory (130) of FIG. 1), and a display (e.g., the display module (160) of FIG. 1).
[0043] According to one embodiment, the system (2) may be a system based on a generative artificial intelligence model. The system (2) may provide images using the generative AI model.
[0044] According to one embodiment, in the system (2), the electronic device may include a user query / response interface (hereinafter, “interface”) as part of its structure or may receive user input through the interface. The user input may be in the form of natural language, images, text, and / or videos, and is not limited to the present disclosure. Along with the user input, context information may be utilized. The context information may include various additional information at the time of the user input. For example, the context information may include information on the application currently being used by the user, information on the user’s location, or history associated with the system (2). The user input may be in the form of a mixture of at least one of natural language, images, sounds, or context information. The user input may be in the form of non-natural language, such as selecting a menu. The electronic device may output to the user a processing result using a generative AI model based on the user input. The processing result may be in the form of an image, natural language, other specific content, or a combination thereof, and may also be provided in the form of an action requested by the user.
[0045] The generative AI model of system (2) can receive user input and coordinate and control each component necessary to carry out the user's intent based on the user's query. Based on the user input received through the interface, the generative AI model can generate prompts suitable for input into a large language model (LLM) or a large multimodal model (LMM). The generative AI model can include an AI component that uses a machine learning algorithm or neural network to develop better prompts over time. User preference data, a prompt library, and prompt examples can be utilized to generate prompts based on user input.
[0046] Generative AI models generally refer to artificial intelligence neural networks that generate new forms of data based on user input. Generative AI models can include models that generate images. Representative examples of these models include generative adversarial networks (GANs) and variational autoencoders (VAEs). Examples include VAEs and diffusion-based generative models that utilize Transformer architectures.
[0047] In one embodiment, an electronic device may provide an AI image using a generative AI model. The AI image may represent an image generated using the generative AI model.
[0048] According to one embodiment, an electronic device may generate an AI image based on data (hereinafter, “data”) (200) for image generation using a generative AI model. The electronic device may generate an AI image based on at least a portion of the data (200) (or corresponding to the data (200)) using the generative AI model. The data (200) may represent an input to the generative AI model. For example, the data (200) may be a prompt. At least a portion of the data (200) may be input to the generative AI model as a prompt.
[0049] According to one embodiment, the data (200) may be input received from a user.
[0050] According to one embodiment, the electronic device may determine data (200) based on at least a portion of input received from a user.
[0051] According to one embodiment, data (200) may include data in various formats. Data (200) may include text, images, voice, video, other formats, or any combination thereof. Data (200) is not limited to the format of input received from a user. For example, data (200) may include voice data for a user's voice input and text converted from the voice data.
[0052] In one embodiment, an electronic device may include a generative AI model. The generative AI model of the electronic device may be an on-device AI model capable of generating AI images without external communication.
[0053] According to one embodiment, the system (2) may further include an external server. For example, the external server may include at least one generative AI model. For example, the external server may include multiple external servers.
[0054] In one embodiment, an electronic device may generate an AI image by linking with (or collaborating with) an external server. The electronic device may transmit an image generation request including data (200) or at least a portion of the data (200) to the external server. The external server may generate an AI image corresponding to the data (200) using the external server's generative AI model, and transmit the AI image corresponding to the data (200) to the electronic device.
[0055] In the present disclosure, the electronic device generating (or providing) an AI image may include generating the AI image using an on-device AI model of the electronic device, generating the AI image by linking with an external server that runs a generative AI model, or a combination thereof.
[0056] According to one embodiment, the electronic device can receive user input including at least one of a drawing input or a voice input.
[0057] According to one embodiment, the user input may include multiple user inputs. The electronic device may receive a first user input including at least one of a first drawing input or a first voice input, a second user input including at least one of a second drawing input or a second voice input, and an Nth user input including at least one of an Nth drawing input or an Nth voice input (N>=2).
[0058] In one embodiment, the electronic device may determine data based on user input and additional information about the user input. The additional information may include at least one of visual information about the time the user input was received or location information about the location where the drawing input was received.
[0059] According to one embodiment, the electronic device may determine first data based on first additional information for the first user input. The first additional information may include at least one of first visual information for the time at which the first user input was received or first location information for the location at which the first drawing input was received. For example, the first data may include first drawing data for the first drawing input, first voice data for the first voice input, first location information, and first visual information. Similarly, the electronic device may determine second data, …, and Nth data.
[0060] Referring to FIG. 2, a portion of the data (200) may include voice data and visual information. For example, if the Nth user input includes the Nth voice input, the Nth data may include the Nth voice data and the Nth visual information.
[0061] According to one embodiment, the electronic device may determine data (200) based on at least a portion of first data, …, or Nth data. For example, the data (200) may include first data to Nth data.
[0062] According to one embodiment, the electronic device can determine at least one object included in a user input (or data (200)). The at least one object included in the user input can represent at least one object to be included in an AI image. The electronic device can determine at least one object based on a user input, such as a drawing input and / or a voice input.
[0063] According to one embodiment, an electronic device may determine an object using an object recognition model. The object recognition model may include a deep learning-based model. For example, the object recognition model may include a neural network model (or algorithm) such as YOLO (You Only Look Once), Faster R-CNN (Region-Based Convolutional Neural Network), SSD (Single Shot Multi-Box Detector), or Mask R-CNN, but is not limited to the present disclosure. The electronic device may determine at least one object included in drawing data (or image) for a user's drawing input using the object recognition model.
[0064] According to one embodiment, an electronic device may determine an object using a context analysis model. The context analysis model may include a deep learning-based model. For example, the context analysis model may include a model based on bidirectional encoder representations from transformers (BERT), a generative pre-trained transformer (GPT), embeddings from language models (ELMo), DeepSpeech, or other transformer architectures, but is not limited to the present disclosure. The electronic device may determine at least one object included in speech data for a user's speech input using the context analysis model. The electronic device may generate text corresponding to the user's utterance represented by the speech data using the context analysis model or a separate speech-to-text (STT) conversion model. The electronic device may determine at least one object included in the text corresponding to the user's utterance using the context analysis model.
[0065] According to one embodiment, a user input may include multiple user inputs. The electronic device may distinguish (or identify, classify, or determine) the user input into multiple user inputs based on at least one object included in the user input. The electronic device may determine at least one object included in continuously received user inputs using at least one of an object recognition model and a context analysis model. The electronic device may distinguish the user input into multiple user inputs, each corresponding to different objects. For example, the electronic device may distinguish the user input into a first user input and a second user input, each corresponding to a first object and a second object, respectively.
[0066] In one embodiment, multiple user inputs may be associated with (or mapped to) visual information regarding the time each user input was received. The visual information may include a timestamp indicating the start and / or end of each user input.
[0067] For example, an electronic device may receive a drawing input of a tree and a voice input describing the tree as user input, and a drawing input of a mountain and a voice input describing the mountain. The electronic device may determine the tree and the mountain as a first object and a second object, respectively, among at least one object included in the user input, using at least one of an object recognition model and a context analysis model. The electronic device may distinguish the user input into a first user input including a drawing input of a tree and a voice input describing the tree, and a second user input including a drawing input of a mountain and a voice input describing the mountain.
[0068] For example, an electronic device may receive a drawing input of a tree and a voice input describing a mountain as user input. The electronic device may use an object recognition model and a context analysis model to determine the tree and the mountain as a first object and a second object, respectively, among at least one object included in the user input. The electronic device may distinguish the user input into a first user input including a drawing input of a tree and a second user input including a voice input describing a mountain.
[0069] Figure 3 is a flowchart of an AI image providing method according to one embodiment.
[0070] According to one embodiment, the operations 310 to 360 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1). For example, the electronic device may include at least one processor (e.g., the processor (120) of FIG. 1), a communication module (e.g., the communication module (190) of FIG. 1), a memory (e.g., the memory (130) of FIG. 1), and a display (e.g., the display module (160) of FIG. 1).
[0071] In one embodiment, a user of an electronic device may verbally describe a drawing while drawing on (or on) the display of the electronic device. Alternatively, the user of the electronic device may verbally describe a drawing while drawing on the display. The user may draw on the display of the electronic device using a pen or a body part (e.g., a hand) connected to the electronic device.
[0072] In one embodiment, a drawing input (e.g., a first drawing input, a second drawing input, etc.) may represent an input in which a user of an electronic device draws a picture on a display of the electronic device.
[0073] In one embodiment, the voice input (e.g., first voice input, second voice input, etc.) may represent an utterance input in which a user of the electronic device describes a picture.
[0074] In operation 310, the electronic device may receive a first user input comprising at least one of a first drawing input or a first voice input.
[0075] According to one embodiment, the electronic device may receive a first user input comprising a first drawing input and a first voice input.
[0076] According to one embodiment, the electronic device can display a first drawing corresponding to a first drawing input through a display at a location where the first drawing input was received.
[0077] In operation 320, the electronic device may determine first data based on a first user input and first additional information about the first user input. The first additional information may include at least one of first visual information about a time at which the first user input was received or first location information about a location at which the first drawing input was received.
[0078] In one embodiment, the first visual information may include a timestamp indicating when the first drawing input and / or the first voice input starts and / or ends. For example, the first visual information may include a timestamp indicating an earlier time between the time the first drawing input starts and the time the first voice input starts, and a timestamp indicating a later time between the time the first drawing input ends and the time the first voice input ends.
[0079] In one embodiment, the electronic device may determine first data associated with first visual information based on a first user input and first additional information about the first user input. For example, the electronic device may map a timestamp of the first visual information to the first data.
[0080] According to one embodiment, the first location information may include at least one coordinate at which the first drawing input is received via a display of the electronic device.
[0081] According to one embodiment, the first data may include at least one of first drawing data for a first drawing input, first voice data for a first voice input, first text generated by converting the first voice data into STT, first location information, or first visual information.
[0082] In operation 330, the electronic device may receive a second user input comprising at least one of a second drawing input or a second voice input.
[0083] According to one embodiment, the electronic device may receive a second user input comprising a second drawing input and a second voice input.
[0084] In operation 340, the electronic device may determine second data based on a second user input and second additional information about the second user input. The second additional information may include at least one of second visual information about a time at which the second user input was received or second location information about a location at which the second drawing input was received.
[0085] According to one embodiment, the second visual information may include a timestamp at which the second drawing input and / or the second voice input started and / or ended.
[0086] In one embodiment, the electronic device may determine second data associated with the second visual information based on the second user input and second additional information about the second user input. For example, the electronic device may map a timestamp of the second visual information to the second data.
[0087] In one embodiment, the second location information may include at least one coordinate at which the second drawing input was received via a display of the electronic device.
[0088] According to one embodiment, the second data may include at least one of second drawing data for a second drawing input, second voice data for a second voice input, second text generated by converting the second voice data into STT, second location information, or second visual information. For example, the second data may include a timetable based on the second drawing data, the second voice data, and visual information mapped to each data.
[0089] In operation 350, the electronic device may generate an AI image based on at least a portion of the first data and the second data using a generative AI model.
[0090] According to one embodiment, the electronic device may generate an AI image based on at least a portion of first data and second data by referencing first visual information and second visual information using a generative AI model. For example, the electronic device may generate an AI image including a first AI object corresponding to first data mapped to the first visual information and a second AI object corresponding to second data mapped to the second visual information.
[0091] According to one embodiment, the electronic device can generate an AI image based on at least a portion of the first data and the second data using an on-device generated AI model of the electronic device.
[0092] In one embodiment, the electronic device may be configured to generate an AI image based on at least a portion of the first data and the second data in conjunction with an external server that drives a generative AI model.
[0093] According to an embodiment, the electronic device may determine a prompt as input to a generative AI model based on at least a portion of the first data and the second data. For example, the prompt may include a timetable based on the first data, the second data, and visual information mapped to each data.
[0094] In Action 360, the electronic device can display AI images through the display.
[0095] According to one embodiment, an electronic device may store metadata including at least a portion of first data and second data in association with an AI image. The metadata may include at least one of drawing data for a drawing input (e.g., a first drawing input, a second drawing input), voice data for a voice input (e.g., a first voice input, a second voice input), text generated by converting voice data into STT, location information (e.g., a first location information, a second location information), or visual information (e.g., a first visual information, a second visual information). For example, the metadata may include drawing data, voice data, and visual information mapped to each data. For example, the metadata may include a timetable based on the drawing data, voice data, and visual information mapped to each data.
[0096] In one embodiment, the electronic device can determine at least one object in the AI image based on metadata stored in association with the AI image.
[0097] According to one embodiment, the electronic device may determine at least one object in an AI image using at least one of an object recognition model and a context analysis model. The metadata may further include information about at least one object in the AI image.
[0098] Although FIG. 3 illustrates a configuration in which a first user input is received at operation 310 and first data is determined at operation 320, this is merely for convenience of explanation, and operations 310 and 320 should not be construed as being performed in a time-series manner, but rather, it should be understood that each operation can be performed in parallel. Similarly, operation 330 for receiving a second user input and operation 340 for determining second data can be performed in parallel.
[0099] FIG. 4 is a diagram for explaining an operation of displaying an AI image according to one embodiment.
[0100] FIG. 4 illustrates screens (41 to 46) displayed on a display (e.g., a display module (160) of FIG. 1) of an electronic device (e.g., an electronic device (101) of FIG. 1). For example, the electronic device may include at least one processor (e.g., a processor (120) of FIG. 1), a communication module (e.g., a communication module (190) of FIG. 1), a memory (e.g., a memory (130) of FIG. 1), and a display.
[0101] The speech bubble in FIG. 4 is illustrated to describe the content of a voice input associated with a drawing input of a user of an electronic device, and may not be displayed on the display of the electronic device. Depending on the embodiment, a graphic object may be displayed to describe the user input, such as the speech bubble in FIG. 4.
[0102] On the screen (41), the electronic device can receive a first user input including a drawing input for drawing a tree (hereinafter, referred to as a first drawing input) and a voice input for describing the tree (hereinafter, referred to as a first voice input). The electronic device can display a first drawing corresponding to the first drawing input, such as a tree, at a location where the first drawing input was received through a display. For example, the first voice input can include object information (referring to FIG. 4 , a tree) and / or color information (referring to FIG. 4 , dark green) associated with the first drawing input.
[0103] On the screen (42), the electronic device can receive a second user input including a drawing input for drawing a mountain (hereinafter, referred to as a second drawing input) and a voice input for describing the mountain (hereinafter, referred to as a second voice input). The electronic device can display a second drawing corresponding to the second drawing input, such as a picture of a mountain, at a location where the second drawing input was received through the display. For example, the second voice input can include descriptive information associated with the second drawing input (referring to FIG. 4, a mountain with white snow).
[0104] On the screen (43), the electronic device can receive a third user input including a drawing input for drawing a cloud (hereinafter, referred to as a third drawing input) and a voice input for describing the cloud (hereinafter, referred to as a third voice input). The electronic device can display a cloud drawing, which is a third drawing corresponding to the third drawing input, at the location where the third drawing input was received through the display.
[0105] On the screen (44), the electronic device can receive a fourth user input including a drawing input for drawing a river (hereinafter, referred to as a fourth drawing input) and a voice input describing the river (hereinafter, referred to as a fourth voice input). The electronic device can display a fourth drawing corresponding to the fourth drawing input, that is, a picture of a river, at the location where the fourth drawing input was received through the display.
[0106] The electronic device may determine first data based on a first user input and first additional information about the first user input. The electronic device may determine first data associated with the first visual information at the time the first user input was received. For example, the electronic device may map a timestamp of the first visual information to the first data. Similarly, the electronic device may determine second data, third data, and fourth data associated with (or mapped to) each visual information.
[0107] The electronic device can generate an AI image based on at least a portion of the aforementioned data (Data 1 to Data 4) using a generative AI model. The electronic device can generate the AI image in response to a user input for generating the AI image. For example, the user can select a button (e.g., the 'generate' button) on the screen (44) to generate the AI image.
[0108] On screen (45), the electronic device may display a standby screen via the display while an AI image is being generated using a generative AI model. For example, the screen (45) may include a loading indicator or a progress indicator. For example, the screen (45) may include a visual effect or percentage indicating the degree of image generation or reception.
[0109] On the screen (46), the electronic device can display an AI image via the display. The AI image can include a dark green tree at a location where a first drawing input was received, a mountain with white snow at a location where a second drawing input was received, clouds at a location where a third drawing input was received, and a river at a location where a fourth drawing input was received.
[0110] FIGS. 5A and 5B are flowcharts of a method for providing objects corresponding to each drawing, according to one embodiment.
[0111] According to one embodiment, the operations below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1). For example, the electronic device may include at least one processor (e.g., the processor (120) of FIG. 1), a communication module (e.g., the communication module (190) of FIG. 1), a memory (e.g., the memory (130) of FIG. 1), and a display (e.g., the display module (160) of FIG. 1).
[0112] Hereinafter, Fig. 5a will be described. According to one embodiment, Fig. 5a is a flowchart of a method for providing a graphic object corresponding to each drawing.
[0113] In operation 510-1, the electronic device may receive a first user input including at least one of a first drawing input or a first voice input. The electronic device may determine first data associated with the first visual information based on the first user input and first additional information for the first user input.
[0114] In operation 520-1, the electronic device can display a first drawing corresponding to the first drawing input through a display at a location where the first drawing input was received.
[0115] According to one embodiment, operations 510-1 and 520-1 may be performed in parallel (or simultaneously).
[0116] In one embodiment, operation 520-1 may be performed after a predetermined delay from operation 510-1. Operations 510-1 and 520-1 may be performed almost simultaneously.
[0117] In operation 530-1, the electronic device can display a first graphic object corresponding to the first drawing based on the first location information by replacing the first drawing through the display.
[0118] In one embodiment, the first graphic object may include a first AI object generated using a pre-stored shape, icon, or generative AI model corresponding to the first drawing.
[0119] In one embodiment, the electronic device may generate a first AI object corresponding to first data using a generative AI model. The electronic device may display the first AI object by replacing the first drawing at the location where the first drawing input was received.
[0120] In one embodiment, the first graphic object may be a graphic object that applies straightening or line smoothing to the first drawing. For example, the electronic device may correct curved or shaky lines of the first drawing to display the first graphic object as a smooth and straight line.
[0121] In operation 540-1, the electronic device may receive a second user input including at least one of a second drawing input or a second voice input. The electronic device may determine second data associated with the second visual information based on the second user input and second additional information about the second user input.
[0122] In operation 550-1, the electronic device can display a second drawing corresponding to the second drawing input through the display at a location where the second drawing input was received.
[0123] According to one embodiment, operations 540-1 and 550-1 may be performed in parallel (or simultaneously).
[0124] In one embodiment, operation 550-1 may be performed after a predetermined delay from operation 540-1. Operations 540-1 and 550-1 may be performed almost simultaneously.
[0125] In operation 560-1, the electronic device can display a second graphic object corresponding to the second drawing based on the second location information by replacing the second drawing through the display.
[0126] In one embodiment, the second graphical object may include a second AI object generated using a pre-stored shape, icon, or generative AI model corresponding to the second drawing.
[0127] In one embodiment, the electronic device may generate a second AI object corresponding to the second data using a generative AI model. The electronic device may display the second AI object by replacing the second drawing at the location where the second drawing input is received.
[0128] In one embodiment, the second graphic object may be a graphic object that has straightening or line smoothing applied to the second drawing. For example, the electronic device may correct curved or shaky lines of the second drawing to display the second graphic object as a smooth and straight line.
[0129] According to one embodiment, operation 310 of receiving a first user input of FIG. 3 may include operation 510-1 of receiving a first user input. Operation 330 of receiving a second user input of FIG. 3 may include operation 540-1 of receiving a second user input.
[0130] According to one embodiment, operation 520-1 of displaying a first drawing may be performed in parallel with operation 320 of FIG. 3 of determining first data. Operation 550-1 of displaying a second drawing may be performed in parallel with operation 340 of FIG. 3 of determining second data.
[0131] According to one embodiment, after the operations of FIG. 5A are performed, the electronic device can generate an AI image (operation 350 of FIG. 3) and display the AI image (operation 360 of FIG. 3).
[0132] Hereinafter, Fig. 5b will be described. According to one embodiment, Fig. 5b is a flowchart of a method for providing an AI image including an AI object corresponding to each drawing.
[0133] In operation 510-2, the electronic device may receive a first user input including at least one of a first drawing input or a first voice input. The electronic device may determine first data associated with the first visual information based on the first user input and first additional information about the first user input.
[0134] In operation 520-2, the electronic device can display a first drawing corresponding to the first drawing input through a display at a location where the first drawing input was received.
[0135] According to one embodiment, operations 510-2 and 520-2 may be performed in parallel (or simultaneously).
[0136] In one embodiment, operation 520-2 may be performed after a predetermined delay from operation 510-2. Operations 510-2 and 520-2 may be performed almost simultaneously.
[0137] In operation 530-2, the electronic device may generate a first AI image including a first AI object corresponding to the first drawing based on at least a portion of the first data using a generative AI model. Before receiving a second user input, the electronic device may generate a first AI image including a first AI object corresponding to the first drawing based on at least a portion of the first data using the generative AI model.
[0138] In one embodiment, the first AI image may include a first AI object at a location where the first drawing input was received.
[0139] In one embodiment, the first AI image may be a screen displayed via the display by operation 520-2 that displays the first drawing, in which the first drawing is replaced with the first AI object.
[0140] According to one embodiment, the first AI image may be a screen displayed through the display by operation 520-2 that displays the first drawing, in which the first drawing is replaced by the first AI object and a graphic effect associated with the first AI object is applied. For example, the first AI image may be a screen displayed by the first AI object with a light and shade effect applied. For example, the first AI image may be a screen displayed by the first AI object with a shading effect applied.
[0141] In operation 530-3, the electronic device can display a first AI image through a display.
[0142] In operation 540-2, the electronic device may receive a second user input including at least one of a second drawing input or a second voice input. The electronic device may determine second data associated with the second visual information based on the second user input and second additional information about the second user input.
[0143] According to one embodiment, the electronic device may receive a second drawing input for (or on) the first AI image.
[0144] In operation 550-2, the electronic device can display a second drawing corresponding to the second drawing input through the display at a location in the first AI image where the second drawing input was received.
[0145] According to one embodiment, operations 540-2 and 550-2 may be performed in parallel (or simultaneously).
[0146] In one embodiment, operation 550-2 may be performed after a predetermined delay from operation 540-2. Operations 540-2 and 550-2 may be performed almost simultaneously.
[0147] In operation 560-2, the electronic device may generate a second AI image including a second AI object corresponding to the second drawing based on at least a portion of the second data using a generative AI model.
[0148] In one embodiment, the second AI image may include a second AI object at a location where the second drawing input was received.
[0149] In one embodiment, the second AI image may be a screen displayed via the display by operation 550-2 that displays the second drawing (i.e., a screen in which the second drawing is displayed in the first AI image), in which the second drawing is replaced with a second AI object.
[0150] In one embodiment, the second AI image may be a screen displayed through the display by operation 550-2, in which the second drawing is replaced by a second AI object, and a graphic effect associated with the second AI object is applied. For example, the second AI image may be a screen displayed by the second AI object with a light and shade effect applied. For example, the second AI image may be a screen displayed by the second AI object with a shading effect applied.
[0151] In operation 560-3, the electronic device can display a second AI image through the display.
[0152] According to one embodiment, operation 310 of receiving a first user input of FIG. 3 may include operation 510-2 of receiving a first user input. Operation 330 of receiving a second user input of FIG. 3 may include operation 540-2 of receiving a second user input.
[0153] According to one embodiment, operation 520-2 of displaying a first drawing may be performed in parallel with operation 320 of FIG. 3 of determining first data. Operation 550-2 of displaying a second drawing may be performed in parallel with operation 340 of FIG. 3 of determining second data.
[0154] FIG. 6 is a drawing for explaining an operation of displaying a graphic object according to one embodiment.
[0155] FIG. 6 illustrates screens (61 to 68) displayed on a display (e.g., a display module (160) of FIG. 1) of an electronic device (e.g., an electronic device (101) of FIG. 1). For example, the electronic device may include at least one processor (e.g., a processor (120) of FIG. 1), a communication module (e.g., a communication module (190) of FIG. 1), a memory (e.g., a memory (130) of FIG. 1), and a display.
[0156] According to one embodiment, FIG. 6 is a drawing for explaining screens (61 to 68) displayed on a display while the operations of FIG. 5a are performed.
[0157] The speech bubble in FIG. 6 is illustrated to describe the content of a voice input associated with a drawing input of a user of an electronic device, and may not be displayed on the display of the electronic device. In some embodiments, a graphic object may be displayed to describe the user input, such as the speech bubble in FIG. 6.
[0158] On the screen (61), the electronic device can receive a first user input including a drawing input for drawing the moon (hereinafter, referred to as a first drawing input) and a voice input for describing the moon (hereinafter, referred to as a first voice input). The electronic device can display a first drawing corresponding to the first drawing input, that is, a picture of the moon, at a location where the first drawing input was received through the display.
[0159] On the screen (62), the electronic device can display a first graphic object corresponding to the first drawing based on the first location information by replacing the first drawing through the display. For example, the first graphic object may be a pre-stored shape corresponding to the first drawing (a circle, as shown in FIG. 6). For example, the first graphic object may be a graphic object to which straightening or line smoothing has been applied to the first drawing.
[0160] On the screen (63), the electronic device can receive a second user input including a drawing input for drawing a hill (hereinafter, referred to as a second drawing input) and a voice input for describing the hill (hereinafter, referred to as a second voice input). The electronic device can display a second drawing corresponding to the second drawing input, that is, a drawing of a hill, at the location where the second drawing input was received through the display.
[0161] On the screen (64), the electronic device can display a second graphic object corresponding to the second drawing based on the second location information by replacing the second drawing through the display. For example, the second graphic object may be a pre-stored shape (a straight line, as shown in FIG. 6 ) corresponding to the second drawing. For example, the second graphic object may be a graphic object to which straightening or line smoothing has been applied to the second drawing.
[0162] On the screen (65), the electronic device can receive a third user input including a drawing input for drawing a tree (hereinafter, referred to as a third drawing input) and a voice input describing the tree (hereinafter, referred to as a third voice input). The electronic device can display a third drawing corresponding to the third drawing input, a tree drawing, at the location where the third drawing input was received through the display.
[0163] On the screen (66), the electronic device can display a third graphic object corresponding to the third drawing based on the third position information by replacing the third drawing through the display. For example, the third graphic object may be a graphic object that has straightening or line smoothing applied to the third drawing.
[0164] An electronic device may determine first data based on a first user input and first additional information about the first user input. The electronic device may determine first data associated with first visual information regarding the time at which the first user input was received. For example, the electronic device may map a timestamp of the first visual information to the first data. Similarly, the electronic device may determine second data and third data associated with (or mapped to) each visual information.
[0165] The electronic device can generate an AI image based on at least a portion of the aforementioned data (Data 1 to Data 3) using a generative AI model. The electronic device can generate the AI image in response to a user input for generating the AI image. For example, the user can select a button (e.g., the "generate" button) on the screen (66) to generate the AI image.
[0166] On screen (67), the electronic device may display a standby screen via the display while the AI image is being generated using the generative AI model. For example, screen (67) may include a loading indicator or progress indicator. For example, screen (67) may include a visual effect or percentage indicating the degree of image generation or reception.
[0167] On the screen (68), the electronic device can display an AI image via the display. The AI image can include a moon at a location where a first drawing input was received, a hill at a location where a second drawing input was received, and a tree at a location where a third drawing input was received.
[0168] FIG. 7 is a diagram for explaining an operation of displaying an AI object according to one embodiment.
[0169] FIG. 7 illustrates screens (71 to 78) displayed on a display (e.g., a display module (160) of FIG. 1) of an electronic device (e.g., an electronic device (101) of FIG. 1). For example, the electronic device may include at least one processor (e.g., a processor (120) of FIG. 1), a communication module (e.g., a communication module (190) of FIG. 1), a memory (e.g., a memory (130) of FIG. 1), and a display.
[0170] According to one embodiment, FIG. 7 is a drawing for explaining screens (71 to 78) displayed on a display while the operations of FIG. 5b are performed.
[0171] The speech bubble in FIG. 7 is illustrated to describe the content of a voice input associated with a drawing input of a user of an electronic device, and may not be displayed on the display of the electronic device. Depending on the embodiment, a graphic object may be displayed to describe the user input, such as the speech bubble in FIG. 7.
[0172] On the screen (71), the electronic device can receive a first user input including a drawing input drawing the sky (hereinafter, referred to as a first drawing input) and a voice input describing the sky (hereinafter, referred to as a first voice input). The electronic device can display a first drawing corresponding to the first drawing input, that is, a picture of the sky, at a location where the first drawing input was received through a display. For example, the first voice input can include description information (see FIG. 7, night sky) associated with the first drawing input. The electronic device can generate a first AI image including a first AI object corresponding to the first drawing using a generative AI model.
[0173] On the screen (72), the electronic device may display a first AI image including a first AI object (see FIG. 7, a night sky) corresponding to the first drawing through the display. Referring to FIG. 7, the first AI image may be a first drawing replaced with a first AI object.
[0174] On the screen (73), the electronic device can receive a second user input including a drawing input (hereinafter, a second drawing input) for drawing a moon on the first AI image (or on the first AI image) and a voice input (hereinafter, a second voice input) for describing the moon. The electronic device can display a second drawing corresponding to the second drawing input, that is, a picture of the moon, at a location in the first AI image where the second drawing input was received through the display. The electronic device can generate a second AI image including a second AI object corresponding to the second drawing using a generative AI model.
[0175] On the screen (74), the electronic device may display a second AI image including a second AI object (see FIG. 7, the moon) corresponding to the second drawing via the display. Compared to the screen (73), the second AI image may be one in which the second drawing is replaced by the second AI object and a graphic effect associated with the second AI object is applied. Referring to FIG. 7, the second AI image may be one in which a light and shade effect is applied by the second AI object.
[0176] On the screen (75), the electronic device can receive a third user input including a drawing input for drawing a hill on the second AI image (hereinafter, referred to as a third drawing input) and a voice input for describing the hill (hereinafter, referred to as a third voice input). The electronic device can display a third drawing corresponding to the third drawing input, that is, a drawing of a hill, at a location in the second AI image where the third drawing input was received through the display. The electronic device can generate a third AI image including a third AI object corresponding to the third drawing using a generative AI model.
[0177] On the screen (76), the electronic device may display a third AI image including a third AI object (see FIG. 7, a hill) corresponding to the third drawing via the display. The third AI image may be a third drawing replaced with a third AI object, as compared to the screen (76).
[0178] On the screen (77), the electronic device can receive a fourth user input including a drawing input of a tree (hereinafter, referred to as a fourth drawing input) for a third AI image and a voice input describing the tree (hereinafter, referred to as a fourth voice input). The electronic device can display a fourth drawing corresponding to the fourth drawing input, that is, a picture of a tree, at a location where the fourth drawing input was received through the display. The electronic device can generate a fourth AI image including a fourth AI object corresponding to the fourth drawing using a generative AI model.
[0179] On the screen (78), the electronic device may display a fourth AI image including a fourth AI object (see FIG. 8, a tree) corresponding to the fourth drawing through the display. The fourth AI image may be, compared to the screen (77), an image in which the fourth drawing is replaced with the fourth AI object.
[0180] Figure 8 is a flowchart of a video providing method according to one embodiment.
[0181] According to one embodiment, the operations 810 to 830 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1). For example, the electronic device may include at least one processor (e.g., the processor (120) of FIG. 1), a communication module (e.g., the communication module (190) of FIG. 1), a memory (e.g., the memory (130) of FIG. 1), and a display (e.g., the display module (160) of FIG. 1).
[0182] As described with reference to FIG. 5B, the electronic device can generate a first AI image based on at least a portion of first data associated with the first visual information, and can generate a second AI image based on at least a portion of second data associated with the second visual information.
[0183] According to one embodiment, operations 810 to 830 may be performed after operations of FIG. 5b for generating the first AI image and the second AI image.
[0184] In operation 810, the electronic device can generate a video including a first frame corresponding to a first AI image and a second frame corresponding to a second AI image in chronological order.
[0185] In one embodiment, the electronic device may generate a video in response to a user input to initiate generation of the video.
[0186] According to one embodiment, the electronic device may generate a video by referencing first visual information and second visual information. The electronic device may generate a video in which the first AI image and the second AI image are arranged in chronological order by referencing the first visual information and the second visual information.
[0187] In one embodiment, the electronic device may receive a video from an external server, the video including a first frame corresponding to a first AI image and a second frame corresponding to a second AI image in chronological order. For example, the external server may be an external server that generated the first AI image and the second AI image. For example, the external server may be an external server that generates a video based on the images.
[0188] At step 820, the electronic device may receive additional user input related to the style of the video. The style of the video may include categories such as the style, mood, and drawing technique of the video (or AI images included in the video). The user may select the style of the video through additional user input to the electronic device.
[0189] In operation 830, the electronic device may use a generative AI model to generate an AI video based on additional user input and the video. The AI video may be a video generated in operation 810 with a style applied that corresponds to the additional user input.
[0190] In one embodiment, the electronic device may generate an AI video in response to a user input to initiate generation of the AI video.
[0191] FIG. 9 is a drawing for explaining an operation of providing a video according to one embodiment.
[0192] FIG. 9 illustrates screens (91, 92) displayed on a display (e.g., a display module (160) of FIG. 1) of an electronic device (e.g., an electronic device (101) of FIG. 1). For example, the electronic device may include at least one processor (e.g., a processor (120) of FIG. 1), a communication module (e.g., a communication module (190) of FIG. 1), a memory (e.g., a memory (130) of FIG. 1), and a display.
[0193] According to one embodiment, operation 810 of FIG. 8 for generating a video may be performed after the operations of FIG. 5b for generating a first AI image and a second AI image.
[0194] In one embodiment, the electronic device may generate a video in response to a user input for generating the video. For example, a user may select a button for generating the video on screen (78) of FIG. 7 or any other screen. Screen (91) may then be displayed.
[0195] On the screen (91), the electronic device may display a standby screen through the display while generating a video including a plurality of frames corresponding to a plurality of AI images (e.g., the first AI image to the fourth AI image described with reference to FIG. 7) in chronological order. For example, the screen (91) may include a loading indicator or a progress indicator. For example, the screen (91) may include a visual effect or percentage indicating the degree of image generation or reception.
[0196] On the screen (92), the electronic device can display the generated video in a playable state through the display.
[0197] According to one embodiment, an electronic device can insert text or sound effects into a video based on metadata. Any description of metadata that overlaps with the above description will be omitted.
[0198] FIG. 10 is a flowchart of a method for providing an edited AI image according to one embodiment.
[0199] According to one embodiment, the operations 1010 to 1040 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1). For example, the electronic device may include at least one processor (e.g., the processor (120) of FIG. 1), a communication module (e.g., the communication module (190) of FIG. 1), a memory (e.g., the memory (130) of FIG. 1), and a display (e.g., the display module (160) of FIG. 1).
[0200] According to one embodiment, operations 1010 to 1040 may be performed after operation 360 of FIG. 3.
[0201] In one embodiment, a user of an electronic device may draw on an AI image displayed on the display of the electronic device while verbally describing the drawing to edit the AI image. Alternatively, the user of the electronic device may draw on an AI image displayed on the display while verbally describing the drawing.
[0202] In operation 1010, the electronic device may receive an editing user input comprising at least one of an editing drawing input or an editing voice input for the AI image.
[0203] According to one embodiment, the electronic device can receive editing user input including editing drawing input and editing voice input for an AI image.
[0204] According to one embodiment, the electronic device can display an edit drawing corresponding to an edit drawing input through a display at a location where the edit drawing input was received.
[0205] In operation 1020, the electronic device may determine edit data based on an edit user input and edit additional information about the edit user input. The edit additional information may include at least one of visual information about the time the edit user input was received or positional information about the location where the edit drawing input was received.
[0206] According to one embodiment, the time information regarding the time at which the editing user input was received may include a timestamp at which the editing drawing input and / or the editing voice input started and / or ended. For example, the time information may include a timestamp that is earlier than the time at which the editing drawing input started and / or the editing voice input started, and a timestamp that is later than the time at which the editing drawing input ended and / or the editing voice input ended.
[0207] In one embodiment, the electronic device may determine edit data associated with visual information based on an editing user input and editing supplementary information for the editing user input. For example, the electronic device may map a timestamp of the visual information to the edit data.
[0208] According to one embodiment, the location information for a location where an editing user input was received may include at least one coordinate where the editing drawing input was received via a display of the electronic device.
[0209] According to one embodiment, the edit data may include at least one of edit drawing data for edit drawing input, edit voice data for edit voice input, edit text generated by converting edit voice data to STT, location information, or visual information.
[0210] According to one embodiment, the electronic device can determine at least one object included in the editing user input. The at least one object included in the editing user input can represent at least one object to be edited in the AI image. The electronic device can determine at least one object based on the editing user input, such as an editing drawing input and / or an editing voice input. The method for determining an object using the object recognition model or context analysis model of FIG. 2 can be similarly applied, and redundant descriptions are omitted.
[0211] According to one embodiment, an editing user input may include a plurality of editing user inputs. The electronic device may distinguish (or identify, classify, or determine) the editing user input into a plurality of editing user inputs based on at least one object included in the editing user input. The electronic device may determine at least one object included in the continuously received editing user input using at least one of an object recognition model and a context analysis model. The electronic device may distinguish the editing user input into a plurality of editing user inputs, each corresponding to a different object. For example, the electronic device may distinguish the editing user input into a first editing user input and a second editing user input, each corresponding to a different first object and a different second object, respectively.
[0212] In one embodiment, a plurality of editing user inputs may be associated with (or mapped to) visual information regarding the time at which each editing user input was received. The visual information may include a timestamp indicating when each editing user input began and / or ended.
[0213] According to one embodiment, when an editing user input includes a plurality of editing user inputs, the electronic device may determine editing data corresponding to each of the plurality of editing user inputs. For example, when the editing user input includes a first editing user input and a second editing user input, the electronic device may determine first editing data associated with time information for a time at which the first editing user input was received based on the first editing user input and first editing additional information for the first editing user input, and similarly determine second editing data.
[0214] In operation 1030, the electronic device may use a generative AI model to generate another AI image based on the edited data and at least a portion of the AI image. The other AI image may represent the edited AI image.
[0215] According to one embodiment, an electronic device can generate another AI image by referencing visual information in edited data using a generative AI model. The electronic device can generate another AI image based on edited data mapped to each visual information.
[0216] According to one embodiment, the electronic device can determine at least one object in the AI image.
[0217] For example, the electronic device may determine at least one object in the AI image based on metadata stored in association with the AI image. The electronic device may determine at least one object in the AI image using at least one of an object recognition model and a context analysis model.
[0218] For example, the electronic device may identify (or reference) information if metadata stored in association with an AI image includes information about at least one object in the AI image.
[0219] For example, an electronic device may use an object recognition model to determine at least one object in an AI image. That is, the electronic device may determine at least one object through object recognition of the AI image itself.
[0220] According to one embodiment, the electronic device can generate another AI image (i.e., an edited AI image) by transforming, based on editing data, at least one object of the AI image, corresponding to at least one object included in the editing user input.
[0221] For example, the electronic device can generate another AI image by converting an object corresponding to a first object among at least one object included in the editing user input, among at least one object of the AI image, based on the editing data. For example, if the editing user input includes an editing drawing input and / or an editing voice input for editing the 'moon', the other AI image can be generated by converting the 'moon' among at least one object of the AI image based on the editing data.
[0222] According to one embodiment, an electronic device may generate another AI image based on at least a portion of edit data, an AI image, and metadata stored in association with the AI image, using a generative AI model. As described with reference to FIG. 3, the electronic device may store metadata, including at least a portion of data (e.g., first data, second data) used to generate the AI image, in association with the AI image. The metadata may include at least one of drawing data for a drawing input (e.g., first drawing input, second drawing input), voice data for a voice input (e.g., first voice input, second voice input), text generated by converting voice data into STT, location information (e.g., first location information, second location information), visual information (e.g., first visual information, second visual information), or at least one object of the AI image. For example, the electronic device may generate another AI image by editing the AI image using data mapped with each visual information by referring to the visual information (timestamp) of the edit data and the metadata.
[0223] In action 1040, the electronic device can display another AI image through the display.
[0224] According to one embodiment, the electronic device may perform the following operations while performing the operations of FIG. 5b.
[0225] According to one embodiment, after operation 530-3 of displaying the first AI image of FIG. 5, the electronic device may receive an editing user input including at least one of an editing drawing input or an editing voice input for the first AI image. The electronic device may determine editing data based on the editing user input and editing additional information for the editing user input. The editing additional information may include at least one of visual information about the time at which the editing user input was received or location information about the location at which the editing drawing input was received. The electronic device may generate another first AI image (i.e., an edited first AI image) corresponding to the first AI image based on the editing data and at least a portion of the first AI image using a generative AI model. Descriptions overlapping with those described above will be omitted.
[0226] FIG. 11 is a flowchart of a method for providing an edited AI image for a stored image according to one embodiment.
[0227] According to one embodiment, the operations 1110 to 1160 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1). For example, the electronic device may include at least one processor (e.g., the processor (120) of FIG. 1), a communication module (e.g., the communication module (190) of FIG. 1), a memory (e.g., the memory (130) of FIG. 1), and a display (e.g., the display module (160) of FIG. 1).
[0228] According to one embodiment, operation 310 of FIG. 3 may include operation 1110, operation 320 may include operation 1120, operation 330 may include operation 1130, operation 340 may include operation 1140, operation 350 may include operation 1150, and operation 360 may include operation 1160.
[0229] In one embodiment, an electronic device may display a stored image stored in memory (or storage) via a display. The stored image may include any image, such as a photograph or picture, previously stored in the memory or storage (e.g., public media storage, individual storage of any application) of the electronic device.
[0230] For example, the electronic device may display an image selected by the user from a gallery or drawing application. For example, the electronic device may display a new image captured by the user using the camera of the electronic device. For example, the electronic device may display an image temporarily stored on the clipboard (e.g., an image copied from any source, such as a web browser application). For example, the electronic device may display a pre-generated AI image. A method for providing an edited AI image for a pre-generated AI image is described with reference to FIG. 10.
[0231] In operation 1110, the electronic device may receive a first user input comprising at least one of a first drawing input or a first voice input for a stored image.
[0232] According to one embodiment, the electronic device can display a first drawing corresponding to a first drawing input through a display at a location on a stored image where the first drawing input was received.
[0233] In operation 1120, the electronic device may determine first data associated with the first visual information based on the first user input and first additional information about the first user input. The first additional information may include at least one of first visual information about the time at which the first user input was received or first location information about the location at which the first drawing input was received.
[0234] In operation 1130, the electronic device may receive a second user input comprising at least one of a second drawing input or a second voice input for the stored image.
[0235] In operation 1140, the electronic device may determine second data associated with the second visual information based on the second user input and second additional information about the second user input. Descriptions overlapping with those described above regarding the first user input and the first data are omitted.
[0236] According to one embodiment, the electronic device can determine at least one object included in the first user input and the second user input. The at least one object included in the first user input and the second user input may represent at least one object to be edited in the stored image. The method for determining an object using the object recognition model or context analysis model of FIG. 2 is similarly applicable, and redundant descriptions are omitted.
[0237] In operation 1150, the electronic device may generate an AI image based on at least a portion of the first data, the second data, and the stored image using a generative AI model. The AI image may represent an edited AI image of the stored image.
[0238] According to one embodiment, the electronic device can generate an AI image based on at least a portion of first data, second data, and stored images by referencing first visual information and second visual information using a generative AI model.
[0239] In one embodiment, the electronic device can determine at least one object in a stored image. For example, the electronic device can determine at least one object in the stored image using an object recognition model.
[0240] According to one embodiment, the electronic device can generate an AI image (i.e., an edited AI image) by converting, based on first data and second data, at least one object of a stored image, an object corresponding to at least one object included in a first user input and a second user input.
[0241] At operation 1160, the electronic device can display an AI image through a display.
[0242] FIG. 12 is a diagram illustrating an operation of providing an edited AI image according to one embodiment.
[0243] FIG. 12 illustrates screens (1201 to 1206) displayed on a display (e.g., a display module (160) of FIG. 1) of an electronic device (e.g., an electronic device (101) of FIG. 1). For example, the electronic device may include at least one processor (e.g., a processor (120) of FIG. 1), a communication module (e.g., a communication module (190) of FIG. 1), a memory (e.g., a memory (130) of FIG. 1), and a display.
[0244] The speech bubble of FIG. 12 is illustrated to describe the content of the edit voice input associated with the edit drawing input of the user of the electronic device, and may not be displayed on the display of the electronic device. In some embodiments, a graphic object may be displayed to describe the user input, such as the speech bubble of FIG. 12.
[0245] On the screen (1201), the electronic device can display an image through the display.
[0246] The electronic device may display an image selected by the user from a gallery or drawing application, a new image taken by the user using the electronic device's camera, an image temporarily saved to the clipboard, or a pre-generated AI image.
[0247] On the screen (1202), the electronic device may receive a first editing user input including an editing drawing input indicating the height of an image (hereinafter, referred to as a first editing drawing input) and an editing voice input describing an aspect ratio (hereinafter, referred to as a first editing voice input). The electronic device may display a first editing drawing corresponding to the first editing drawing input through a display at a location where the first editing drawing input was received. For example, the electronic device may receive a first editing voice input describing changing the aspect ratio to 3:2.
[0248] On the screen (1203), the electronic device may receive a second editing user input including an editing drawing input indicating the moon in the image (hereinafter, referred to as a second editing drawing input) and an editing voice input describing the size of the moon (hereinafter, referred to as a second editing voice input). The electronic device may display a second editing drawing corresponding to the second editing drawing input through the display at the location where the second editing drawing input was received. For example, the electronic device may receive a second editing voice input describing reducing the size of the moon.
[0249] On the screen (1204), the electronic device may receive a third editing user input including an editing drawing input for adding a person to an image (hereinafter, referred to as a third editing drawing input) and an editing voice input describing the person (hereinafter, referred to as a third editing voice input). The electronic device may display a third editing drawing corresponding to the third editing drawing input through the display at the location where the third editing drawing input was received. For example, the electronic device may receive a third editing voice input describing a person walking toward the moon.
[0250] The electronic device may determine first edit data associated with the visual information received by the first edit user input based on the first edit user input and first additional information about the first edit user input. Similarly, the electronic device may determine second edit data and third edit data.
[0251] On screen (1204), the electronic device may generate an edited AI image in response to a user input for generating the edited AI image. For example, the user may select a button (e.g., a 'regenerate' button) on screen (1204) to generate the edited AI image.
[0252] In one embodiment, when editing a pre-generated AI image, the electronic device may use a generative AI model to generate another AI image (i.e., an edited AI image) based on the aforementioned edit data (the first edit data to the third edit data) and at least a portion of the AI image.
[0253] The electronic device can determine at least one object of the AI image. For example, the electronic device can determine at least one object of the AI image based on metadata stored in association with the AI image. For example, if the metadata stored in association with the AI image includes information about at least one object of the AI image, the electronic device can verify the information. For example, the electronic device can determine at least one object of the AI image using an object recognition model. The electronic device can generate another AI image by converting, based on editing data, an object corresponding to at least one object included in an editing user input (a first editing user input to a third editing user input) among at least one object of the AI image. Descriptions that overlap with the method for editing the AI image of FIG. 10 are omitted.
[0254] In one embodiment, when editing an image such as a photograph or drawing (hereinafter, a non-AI image) that is not a pre-generated AI image, the electronic device may use a generative AI model to generate an edited AI image based on the aforementioned editing data (first editing data to third editing data) and at least a portion of the non-AI image.
[0255] The electronic device can determine at least one object in a non-AI image. For example, the electronic device can determine at least one object in the non-AI image using an object recognition model. The electronic device can generate an edited AI image by converting, based on editing data, at least one object corresponding to an editing user input (a first editing user input to a third editing user input), among the at least one object in the non-AI image. Any description that overlaps with the method for editing a stored image of FIG. 11 will be omitted.
[0256] On screen (1205), the electronic device can display a standby screen through the display while an edited AI image is generated using a generative AI model.
[0257] On the screen (1206), the electronic device can display an edited AI image corresponding to an image (e.g., an AI image, a non-AI image) via the display. The edited AI image can have an aspect ratio of the image on the screen (1201) changed to 3:2, include a moon smaller than the moon in the image on the screen (1201) at a location where a second editing drawing input is received, and include a person walking toward the moon at a location where a third editing input is received.
[0258] According to one embodiment, an electronic device (101) comprises: a display; at least one processor (120); And a memory storing instructions executable by at least one processor (120), wherein when the instructions are executed by the at least one processor (120), the electronic device (101) causes at least: receiving a first user input including at least one of a first drawing input or a first voice input, and determining first data based on the first user input and first additional information about the first user input, wherein the first additional information includes at least one of first visual information about a time at which the first user input was received or first location information about a location at which the first drawing input was received; receiving a second user input including at least one of a second drawing input or a second voice input, and determining second data based on the second user input and second additional information about the second user input, wherein the second additional information includes at least one of second visual information about a time at which the second user input was received or second location information about a location at which the second drawing input was received; generating an AI image based on at least a portion of the first data and the second data using a generative artificial intelligence (AI) model. You can create and display AI images through the display.
[0259] According to one embodiment, when the instructions are executed by at least one processor (120), the electronic device (101) may cause at least: to receive a first user input comprising a first drawing input and a first voice input, and determine first data associated with first visual information based on the first user input and first additional information about the first user input, and to receive a second user input comprising a second drawing input and a second voice input, and determine second data associated with the second visual information based on the second user input and second additional information about the second user input.
[0260] According to one embodiment, when the instructions are executed by at least one processor (120), the electronic device (101) may cause at least: a generative AI model to generate an AI image based on at least a portion of first data and second data by referencing first visual information and second visual information.
[0261] According to one embodiment, when the instructions are executed by at least one processor (120), the electronic device (101) may store metadata in association with an AI image, the metadata including at least a portion of the first data and the second data.
[0262] According to one embodiment, the metadata may include at least one of first drawing data for a first drawing input, first voice data for a first voice input, first text generated by converting the first voice data into speech-to-text (STT), first location information, or first visual information.
[0263] According to one embodiment, when the instructions are executed by at least one processor (120), the electronic device (101) may cause at least: a first drawing corresponding to a first drawing input to be displayed at a location where the first drawing input was received through a display.
[0264] According to one embodiment, when the instructions are executed by at least one processor (120), the electronic device (101) may cause at least: displaying a first graphical object corresponding to the first drawing by replacing the first drawing through the display based on the first location information, wherein the first graphical object includes a pre-stored shape, icon, or first AI object generated using a generative AI model corresponding to the first drawing.
[0265] According to one embodiment, when the instructions are executed by at least one processor (120), the electronic device (101) may cause at least: prior to receiving a second user input, to generate a first AI image including a first AI object corresponding to a first drawing based on at least a portion of the first data using a generative AI model, and to display the first AI image through a display.
[0266] According to one embodiment, when the instructions are executed by at least one processor (120), the electronic device (101) may cause at least: display a second drawing corresponding to a second drawing input via a display at the location where the second drawing input was received on the first AI image; generate a second AI image including a second AI object corresponding to the second drawing based on at least a portion of the second data using a generative AI model; and display the second AI image via the display.
[0267] According to one embodiment, when the instructions are executed by at least one processor (120), the electronic device (101) may generate a video including at least: a first frame corresponding to a first AI image and a second frame corresponding to a second AI image in chronological order.
[0268] According to one embodiment, when the instructions are executed by at least one processor (120), the electronic device (101) may cause at least: receive additional user input associated with a style of the video, and generate an AI video based on the additional user input and the video using a generative AI model.
[0269] According to one embodiment, when the instructions are executed by at least one processor (120), the electronic device (101) may at least: insert text or sound effects into a video based on metadata.
[0270] According to one embodiment, when the instructions are executed by at least one processor (120), the electronic device (101) may cause at least: receive an editing user input including at least one of an editing drawing input or an editing voice input for an AI image; determine editing data based on the editing user input and editing additional information for the editing user input, wherein the editing additional information includes at least one of visual information about a time at which the editing user input was received or location information about a location at which the editing drawing input was received; generate another AI image based on at least a portion of the editing data and the AI image using a generative AI model; and display the other AI image through a display.
[0271] According to one embodiment, when the instructions are executed by at least one processor (120), the electronic device (101) may cause at least: determine at least one object of an AI image based on metadata or an object recognition model, and use the generative AI model to generate another AI image such that at least a portion of the at least one object is changed into at least one other object corresponding to an editing user input.
[0272] According to one embodiment, when the instructions are executed by at least one processor (120), the electronic device (101) causes at least: display a stored image stored in a memory (130) through a display, receive a first user input including at least one of a first drawing input or a first voice input for the stored image, determine first data based on the first user input and first additional information for the first user input, wherein the first additional information includes at least one of first visual information for a time at which the first user input was received or first location information for a location at which the first drawing input was received, receive a second user input including at least one of a second drawing input or a second voice input for the stored image, determine second data based on the second user input and second additional information for the second user input, wherein the second additional information includes at least one of second visual information for a time at which the second user input was received or second location information for a location at which the second drawing input was received, and, using a generative AI model, generate the first data, the second data, and the stored image. An AI image can be generated based on at least a portion of an image, and the AI image can be displayed through a display.
[0273] According to one embodiment, a method performed by an electronic device (101), wherein the electronic device (101) comprises: a display; at least one processor (120); and a memory (130) storing instructions executable by the at least one processor (120), the method comprising: receiving a first user input including at least one of a first drawing input or a first voice input; determining first data based on the first user input and first additional information about the first user input, wherein the first additional information includes at least one of first visual information about a time at which the first user input was received or first location information about a location at which the first drawing input was received; receiving a second user input including at least one of a second drawing input or a second voice input; determining second data based on the second user input and second additional information about the second user input, wherein the second additional information includes at least one of second visual information about a time at which the second user input was received or second location information about a location at which the second drawing input was received; The method may include an operation of generating an AI image based on at least a portion of first data and second data using a generative AI (artificial intelligence) model; and an operation of displaying the AI image through a display.
[0274] According to one embodiment, in a method performed by an electronic device (101), a first user input may include a first drawing input and a first voice input, a second user input may include a second drawing input and a second voice input, an operation of determining first data may include an operation of determining first data associated with first visual information based on the first user input and first additional information about the first user input, and an operation of determining second data may include an operation of determining second data associated with second visual information based on the second user input and second additional information about the second user input.
[0275] According to one embodiment, in a method performed by an electronic device (101), the operation of generating an AI image may include an operation of generating an AI image based on at least a portion of first data and second data by referring to first visual information and second visual information using a generative AI model.
[0276] According to one embodiment, the method performed by the electronic device (101) may further include an operation of storing metadata including at least a portion of the first data and the second data in association with the AI image.
[0277] According to one embodiment, a computer-readable recording medium may store a program for executing a method performed by an electronic device (101) in combination with hardware.
[0278] The embodiments described above may be implemented using hardware components, software components, and / or a combination of hardware components and software components. For example, the devices, methods, and components described in the embodiments may be implemented using a general-purpose computer or a special-purpose computer, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and software applications running on the operating system. Furthermore, the processing device may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.
[0279] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may, independently or collectively, command the processing device. The software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave, for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on a computer-readable recording medium.
[0280] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., alone or in combination, and the program commands recorded on the medium may be those specially designed and configured for the embodiment or may be known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc.
[0281] The hardware device described above may be configured to operate as one or more software modules to perform the operations of the embodiment, and vice versa.
[0282] Although the embodiments have been described with limited drawings, those skilled in the art will appreciate that various technical modifications and variations can be applied based on the described embodiments. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.
[0283] Therefore, other implementations, embodiments and equivalents to the claims also fall within the scope of the claims described below.
Claims
1. In an electronic device (101), display; at least one processor (120); and It includes a memory (130) that stores instructions executable by at least one processor (120), When the above instructions are executed by the at least one processor (120), the electronic device (101) causes at least: Receiving a first user input comprising at least one of a first drawing input or a first speech input, Determining first data based on the first user input and first additional information about the first user input, wherein the first additional information includes at least one of first visual information about a time at which the first user input was received or first location information about a location at which the first drawing input was received; receiving a second user input comprising at least one of a second drawing input or a second speech input; Determining second data based on the second user input and second additional information about the second user input, wherein the second additional information includes at least one of second visual information about the time at which the second user input was received or second location information about the location at which the second drawing input was received; Using a generative AI (artificial intelligence) model, an AI image is generated based on at least a portion of the first data and the second data, Display the AI image through the above display To do, Electronic device (101).
2. In paragraph 1, When the above instructions are executed by the at least one processor (120), the electronic device (101) causes at least: Receiving the first user input including the first drawing input and the first voice input, Based on the first user input and the first additional information for the first user input, determining the first data associated with the first visual information; Receiving the second user input including the second drawing input and the second voice input, Based on the second user input and the second additional information for the second user input, the second data associated with the second visual information is determined, Using the generative AI model, the AI image is generated based on at least a portion of the first data and the second data by referring to the first visual information and the second visual information. To do, Electronic device (101).
3. In either of paragraphs 1 and 2, When the above instructions are executed by the at least one processor (120), the electronic device (101) causes at least: Storing metadata including at least a portion of the first data and the second data in association with the AI image, wherein the metadata includes at least one of first drawing data for the first drawing input, first voice data for the first voice input, first text generated by converting the first voice data into STT (speech-to-text), the first location information, or the first visual information. To do, Electronic device (101).
4. In any one of paragraphs 1 to 3, When the above instructions are executed by the at least one processor (120), the electronic device (101) causes at least: Displaying a first drawing corresponding to the first drawing input through the display at the location where the first drawing input was received To do, Electronic device (101).
5. In any one of paragraphs 1 to 4, When the above instructions are executed by the at least one processor (120), the electronic device (101) causes at least: Displaying a first graphic object corresponding to the first drawing by replacing the first drawing through the display based on the first location information, wherein the first graphic object includes a pre-stored shape, icon, or first AI object generated using the generative AI model corresponding to the first drawing. To do, Electronic device (101).
6. In any one of paragraphs 1 to 5, When the above instructions are executed by the at least one processor (120), the electronic device (101) causes at least: Before receiving the second user input, using the generative AI model, a first AI image including a first AI object corresponding to the first drawing is generated based on at least a portion of the first data, Displaying the first AI image through the above display To do, Electronic device (101).
7. In any one of paragraphs 1 to 6, When the above instructions are executed by the at least one processor (120), the electronic device (101) causes at least: Displaying a second drawing corresponding to the second drawing input through the display at the location where the second drawing input was received on the first AI image, Using the generative AI model, a second AI image including a second AI object corresponding to the second drawing is generated based on at least a portion of the second data, Displaying the second AI image through the above display To do, Electronic device (101).
8. In any one of paragraphs 1 to 7, When the above instructions are executed by the at least one processor (120), the electronic device (101) causes at least: Generate a video including a first frame corresponding to the first AI image and a second frame corresponding to the second AI image in chronological order To do, Electronic device (101).
9. In any one of paragraphs 1 to 8, When the above instructions are executed by the at least one processor (120), the electronic device (101) causes at least: Receive additional user input related to the style of the above video, Generate AI video based on the additional user input and the video using the generative AI model To do, Electronic device (101).
10. In any one of paragraphs 1 to 9, When the above instructions are executed by the at least one processor (120), the electronic device (101) causes at least: Insert text or sound effects into the video based on the above metadata To do, Electronic device (101).
11. In any one of paragraphs 1 to 10, When the above instructions are executed by the at least one processor (120), the electronic device (101) causes at least: Receiving an editing user input including at least one of an editing drawing input or an editing voice input for the AI image; Determining edit data based on the above editing user input and editing additional information for the editing user input, wherein the editing additional information includes at least one of visual information about the time at which the editing user input was received or positional information about the location at which the editing drawing input was received. Using the generative AI model, another AI image is generated based on the edited data and at least a portion of the AI image, Display the other AI image through the above display To do, Electronic device (101).
12. In any one of paragraphs 1 to 11, When the above instructions are executed by the at least one processor (120), the electronic device (101) causes at least: determining at least one object in said AI image based on metadata or an object recognition model; Using the generative AI model, generate another AI image such that at least a part of the at least one object is changed into at least one other object corresponding to the editing user input. To do, Electronic device (101).
13. In any one of paragraphs 1 to 12, When the above instructions are executed by the at least one processor (120), the electronic device (101) causes at least: Displaying the stored image stored in the memory (130) through the above display, Receiving the first user input comprising at least one of the first drawing input or the first voice input for the stored image, Determining first data based on the first user input and the first additional information for the first user input, wherein the first additional information includes at least one of first visual information for the time at which the first user input was received or first location information for the location at which the first drawing input was received; Receiving said second user input comprising at least one of said second drawing input or said second voice input for said stored image, Determining second data based on the second user input and the second additional information for the second user input, wherein the second additional information includes at least one of second visual information for the time at which the second user input was received or second location information for the location at which the second drawing input was received; Using the generative AI model, the AI image is generated based on at least a portion of the first data, the second data, and the stored image, Display the AI image through the above display To do, Electronic device (101).
14. In a method performed by an electronic device (101), The above electronic device (101) display; at least one processor (120); and It includes a memory (130) that stores instructions executable by at least one processor (120), An action of receiving a first user input comprising at least one of a first drawing input or a first speech input; An operation of determining first data based on the first user input and first additional information about the first user input, wherein the first additional information includes at least one of first visual information about a time at which the first user input was received or first location information about a location at which the first drawing input was received; An action of receiving a second user input comprising at least one of a second drawing input or a second speech input; An operation of determining second data based on the second user input and second additional information about the second user input, wherein the second additional information includes at least one of second visual information about a time at which the second user input was received or second location information about a location at which the second drawing input was received; An operation of generating an AI image based on at least a portion of the first data and the second data using a generative AI (artificial intelligence) model; and An action of displaying the AI image through the above display Including, method.
15. A computer-readable recording medium storing a program for executing the method of claim 14 in combination with hardware.
Citation Information
Patent Citations
Voice drawing system and drawing method thereof
CN111613218A
Voice drawing method and device and computer equipment
CN114995729A
Method for generating multiple drawing images based on artificial intelligence, electronic equipment and medium
CN116342739A
Picture drawing support device, method, and program
JP2014186372A
A method for providing information about the responsiveness to lipid-lowering therapy in individuals with familial hypercholesterolemia
KR102526094B1