Method for converting image and electronic device supporting same

WO2026205875A1PCT designated stage Publication Date: 2026-10-01SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/004348
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-04-17
Filing Date
2026-03-18
Publication Date
2026-10-01

Smart Images

  • Figure KR2026004348_01102026_PF_FP_ABST
    Figure KR2026004348_01102026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device disclosed in the present document may comprise: a display; a memory; and at least one processor including processing circuitry. The electronic device may receive a first image. The electronic device may determine a first prompt regarding conversion of the first image. The electronic device may receive text related to modification of the first image. The electronic device may determine a first criterion among a plurality of criteria included in the first prompt on the basis of the text. The electronic device may generate a second prompt by modifying the first criterion. The electronic device may generate a second image by inputting the first image and the second prompt to an image generation model.
Need to check novelty before this filing date? Find Prior Art

Description

Method for converting images and electronic device supporting the same

[0001] The embodiments disclosed in this document relate to a method for converting images and an electronic device supporting the same.

[0002] Electronic devices (or portable electronic devices, mobile devices) such as smartphones or tablet PCs can perform various functions. For example, electronic devices can run applications to perform various functions such as making calls, playing music or videos, taking notes, or drawing.

[0003] Recently, various functions utilizing artificial intelligence (AI) are being applied. Vision generative AI technology can generate high-quality images based on sketch input or text prompts. When a user draws on a display using a mouse, stylus pen, or a part of their body (e.g., hand) (sketch input), the electronic device can input the sketch image (image resulting from the sketch input) into a vision generative AI model to generate a new, high-quality image. Alternatively, the electronic device can receive the sketch image and generate an animation consisting of multiple frames.

[0004] For example, electronic devices can use Stable Diffusion to convert sketch images into high-quality images. Stable Diffusion is a deep learning model (text-to-image AI model) that generates high-quality images based on text prompts. Based on text-based descriptions, Stable Diffusion can generate detailed and realistic images. Stable Diffusion can convert simple sketches or doodles drawn by users into high-quality digital images. Additionally, Stable Diffusion can display a user interface for performing various image processing tasks, such as image editing, inpainting, and outpainting.

[0005] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.

[0006] An electronic device according to one embodiment may include at least one processor comprising a display, memory, and processing circuitry. The memory may store instructions that cause the electronic device to perform the following operations when executed individually or collectively by the at least one processor. The electronic device may receive a first image. The electronic device may determine a first prompt regarding the transformation of the first image. The electronic device may receive text regarding the modification of the first image. The electronic device may determine a first criterion among a plurality of criteria included in the first prompt based on the text. The electronic device may generate a second prompt by modifying the first criterion. The electronic device may generate a second image by inputting the first image and the second prompt into an image generation model.

[0007] An image conversion method according to one embodiment may be performed in an electronic device. The image conversion method may include the operation of receiving a first image, the operation of determining a first prompt regarding the conversion of the first image, the operation of receiving text related to the modification of the first image, the operation of determining a first criterion among a plurality of criteria included in the first prompt based on the text, the operation of modifying the first criterion to generate a second prompt, and the operation of inputting the first image and the second prompt into an image generation model to generate a second image.

[0008] A computer-readable storage medium according to one embodiment may store instructions executable by a processor. When the instructions are executed, the processor of an electronic device may perform the following operations. The processor may receive a first image. The processor may determine a first prompt regarding the transformation of the first image. The processor may receive text related to the modification of the first image. Based on the text, the processor may determine a first criterion among a plurality of criteria included in the first prompt. The processor may modify the first criterion to generate a second prompt. The processor may input the first image and the second prompt into an image generation model to generate a second image.

[0009] FIG. 1 is a block diagram of an electronic device in a network environment according to various embodiments.

[0010] FIG. 2 shows a configuration diagram of an electronic device according to one embodiment.

[0011] FIG. 3 is a configuration diagram of an image conversion unit according to one embodiment.

[0012] FIG. 4 is a flowchart of an image conversion method according to one embodiment.

[0013] FIG. 5 is a flowchart of an image conversion method that displays an intermediate image according to one embodiment.

[0014] FIG. 6a is an example of an LLM prompt for determining a first criterion according to one embodiment.

[0015] FIG. 6b shows the selection of a first criterion in a first prompt according to one embodiment.

[0016] FIG. 7a is an example of an LLM prompt for determining a second criterion according to one embodiment.

[0017] FIG. 7b is an example showing the determination of the second standard according to one embodiment.

[0018] FIG. 7c is an example showing the generation of a second prompt according to one embodiment.

[0019] FIG. 8 shows a user interface related to image conversion according to one embodiment.

[0020] FIG. 9 shows a user interface related to image style conversion according to one embodiment.

[0021] FIG. 10 shows the attributes of a plurality of standards according to one embodiment.

[0022] FIG. 11 is an image conversion method using a plurality of paths according to one embodiment.

[0023] In relation to the description of the drawings, the same or similar reference numerals may be used for identical or similar components.

[0024] Hereinafter, various embodiments of this document are described with reference to the accompanying drawings. However, this is not intended to limit the technology described in this document to specific embodiments and should be understood to include various modifications, equivalents, and / or alternatives to the embodiments of this document. In relation to the description of the drawings, similar reference numerals may be used for similar components.

[0025] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (or display) (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).

[0026] The processor (120) can control at least one other component (e.g., hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., program (140)), and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., sensor module (176) or communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., central processing unit or application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., graphics processing unit, neural processing unit (NPU), image signal processor, sensor hub processor, or communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.

[0027] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence is performed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.

[0028] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).

[0029] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0030] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0031] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.

[0032] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.

[0033] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).

[0034] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0035] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0036] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0037] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that the user can perceive through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.

[0038] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0039] The power management module (188) can manage the power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).

[0040] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0041] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).

[0042] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) can support a Peak data rate (e.g., 20 Gbps or more) for realizing eMBB, loss coverage (e.g., 164 dB or less) for realizing mMTC, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for realizing URLLC.

[0043] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).

[0044] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.

[0045] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.

[0046] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In one embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0047]

[0048] FIG. 2 shows a configuration diagram of an electronic device according to one embodiment. FIG. 2 illustrates configurations related to image conversion classified according to function as examples, but is not limited thereto.

[0049] Referring to FIG. 2, an electronic device (201) (e.g., the electronic device (101) of FIG. 1) can convert a sketch image (hereinafter, first image) based on a sketch input into a high-quality image (hereinafter, second image) using a vision generative AI model.

[0050] According to one embodiment, the electronic device (201) may include a control unit (210), a prompt determination unit (220), a prompt generation unit (230), an image conversion unit (240), and a prompt database (250). The operation of the control unit (210), the prompt determination unit (220), the prompt generation unit (230), and the image conversion unit (240) may be the operation of a processor included in the electronic device (201) (e.g., the processor (120) of FIG. 1). The prompt database (250) may be part of the memory of the electronic device (201) (e.g., the memory (130) of FIG. 1).

[0051] According to one embodiment, the control unit (210) can control the prompt determination unit (220), the prompt generation unit (230), and the image conversion unit (240) in relation to image conversion. The control unit (210) can control the transmission of signals or data between the prompt determination unit (220), the prompt generation unit (230), and the image conversion unit (240). The control unit (210) can process user input related to image conversion and transmit it to the prompt determination unit (220), the prompt generation unit (230), or the image conversion unit (240).

[0052] According to one embodiment, the control unit (210) receives a first image of a user and can process tasks that occur during the process of determining and generating a prompt for generating a second image and the image conversion process. The control unit (210) can control the generation of a new prompt based on user input and can perform control to generate a second image based on the input sketch image and the modified prompt.

[0053] According to one embodiment, the prompt determination unit (220) can determine a prompt (hereinafter, a first prompt) corresponding to a style (conversion style) selected by a default setting or user input. The prompt determination unit (220) can determine a first prompt corresponding to a style selected by user input from among a plurality of prompts stored in the prompt database (250). The determined first prompt can be used as an input to the image conversion unit (240). Alternatively, the first prompt can be changed to a second prompt in which other elements or criteria are additionally reflected in the prompt generation unit (230).

[0054] According to one embodiment, the first prompt may include a plurality of criteria related to image conversion (or a plurality of image conversion criteria, a plurality of rules). The plurality of criteria may be selected by considering data learned in the image conversion unit (240) (e.g., stable diffusion).

[0055] According to one embodiment, the prompt determining unit (220) can determine a first prompt corresponding to a style of image transformation selected by the user (e.g., watercolor, illustration, pop art). When "watercolor" is determined as the style for transforming the first image by user input, the prompt determining unit (220) can determine the first prompt by loading a prompt corresponding to "watercolor" from among a plurality of prompts stored in the prompt database (250). For example, the first prompt may include natural language in text form such as "transform this sketch into a vibrant watercolor painting" or "soft pastel hues".

[0056] The prompt generation unit (230) can generate a second prompt by changing the first prompt. The prompt generation unit (230) can determine the second prompt by changing the first prompt based on text separately entered by the user regarding image conversion (or text reflecting the user's request regarding the direction of image modification) (hereinafter, modification text).

[0057] According to one embodiment, the modified text may be text-based data entered by a user in connection with the generation of a second image. The modified text may reflect the user's intention regarding image conversion. The modified text may be entered through a field of a user interface displayed on a display. For example, the modified text may be "rougher."

[0058] According to one embodiment, the second prompt may be a prompt in which at least some of the multiple criteria (or multiple rules) included in the first prompt are replaced or deleted, or a separate criterion (or rule) is added. The second prompt may be transmitted to an image conversion unit (240) and used to generate a second image. The second prompt may include a limited number of words, phrases, or sentences that an AI model included in the image conversion unit (240) can understand.

[0059] According to one embodiment, the prompt generation unit (230) may use a separate language AI model to generate a second prompt. For example, a language model such as a large language model (LLM) may be used (see FIGS. 6a to 7c).

[0060] The image conversion unit (or sketch-to-image conversion module) (240) can generate a high-quality image by inputting a first image and a prompt to a vision generative AI model. For example, the image conversion unit (240) may include a lightweight vision transformer model such as MobileViT suitable for mobile environments. Alternatively, the image conversion unit (240) may be an AI model such as stable diffusion, which is a sketch-to-image conversion model, or a model such as GAN, MobileNet, or StyleGAN2-ADA. The image conversion unit (240) may apply quantization techniques and model compression techniques for real-time processing of the image. Additional information regarding the image conversion unit (240) may be provided through FIG. 3.

[0061] The prompt database (or conversion prompt database) (250) can store various prompts related to image conversion. The prompt database (250) can store prompts for each of the various image conversion styles (e.g., watercolor, illustration, pop art) applicable to the electronic device (201). Alternatively, the prompt database (250) can store a second prompt newly generated by reflecting the modified text.

[0062] According to one embodiment, each prompt stored in the prompt database (250) may include a positive criterion, a negative criterion, or a positive / negative criterion. A positive criterion may represent a criterion that must be expressed or reflected in relation to image transformation, and a negative criterion may represent a criterion that must not be expressed or reflected in relation to image transformation. A positive / negative criterion may be a criterion that is used positively or negatively according to specified conditions.

[0063] According to one embodiment, the criteria for image conversion included in the prompt may consist of a limited number of words, phrases, and sentences. The limited number of words, phrases, and sentences may be selected based on training data used when training the image conversion unit (240) or a part of the image conversion unit (240) (e.g., vision / text encoder).

[0064] According to one embodiment, the prompts stored in the prompt database (250) may have various attributes. Additional information regarding the attributes of the prompts stored in the prompt database (250) may be provided through FIG. 10.

[0065]

[0066] FIG. 3 is a configuration diagram of an image conversion unit according to one embodiment. FIG. 3 is exemplary and is not limited thereto.

[0067] Referring to FIG. 3, the image conversion unit (or sketch-to-image conversion module) (240) can generate a high-quality image using an image from a sketch input and a text-based prompt (a prompt reflected for image conversion). The image conversion unit (240) can generate images of various styles and formats. The image conversion unit (240) can be implemented with a plurality of artificial neural networks. For example, the image conversion unit (240) may be a Stable Diffusion operating in a mobile environment.

[0068] According to one embodiment, the image conversion unit (240) may include a vision encoder (310), a text encoder (320), a control unit (330), an image conversion AI model (340), and an automatic encoder / decoder (350).

[0069] The vision encoder (310) can extract a content feature map from a first image input by a user. The content features may include information such as the shape of objects included in the first image, patterns, and arrangement of objects. The vision encoder (310) can be implemented as CLIP (Contrastive Language-Image Pre-training), BLIP (Bootstrapping Language-Image Pre-training), or a CNN model (e.g., VGGNet-16).

[0070] The text encoder (320) can extract a style feature map by encoding a text-type conversion prompt (a first prompt or a second prompt). The text encoder (320) can be implemented with CLIP, BLIP, or ViT-L / 14.

[0071] The control unit (330) can receive a content feature map from the vision encoder (310) and a style feature map from the text encoder (320). The control unit (330) can control a stable diffusion model using the content feature map and the style feature map. The control unit (330) can generate a combined feature map by combining the content feature map and the style feature map. The control unit (330) can use a control AI model such as AdaIN (Adaptive Instance Normalization) or ControlNet.

[0072] The image transformation AI model (Stable diffusion model, or UNet) (340) can receive a combined feature map from the control unit (330). The image transformation AI model (340) can gradually remove noise using the combined feature map. Through this, the image transformation AI model (340) can generate a latent representation of a new image.

[0073] According to one embodiment, the image conversion AI model (340) can control noise prediction at each denoising step through classifier-free guidance (CFG) or denoising strength.

[0074] For example, the image transformation AI model (340) can predict noise at each denoising step through classifier-free guidance (CFG). The image transformation AI model (340) can interpolate conditional and unconditional predictions based on content feature maps and style feature maps, and can control the value used for interpolation with CFG. If CFG is low, more creative and unpredictable images can be generated. Conversely, if CFG is high, images that strictly follow prompts can be generated.

[0075] As another example, the image conversion AI model (340) can predict noise at each denoising step through denoising strength. Denoising strength is a variable that controls the noise removal strength at each denoising step, and if the denoising strength is low, an image close to the shape of a sketch is generated, and if the denoising strength is high, a new image different from the original can be generated.

[0076] The autoencoder / decoder (or fecal autoencoder / decoder) (350) can compress the image into a low-dimensional latent space. The autoencoder / decoder can restore details of the latent representation generated by the image transformation AI model (340) and generate a final transformed image.

[0077]

[0078] FIG. 4 is a flowchart of an image conversion method according to one embodiment.

[0079] Referring to FIGS. 1 and FIGS. 4, in operation 410, the processor (120) may receive a first image. The first image may be an image generated by a sketch input in which a user draws on a display (or display module) (160) using a mouse, an electronic pen (stylus pen), or a part of the body (e.g., hand). Alternatively, the first image may be a pre-stored image of the user (e.g., a photo taken through a camera, a photo selected in a gallery app, a drawing sketched in a drawing app), or an image processed from a sketch image.

[0080] In operation 420, the processor (120) may determine a first prompt corresponding to the transformation style of the first image. The first prompt may include a plurality of criteria (or a plurality of rules) related to the transformation of the image. The first prompt may be processed by an AI model related to the transformation of the image (e.g., the image transformation unit (240) of FIG. 3).

[0081] According to one embodiment, the processor (120) may determine the first prompt by a preset default setting without separate user input, or determine the first prompt by a conversion style (e.g., watercolor, illustration, pop-art) selected by user input.

[0082] According to one embodiment, the processor (120) may display touch buttons (e.g., watercolor, illustration, pop-art) for a conversion style selectable by user input while a sketch image is displayed. A first prompt corresponding to the conversion style selected by user input may be loaded from a database (e.g., the prompt database (250) of FIG. 2).

[0083] According to one embodiment, the first prompt may include a plurality of criteria related to image conversion. For example, in the case of a watercolor style, it may include positive criteria and negative criteria as follows.

[0084]

[0085] [Prompt 1 - positives]

[0086] 1. transform this sketch into a vibrant watercolor painting,

[0087] 2. soft pastel hues,

[0088] 3. flowing brush strokes,

[0089] 4. wet-on-wet technique,

[0090] 5. delicate color blending,

[0091] 6. white paper texture visible,

[0092] 7. Slight color bleeding at edges

[0093]

[0094] [Prompt 1 - negatives]

[0095] 1. plastic-like texture,

[0096] 2. neon colors,

[0097] 3. overly saturated colors,

[0098] 4. geometric patterns,

[0099] 5. pixel art,

[0100] 6. clean edges,

[0101] 7. precise outlines,

[0102] 8. High contrast

[0103]

[0104] Here, positive criteria may represent criteria that must be expressed or reflected in relation to image transformation, and negative criteria may represent criteria that must not be expressed or reflected in relation to image transformation.

[0105] According to one embodiment, the processor (120) can input a first image and a first prompt into an image conversion model (e.g., an image generation unit of FIG. 2) to generate an intermediate image and display it on a display (160). The intermediate image can be displayed in correspondence with a style selected by the user before the modified text is reflected. After checking the intermediate image, the user can convert the image more precisely through the modified text.

[0106] In operation 430, the processor (120) may receive modification text related to the transformation of the first image. The processor (120) may receive modification text (e.g., text reflecting a user's request regarding the direction of modification of the image) via keyboard input or voice input. The modification text may include natural language and may include specified words, phrases, or sentences. For example, the modification text may be "a bit rougher."

[0107] In operation 440, the processor (120) can determine a first criterion (or first criterion group, first rule group) that needs to be modified among a plurality of criteria included in the first prompt based on the modification text. The first criterion may be a criterion that needs to be modified to reflect an effect corresponding to the modification text among the plurality of criteria included in the first prompt.

[0108] According to one embodiment, the processor (120) may measure the text relevance between the modified text and a plurality of criteria and determine a first criterion based on the measured text relevance. Each of the plurality of criteria may be written in text or a specified format and stored in a database. Each of the plurality of criteria may be part of a pool of criteria stored in the database. Each criterion included in the pool of criteria may be matched with one or more specified tags (e.g., #brush, #soft, #border, #color, #vibrant). The processor (120) may select a tag corresponding to the modified text based on text similarity. The processor (120) may select a criterion corresponding to the selected tag as the first criterion.

[0109] According to one embodiment, the processor (120) may use a text encoder such as BERT (Bidirectional Encoder Representations from Transformers) to measure text relevance. For example, the processor (120) may convert the modified text and a plurality of criteria into embeddings through BERT and calculate the cosine similarity between the embeddings. The processor (120) may determine the conversion criterion with high cosine similarity as the first criterion.

[0110] According to one embodiment, the processor (120) may determine a conversion criterion with a cosine similarity close to -1 as a first criterion to find a conversion criterion that is related to the modified text but is semantically opposite.

[0111] According to one embodiment, the processor (120) may determine a conversion criterion with a cosine similarity close to 0 as a first criterion to find a conversion criterion that needs to be deleted because it has no relation to the modified text.

[0112] According to one embodiment, the processor (120) may determine a first criterion using a separate LLM. Additional information may be provided through FIGS. 6a and FIGS. 6b.

[0113] In operation 450, the processor (120) can determine a second prompt by changing the first criterion to a second criterion (or second criterion group, second rule group). The processor (120) can replace or delete the first criterion with another criterion by reflecting the modified text. Or the processor (120) can add a separate criterion by reflecting the modified text.

[0114] According to one embodiment, the processor (120) may determine a second criterion using a pre-stored text pool. The text pool may include words, phrases, and sentences. The text pool may be determined based on training data used when training an image conversion model (e.g., the image generation unit of FIG. 2).

[0115] According to one embodiment, the processor (120) may determine a second prompt using a separate reference-based algorithm, an AI model, or an LLM. Additional information regarding the LLM prompt when an LLM is used may be provided through FIG. 7a and FIG. 7b.

[0116] In operation 460, the processor (120) can generate a second image by inputting the first image and the second prompt into an image conversion model (e.g., the image generation unit of FIG. 2). The second image may be a high-quality image generated by reflecting the modified text (reflecting the user's modification intention). Additional information regarding the generation of the second image may be provided through the drawings below.

[0117]

[0118] FIG. 5 is a flowchart of an image conversion method that displays an intermediate image according to one embodiment.

[0119] Referring to FIG. 1 and FIG. 5, operations 510 and 520 in FIG. 5, respectively, may be the same or similar to operations 410 and 420 in FIG. 4.

[0120] In operation 525, the processor (120) inputs the first image and the first prompt into an image conversion model (e.g., the image generation unit of FIG. 2) to generate an intermediate image and display it on the display (160). The intermediate image may be displayed in correspondence with the style selected by the user before the modified text is reflected. After checking the intermediate image, the user can convert the image more precisely through the modified text.

[0121] Each of operations 530 to 560 in FIG. 5 may be the same or similar to operations 430 to 460 in FIG. 4.

[0122] In operation 565, the processor (120) may store the second prompt used to generate the second image in a database (e.g., the prompt database (250) of FIG. 2). The processor (120) may store the second prompt automatically or upon separate user input (see FIG. 8 and FIG. 9).

[0123]

[0124] FIG. 6a is an example of an LLM prompt for determining a first criterion according to one embodiment. FIG. 6a is exemplary and is not limited thereto.

[0125] Referring to FIG. 6a, the LLM prompt (601) may be text containing a description in a form similar to human speech (e.g., natural language). The LLM prompt (601) may include criteria, output conditions, and output requests necessary for image transformation and determination of a first criterion. The LLM prompt (601) may be input into and processed by an artificial intelligence model (LLM (large language model)).

[0126] According to one embodiment, the LLM prompt (601) may include a first part (610), a second part (620), and a third part (630).

[0127] The first part (610) may include text describing output conditions and output requests. The first part (610) may describe matters that the artificial intelligence model must verify or review in determining the first criterion. The content of the first part (610) may vary depending on whether the multiple criteria included in the second part (620) are positive criteria or negative criteria.

[0128] The second part (620) may include multiple criteria related to image conversion. If the first prompt is determined by a conversion style (e.g., watercolor, illustration, pop-art) selected by user input, the second part (620) may include multiple criteria included in the first prompt. Additionally, the second part (620) may include information indicating whether the multiple criteria are positive criteria or negative criteria.

[0129] The third part (630) may include modified text. The modified text may include natural language and may include specified words, phrases, or sentences. For example, the modified text may be "a bit rougher." The modified text may be added by user input (keyboard input, voice input).

[0130]

[0131] FIG. 6b illustrates the selection of a first criterion in a first prompt according to one embodiment. FIG. 6b is exemplary and is not limited thereto.

[0132] Referring to FIG. 6b, the first prompt (650) may include multiple criteria regarding style. For example, in the case of a watercolor style, it may include first to seventh positive criteria (651) and first to eighth negative criteria (652).

[0133] When the LLM prompt (601) of FIG. 6a is input into the LLM, a first criterion (660) corresponding to the correction text may be output. The first criterion (660) may include a positive correction criterion (positive first criterion) (661) and a negative correction criterion (negative first criterion) (662).

[0134] For example, in a watercolor style, the correction text "a bit rougher" can be entered. The positive correction criteria (positive first criterion) (661) can be determined as 3. flowing brush strokes, 5. delicate color blending, and 7. slight color bleeding at edges. The negative correction criteria (negative first criterion) (662) can be determined as 3. overly saturated colors and 8. high contrast.

[0135] Additional information regarding the process of generating a second prompt based on a selected first criterion may be provided through the following FIGS. 7a and 7b.

[0136]

[0137] FIG. 7a is an example of an LLM prompt for determining a second criterion according to one embodiment.

[0138] Referring to FIG. 7a, the LLM prompt (701) may be text containing a description in a form similar to human speech (e.g., natural language). The LLM prompt (701) may include criteria, output conditions, and output requests necessary for image transformation and determination of a second criterion. The LLM prompt (701) may be input into and processed by an artificial intelligence model (LLM (large language model)).

[0139] According to one embodiment, the LLM prompt (701) may include a first part (710), a second part (720), and a third part (730).

[0140] The first part (710) may include text describing output conditions and output requests. The first part (710) may describe matters that the artificial intelligence model must verify or review in determining the second criterion.

[0141] The second part (720) may include a first criterion that requires correction among a plurality of criteria related to image transformation. The second part (720) may include positive correction criteria (positive first criterion) (661) (e.g., 3. flowing brush strokes, 5. delicate color blending, 7. slight color bleeding at edges) determined through FIGS. 6A and 6B, and negative correction criteria (negative first criterion) (662) (e.g., 3. overly saturated colors, 8. high contrast).

[0142] The third part (730) may include modified text. The modified text may include natural language and may include specified words, phrases, or sentences. For example, the modified text may be "a bit rougher." The modified text may be added by user input (keyboard input, voice input).

[0143]

[0144] FIG. 7b is an example illustrating the determination of a second criterion according to one embodiment. FIG. 7b is exemplary and is not limited thereto.

[0145] Referring to FIG. 7b, the first criterion (660) may include criteria that require correction by reflecting the correction text. The first criterion (660) may include a positive correction criterion (positive first criterion) (661) and a negative correction criterion (negative first criterion) (662).

[0146] For example, in a watercolor style, the correction text "a bit rougher" can be entered. The positive correction criteria (positive first criterion) (661) can be determined as 3. flowing brush strokes, 5. delicate color blending, and 7. slight color bleeding at edges. The negative correction criteria (negative first criterion) (662) can be determined as 3. overly saturated colors and 8. high contrast.

[0147] The second criterion (750) may be determined by changing each item of the first criterion (660) to reflect the modified text. The second criterion (750) may include a positive modification criterion (positive second criterion) (751) and a negative modification criterion (negative second criterion) (752). The second criterion (750) may be created by replacing or deleting some of the criteria included in the first criterion (660) with other criteria. Criteria not included in the first criterion (660) may be added to the second criterion (750).

[0148] For example, in a watercolor style, the correction text "a bit rougher" may be entered. The positive correction criteria (positive second criteria) (751) may be determined as 3. dry brush technique, 5. intentional hard edges, 7. subtle color bleeding at edges. The negative correction criteria (negative second criteria) (752) may be determined as 3. vibrant colors, 8. [deleted], 9. flowing brush strokes.

[0149]

[0150] FIG. 7c is an example showing the generation of a second prompt according to one embodiment.

[0151] Referring to FIG. 1 and FIG. 7c, the first prompt (760) may include multiple criteria in relation to style. For example, in the case of a watercolor style, it may include first to seventh positive criteria (761) and first to eighth negative criteria (762).

[0152] The processor (120) can generate a second prompt (770) by reflecting the modified text in the first prompt (760).

[0153] For example, as shown in FIGS. 6a and 6b, a first criterion can be selected first, and after changing from the first criterion to the second criterion as shown in FIGS. 7a and 7b, this can be reflected back into the first prompt (760) to generate the second prompt (770).

[0154] As another example, the processor (120) can input a first prompt (760) and a modified text into an artificial intelligence model to generate a second prompt (770) that reflects the modified text.

[0155] The second prompt (770) may be created by replacing or deleting some of the criteria in the first prompt (760) with other criteria. Alternatively, the second prompt (770) may be created by adding new criteria that are not in the first prompt (760).

[0156] For example, in the case of the watercolor style, when the modified text "rougher" is entered, the 3rd positive criterion, the 5th positive criterion, and the 7th positive criterion may be replaced with other criteria. The 8th negative criterion, which does not correspond to the modified text "rougher", may be deleted.

[0157] According to one embodiment, the second prompt (770) may be generated by adding a separate criterion not included in the first prompt (760). A new ninth negative criterion corresponding to the modified text "rougher" may be added to the second prompt (770).

[0158] According to one embodiment, the processor (120) may generate a single LLM prompt (hereinafter referred to as an integrated LLM prompt) to perform an operation to determine a first criterion and an operation to determine a second criterion at once. The integrated LLM prompt may be a combination of the LLM prompt (601) of FIG. 6a and the LLM prompt (701) of FIG. 7a, and may be used to determine the second criterion through a single LLM inference process (a process of inputting and processing into a single LLM).

[0159]

[0160] FIG. 8 shows a user interface related to image conversion according to one embodiment.

[0161] Referring to FIGS. 1 and FIGS. 8, in the first user interface (801), the processor (120) can acquire a first image (810). The first image (810) can be generated by the user's sketch input. For example, the processor (120) can run an application that allows the user to draw. The user's input can be received in a designated drawing area (805), and the first image (810) can be generated. Alternatively, the first image (810) may be an image previously stored by the user's sketch input, or an image processed from an existing sketch image.

[0162] According to one embodiment, the processor (120) may display an image generation button (or image conversion button) (815) for generating a second image. Additionally, the processor (120) may display a field (818) for receiving modification text that may be reflected during the process of generating the second image.

[0163] In the second user interface (802), when user input (e.g., touch input) occurs on the image creation button (815), the processor (120) may display buttons (or options) regarding the transformation style (820) of the first image (810). The transformation style (820) may indicate a classification of the transformation method or transformation effect of the first image (810). For example, the transformation style (820) may include watercolor, illustration, and pop-art.

[0164] According to one embodiment, the conversion style (820) may be automatically displayed during the process of inputting the first image (810). For example, the processor (120) may detect features of the first image (810) and display a recommended conversion style (820).

[0165] In the third user interface (803), when one of the conversion styles (820) is selected by user input, the processor (120) may display an intermediate image (or temporary image) (830) that has changed the first image to the selected conversion style (e.g., watercolor). The intermediate image (830) may be generated by inputting the first image (810) and the first prompt corresponding to the conversion style (e.g., watercolor) selected by user input into an image conversion model (e.g., the image conversion unit (240) of FIG. 3).

[0166] According to one embodiment, the processor (120) may display a field (818) for inputting modification text. The field (818) may input modification text regarding criteria, style, format, or effect determined by the user related to the modification of the intermediate image (830). For example, the modification text may be text such as "rougher." The modification text may be received via keyboard or voice input.

[0167] In the fourth user interface (804), the processor (120) can display a second image (840) that reflects the modified text in the intermediate image (830). The processor (120) can generate a second prompt by modifying the criteria included in the first prompt based on the modified text. The second prompt and the first image (810) can be input into an image conversion unit (e.g., the image conversion unit of FIG. 2 and FIG. 3) to generate the second image (840).

[0168] According to one embodiment, the processor (120) may display a button (845) for saving the second prompt used to generate the second image (840). When the second prompt is saved, the user can easily apply the same style as the second image (840) to another sketch image (see FIG. 9).

[0169]

[0170] FIG. 9 illustrates a user interface related to image style conversion according to one embodiment. FIG. 9 is exemplary and is not limited thereto.

[0171] Referring to FIG. 9, in the user interface (901), the processor (120) can acquire a first image (910). The first image (910) can be generated by the user's sketch input. The user's input can be received in a designated drawing area (905), and the first image (910) can be generated.

[0172] According to one embodiment, the processor (120) may display an image generation button (or image conversion button) (915) for generating a second image. Additionally, the processor (120) may display a field (918) for receiving modification text that may be reflected during the process of generating the second image.

[0173] When user input (e.g., touch input) occurs on the image creation button (915), the processor (120) may display buttons (920a, 920b, 925) regarding the transformation style (930) of the first image (910). The transformation style (930) may indicate a classification of the transformation method or transformation effect of the first image (910). For example, the transformation style (920) may include watercolor, illustration, and pop-art.

[0174] The processor (120) may display a conversion style (920) reflecting a history of previous use. For example, if there is a history of a second image being generated with "rougher" reflected as modification text in watercolor, a "watercolor and rougher" button (920b) may be displayed separately from the "watercolor" button (920a). The "watercolor and rougher" button (920b) may be in a form where "watercolor" and "rougher" are partially overlapping.

[0175] If the user selects the "watercolor and coarser" button (920b), the previously used second prompt and the first image (910) currently being displayed are input into the image conversion model to generate a new second image.

[0176]

[0177] FIG. 10 shows the attributes of a plurality of standards according to one embodiment.

[0178] Referring to FIG. 2 and FIG. 10, the prompt database (250) may store a list of attributes of a plurality of criteria (hereinafter, attribute list) (1001). The attribute list (1001) may include a prompt element (1010), an assigned prompt (or style name) (1020), a positive / negative attribute indicator (1030), and an adjustable attribute indicator (1040).

[0179] The prompt element (or transformation criterion) (1010) may be the name or identifier of each of the criteria included in the first prompt or the second prompt. The assigned prompt (1020) may represent the style classification or style category containing each criterion.

[0180] The positive / negative attribute indicator (P / N) (1030) indicates whether each criterion is a positive criterion or a negative criterion.

[0181] The adjustable attribute indicator (P / A) (1040) indicates whether each criterion is allowed to be changed by the modification text or not. If the conversion criterion is not editable, it may not be modified (changed or removed) during the process of generating the second prompt.

[0182] According to one embodiment, the prompt element (1010) may not have an assigned prompt (1020) specified, or both positive and negative may be possible in the positive / negative attribute indicator (P / N) (1030). The attribute list (1001) may include information on whether grouping among the criteria (gathering or grouping transformation criteria to generate a single prompt) is possible or not.

[0183] According to one embodiment, each prompt element (1010) may include one or more designated tags (e.g., #brush, #soft, #border, #color, #vibrant, etc.). Tags may be used to group with other criteria, or to change or replace criteria. For example, it may be configured so that only prompt elements (1010) belonging to the tag "#soft" are allowed to group, or only prompt elements (1010) belonging to the tag "#brush" are allowed to replace.

[0184]

[0185] FIG. 11 is an image conversion method using a plurality of paths according to one embodiment. The operation of each component of FIG. 11 may be the operation of the processor 120 of FIG. 1.

[0186] Referring to FIG. 11, when generating a second conversion prompt by modifying the first conversion prompt based on the modified text, the method determining unit (1110) can determine the level of image conversion related to the modified text. If the degree of image conversion by the modified text is at a relatively small first level (e.g., grayscale processing, blurring, sharpening, brightness / contrast / saturation change), the method determining unit (1110) does not generate a second prompt, inputs the first image and the first prompt into the first AI model (1120), and then generates a second image using a post-processing filter (1125).

[0187] When the modified text is at a second level with a relatively large degree of image conversion, the method determining unit (1110) can convert the first prompt stored in the prompt database (1135) into a second prompt reflecting the modified text, and input the second prompt and the first image into the second AI model (1130) to generate a second image.

[0188] In FIG. 11, the first AI model (1120) and the second AI model (1130) are shown as being distinct from each other, but this is not limited thereto. The first AI model (1120) and the second AI model (1130) may be the same single image conversion model.

[0189]

[0190] Stable Diffusion requires significant resources (computational load, memory) and allows for limited customization. Additionally, when using the computationally intensive sketch-to-image conversion module, users cannot customize the conversion direction because it utilizes pre-defined prompts or models pre-trained with specific image data for model optimization.

[0191] The present disclosure may provide an electronic device that converts a sketch image into a high-quality image in a mobile environment based on modified text that reflects the user's intent.

[0192] The technical problems to be solved in this disclosure are not limited to those mentioned above, and other technical problems not mentioned will be clearly understood by those skilled in the art to which this disclosure pertains.

[0193] An electronic device according to embodiments disclosed in this document can acquire a transformed image according to the user's intent using modified text. The electronic device can generate a high-quality image suitable for a mobile device without compromising the attributes of the transformation criteria included in the sketch-to-image transformation model.

[0194] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs.

[0195]

[0196] An electronic device (101;201) according to one embodiment may include a display (160), a memory (130), and at least one processor (120) including processing circuitry. The memory (130) may store instructions that cause the electronic device (101;201) to perform the following operations when executed individually or collectively by the at least one processor (120). The electronic device (101;201) may receive a first image. The electronic device (101;201) may determine a first prompt regarding the transformation of the first image. The electronic device (101;201) may receive text regarding the modification of the first image. The electronic device (101;201) may determine a first criterion among a plurality of criteria included in the first prompt based on the text. The electronic device (101;201) may generate a second prompt by modifying the first criterion. The electronic device (101;201) can generate a second image by inputting the first image and the second prompt into an image generation model.

[0197] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (120), the electronic device (101;201) may generate the first image based on sketch inputs generated through the display (160).

[0198] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (120), the electronic device (101;201) may display a plurality of options for determining a style related to the conversion of the first image on the display (160), and when one of the plurality of options is selected by user input, the first prompt corresponding to the selected option may be determined.

[0199] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (120), the electronic device (101;201) may load the first prompt corresponding to the selected option among a plurality of prompts stored in a database related to image conversion.

[0200] According to one embodiment, the plurality of criteria may have one of a positive attribute, a negative attribute, or a positive / negative attribute.

[0201] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (120), the electronic device (101;201) may determine a second criterion by using a first large language model (LLM) to replace or delete the first criterion or add a separate criterion.

[0202] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (120), the electronic device (101;201) may generate the second prompt by reflecting the second criterion in the first prompt.

[0203] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (120), the electronic device (101;201) may determine the first criterion using a second large language model (LLM).

[0204] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (120), the electronic device (101;201) may input the first prompt into a third large language model (LLM) to generate the second prompt.

[0205] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (120), the electronic device (101;201) may generate an intermediate image based on the first prompt and the first image and display it on the display (160) before receiving the text.

[0206] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (120), the electronic device (101;201) may display a user interface on the display (160) including a drawing area for receiving the first image and a field for inputting the text.

[0207] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (120), the electronic device (101;201) may display a plurality of options for determining a style related to the transformation of the first image on the display (160). One of the plurality of options may be displayed overlapping with the text.

[0208] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (120), the electronic device (101;201) may store the second prompt in a database.

[0209] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (120), the electronic device (101;201) may determine the first criterion using text similarity between the text and tags set on each of the plurality of criteria.

[0210] According to one embodiment, the image generation model may be a vision generative AI model.

[0211] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (120), the electronic device (101;201) may analyze the text to determine a level related to image conversion, and if the level is greater than or equal to a specified reference value, generate the second prompt.

[0212] According to one embodiment, when the instructions are executed individually or collectively by the at least one processor (120), the electronic device (101;201) may determine a post-processing filter corresponding to the text when the level is less than a specified reference value, input the first image and the first prompt into the image generation model to generate a third image, and apply the post-processing filter to the third image.

[0213] An image conversion method according to one embodiment may be performed in an electronic device (101; 201). The image conversion method may include the operation of receiving a first image, the operation of determining a first prompt regarding the conversion of the first image, the operation of receiving text related to the modification of the first image, the operation of determining a first criterion among a plurality of criteria included in the first prompt based on the text, the operation of modifying the first criterion to generate a second prompt, and the operation of inputting the first image and the second prompt into an image generation model to generate a second image.

[0214] According to one embodiment, the operation of generating the second prompt may include the operation of determining a second criterion by using a first large language model (LLM) to replace or delete the first criterion or to add a separate criterion.

[0215]

[0216] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., first) component is referred to as “coupled” or “connected” to another (e.g., second) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., by wire), wirelessly, or through a third component.

[0217] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0218] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0219] According to one embodiment, the method according to the various embodiments disclosed herein may be provided as included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0220] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device, Display; memory; and It includes at least one processor comprising processing circuitry, and When the above memory is executed individually or collectively by the at least one processor, the electronic device, Receive the first image, Determining a first prompt regarding the transformation of the first image above, and Receive text related to the modification of the first image above, and Based on the above text, determine the first criterion among the plurality of criteria included in the first prompt, and Modify the above first criterion to generate a second prompt, and An electronic device that stores instructions for generating a second image by inputting the first image and the second prompt into an image generation model.

2. In paragraph 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that generates the first image based on sketch input generated through the above display.

3. In paragraph 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, A plurality of options for determining a style related to the transformation of the first image are displayed on the above display, and An electronic device that determines the first prompt corresponding to the selected option when one of the above plurality of options is selected by user input.

4. In paragraph 3, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that loads the first prompt corresponding to the selected option among a plurality of prompts stored in a database related to video conversion.

5. In paragraph 1, the plurality of standards are An electronic device characterized by having one of a positive attribute, a negative attribute, or a positive / negative attribute.

6. In paragraph 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that uses a first large language model (LLM) to determine a second criterion by replacing or deleting the first criterion or adding a separate criterion.

7. In paragraph 6, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that generates the second prompt by reflecting the second criterion in the first prompt.

8. In paragraph 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that determines the first criterion using a second large language model (LLM).

9. In paragraph 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that inputs the first prompt into a third large language model (LLM) to generate the second prompt.

10. In paragraph 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that generates an intermediate image based on the first prompt and the first image and displays it on the display before receiving the text above.

11. In paragraph 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that displays a user interface on the display, the user interface including a drawing area for receiving the first image and a field for inputting the text.

12. In paragraph 11, when the instructions are executed individually or collectively by the at least one processor, the electronic device, A plurality of options for determining a style related to the transformation of the first image are displayed on the above display, and An electronic device characterized in that one of the above multiple options is displayed overlapping with the text.

13. In paragraph 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device for determining the first criterion using text similarity between the text and tags set for each of the plurality of criteria.

14. In paragraph 1, when the instructions are executed individually or collectively by the at least one processor, the electronic device, By analyzing the above text, determine the level related to image conversion, and An electronic device that generates the second prompt when the above level is greater than or equal to a specified threshold value.

15. A method for converting an image performed on an electronic device, Operation of receiving the first image; An operation to determine a first prompt regarding the conversion of the first image above; The operation of receiving text related to the modification of the first image above; An operation to determine a first criterion among a plurality of criteria included in the first prompt based on the above text; The operation of modifying the above-mentioned first standard to generate a second prompt; and A method including the operation of generating a second image by inputting the first image and the second prompt into an image generation model.