Electronic device and control method therefor

By modifying user prompts and optimizing image quality parameters, the electronic device addresses delays in generating high-quality images, ensuring quick and efficient image display.

WO2026054270A1PCT designated stage Publication Date: 2026-03-12SAMSUNG ELECTRONICS CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing electronic devices experience significant delays in generating and displaying high-quality images due to the time required for image upscaling and transmission from external servers, leading to a degraded user experience, especially when multiple iterations are needed.

Method used

An electronic device with a processor and memory that modifies user prompts based on identified image features, generates candidate images locally or remotely, and selects high-quality images for immediate display, reducing delays by optimizing image quality parameters and transmission.

Benefits of technology

The solution enables rapid generation and display of high-quality images, enhancing user experience by minimizing wait times and allowing for efficient selection of desired images without excessive delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025009381_12032026_PF_FP_ABST
    Figure KR2025009381_12032026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device and a control method therefor are provided. The electronic device comprises: a memory for storing one or more instructions; and at least one processor. When executed individually or collectively by the at least one processor, the instructions instruct the electronic device to: receive an input prompt; identify, on the basis of the input prompt, feature information in a first portion of a first generated image; modifies the input prompt on the basis of the feature information so as to obtain a modified prompt; obtain, on the basis of an input text, a plurality of candidate images of a first image quality, corresponding to the modified prompt, by using a first generative AI model for generating an output image; display a UI including the plurality of candidate images on a display; and, when a selected candidate image is identified among the plurality of candidate images, obtain a second generated image of a second image quality that corresponds to the selected candidate image, wherein second image-quality parameters of the second generated image are higher than first image-quality parameters of the first generated image.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device and method of controlling the same

[0001] The present disclosure relates to an electronic device and a method for controlling the same, and more particularly, to an electronic device and a method for controlling the same for obtaining a generated image based on a prompt input by a user.

[0002] Recently, generative AI models have been used to generate a variety of images. Electronic devices, such as refrigerators, can obtain high-quality generated images based on prompts by performing a process via an external server.

[0003] In the past, a prompt was obtained based on user input, the prompt was transmitted to an external server, the server obtained a generated image corresponding to the prompt using a generative AI model, and an electronic device received the generated image from the server, performed upscaling, and then provided a high-quality generated image to the user.

[0004] The process of upscaling an image on an electronic device after it has been generated and transmitted can take a significant amount of time before users can view the resulting high-quality image. For example, the server can take several seconds to generate a medium-sized image, further time is required to transmit the image, and an additional several seconds are required for the electronic device to upscale the image. This results in users having to wait excessively long periods of time before seeing the image they want to set as their wallpaper, significantly degrading the user experience. Furthermore, if the generated image is not approved by the user, the process can take a significant amount of time to repeat.

[0005] Therefore, other devices and methods are needed to obtain high-quality generated images.

[0006] According to an embodiment of the present disclosure, an electronic device includes a memory storing one or more instructions and at least one processor, wherein the one or more instructions, when individually or collectively executed by the at least one processor, cause the electronic device to receive an input prompt, identify feature information within a first portion of a first generated image based on the input prompt, modify the input prompt based on the feature information to obtain a modified prompt, obtain a plurality of candidate images of a first quality corresponding to the modified prompt using a first generative AI model for generating an output image based on input text, and display a UI including the plurality of candidate images through a display, and when a selected candidate image among the plurality of candidate images is identified, obtain a second generated image of a second quality corresponding to the selected candidate image, wherein a second quality parameter of the second generated image may be higher than a first quality parameter of the first generated image.

[0007] In this case, the one or more instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to modify the input prompt to include text that instructs the electronic device to modify a first area corresponding to a first portion of the first generated image to increase a third quality parameter in the first area, and to modify a second area corresponding to a second portion surrounding the first portion of the first generated image to decrease a fourth quality parameter in the second area.

[0008] Meanwhile, the one or more instructions, when individually or collectively executed by the at least one processor, may modify the input prompt to include text that instructs the electronic device to generate the second generated image of the second quality and to compress the second generated image based on information obtained from a network via the at least one memory or communication interface.

[0009] Meanwhile, the electronic device may further include a communication interface, wherein the first generative AI model is stored in a storage of an external server, and the one or more instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit the modified prompt to the external server through the communication interface so that the external server inputs the modified prompt into the first generative AI model, and receive the plurality of candidate images from the external server while the external server processes the modified prompt through the first generative AI model.

[0010] In this case, when the electronic device requests the plurality of candidate images, the electronic device causes the external server to generate a plurality of images corresponding to the plurality of candidate images through at least one neural network model for increasing at least one image quality parameter, and the one or more instructions, when individually or collectively executed by the at least one processor, cause the electronic device to receive the second generated image from the external server through the communication interface when the selected candidate image is identified.

[0011] In this case, the external server obtains the plurality of candidate images and a seed value representing each of the plurality of candidate images using a generative AI model for generating an image corresponding to the text, and the one or more instructions, when individually or collectively executed by the at least one processor, can cause the electronic device to receive a second image quality generated image corresponding to the selected candidate image obtained using the seed value corresponding to the selected candidate image.

[0012] In this case, the communication interface may further be included, and the first generative AI model may be stored in the memory, and the one or more instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to input the modified prompt to the first generative AI model to obtain a plurality of candidate images, and to transmit the plurality of candidate images and the plurality of seed values ​​corresponding to the plurality of candidate images to an external server through the communication interface while displaying a UI including the plurality of candidate images.

[0013] In this case, the one or more instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive the second generated image from the external server through the communication interface based on transmitting a seed value corresponding to the selected candidate image to the external server when the selected candidate image is identified.

[0014] In this case, the camera may be further included, and the one or more instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to take an internal image of the electronic device through the camera, obtain information related to an object in the internal image, and integrate the information related to the object into the input prompt to obtain the modified prompt.

[0015] Meanwhile, the one or more instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to set the second generated image as a background screen of the electronic device.

[0016] A control method of an electronic device according to the present disclosure includes the steps of receiving an input prompt, identifying feature information within a first portion of a first generated image based on the input prompt, modifying the input prompt based on the feature information to obtain a modified prompt, using a first generative AI model for generating an output image based on input text to obtain a plurality of candidate images of a first quality corresponding to the modified prompt, displaying a UI including the plurality of candidate images through a display, and when a selected candidate image among the plurality of candidate images is identified, obtaining a second generated image of a second quality corresponding to the selected candidate image, wherein a second quality parameter of the second generated image may be higher than a first quality parameter of the first generated image.

[0017] In this case, the step of obtaining the modified prompt may modify the input prompt to include text instructing to modify a first area corresponding to a first portion of the first generated image to increase a third quality parameter in the first area, and to modify a second area corresponding to a second portion surrounding the first portion of the first generated image to decrease a fourth quality parameter in the second area.

[0018] Meanwhile, the step of obtaining the modified prompt may include a step of modifying the input prompt to include text instructing to generate the second generated image of the second quality and to compress the second generated image based on information obtained from a network via at least one memory or communication interface of the electronic device.

[0019] Meanwhile, the first generative AI model may be stored in a storage of an external server, and the step of obtaining the plurality of candidate images may include the step of transmitting the modified prompt to the external server via a communication interface so that the external server inputs the modified prompt into the first generative AI model, and the step of receiving the plurality of candidate images from the external server while the external server processes the modified prompt through the first generative AI model.

[0020] In this case, when the electronic device requests the plurality of candidate images, the electronic device causes the external server to generate a plurality of images corresponding to the plurality of candidate images through at least one neural network model for increasing at least one image quality parameter, and the step of obtaining the second generated image may include receiving the second generated image from the external server through the communication interface when the selected candidate image is identified.

[0021] The above and other aspects, features and advantages of the present invention will become more apparent by describing embodiments of the present invention below with reference to the attached drawings.

[0022] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.

[0023] FIG. 1 is a diagram illustrating a system for obtaining a generated image according to one embodiment;

[0024] FIG. 2 is a block diagram illustrating a configuration of an electronic device according to one embodiment;

[0025] FIG. 3 is a block diagram illustrating a configuration of an electronic device for obtaining an image corresponding to a prompt according to one embodiment;

[0026] FIG. 4 is a sequence diagram illustrating an embodiment for obtaining an image corresponding to a prompt according to one embodiment;

[0027] FIG. 5 is a diagram illustrating a UI for obtaining a prompt according to one embodiment;

[0028] FIG. 6 is a flowchart illustrating a method for modifying a prompt according to one embodiment;

[0029] FIG. 7 is a diagram illustrating a UI including multiple candidate images according to one embodiment;

[0030] FIG. 8 is a sequence diagram illustrating an embodiment for obtaining an image corresponding to a prompt according to another embodiment, and

[0031] FIG. 9 is a flowchart illustrating a method of controlling an electronic device for obtaining an image corresponding to a prompt according to one embodiment.

[0032] The embodiments described in this specification and the configurations illustrated in the drawings are merely examples of embodiments, and various modifications may be made without departing from the scope and spirit of this specification.

[0033] The present embodiments may be modified and have various embodiments, and the embodiments are illustrated in the drawings and described in detail in the detailed description. This is not intended to limit the scope of the embodiments, but should be understood to encompass various modifications, equivalents, and / or alternatives of the embodiments of the present disclosure. In connection with the description of the drawings, similar reference numerals may be used for similar components.

[0034] The following examples may be modified in various ways, and the scope of the technical concepts of the present disclosure is not limited to these examples. Rather, these examples are provided to further faithfully and completely convey the technical concepts of the present disclosure to those skilled in the art.

[0035] The terminology used in this disclosure is for the purpose of describing embodiments only and is not intended to limit the scope of the rights. Singular expressions include plural expressions unless the context clearly dictates otherwise.

[0036] Expressions such as "has," "can have," "includes," or "may include" indicate the presence of a feature (e.g., a number, function, operation, or component such as a part) and do not exclude the presence of additional features.

[0037] Expressions such as "A or B," "at least one of A and / or B," or "one or more of A and / or B" can include all possible combinations of the items listed together. For example, "A or B," "at least one of A and B," or "at least one of A or B" can all refer to (1) including at least one A, (2) including at least one B, or (3) including both at least one A and at least one B.

[0038] The expressions “first,” “second,” “first,” or “second,” etc., used in this disclosure can describe various components, regardless of order and / or importance, and are only used to distinguish one component from another, but do not limit the components.

[0039] When it is said that a component (e.g., a first component) is “(operatively or communicatively) coupled with / to” or “connected to” another component (e.g., a second component), it should be understood that said component may be directly coupled to said other component, or may be coupled via another component (e.g., a third component).

[0040] On the other hand, when it is said that a component (e.g., a first component) is "directly connected" or "directly connected" to another component (e.g., a second component), it can be understood that no other component (e.g., a third component) exists between said component and said other component.

[0041] The expression "configured to" as used in the present disclosure may be used interchangeably with, for example, "suitable for," "having the capacity to," "designed to," "adapted to," "made to," or "capable of." The term "configured to" may not necessarily mean only "specifically designed to" in terms of hardware.

[0042] Instead, in some contexts, the phrase "a device configured to" may mean that the device, in conjunction with other devices or components, is "capable of" performing A, B, and C. For example, the phrase "a processor configured (or set) to perform A, B, and C" may refer to a dedicated processor (e.g., an embedded processor) for performing those operations, or a general-purpose processor (e.g., a CPU or application processor) that can perform those operations by executing one or more software programs stored in a memory device.

[0043] In the embodiments, a 'module' or 'part' performs at least one function or operation, and may be implemented as hardware or software, or as a combination of hardware and software. Furthermore, a plurality of 'modules' or 'parts' may be integrated into at least one module and implemented as at least one processor, except for a 'module' or 'part' that needs to be implemented as a specific hardware.

[0044] The various elements and areas in the drawings are schematically drawn. The technical concept of the present invention is not limited by the relative sizes or spacings drawn in the attached drawings.

[0045] In the present disclosure, the meaning of an electronic device (100) “providing” information may include not only displaying information through an output device (e.g., a display (120)) included in the electronic device (100), but also transmitting information to a user terminal that is in communication with the electronic device (100) and displaying the information through a display of the user terminal.

[0046] Throughout this disclosure, "image quality" or "image quality parameter" may refer not only to resolution, but may also include, for example, image detail, color depth, contrast ratio, frame rate, or noise. "Increasing a quality parameter" or "higher quality" may include, for example, increasing resolution, increasing detail, increasing color depth, increasing contrast ratio, increasing frame rate, or reducing noise. Conversely, "reducing a quality parameter" or "lower quality" may include, for example, decreasing resolution, decreasing detail, decreasing color depth, decreasing contrast ratio, decreasing frame rate, or increasing noise.

[0047] The present disclosure will be described in detail with reference to the drawings.

[0048] FIG. 1 is a diagram illustrating a system for acquiring a generated image according to one embodiment. As illustrated in FIG. 1, the system may include an electronic device (100) and an external server (200). Here, the electronic device (100) may be implemented as a refrigerator as illustrated in FIG. 1, but this is only one embodiment, and may be implemented as a home appliance such as a TV, a washing machine, an air conditioner, or a cooking appliance, or may be implemented as a user terminal such as a smart phone, a tablet PC, or a laptop PC, for example. In addition, the external server (200) may be implemented as a single server, but this is only one embodiment, and of course, may be implemented as a plurality of servers.

[0049] The electronic device (100) may receive user input to obtain a prompt. Here, the prompt may refer to an input for initiating an interaction with a generative AI model, and may be in the form of text. In one or more embodiments, the user input may be a user voice, but this is only one embodiment, and may be implemented as, for example, a user touch for entering text through a text window, or a user touch for selecting a UI element in a UI for generating a prompt, but is not limited thereto.

[0050] The electronic device (100) can modify the prompt to quickly and efficiently obtain the generated image. In one or more embodiments, the electronic device (100) can identify features of the generated image based on the prompt and modify the prompt based on the identified features of the generated image. In one or more embodiments, the electronic device (100) can modify the prompt based on information about the electronic device (100) (e.g., performance information) and network information. In one or more embodiments, the electronic device (100) can modify the prompt based on information about objects associated with the electronic device (100) (e.g., objects stored within the electronic device (100).

[0051] The electronic device (100) may transmit the modified prompt to an external server (200). The external server (200) may include a generative AI model and at least one neural network model for image quality improvement. The external server (200) may input the prompt received from the electronic device (100) into the generative AI model to obtain multiple candidate images. The candidate images may be low-quality images.

[0052] An external server (200) can transmit a plurality of generated candidate images to an electronic device (100). The external server (200) can input the plurality of candidate images into at least one neural network model to obtain a high-quality generated image.

[0053] The electronic device (100) can display a UI including multiple candidate images and select one of the multiple candidate images based on user input. The electronic device (100) can transmit information about the selected candidate image to an external server (200).

[0054] The external server (200) can identify a generated image corresponding to a selected candidate image among a plurality of generated generated images and transmit the generated image to the electronic device (100). For example, the external server (200) can acquire a plurality of generated images in advance and transmit a high-quality generated image corresponding to a candidate image selected by the user to the electronic device (100) without a separate delay time. As a result, the user can acquire a high-quality generated image corresponding to the prompt more quickly.

[0055] In the embodiment, an external server (200) is described as acquiring multiple candidate images using a generative AI model. This is merely one embodiment, and the electronic device (100) can also acquire multiple candidate images using a generative AI model. This will be described in detail later with reference to FIG. 8.

[0056] FIG. 2 is a block diagram illustrating a configuration of an electronic device according to one embodiment. As illustrated in FIG. 2, the electronic device (100) may include a communication interface (110), a display (120), a camera (130), an input interface (140), a memory (150), a functional unit (160), and a processor (170). The configuration illustrated in FIG. 2 is an embodiment, and some configurations may be added depending on the implementation of the electronic device (100).

[0057] The communication interface (110) can communicate with an external server (200) or an external terminal device. The communication interface (110) can transmit a prompt to the external server (200) and receive a plurality of candidate images corresponding to the prompt from the external server (200). The communication interface (110) can transmit information about a selected candidate image to the external server (200) and receive a generated image corresponding to the selected candidate image from the external server (200).

[0058] The communication interface (110) can communicate with various external devices using various wireless communication technologies or mobile communication technologies. Wireless communication technologies may include, for example, Bluetooth, Bluetooth Low Energy, CAN communication, Wi-Fi, Wi-Fi Direct, ultrawide band (UWB), Zigbee, infrared Data Association (IrDA), or near field communication (NFC). Mobile communication technologies may include, for example, 3GPP, Wi-Max, LTE (Long Term Evolution), 5G, etc.

[0059] The display (120) can provide various information. The display (120) can display a UI for generating a prompt and a UI including multiple candidate images. The display (120) can provide a generated image corresponding to a candidate image selected by the user. In this case, the generated image can be set as the background image.

[0060] If the electronic device (100) is implemented as a refrigerator, the display (120) may be placed on some of the multiple doors of the refrigerator.

[0061] The camera (130) is configured to capture an image of a subject and generate an image, wherein the captured image may include both a moving image and a still image. Throughout the present disclosure, "image" may include both an image output on a display (120) and an image frame captured by the camera (110).

[0062] When the electronic device (100) is implemented as a refrigerator, the camera (110) can capture images of the storage compartment inside the main body of the refrigerator and the door bin (or door basket, pantry) area of ​​the door. The camera (130) can be installed in the upper area inside the main body to capture images of at least a portion of the inside of the main body and the door bin area of ​​the door. The camera (130) can be installed in the side area or the lower area inside the main body to capture images of at least a portion of the inside of the main body. The camera (130) can be installed on the outside of the refrigerator to capture images of the outside of the refrigerator. The camera (130) can be implemented as a single camera, or can be implemented as a plurality of cameras depending on the embodiment.

[0063] The processor (170) can obtain information about an object contained within the electronic device (100) from an image captured by the camera (130).

[0064] The input interface (140) may include, for example, a button, a lever, a switch, a touch interface, or a microphone. In this case, the touch interface may be implemented in a manner of receiving input by the user's touch on the display (120) screen of the electronic device (100).

[0065] In particular, the input interface (140) can receive, for example, user input for obtaining a prompt, user input for selecting one of a plurality of candidate images.

[0066] The memory (150) may store an operating system (OS) for controlling the overall operation of the components of the electronic device (100) and instructions or data related to the components of the electronic device (100). The memory (150) may include various modules for obtaining a generated image corresponding to a prompt. When an event for obtaining a generated image corresponding to a prompt occurs, the electronic device (100) may load data for various modules stored in the non-volatile memory to perform various operations into the volatile memory. Here, loading means an operation of calling and storing data stored in the non-volatile memory into the volatile memory so that the processor (170) can access it.

[0067] The memory (150) may be implemented as a non-volatile memory (e.g., hard disk, SSD (Solid state drive), flash memory), volatile memory (which may also include memory within the processor (170)), etc.

[0068] According to one embodiment, the memory (150) may store a generative AI model for generating a plurality of candidate images corresponding to a prompt.

[0069] The functional unit (160) can perform various functions of the electronic device (100). For example, if the electronic device (100) is implemented as a refrigerator, the functional unit (160) may include a configuration for refrigerating or freezing food. As another example, if the electronic device (100) is implemented as a washing machine, the functional unit (160) may include a configuration for washing laundry.

[0070] The processor (170) can control the electronic device (100) according to at least one instruction stored in the memory (150).

[0071] The processor (170) may include one or more processors. The one or more processors may include one or more of a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), an Accelerated Processing Unit (APU), a Many Integrated Core (MIC), a Digital Signal Processor (DSP), a Neural Processing Unit (NPU), a hardware accelerator, or a machine learning accelerator. The one or more processors may control one or any combination of other components of the electronic device, and may perform operations related to communication or data processing. The one or more processors may execute one or more programs or instructions stored in a memory. For example, the one or more processors may perform a method according to an embodiment by executing one or more instructions stored in a memory.

[0072] When a method according to an embodiment includes a plurality of operations, the plurality of operations may be performed by one processor or by a plurality of processors. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first processor, or the first operation and the second operation may be performed by a first processor (e.g., a central processing unit (CPU)) and the third operation may be performed by a second processor (e.g., a graphics processing unit (GPU) or a neural processing unit (NPU)). For example, according to an embodiment, an operation of identifying a corner in a handwritten image using a neural network model or correcting a space in a handwritten image may be performed by a processor that performs parallel operations, such as a GPU or an NPU, and an operation of generating / editing a floor plan image or a post-processing operation may be performed by a CPU.

[0073] One or more processors may be implemented as a single core processor including one core, or may be implemented as one or more multicore processors including multiple cores (e.g., homogeneous multicores or heterogeneous multicores). When one or more processors are implemented as a multicore processor, each of the multiple cores included in the multicore processor may include internal processor memory, such as cache memory or on-chip memory, and a common cache shared by the multiple cores may be included in the multicore processor. Each of the multiple cores (or some of the multiple cores) included in the multicore processor may independently read and execute program instructions for implementing a method according to an embodiment, or all (or some) of the multiple cores may be linked to read and execute program instructions for implementing a method according to an embodiment.

[0074] When a method according to an embodiment includes a plurality of operations, the plurality of operations may be performed by one core among the plurality of cores included in a multi-core processor, or may be performed by the plurality of cores. For example, when a first operation, a second operation, and a third operation are performed by a method according to an embodiment, the first operation, the second operation, and the third operation may all be performed by a first core included in the multi-core processor, or the first operation and the second operation may be performed by a first core included in the multi-core processor, and the third operation may be performed by a second core included in the multi-core processor.

[0075] In embodiments, the processor (170) may mean a system on a chip (SoC) in which one or more processors and other electronic components are integrated, a single-core processor, a multi-core processor, or a core included in a single-core processor or a multi-core processor, wherein the core may be implemented as a CPU, a GPU, an APU, a MIC, a DSP, an NPU, a hardware accelerator, or a machine learning accelerator, but the embodiments are not limited thereto.

[0076] The processor (170) receives a prompt for generating an image by executing at least one instruction stored in the memory (150), identifies information about a portion of the generated image based on the prompt, modifies the prompt based on information about a feature portion of the image, obtains a plurality of candidate images of a first quality corresponding to the modified prompt using a generative AI model for generating an image corresponding to text, displays a UI including the plurality of candidate images through the display (120), and, when one of the plurality of candidate images is selected, obtains a generated image of a second quality corresponding to the selected candidate image. Here, the second quality is higher than the first quality.

[0077] In one or more embodiments, the processor (170) may modify the prompt to include text that causes the region corresponding to the feature of the image to be generated in high quality and the region corresponding to the periphery of the image to be generated in low quality.

[0078] In one embodiment, the prompt may be modified to include text that allows the user to determine the second quality and compression of the generated image based on information about the electronic device (100) and network information.

[0079] In one or more embodiments, the generative AI model is stored in an external server (200), the processor (170) transmits a modified prompt to the external server (200) via the communication interface (110), and when the external server (200) inputs the modified prompt into the generative AI model to obtain a plurality of candidate images, the plurality of candidate images can be received from the external server (200).

[0080] In one or more embodiments, the external server (200) may include at least one neural network model for improving image quality. Here, the external server (200) transmits a plurality of candidate images and then acquires a plurality of images corresponding to the plurality of candidate images using the at least one neural network model, and when one of the plurality of candidate images is selected, the processor (170) may receive a second image of a generated image corresponding to the selected candidate image from the plurality of images from the external server (200).

[0081] In one or more embodiments, an external server (200) may obtain a plurality of candidate images and seed values ​​representing each of the plurality of candidate images using a generative AI model for generating images corresponding to text. The processor (170) may receive a second image quality generated corresponding to the selected candidate image, obtained using the seed values ​​corresponding to the selected candidate image.

[0082] In one or more embodiments, the memory (150) may store a generative AI model. Here, the processor (170) may input a modified prompt into the generative AI model to obtain a plurality of candidate images of a first quality corresponding to the modified prompt, and may transmit the plurality of candidate images and seed values ​​representing each of the plurality of candidate images to an external server (200) through the communication interface (110) while displaying a UI including the plurality of candidate images through the display (120). When one of the plurality of candidate images is selected, the processor (170) may receive a generated image obtained through the seed value corresponding to the selected candidate image from the external server (200) through the communication interface (110).

[0083] In one or more embodiments, the processor (170) may capture an internal image of the electronic device (100) via the camera (130), obtain information about an object included in the internal image, and modify the prompt by adding the information about the object to the prompt.

[0084] In one or more embodiments, the processor (170) may set the generated image as the background screen of the electronic device (100).

[0085] FIG. 3 is a block diagram illustrating a configuration of an electronic device for acquiring an image corresponding to a prompt according to one embodiment. As illustrated in FIG. 3, the electronic device (100) may include a prompt acquisition module (310), a prompt modification module (320), a feature identification module (330), a device information acquisition module (340), a network information acquisition module (350), an object information acquisition module (360), a candidate image acquisition module (370), and a generated image acquisition module (380). The plurality of modules disclosed in FIG. 3 may be implemented in software, but this is only one embodiment, and may also be implemented in a combination of software and hardware. According to one embodiment, some of the plurality of modules illustrated in FIG. 3 may not be included, and new modules may be further included.

[0086] The prompt acquisition module (310) may acquire a prompt for generating an image. According to the present disclosure, a "prompt" may mean an input for initiating an interaction with a generative AI model. The prompt may be a text input or a voice input including one or more texts and / or one or more sentences. In one embodiment, the prompt may include natural language text. The natural language text may include various information that the generative AI model may utilize to generate a response to a user inquiry or to control the electronic device (100), such as context, intent, task, and constraints. The prompt may be referred to as a substitute for various expressions representing the same / similar concept. The prompt may be replaced with expressions such as, for example, “input”, “user input”, “input phrase”, “user command”, “directive”, “starting sentence”, “task query”, “trigger sentence”, “message”, etc., and is not limited to the examples described above. The user voice input into the first electronic device (100) may also be a prompt, but for the convenience of explanation, the user voice input initially and the prompt generated by additional information are distinguished.

[0087] In one or more embodiments, the prompt acquisition module (310) may acquire a prompt based on a user's voice input via a microphone. The prompt acquisition module (310) may receive a user's voice via a microphone, perform voice recognition on the user's voice, and acquire a prompt (or text) corresponding to the user's voice.

[0088] In one or more embodiments, the prompt acquisition module (310) may acquire text entered as a prompt through a text input UI displayed on the electronic device (100) or a text input UI displayed on a user terminal connected to the electronic device (100).

[0089] In one or more embodiments, the prompt acquisition module (310) may display a UI for selecting attributes (e.g., category, style, color, etc.) of an image that the user wishes to generate, and may acquire a prompt based on the attributes of the image selected through the UI. For example, if the "cooking" category and the "neat" style are selected through the UI for selecting an image category (e.g., cooking, food, etc.) and an image style (e.g., neat, colorful, etc.), the prompt acquisition module (310) may acquire a prompt such as "Generate a cooking image in a neat style" based on the attributes of the selected image. Here, the prompt acquisition module (310) may acquire the prompt by inputting information about the attributes of the selected image into a trained neural network model or a template prompt.

[0090] The prompt modification module (320) can modify (or correct, update, etc.) the acquired prompt to obtain an optimal image. The prompt modification module (320) can modify the acquired prompt based on various information acquired through the feature identification module (330), the device information acquisition module (340), the network information acquisition module (350), and the object information acquisition module (360).

[0091] In one or more embodiments, the prompt modification module (320) can obtain information about features of the generated image obtained through the feature identification module (330).

[0092] The feature identification module (330) can identify information about features of a generated image based on a prompt. Here, the generated image may be an image generated by inputting the prompt into a neural network model including a generative AI model. The features of the generated image may be a key area or a target object included in the generated image.

[0093] The feature identification module (330) can identify a target word (or main keyword) corresponding to an image that the user wants to create among a plurality of words included in a prompt, thereby identifying a feature of the generated image. In one embodiment, the feature identification module (330) can identify the target word by inputting a prompt to a neural network model for identifying the target word. In one embodiment, the feature identification module (330) can identify the target word by identifying an object word among a plurality of words. The feature identification module (330) can identify a feature of the generated image based on the identified target word.

[0094] The prompt modification module (320) can modify the prompt based on the features of the generated image identified by the feature identification module (330). The prompt modification module (320) can modify the prompt to include text that causes a region corresponding to the features of the generated image to be generated in high quality and a region corresponding to the periphery of the generated image to be generated in low quality. For example, if the prompt is "Draw a dish that can be made with the food in the refrigerator," the feature identification module (320) can identify the target word "cooking" as the feature of the generated image. The prompt modification module (320) can modify the prompt to include text that says "In each image, make the cooking part clear and make the rest blurry."

[0095] In one or more embodiments, the prompt modification module (320) can obtain information about the electronic device (100) and network information through the device information acquisition module (340) and the network information acquisition module (350).

[0096] The device information acquisition module (340) can acquire information about the electronic device (100). Here, the information about the electronic device (100) can include identification information and performance information of the electronic device (100). The identification information of the electronic device (100) can include information about the product name, product number, and manufacturer of the electronic device. The performance information of the electronic device (100) can include usage information of the processor (170), usage information of the memory (150), battery information of the electronic device (100), display information of the electronic device (100), etc.

[0097] The device information acquisition module (340) can acquire information about the electronic device (100) by reading out information about the electronic device (100) stored in the memory (150) of the electronic device (100). The device information acquisition module (340) can receive information about the electronic device (100) through an external server (200).

[0098] The network information acquisition module (350) can acquire information on the network performance of the electronic device (100). Here, the information on the network performance may include at least one of bandwidth information, delay time information, jitter information, packet loss information, throughput information, and QoS (Quality of Service) information. The network information acquisition module (350) can acquire information on delay time and packet loss through a ping test, can acquire information on bandwidth (download / upload time, etc.) through a speed test, and can evaluate the throughput and bandwidth of the network using iPerf. This is one embodiment, and network information can be acquired through various methods.

[0099] The prompt modification module (320) can modify the prompt based on information about the electronic device (100) and network information acquired through the device information acquisition module (340) and the network information acquisition module (350). For example, the prompt modification module (320) can modify the prompt to include information about the electronic device (100) and network information along with text such as "Determine the quality level based on the given network speed, and determine whether to compress the image based on the device specifications."

[0100] In one or more embodiments, the prompt modification module (320) may obtain information about an object contained within the electronic device (100) through the object information acquisition module (360). For example, if the electronic device (100) is a refrigerator, the object may include food, dishes, etc. stored in the refrigerator.

[0101] The object information acquisition module (360) can acquire information about an object included in the electronic device (100) based on an internal image of the electronic device (100) captured by a camera. In one embodiment, the object information acquisition module (360) can acquire information about an object included in the electronic device (100) by inputting the internal image into a neural network model trained to identify an object included in the image.

[0102] The prompt modification module (320) can modify the prompt modification module (320) based on information about the object acquired through the object information acquisition module (360). For example, if the prompt is "Draw a dish that can be made with the food in the refrigerator" and the object is "kimchi," the prompt modification module (320) can modify the prompt to "Draw a dish that can be made with kimchi" based on information about the object.

[0103] The candidate image acquisition module (370) can acquire multiple candidate images corresponding to the modified prompt using a generative AI model (385).

[0104] Here, the generative AI model (385) is an artificial intelligence model that generates new content based on input data, and may be a model that generates an image (or video) by inputting text. According to one embodiment, the generative AI model (385) may be a neural network model trained to input a prompt and generate a plurality of candidate images corresponding to the prompt. Here, the plurality of candidate images may be images of a first quality (or low quality). The first quality may be a low quality with a resolution of 480p (640x480 pixels) or lower.

[0105] In one or more embodiments, the generative AI model (385) may be stored in an external server (200). For example, the candidate image acquisition module (370) may transmit a prompt obtained from the prompt modification module (320) to the external server (200). The candidate image acquisition module (370) may input a prompt from the external server (200) into the generative AI model (385) and receive a plurality of candidate images obtained.

[0106] Here, the external server (200) can adjust the image quality of the plurality of candidate images depending on the network speed and the performance of the electronic device. For example, if the network speed is above a threshold or the performance of the electronic device (100) is high, the external server (200) can transmit a plurality of candidate images of the first image quality. If the network speed is below the threshold or the performance of the electronic device (100) is low, the external server (200) can transmit a plurality of compressed candidate images.

[0107] In one or more embodiments, the generative AI model (385) may be stored in the memory (150) of the electronic device (100). For example, the candidate image acquisition module (370) may input a prompt obtained from the prompt modification module (320) into the generative AI model (385) to obtain multiple candidate images.

[0108] When generating multiple candidate images, the generative AI model (385) can obtain seed values ​​corresponding to each of the multiple candidate images. Here, the seed value is a value representing the candidate image, and may be a value that allows the neural network model to generate the same output for the same input.

[0109] If the image desired by the user is not found among multiple candidate images, the candidate image acquisition module (370) can request a new candidate image from an external server (200).

[0110] The generated image acquisition module (380) can acquire a generated image selected by the user from among a plurality of candidate images. Here, the generated image may be an image of second quality (or high definition). For example, the second quality may be an image having a resolution of 1080p (1920x1080 pixels), 4K (3840x2160 pixels), or higher.

[0111] The generated image acquisition module (380) may display a UI including multiple candidate images. When a user input for selecting one of the multiple candidate images is received, the generated image acquisition module (380) may transmit information (e.g., a seed value) about the selected candidate image to an external server (200).

[0112] Here, the external server (200) can obtain second high-quality generated images corresponding to the plurality of candidate images while the user selects one of the plurality of candidate images (or while a UI including the plurality of candidate images is displayed). For example, the external server (200) can generate a plurality of candidate images and obtain a plurality of generated images corresponding to the plurality of candidate images. The external server (200) stores the obtained plurality of generated images. When information regarding a selected candidate image among the plurality of candidate images is received, the external server (200) can transmit a generated image corresponding to the selected candidate image to the electronic device (100). This allows the user to obtain a high-quality generated image more quickly.

[0113] As illustrated in FIG. 3, the external server (200) can obtain a second image of a generated image corresponding to a candidate image using at least one neural network model (390). In one or more embodiments, the at least one neural network model (390) is a model for improving the image quality of an input image, and may include an upscaling model (391) and a refinement model (393).

[0114] The upscaling model (391) may be an artificial intelligence model used to improve the image quality by converting a low-quality (or low-resolution) image (or video) to a higher-quality (or higher-resolution) image. The upscaling model (391) may be trained to increase the resolution of an input image. For example, the upscaling model (391) may be one of, but is not limited to, a Super-Resolution Convolutional Neural Network (SRCNN), an Enhanced Deep Super-Resolution Network (EDSR), or an Enhanced Super-Resolution Generative Adversarial Network (ESRGAN).

[0115] The refinement model (393) may be an artificial intelligence model used to refine the input image to make it more precise and accurate. The refinement model (393) may be trained to enhance the details of the input image. The refinement model (393) may be an artificial intelligence model based on a Generative Adversarial Network (GAN), but is not limited thereto.

[0116] The external server (200) can obtain information about a generated image of a second quality with improved resolution by inputting a candidate image into an upscaling model (391), and can obtain a generated image of a second quality with improved detail by inputting the generated image of the second quality into a refining model (393). In FIG. 3, the operation order of the upscaling model (391) and the refining model (393) can be changed. For example, the external server (200) can obtain information about a generated image of a first quality with improved detail by inputting a candidate image into a refining model (393), and can obtain a generated image of a second quality with improved detail by inputting the generated image of the first quality with improved detail into the upscaling model (391).

[0117] In FIG. 3, at least one neural network model (390) is described as including an upscaling model (391) and a refinement model (393), but this is only one embodiment, and in addition to the upscaling model (391) and the refinement model (393), additional neural network models (e.g., a noise removal model, etc.) may be included.

[0118] In Fig. 3, the neural network model (390) is described as including multiple neural network models (391, 393), but this is only one embodiment, and may be implemented as a single neural network model (i.e., a neural network model in which an upscaling model (391) and a refinement model (393) are combined) for improving the image quality.

[0119] The generated image acquisition module (380) can receive a generated image of second quality corresponding to the selected candidate image from an external server (200). The generated image acquisition module (380) can provide the generated image of second quality through a display (120) and store it in a memory (150). The generated image acquisition module (380) can set the acquired generated image as a background screen.

[0120] Figure 4 is a sequence diagram illustrating an embodiment for obtaining an image corresponding to a prompt, according to one embodiment. In the following embodiments, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0121] The electronic device (100) may obtain a prompt (S405). Here, the prompt may be obtained based on a user input for obtaining a generated image. In one or more embodiments, the electronic device (100) may receive a user voice input through a microphone. The electronic device (100) may obtain text corresponding to the user voice acquired through Automatic Speech Recognition (ASR) as a prompt. In one or more embodiments, the electronic device (100) may display a UI for text input on the display (120). The electronic device (100) may obtain text input through the UI for text input as a prompt. In one or more embodiments, the electronic device (100) may obtain text acquired by a user terminal as a prompt. Here, the user terminal may obtain text based on the user voice acquired through a microphone or obtain text through the UI for text input. In one or more embodiments, the electronic device (100) may display a UI for obtaining a prompt. Here, the UI for obtaining a prompt may include a UI element for selecting a keyword to be entered into the prompt. For example, as illustrated in FIG. 5, the electronic device (100) may display a UI (500) including UI elements for selecting a keyword entered into the prompt, such as a category (510) of the generated image, a style (520) of the generated image, a background screen (530) of the generated image, etc. When a keyword to be entered into the prompt is selected through the UI (500), the electronic device (100) may input the selected keyword into a pre-stored template prompt to generate a prompt.For example, if the template prompt "Generate ZZZ images with XXX style and YYY wallpaper" is saved and "Cooking category", "Clean style", and "Indoor wallpaper" are selected through the UI (500), the electronic device (100) can obtain a prompt such as "Generate cooking images with an indoor wallpaper with a clean style."

[0122] The electronic device (100) can modify the prompt (S410). The electronic device (100) can modify the prompt based on at least one of the features of the generated image, information about the electronic device (100), network information, and information about the object. This will be described with reference to FIG. 6.

[0123] FIG. 6 is a flowchart illustrating a method for modifying a prompt according to one embodiment.

[0124] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0125] According to one embodiment, S610 to S670 may be understood to be performed in a processor (e.g., processor (170) of FIG. 2) of an electronic device (e.g., electronic device (100) of FIG. 2).

[0126] The electronic device (100) can obtain a prompt (S610). For additional implementation details, refer to the description of step S405 of FIG. 4.

[0127] Here, the electronic device (100) can identify a feature of the generated image (S620). The electronic device (100) can identify a target word (or a main keyword) corresponding to an image that the user wants to generate among a plurality of words included in the prompt to identify the feature of the generated image. In one embodiment, the electronic device (100) can input the input prompt into a trained neural network model to identify the target word. In another embodiment, the electronic device (100) can identify a keyword selected through a category UI element (510) among a plurality of UI elements included in the UI (500) as the target word. In another embodiment, the electronic device (100) can identify the target word by identifying an object among a plurality of words included in the prompt.

[0128] The electronic device (100) can identify information about the electronic device (100) (S630). Here, the information about the electronic device (100) may include identification information of the electronic device (100) and performance information of the electronic device (100). Here, the identification information of the electronic device (100) may include information about the product name, product number, and manufacturer of the electronic device. The performance information of the electronic device (100) may include usage information of the processor (170), usage information of the memory (150), battery information of the electronic device (100), display information of the electronic device (100), etc.

[0129] The electronic device (100) can obtain network information (S640). Here, the network performance information may include at least one of bandwidth information, delay time information, jitter information, packet loss information, throughput information, and QoS (Quality of Service) information. The electronic device (100) can obtain information on delay time and packet loss through a ping test, obtain information on bandwidth (download / upload time, etc.) through a speed test, and evaluate the network throughput and bandwidth using iPerf.

[0130] The electronic device (100) can identify whether the prompt indicates that information about an object should be acquired (S650). The electronic device (100) can identify whether a word related to an object stored within the electronic device (100) is included among a plurality of words included in the prompt. For example, if a prompt such as "Generate an image with food in the refrigerator," the electronic device (100) can identify that the word related to an object stored within the electronic device (100), such as "food in the refrigerator," is included among the plurality of words included in the prompt. The electronic device (100) can identify that the prompt indicates that information about an object should be acquired.

[0131] When it is identified that information about an object needs to be acquired (S650-Y), the electronic device (100) can acquire information about the object (S660). The electronic device (100) can input an internal image of the electronic device (100) captured by the camera (130) into a trained neural network model to acquire information about the object included in the internal image. The electronic device (100) can acquire information about the object input by the user when storing the object within the electronic device (100).

[0132] The electronic device (100) may modify the prompt (S670). In one or more embodiments, the electronic device (100) may modify the prompt to include text that instructs the electronic device (100) to generate a high-quality image corresponding to a feature of the image and to generate a low-quality image corresponding to a periphery of the image based on the feature of the image. For example, if the prompt "Draw a dish that can be made with the food in the refrigerator" is input, the electronic device (100) may identify the feature of the image as "dish" and modify the prompt to include text that instructs the electronic device (100) to determine the second quality and whether to compress the generated image based on information about the electronic device (100) and network information. For example, the electronic device (100) may modify the prompt to include text such as "Decide the quality level of the image based on the identified network speed (x bps) and decide whether to compress the image based on the electronic device specifications (product XYY)." In one or more embodiments, the electronic device (100) may modify the prompt to add information about the object. For example, if the prompt "Draw a dish that can be made with the food in the refrigerator" is input, the electronic device (100) may modify the prompt based on the information about the object, such as "Draw a dish that can be made with the kimchi in the refrigerator."

[0133] Referring back to FIG. 4, the electronic device (100) can transmit the modified prompt to an external server (200) (S420).

[0134] The external server (200) can obtain multiple candidate images based on the modified prompt (S425). The external server (200) can obtain multiple candidate images by inputting the modified prompt into the generative AI model. For example, if the prompt "Draw a dish that can be made with kimchi in the refrigerator" is obtained, the external server (200) can generate multiple candidate images including dishes related to kimchi using the generative AI model. Here, the generated multiple candidate images may be low-quality images. The external server (200) can obtain the multiple candidate images along with seed values ​​representing the multiple candidate images.

[0135] The external server (200) can adjust the image quality of multiple candidate images depending on the network speed and the performance of the electronic device. For example, if the network speed is above a threshold or the performance of the electronic device (100) is high, the external server (200) can transmit multiple candidate images of the first image quality. If the network speed is below the threshold or the performance of the electronic device (100) is low, the external server (200) can transmit multiple compressed candidate images.

[0136] An external server (200) can transmit multiple candidate images to an electronic device (100) (S430).

[0137] The electronic device (100) may display a UI including a plurality of candidate images (S440). For example, as illustrated in FIG. 7, the electronic device (100) may display a UI (700) including a plurality of low-quality candidate images (710 to 740). Here, if the image desired by the user is not among the candidate images, the electronic device (100) may include a "See more candidate images" UI element (750) to generate additional candidate images. For example, when the "See more candidate images" UI element (750) is selected, the electronic device (100) may transmit a request signal to an external server (200) to generate additional candidate images.

[0138] While the electronic device (100) provides a plurality of candidate images, the external server (200) can obtain generated images corresponding to the plurality of candidate images (S445). For example, the external server (200) can obtain generated images corresponding to the plurality of candidate images regardless of the user's selection of candidate images. Here, the external server (200) can obtain generated images with improved image quality using at least one neural network model. In one or more embodiments, the at least one neural network model is a model for improving the image quality of an input image, and may include an upscaling model and a refinement model. The external server (200) can improve the image quality of the candidate images using a seed value representing the plurality of candidate images.

[0139] The electronic device (100) can select one of a plurality of candidate images based on a user input (S450). For example, if a user input for selecting one candidate image is detected through the UI illustrated in FIG. 7, the electronic device (100) can select one of the plurality of candidate images based on the user input.

[0140] The electronic device (100) can transmit information about the selected candidate image to an external server (200) (S455). Here, the information about the selected candidate image may be identification information of the selected candidate image or a seed value corresponding to the selected candidate image.

[0141] The external server (200) may transmit a generated image corresponding to the selected candidate image to the electronic device (100) (S460). For example, the external server (200) may identify a generated image corresponding to the selected candidate image among a plurality of generated images based on information received from the electronic device (100). The external server (200) may transmit the identified generated image to the electronic device.

[0142] The electronic device (100) can provide the received generated image (S465). The electronic device (100) can provide the generated image through the display (120) and can provide the generated image to an external user terminal. In one or more embodiments, the electronic device (100) can automatically set the generated image as the background screen.

[0143] FIG. 8 is a sequence diagram illustrating an embodiment for obtaining an image corresponding to a prompt according to another embodiment. In the following embodiments, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel. For additional implementation details of operations S805 and S810 of FIG. 8, reference may be made to the description of operations S405 and S410 of FIG. 4.

[0144] The electronic device (100) can acquire multiple candidate images using the modified prompt (S815). The electronic device (100) can acquire multiple candidate images by inputting the modified prompt into the generative AI model. Here, the generated multiple candidate images may be low-quality images. The electronic device (100) can acquire seed values ​​representing the multiple candidate images together with the multiple candidate images.

[0145] The electronic device (100) can transmit information about multiple candidate images to an external server (200) (S820). Here, the information about the multiple candidate images may be seed values ​​representing the multiple candidate images, but is not limited thereto.

[0146] The electronic device (100) can display a UI including multiple candidate images (S825). For example, as illustrated in FIG. 7, the electronic device (100) can display a UI (700) including multiple low-quality candidate images (710 to 740).

[0147] While the electronic device (100) provides multiple candidate images, the external server (200) can obtain generated images corresponding to the multiple candidate images (S830). Here, the external server (200) can obtain generated images corresponding to the multiple candidate images using a seed value. The external server (200) can obtain generated images with improved image quality using at least one neural network model.

[0148] The electronic device (100) can select one of a plurality of candidate images based on a user input (S840). For example, if a user input for selecting one candidate image is detected through the UI illustrated in FIG. 7, the electronic device (100) can select one of the plurality of candidate images based on the user input.

[0149] The electronic device (100) can transmit information about the selected candidate image to an external server (200) (S845). Here, the information about the selected candidate image may be identification information of the selected candidate image or a seed value corresponding to the selected candidate image.

[0150] An external server (200) can transmit a generated image corresponding to the selected candidate image to the electronic device (100) (S850).

[0151] The electronic device (100) can provide the received generated image (S855). The electronic device (100) can provide the generated image through the display (120) and can provide the generated image to an external user terminal. In one or more embodiments, the electronic device (100) can automatically set the generated image as the background screen.

[0152] As described in FIGS. 4 to 8, since the image enhancement work is performed on the external server (200), the external server (200) can perform the image enhancement work, which was previously impossible on the electronic device (100), in a short time of several hundred milliseconds. For example, the processing time can be significantly reduced. The computational burden on the electronic device (100) is reduced, which contributes to extending the battery life and preserving the performance of the electronic device (100), and allows more resources to be allocated to other tasks. The external server (200) can perform the image enhancement work to obtain a generated image of even better quality. The user experience can be improved due to the fast image processing and improved image quality. In addition, the optimal generated image can be provided by reflecting the user prompt, the network environment, and information on the electronic device.

[0153] FIG. 9 is a flowchart illustrating a method of controlling an electronic device for obtaining an image corresponding to a prompt according to one embodiment.

[0154] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0155] According to one embodiment, S910 to S970 may be understood to be performed in a processor (e.g., processor (170) of FIG. 2) of an electronic device (e.g., electronic device (100) of FIG. 2).

[0156] The electronic device (100) receives a prompt for generating an image (S910).

[0157] The electronic device (100) identifies information about a portion of the generated image based on the prompt (S920).

[0158] The electronic device (100) modifies the prompt based on information about the features of the image (S930). In one or more embodiments, the electronic device (100) may modify the prompt to include text that causes the area corresponding to the features of the image to be generated in high definition, and the area corresponding to the periphery of the image to be generated in low definition.

[0159] The electronic device (100) obtains a plurality of candidate images of a first image quality corresponding to the modified prompt using a generative AI model for generating an image corresponding to text (S940). In one or more embodiments, the electronic device (100) may transmit the modified prompt to an external server (200), and when the external server (200) inputs the modified prompt into the generative AI model to obtain a plurality of candidate images, the electronic device (100) may receive the plurality of candidate images from the external server (200). The external server (200) may obtain a plurality of candidate images and a seed value representing each of the plurality of candidate images using the generative AI model for generating an image corresponding to text.

[0160] In one or more embodiments, the electronic device (100) may input a modified prompt into a generative AI model to obtain a plurality of candidate images of a first quality corresponding to the modified prompt. While displaying a UI including the plurality of candidate images, the electronic device (100) may transmit the plurality of candidate images and seed values ​​representing each of the plurality of candidate images to an external server (200).

[0161] After transmitting a plurality of candidate images, the external server (200) can obtain a plurality of images corresponding to the plurality of candidate images by using at least one neural network model for image quality improvement.

[0162] The electronic device (100) provides a UI including multiple candidate images (S950).

[0163] The electronic device (100) selects one of a plurality of candidate images (S960).

[0164] The electronic device (100) obtains a generated image of a second quality corresponding to the selected candidate image (S970). Here, the second quality may be higher than the first quality. In one or more embodiments, when one of the plurality of candidate images is selected, the electronic device (100) may receive a generated image of a second quality corresponding to the selected candidate image from the external server (200). The electronic device (100) may receive a generated image of a second quality corresponding to the selected candidate image using a seed value corresponding to the selected candidate image.

[0165] In one or more embodiments, the electronic device (100) may modify the prompt to include text that allows the electronic device (100) to determine the second quality and whether to compress the generated image based on information about the electronic device (100) and network information.

[0166] In one or more embodiments, the electronic device (100) may capture an internal image of the electronic device (100) via a camera (130) and obtain information about an object included in the internal image. In addition, the electronic device (100) may modify the prompt by adding information about the object to the prompt.

[0167] In one or more embodiments, the electronic device (100) may set the generated image as the background screen of the electronic device (100).

[0168] The method according to various embodiments may be provided as a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product (e.g., a downloadable app) may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0169] The method according to various embodiments may be implemented as software including instructions stored in a machine-readable storage medium that can be read by a machine (e.g., a computer). The device is a device that can call instructions stored in the storage medium and operate according to the called instructions, and may include an electronic device (e.g., a TV) according to embodiments.

[0170] A device-readable storage medium may be provided in the form of a non-transitory storage medium. Here, the term "non-transitory storage medium" simply means a tangible device that does not contain signals (e.g., electromagnetic waves). This term does not distinguish between cases where data is permanently stored in the storage medium and cases where data is temporarily stored. For example, a "non-transitory storage medium" may include a buffer in which data is temporarily stored.

[0171] When the above instruction is executed by the processor, the processor may perform the function corresponding to the instruction directly or by using other components under the control of the processor. The instruction may include code generated or executed by a compiler or interpreter.

[0172] Although the preferred embodiments have been illustrated and described above, the present disclosure is not limited to the embodiments, and various modifications may be made by a person having ordinary skill in the art to which the present disclosure pertains without departing from the gist of the present disclosure as claimed in the claims. Furthermore, such modifications should not be understood individually from the technical idea or prospect of the present disclosure.

Claims

1. In electronic devices, memory that stores one or more instructions; and Contains at least one processor, The one or more instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Receive an input prompt, Identifying feature information within a first portion of a first generated image based on the above input prompt, Modify the input prompt based on the above characteristic information to obtain a modified prompt, Obtaining a plurality of candidate images of the first quality corresponding to the modified prompt using a first generative AI model for generating an output image based on input text, Displaying a UI including the plurality of candidate images through a display, When a selected candidate image is identified among the plurality of candidate images, a second generated image of a second quality corresponding to the selected candidate image is obtained. An electronic device wherein the second quality parameter of the second generated image is higher than the first quality parameter of the first generated image.

2. In paragraph 1, The one or more instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Modifying a first area corresponding to a first part of the first generated image to increase a third image quality parameter in the first area, To reduce the fourth quality parameter in the second region by modifying the second region corresponding to the second region surrounding the first portion of the first generated image, An electronic device that modifies said input prompt to include the text to be directed.

3. In paragraph 1, The one or more instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Generating the second generated image of the second image quality, Compress the second generated image based on information obtained from a network via at least one of the above memory or communication interfaces; An electronic device that modifies said input prompt to include the text to be directed.

4. In paragraph 1, Including further communication interfaces, The above first generative AI model is stored in the storage of an external server, The one or more instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: The external server transmits the modified prompt to the external server via the communication interface so that the external server inputs the modified prompt into the first generative AI model, An electronic device that receives the plurality of candidate images from the external server while the external server processes the modified prompt through the first generative AI model.

5. In paragraph 4, When the electronic device requests the plurality of candidate images, the electronic device causes the external server to generate a plurality of images corresponding to the plurality of candidate images through at least one neural network model for increasing at least one image quality parameter, and the one or more instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: An electronic device that, when the selected candidate image is identified, receives the second generated image from the external server through the communication interface.

6. In paragraph 5, The above external server is, Obtaining the plurality of candidate images and seed values ​​representing each of the plurality of candidate images using a generative AI model for generating images corresponding to text, The one or more instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: An electronic device that receives a second image of a quality corresponding to the selected candidate image obtained using a seed value corresponding to the selected candidate image.

7. In paragraph 1, Including further communication interfaces, The above first generative AI model is stored in the memory, The one or more instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: By inputting the above modified prompt into the first generative AI model, multiple candidate images are obtained, An electronic device that transmits the plurality of candidate images and the plurality of seed values ​​corresponding to the plurality of candidate images to an external server through the communication interface while displaying a UI including the plurality of candidate images.

8. In paragraph 7, The one or more instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: An electronic device that, when the selected candidate image is identified, transmits a seed value corresponding to the selected candidate image to the external server and receives the second generated image from the external server through the communication interface.

9. In paragraph 1, Including more cameras, The one or more instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Taking an internal image of the electronic device through the camera, Obtain information related to objects within the above internal image, An electronic device that integrates information related to the object into the input prompt to obtain the modified prompt.

10. In paragraph 1, The one or more instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: An electronic device that sets the second generated image as the background screen of the electronic device.

11. In a method for controlling an electronic device, Step of receiving an input prompt; A step of identifying feature information within a first portion of a first generated image based on the input prompt; A step of modifying the input prompt based on the above characteristic information to obtain a modified prompt; A step of obtaining a plurality of candidate images of a first quality corresponding to the modified prompt using a first generative AI model for generating an output image based on input text; A step of displaying a UI including the plurality of candidate images through a display; and When a selected candidate image is identified among the plurality of candidate images, a step of obtaining a second generated image of a second quality corresponding to the selected candidate image is included; A control method wherein the second image quality parameter of the second generated image is higher than the first image quality parameter of the first generated image.

12. In paragraph 11, The steps for obtaining the above modified prompt are: Modifying a first area corresponding to a first part of the first generated image to increase a third image quality parameter in the first area, To reduce the fourth quality parameter in the second region by modifying the second region corresponding to the second region surrounding the first portion of the first generated image, A control method for modifying the above input prompt to include the text to be directed.

13. In paragraph 11, The steps for obtaining the above modified prompt are: Generating the second generated image of the second image quality, Compress the second generated image based on information obtained from a network via at least one memory or communication interface of the electronic device; A control method comprising the step of modifying said input prompt to include instructive text; 14. In paragraph 11, The above first generative AI model is stored in the storage of an external server, The step of obtaining the above multiple candidate images is: A step of transmitting the modified prompt to the external server through a communication interface so that the external server inputs the modified prompt into the first generative AI model; and A control method comprising: receiving the plurality of candidate images from the external server while the external server processes the modified prompt through the first generative AI model; 15. In paragraph 14, When the electronic device requests the plurality of candidate images, the electronic device causes the external server to generate a plurality of images corresponding to the plurality of candidate images through at least one neural network model for increasing at least one image quality parameter, and the step of obtaining the second generated image is A control method for receiving the second generated image from the external server through the communication interface when the selected candidate image is identified.

Citation Information

Patent Citations

  • Neural image compression with controllable spatial bit allocation

    US20230156207A1

  • Method of on-device generation and supplying wallpaper stream and computing device implementing the same

    WO2022075533A1

  • KR20240069069A