ELECTRONIC DEVICES

VN126590APending Publication Date: 2026-07-01SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
VN · VN
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2024-08-30
Publication Date
2026-07-01

AI Technical Summary

Technical Problem

Existing electronic devices struggle to efficiently allow users to select and edit specific parts of an image, limiting the ease of image processing and editing.

Method used

The electronic device employs a processor with a display, processing circuit, and memory, using AI models to identify objects in images, generate candidate and detailed prompts, and perform image editing based on user input.

Benefits of technology

This solution enables users to easily select and edit specific objects within images, facilitating various editing options and improving the overall image processing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure VN1202603137_0
    Figure VN1202603137_0
Patent Text Reader

Abstract

The invention relates to an electronic device. The electronic device may include: a display unit; a processor including processing circuitry; and memory for storing instructions. The instructions, when executed individually or collectively by the processor, cause the electronic device to: identify the object of the input image and one or more objects associated with the object; generate candidate statement parts for the object and one or more objects through an artificial intelligence (AI) model; display the candidate statement parts and the input image through the display unit; generate an instruction to correct the image based on user input for at least one of the candidate statement parts; and generate an output image by performing image correction on the input image according to the instruction generated through the AI ​​image model. The candidate statement parts may be used to indicate editable objects in the input image.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic devices and methods for image processing

[0001] The descriptions below relate to electronic devices and methods for image processing.

[0002] Electronic devices can perform image processing on images. The electronic devices can perform image processing in response to user input. Generative artificial intelligence (AI) can be used for image processing.

[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.

[0004] In embodiments, an electronic device is provided. The electronic device may include a display, a processor including processing circuitry, and a memory storing instructions. The instructions, when individually or collectively executed by the processor, may cause the electronic device to identify an object in an input image and one or more objects associated with the object, generate candidate prompt portions for the object and the one or more objects through an artificial intelligence (AI) model, display the candidate prompt portions and the input image through the display, generate a prompt for image editing based on a user input for at least one of the candidate prompt portions, and perform image editing on the input image according to the generated prompt through an image AI model, thereby generating an output image. The candidate prompt portions may be used to indicate editable objects in the input image.

[0005] In embodiments, an electronic device is provided. The electronic device may include a display, a processor including a processing circuit, and a memory storing instructions. The instructions, when individually or collectively executed by the processor, may cause the electronic device to generate text for an input image through an artificial intelligence (AI) model, display the input image and the text through the display, identify an object corresponding to the target candidate prompt portion among candidate prompt portions of the text in response to a first user input for a target candidate prompt portion, generate detailed prompt portions associated with the object through the AI ​​model, display the detailed prompt portions associated with the object through the display, and generate a prompt for image editing based on a second user input for indicating an execution prompt among the detailed prompt portions. The candidate prompt portions of the text may respectively correspond to editable objects within the input image.

[0006] Figure 1 is a block diagram of an electronic device within a network environment.

[0007] Figure 2 shows an example of generating a prompt using an AI (artificial intelligence) model.

[0008] Figure 3 shows examples of candidate prompt parts according to the input image.

[0009] Figure 4 shows examples of detailed prompt parts according to the selection of candidate prompt parts.

[0010] Figure 5 shows examples of editing actions.

[0011] Figure 6 shows an example of a candidate prompt portion.

[0012] Figure 7 shows an example of image processing using an AI model.

[0013] Figures 8a to 8d illustrate examples of scenarios of image editing functions through interactive services.

[0014] Figure 9 illustrates the operation flow of an electronic device for editing an image using a prompt generated through an AI model.

[0015] Figure 10 illustrates the operational flow of an electronic device for generating a prompt for image editing.

[0016] Figure 11 illustrates an operation flow of an electronic device for editing an image including an object and a background.

[0017] The terms used in this disclosure are used only to describe specific embodiments and may not be intended to limit the scope of other embodiments. The singular expression may include plural expressions unless the context clearly indicates otherwise. Terms used herein, including technical or scientific terms, may have the same meaning as commonly understood by those of ordinary skill in the art described in this disclosure. Terms defined in general dictionaries among the terms used in this disclosure may be interpreted as having the same or similar meaning in the context of the relevant technology, and shall not be interpreted in an idealized or overly formal sense unless explicitly defined in this disclosure. In some cases, even if a term is defined in this disclosure, it cannot be interpreted to exclude embodiments of the present disclosure.

[0018] The various embodiments of the present disclosure described below illustrate a hardware-based approach as an example. However, since the various embodiments of the present disclosure include techniques utilizing both hardware and software, the various embodiments of the present disclosure do not exclude a software-based approach.

[0019] In the following description, terms referring to signals (e.g., signal, information, message, signaling), terms referring to data types (e.g., list, set, subset), terms for operational states (e.g., step, operation, procedure), terms referring to data (e.g., packet, user stream, information, bit, symbol, codeword), terms referring to resources (e.g., symbol, slot, subframe, radio frame, subcarrier, resource element (RE), resource block (RB), bandwidth part (BWP), occasion), terms referring to channels, terms referring to network entities, terms referring to components of devices, etc. are examples for convenience of description. Therefore, the present disclosure is not limited to the terms described below, and other terms having equivalent technical meanings may be used.

[0020] In the following description, terms referring to input data to an AI (artificial intelligence) model (e.g., signal, information, input data, prompt, candidate prompt, prompt phrase, input text, input object), information for indicating an object (e.g., prompt part, prompt area, prompt target, candidate prompt part, prompt object, object indicator, object indicator information, object input information, masking area indicator information, masking information), terms referring to components of a device, etc. are examples for convenience of description. Therefore, the present disclosure is not limited to the terms described below, and other terms having equivalent technical meanings may be used.

[0021] In addition, in the present disclosure, expressions such as "more than" or "less than" may be used to determine whether a specific condition is satisfied or fulfilled, but this is merely a description for expressing an example and does not exclude descriptions such as "more than" or "less than." A condition described as "more than" may be replaced with "more than," a condition described as "less than" may be replaced with "less than," and a condition described as "more than and less than" may be replaced with "more than and less than." In addition, hereinafter, "A" to "B" mean at least one of elements from A (including A) to B (including B). hereinafter, "C" and / or "D" mean at least one of "C" or "D," that is, including {"C", "D", "C" and "D"}.

[0022] Figure 1 is a block diagram of an electronic device within a network environment.

[0023] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0024] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculation, the processor (120) may store a command or data received from another component (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the command or data stored in the volatile memory (132), and store the resulting data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor)) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0025] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0026] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0027] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0028] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0029] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0030] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0031] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0032] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0033] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0034] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0035] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0036] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0037] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).

[0038] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0039] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0040] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimizing terminal power and connecting multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0041] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0042] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0043] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0044] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0045] The introduction of generative artificial intelligence (AI) enables various image editing, allowing users to edit images more conveniently. By inputting a prompt for image editing into an AI model, the user can obtain a desired output image. Meanwhile, it is not easy to specify only the portion of an image to be edited within an electronic device, and the designation of image processing functions for that portion has been limited. In various embodiments of the present disclosure, a technique for generating a prompt for image editing is provided. The user can more easily select the portion of an image to be edited and perform various editing operations. For example, through text selection, touch input to an area of ​​an image, text, multi-modal input, and / or voice input, the electronic device can automatically generate a prompt for editing a desired portion of an image.

[0046] Figure 2 illustrates an example of generating a prompt using an artificial intelligence (AI) model. An electronic device (e.g., electronic device (101)) can generate a prompt.

[0047] Referring to FIG. 2, the electronic device (101) can obtain an input image (210). The input image (210) can be input to an AI model (250). The electronic device (101) can analyze the input image (210) through the AI ​​model (250). The AI ​​model (250) can be an AI model configured to generate a prompt (240) for image processing. The AI ​​model (250) can be configured to output the prompt (240) based on input data (e.g., input image (210), user input). The electronic device (101) can utilize the AI ​​model (250). For example, the electronic device (101) can input input data (e.g., input image (210), selected object) to the AI ​​model (250) implemented in the form of an on-device within the electronic device (101). The electronic device (101) can obtain output data (e.g., candidate prompt portions (220), detailed prompt portions (230)) corresponding to input data. For another example, the electronic device (101) can utilize an AI model (250) located in another electronic device (e.g., electronic device (102), electronic device (104), server (108)). The electronic device (101) can transmit input data (e.g., input image, selected object) to the other electronic device and receive results (e.g., candidate prompt portions (220), detailed prompt portions (230)) from the AI ​​model (250). For another example, the electronic device (101) can utilize an AI model implemented on-device and an AI model located externally together. Hereinafter, operations of the electronic device (101) are described based on the AI ​​model (250) of the electronic device (101), but embodiments of the present disclosure are not limited thereto.In one embodiment, at least a portion of the AI ​​model (250) is located in another electronic device (e.g., electronic device (102), electronic device (104), server (108)), and the electronic device (101) can receive data from the other electronic device.

[0048] An AI model (250) for generating a prompt (240) can be generated and / or updated through learning. Learning of the AI ​​model (250) can be performed on the electronic device (101) itself (e.g., the auxiliary processor (123) of FIG. 1 ) or can be performed through a separate server (e.g., the server (108)). In one embodiment, learning of the AI ​​model (250) can be performed based on user input received during the process of generating the input image (210), candidate prompt portions (220), and / or detailed prompt portions (230). For example, the AI ​​model (250) can identify the type of object or the method of image processing preferred by a user who desires image editing through statistics of the detailed prompt portions (230) related to the generation of the final prompt (240). Learning of the AI ​​model (250) can be performed through the identified type and method. As a non-limiting example, supervised learning of the AI ​​model (250) may be performed depending on whether the generated prompt (240) is actually input into the image AI model or the prompt is rewritten. In one embodiment, learning of the AI ​​model (250) may be performed based on image analysis. For example, if the resolution of the input image (250) or the resolution of a major object within the input image (250) is lower than a threshold (e.g., an average of previously input images, a resolution of an image for which resolution adjustment has been previously requested, a resolution determined by the type of the input image (250), the AI ​​model (250) may generate a prompt (240) that includes a command requesting an increase in resolution (e.g., text indicating upscaling). The threshold may be set through learning. For example, if the saturation or brightness of the input image (250) is higher than the threshold, the AI ​​model (250) may generate a prompt (240) that includes a command requesting a decrease in saturation or brightness. The above threshold can be set through learning.

[0049] The electronic device (101) can identify an object included in the input image (210). The electronic device (101) can identify a target object. The target object refers to an object that can be recognized as a major area within the input image (210). For example, the electronic device (101) can identify the target object from the input image (210) through an AI algorithm. For example, the electronic device (101) can identify an object that occupies the largest portion of the central area among the objects of the input image (210) as the target object. For example, the electronic device (101) can identify the target object among the objects of the input image (210) according to a category of the user's editing history. For example, if the user of the electronic device (101) has a greater history of editing flowers than other objects (e.g., people, backgrounds), the electronic device can identify the flower or an object related to the flower as the target object. The electronic device (101) can generate a segmentation area (or may be referred to as a masking area) associated with the target object. The segmentation area represents an area that includes the target object.

[0050] The electronic device (101) can identify one or more objects related to the target object. For example, the electronic device (101) can identify an object located within a certain range from the target object. For example, the electronic device (101) can identify an object of the same type as the target object (e.g., an object representing a person when the target object is a person). For example, the electronic device (101) can identify an object having similar properties (e.g., the same color series) to the target object. The electronic device (101) can generate a segmentation area associated with the object. The segmentation area represents an area including the object.

[0051] The electronic device (101) may generate a candidate prompt portion for the target object. For example, the candidate prompt portion may include text describing the target object. As another example, the candidate prompt portion may be an icon or image portion representing the target object. The electronic device (101) may generate a candidate prompt portion for an object related to the target object. For example, the candidate prompt portion may include text describing the target object. As another example, the candidate prompt portion may be an icon or image portion representing the target object. The electronic device (101) may generate candidate prompt portions (220) from the input image (210) through the AI ​​model (250). The candidate prompt portions (220) may be used to determine the prompt (240) to be ultimately generated. According to one embodiment, the candidate prompt portions (220) may represent editable objects. For example, the input image (210) may include a plurality of objects. The above-described plurality of objects may be objects that can be edited through image processing. The above-described plurality of objects may each correspond to candidate prompt portions (220). The electronic device (101) may display the candidate prompt portions (220). The electronic device (101) may display the candidate prompt portions (220) through a display (e.g., a display module (160)) to inform the user of the editable objects. The electronic device (101) may receive a user input that indicates at least one of the candidate prompt portions (220). For example, the user input may indicate a specific object in the input image (210). The electronic device (101) may input the candidate prompt portion corresponding to the user input into the AI ​​model (250).The electronic device (101) can generate detailed prompt portions (230) through the AI ​​model (250). The detailed prompt portions (230) can be associated with the specific object indicated by the user input.

[0052] The electronic device (101) may display detailed prompt portions (230). The detailed prompt portions (230) may indicate detailed objects (or detailed items) related to the specific object. For example, if the specific object is a person, the detailed objects may include earrings, clothes, a hat, and / or glasses. At least one of the detailed prompt portions (230) may indicate an editing action. For example, if the specific object is a person, the detailed prompt portion may indicate a change in the person's facial expression (e.g., smiling, crying, frowning). As another example, if the specific object is a tree, the detailed prompt portion may indicate a change in the color of the leaves (e.g., green, brown, yellow, ochre). For example, the electronic device (101) may display primary detailed prompt portions (230-1). The electronic device (101) can receive a user input for at least one of the first detailed prompt parts (230-1). Through the user input, the electronic device (101) can input information about a specific detailed item among the detailed items of a specific object into the AI ​​model (250). The electronic device (101) can obtain second detailed prompt parts for the specific detailed item from the AI ​​model (250). In this way, the AI ​​model (250) can also generate n-th detailed prompt parts (230-n) having a hierarchical structure in response to at least one user input.

[0053] The electronic device (101) can determine a detailed prompt portion in response to user inputs to guide editing in a desired direction by the user. The electronic device (101) can receive a user input for a detailed prompt portion indicating an editing action. The electronic device (101) can generate a prompt (240) to be ultimately used for image editing in response to the user input. The electronic device (101) can generate the prompt (240) through an AI model (250). For example, the electronic device (101) can input information about the detailed prompt portion indicated by the user input into the AI ​​model (250). The AI ​​model (250) can generate the prompt (240) using candidate prompt portions and detailed prompt portions (e.g., 1st detailed prompt portion, 2nd detailed prompt portion, ..., nth detailed prompt portion) selected by the user of the electronic device (101).

[0054] A prompt (240) generated through an AI model (250) may include commands to be input into an AI model (hereinafter, an image AI model) for generating an output image through image editing. For example, the prompt (240) may include not only content about an object, a subject, and / or a background to be input into the image AI model, but also a portion (e.g., text) regarding technical elements for image processing. According to one embodiment, a portion of the prompt (240) may include text commands for image processing (e.g., high dynamic range (HDR), color, lighting, resolution). The above command may be provided not only through direct text (e.g., HDR, color-#99FF66 (color corresponding to R of RGB is 153, G is 255, B is 102), resolution set to QHD), but also through predefined keywords for the image AI model (e.g., fhd - resolution control, qhd - resolution control, uhd - resolution control, HDR - HDR image processing, color_xxx - color control with xxx). For example, a part of the prompt (240) may include 'HDR'. The image AI model may provide an output image in which HDR processing is performed on the input image (210) through the part. For example, a part of the prompt (240) may include "eye_light_xx". The image AI model may provide an output image in which the brightness of the area including the eye in the input image (210) is 'xx' through the part. For example, a part of the prompt (240) may include "resolution_UHD". The image AI model can provide an output image having a resolution corresponding to UHD (e.g., 3840 x 2160) by adjusting the resolution of the input image (210) through the above part (e.g., upscaling).Additionally, according to one embodiment, a portion of the prompt (240) may include a command of an element representing a means of expression of an image (e.g., oil painting, illustration, photograph, 3D rendering). Depending on the means of expression indicated through the command, the image AI model may provide an output image by reconstructing the input image (210) in a style according to the indicated means of expression. The command may be provided not only through direct text (e.g., oil painting, illustration, photograph, 3D rendering), but also through predefined identification numbers for the image AI model (e.g., 1-oil painting, 2-illustration, 3-photography, 4-3D rendering), or keywords (e.g., oil-oil painting, ill-illustration, pho-photography, 3d-3D rendering).

[0055] In addition to the examples described above, an image on which high-quality image processing has been performed may be output. For example, the prompt (240) may include words such as 'HDR', 'UHD', and '64K'. For other examples, the prompt (240) may include 'highly detailed', 'studio lighting', 'professional', 'vivid color (or vivid intense color)', or 'bokeh'. Depending on the words included in the prompt (240), an output image may be provided. For example, a prompt (240) such as "The child is riding a horse and a man in a hat is helping the child. The child is smiling brightly. It's outdoors with mountains in the back. HDR, highly detailed, Studio Lighting, Professional." may be input to the image AI model. This may provide an output image that includes a child riding a horse and a man wearing a hat assisting the child, and that reflects the 'studio lighting' effect and the 'expert' effect, with high resolution and high detail processing. The 'studio lighting' effect and / or the 'expert' effect may represent image processing according to settings predefined in the image AI model.

[0056] Hereinafter, in various embodiments of the present disclosure, terms expressed as candidate prompt portions may be referred to as object keywords, key phrases, object identification words, object identification phrases, candidate portions, candidate prompt data, object prompt portions, first prompt portions, candidate wording portions, candidate contexts, first contexts, and / or equivalent technical terms in terms of selecting editable objects.

[0057] Hereinafter, terms expressed as detailed prompt portions in various embodiments of the present disclosure may be referred to as detailed candidate prompt portions, detailed keywords, detailed key phrases, detailed object identification words, detailed object identification phrases, detailed candidate portions, detailed prompt data, detailed object prompt portions, second prompt portions, detailed wording portions, detailed contexts, second contexts, and / or equivalent technical terms.

[0058] Hereinafter, in various embodiments of the present disclosure, terms expressed as descriptive text to indicate a description of an image may be referred to as, in addition to descriptive text, descriptive information, descriptive information, descriptive portion, text portion, descriptive portion, descriptive object, visual object, visual portion, descriptive object, and / or equivalent technical terms.

[0059] The processor (120) of the present disclosure may include various processing circuits and / or multiple processors. For example, the term "processor" as used herein, including in the claims, may include various processing circuits including at least one processor, one or more of which may be individually and / or collectively configured to perform the various functions(s) described in the present disclosure. As used herein, when "processor," "at least one processor," and "one or more processors" are described as being configured to perform various functions, these terms may include, for example, without limitation, situations where a single processor performs some of the recited functions, situations where other processor(s) perform other of the recited functions, situations where a single processor can perform all of the recited functions, and / or combinations of processors performing in a distributed manner. Additionally, instructions (or program commands) for various functions(s) in the present disclosure, when executed by a processor, may cause an electronic device (e.g., electronic device (101)) to execute the various functions(s).

[0060] FIG. 3 illustrates examples of candidate prompt portions (e.g., candidate prompt portions (220)) according to an input image (e.g., input image (210)).

[0061] Referring to FIG. 3, an electronic device (e.g., electronic device (101)) may obtain an input image (300). The input image (300) may include a background area (310). The input image (300) may include a plurality of objects. For example, the input image (300) may include a first object (321), a second object (322), a third object (323), a fourth object (324), a fifth object (325), and / or a sixth object (326). In order to easily edit the input image (300), it is required that a user easily recognize editable objects within the input image (300). The electronic device (101) may display candidate prompt portions corresponding to the editable objects. The candidate prompt portions may include a description, a name, and / or an image for each editable object. As the above candidate prompt portions are displayed, the user of the electronic device (101) can select an object to edit with just one input.

[0062] The electronic device (101) can identify a target object from the input image (300). For example, the electronic device (101) can identify a first object (321) as the target object. The electronic device (101) can generate a first candidate prompt portion (351) representing the first object (321). The electronic device (101) can display the first candidate prompt portion (351). For example, the first object (321) can be a child in the input image (300). The electronic device (101) can display the first candidate prompt portion (351) including the text of 'child on horseback' through a display (e.g., a display module (160)).

[0063] The electronic device (101) can identify one or more objects associated with the first object (321). For example, the electronic device (101) can identify objects (e.g., a second object (322) and a third object (323)) located within a certain range from the first object (321). The electronic device (101) can generate a second candidate prompt portion (352) indicating the second object (322). The electronic device (101) can generate a third candidate prompt portion (353) indicating the third object (323). The electronic device (101) can display the second candidate prompt portion (352). The electronic device (101) can display the third candidate prompt portion (353). For example, the second object (322) can be a horse in the input image (300). The third object (323) can be an adult in the input image (300). The electronic device (101) can display a second candidate prompt portion (352) including the text of 'horse' through a display (e.g., a display module (160)). The electronic device (101) can display a third candidate prompt portion (353) including the text of 'man wearing a hat' through a display (e.g., a display module (160)).

[0064] The input image (300) may further include other objects (e.g., a fourth object (324), a fifth object (325), and a sixth object (326)) in addition to the first object (321), the second object (322), and the third object (323). The electronic device (101) may display a button (354) to select one of the other objects or to request a candidate prompt portion for at least one of the other objects. In response to a user input to the button (354), the electronic device (101) may display additional candidate prompt portions in addition to the first candidate prompt portion (351), the second candidate prompt portion (352), and the third candidate prompt portion (353). For example, the electronic device (101) may first display a first candidate prompt portion (351) corresponding to a first object (321), which is a main object, and additionally display a second candidate prompt portion (352) corresponding to a second object (322) and a third candidate prompt portion (353) corresponding to a third object (323).

[0065] The electronic device (101) may display candidate prompt portions to inform the user of editable objects. The electronic device (101) may receive a user input indicating at least one of the candidate prompt portions. The electronic device (101) may generate detailed prompt portions indicating editable options based on an object (e.g., a first object (321)) identified by the user input. The detailed prompt portions are described with reference to FIG. 4. Meanwhile, the user of the electronic device (101) may wish to edit an object other than an object (e.g., a first object (321), a second object (322), a third object (323)) provided by the displayed candidate prompt portions (e.g., a first candidate prompt portion (351), a second candidate prompt portion (352), and a third candidate prompt portion (353)). For example, the electronic device (101) may receive a user input. The user input may indicate another object (e.g., a fourth object (324), a fifth object (325), or a sixth object (326)) within the input image (300). For example, the user input may include a text input, a multi-modal input, or a voice input indicating the other object. The electronic device (101) may generate detailed prompt portions indicating editable options based on the object (e.g., the fourth object (324)) identified by the user input.

[0066] Figure 4 illustrates examples of detailed prompt parts according to the selection of the candidate prompt part. Figure 4 describes a situation in which the first candidate prompt part (351) is selected among the candidate prompt parts of Figure 3.

[0067] Referring to FIG. 4, the electronic device (101) may display an input image (300) and a plurality of candidate prompt portions through a display (e.g., a display module (160)). For example, the plurality of candidate prompt portions may include a first candidate prompt portion (351) including the text of 'child on horseback', a second candidate prompt portion (352) including the text of 'horse', and a third candidate prompt portion (353) including the text of 'man in hat'. The electronic device (101) may receive a user input (410) on the first candidate prompt portion (351). Although FIG. 4 illustrates a user input (410) including a touch input for the first candidate prompt portion (351), embodiments of the present disclosure are not limited thereto. In addition to a touch input, an input for selecting the first candidate prompt portion (351) may be an input through a separate input means (e.g., a controller, a keyboard, a pen), a multi-modal input, and / or a voice input.

[0068] The electronic device (101) may, in response to a user input (410), identify an object region (e.g., an object region (400a)) including a first object (321) corresponding to a first candidate prompt portion (351). The object region (400a) may include a segmentation region corresponding to the first object (321). The segmentation region refers to a region corresponding to the first object (321) within the input image (300). The object region (400a) may be set as a region of interest (ROI). The electronic device (101) may perform an analysis on the object region (400a). Through the analysis, the electronic device (101) may generate detailed prompt portions representing one or more objects included in the object region (400a). The one or more objects may include detailed items of the first object (321). The electronic device (101) can display the above detailed prompt parts through a display (e.g., a display module (160)).

[0069] The electronic device (101) can magnify the object area (400a). The electronic device (101) can display an image (400b) in which the object area (400a) is magnified. The image (400b) can include an object area having a magnification that is greater than the magnification of the object area (400a) in the input image (300). The image (400b) can include details of the first object (321). For example, the image (400b) can include a child's face (421), a helmet (422), and / or clothes (423). The electronic device (101) can provide editable details of the first object (321) to enable detailed editing of the first object (321). The electronic device (101) can display detail prompt portions corresponding to the details. For example, the electronic device (101) may display a first detailed prompt portion (451) corresponding to the child's face (421). The electronic device (101) may display a second detailed prompt portion (452) corresponding to the helmet (422). The electronic device (101) may display a third detailed prompt portion (453) corresponding to the clothes (423).

[0070] In addition to indicating editable details for a specific object, the detailed prompt portion may include an editing action for the specific object. For example, the electronic device (101) may display a fourth detailed prompt portion (454). The fourth detailed prompt portion (454) may indicate image editing (e.g., highlighting) of the first object (321) (e.g., an eye). The fourth detailed prompt portion (454) may include text having "Emphasis only on the eye." While the "highlight" processing is described as an example in FIG. 4 , embodiments of the present disclosure are not limited thereto. Any detailed prompt portion indicating a function for image editing may be understood as an embodiment of the present disclosure. For example, if the specific object is a blurred object, the detailed prompt portion may indicate deblurring. If the specific object is the sky, the detailed prompt portion may indicate cloud removal or color correction. The editing action provided through the detailed prompt portion may vary depending on the type of object selected in the previous step. For a specific example of editing behavior, reference may be made to Figure 5.

[0071] FIG. 5 illustrates examples of editing actions. Referring to FIG. 5, in response to a user input for the first detailed prompt portion (451), the electronic device (101) may display candidate editing actions for the child's face (421) as the detailed prompt portion. For example, the candidate editing actions may include 'brighten the child's face' (551a), 'make the child's face clear' (551b), 'make the child's face bright' (551c), 'make the child's face smile' (551d), and / or 'make the child's face's eyes wide' (551e). In response to a user input for the second detailed prompt portion (452), the electronic device (101) may display candidate editing actions for the helmet (422) as the detailed prompt portion. For example, the candidate edit actions may include 'remove helmet' (552a), 'change helmet to hat' (552b), and / or 'change helmet color' (552c).

[0072] The electronic device (101) may receive user input for selecting other objects in addition to the presented details. For example, the electronic device (101) may receive user input for the sky (510) among the background areas of the input image (300). The electronic device (101) may display candidate edit actions for the sky (510) as a part of the detail prompt. The candidate edit actions may include 'make the sky blue' (511), 'add clouds to the sky' (512), and / or 'add birds to the sky' (513). For example, the electronic device (101) may display candidate edit actions for other people (520) (e.g., the fourth object (324)) in the input image (300). The candidate edit actions may include 'remove other people (521)'. For example, the electronic device (101) may display candidate edit actions for a background figure (530) (e.g., background area (310)) within an input image (300). The candidate edit actions may include 'blur background' (531) and / or 'change background to grassland' (532).

[0073] In this way, by displaying detailed items and / or editing actions after the first object (321), the electronic device (101) can enable the user to make step-by-step selections. In response to a user input indicating at least one of the detailed prompt portions, the electronic device (101) can perform the following actions. For example, the electronic device (101) can display texts of secondary detailed items that describe the selected detailed item in more detail. As another example, the electronic device (101) can perform the editing action indicated by the detailed prompt portion. For example, the electronic device (101) can perform image processing to emphasize the border of the eye of the first object (321) in the object area (400a) in the input image (300). Depending on the user's selection, the editing target can be made more specific or the prompt including the editing command can be updated. When an object or detailed item is selected, the selected area can be enlarged. Additionally, detailed prompt portions that are described in more detail can be displayed in the selected area. By repeating this hierarchical structure, the electronic device (101) can provide the user with more detailed image editing functions. For example, but not limited to, editing is possible not only through the user's input, but also through touch, voice, and / or separate text input.

[0074] The electronic device (101) can generate a prompt (e.g., prompt (240)), which is a final command, through the user's selections. The electronic device (101) can generate the prompt (240) including prompt parts based on the user's inputs through an AI model (e.g., AI model (250)). The electronic device (101) can obtain an output image through the input image (300), the prompt (240), and the area of ​​the object to be edited (e.g., the first object (321), and if the detailed item is determined, the detailed item (e.g., helmet (422)). The electronic device (101) can input the input image (300), the prompt (240), and the area of ​​the object into the image AI model. The electronic device (101) can obtain an output image from the image AI model. The output image can represent the result of image processing performed according to the prompt (240).

[0075] Figure 6 shows an example of a candidate prompt portion.

[0076] Referring to FIG. 6, the electronic device (101) can display a screen (600). The screen (600) can include an input image (300). The electronic device (101) can display the screen (600) through a display (e.g., a display module (160)). The input image (300) can include a plurality of objects. For example, the input image (300) can include a first object (321), a second object (322), a third object (323), a fourth object (324), a fifth object (325), and / or a sixth object (326). In order to easily edit the input image (300), it is required that the user can easily recognize editable objects within the input image (300).

[0077] The electronic device (101) may display candidate prompt portions corresponding to the editable objects. According to one embodiment, the electronic device (101) may display the candidate prompt portions in the form of sentences including multiple words. Each word may represent an object of the input image (300). The screen (600) may include a control area (640). The electronic device (101) may display candidate prompt portions composed of sentences on the control area (640). For example, the electronic device (101) may display the text 'A boy and a man riding a horse.' In the text, the word 'horse' may correspond to the first candidate prompt portion (651). The first candidate prompt portion (651) may indicate that it is an editable object through additional markings (e.g., underlining, bolding, a different color than the entire sentence). The first candidate prompt portion (651) may correspond to the second object (322). In the above text, the word 'child' may correspond to the second candidate prompt portion (652). The second candidate prompt portion (652) may indicate that it is an editable object through additional markings (e.g., underlining, bolding, a different color from the entire sentence). The second candidate prompt portion (652) may correspond to the first object (321). In the above text, the word 'male' may correspond to the third candidate prompt portion (653). The third candidate prompt portion (653) may indicate that it is an editable object through additional markings (e.g., underlining, bolding, a different color from the entire sentence). The third candidate prompt portion (653) may correspond to the third object (323).

[0078] A user of the electronic device (101) may wish to edit images of other objects (e.g., a fourth object (324), a fifth object (325), a sixth object (326)). The electronic device (101) may display a button (660) for adding objects within the control area (640). In response to a user input to the button (660), the electronic device (101) may display an additional candidate prompt portion (e.g., “chair in front of man” corresponding to the sixth object (326). The electronic device (101) may select an object to be edited not only through a simple touch input, but also through a separate text input or voice input. For example, the electronic device (101) may display an input field (630). The electronic device (101) may input an object to be edited (e.g., “person wearing glasses”) through text input on the input field (630). The electronic device (101) can identify the fifth object (325) in response to the text input. Although not illustrated in FIG. 6, the electronic device (101) can display detailed prompt portions corresponding to detailed items of the fifth object (325). The electronic device (101) can perform image editing on the fifth object (325) in response to a user input for at least one of the detailed prompt portions.

[0079] Figure 7 illustrates an example of image processing using an AI model. The AI ​​model may be distinguished from the AI ​​model for generating prompts (e.g., prompt (240)) described in Figures 2 through 6. The AI ​​model may be referred to as an image AI model in the context of image editing, rather than generating prompts (240). In one embodiment, the AI ​​model may include a generative AI model.

[0080] Referring to FIG. 7, the electronic device (101) can generate an output image (760) through an image AI model (750). The electronic device (101) can generate an output image (760) through the image AI model (750) that uses an input image (210) (or an input image (300)), a prompt (240), and a segmentation area (715) as input data. The electronic device (101) can generate a prompt (240) for image editing through user inputs, as described with reference to FIGS. 2 to 6. For example, the electronic device (101) can generate the prompt (240) through an AI model (250) that uses an input image (210) (or an input image (300)) as input data. In the process of selecting an object using the AI ​​model (250), the electronic device (101) can receive one or more user inputs. For example, the electronic device (101) may receive a user input indicating 'a child riding a horse' (e.g., a user input for the first candidate prompt portion (351)), a user input indicating 'a child's face' (e.g., a first detailed prompt portion (451)), and a user input for 'a child's face with a smiling expression' (551d), as shown in FIG. 4. The electronic device (101) may generate a prompt (240) including the prompt portions obtained through the user inputs. For example, the electronic device (101) may generate text such as 'Change the face of the child riding a horse to a smiling expression' as the prompt (240). The electronic device (101) may edit an image according to the prompt (240) generated for editing.

[0081] The electronic device (101) can specify an object to be edited within the input image (300) through the user inputs. The more detailed items are selected through the user input, the more detailed the editing target can be. The electronic device (101) can determine a segmentation area (715) on which image processing is to be performed through the user inputs. For example, the electronic device (101) can specify a first object (321) representing a child. In addition, for example, the electronic device (101) can specify a child's face (421) in the first object (321). The child's face (421) can be determined as the segmentation area (715). The image AI model (750) can perform image processing according to the prompt (240) on the segmentation area (715) of the input image (210).

[0082] Figures 8a to 8d illustrate examples of scenarios of image editing functions through interactive services.

[0083] Referring to FIG. 8A, the electronic device (101) can display a screen (800). The screen (800) can include an input image (300). The electronic device (101) can display the screen (800) through a display (e.g., a display module (160)). The electronic device (101) can display a description text (805) for the input image (300). For example, the electronic device (101) can identify a first object (321) ('child') as a target object in the input image (300). The electronic device (101) can generate a description text (805) associated with the first object (321). The electronic device (101) can display the description text (805) through a display (e.g., a display module (160)). For example, the description text (805) can indicate 'the child is riding a horse'.

[0084] An electronic device (101) can receive a user input. Depending on which object the user input indicates within the input image (300), the electronic device (101) can display a specific screen. The electronic device (101) can determine which object the user input is intended to target and / or which editing is desired, and display a screen according to the determined result. Hereinafter, according to one embodiment, when a user input for a target object (e.g., a first object (321)) is received, examples of a screen displayed next to the screen (800) are described through FIG. 8B. According to another embodiment, when a user input for another object (e.g., a fourth object (324)) is received, examples of a screen displayed next to the screen (800) are described through FIG. 8C. According to yet another embodiment, when a user input for an object related to the target object (e.g., a third object (323)) is received, examples of a screen displayed next to the screen (800) are described through FIG. 8D.

[0085] Referring to FIG. 8B, the electronic device (101) can display a screen (830a). On the screen (830a), the electronic device (101) can receive a user input (810) pointing to a first object (321). In response to the user input (810), the electronic device (101) can display a screen (830b). The screen (830b) can include a first object (321) with a larger magnification than the first object (321) on the screen (830a). The electronic device (101) can determine an object area displayed through the screen (830b) as a region of interest (ROI). The electronic device (101) can receive a user input (830) pointing to a face on the first object (321) on the screen (830b). The electronic device (101) may display an inquiry text (815) in response to a user input (830). The inquiry text (815) may indicate an editing option for the face of the first object (321). The electronic device (101) may display the inquiry text (815) through a display (e.g., a display module (160)). For example, the inquiry text (815) may indicate 'Shall I make your face smile?'. The electronic device (101) may receive a confirmation response or a rejection response for the inquiry text (815). For example, the electronic device (101) may receive a user input for the face of the first object (321) as the confirmation response. For example, the electronic device (101) may obtain a rejection response if there is no additional input for a certain period of time.

[0086] The electronic device (101) can display a screen (830c). The screen (830c) can be used to provide options for editing other images. The electronic device (101) can display descriptive text (835) corresponding to at least a portion of the inquiry text (815). The electronic device (101) can receive user input (840) on the inquiry text (815). In response to the user input (840), the electronic device (101) can display candidate prompt portions for editing. The candidate prompt portions can include a first candidate prompt portion (831), a second candidate prompt portion (832), a third candidate prompt portion (833), and a fourth candidate prompt portion (834). The first candidate prompt portion (831) can include the text 'Brightly'. The first candidate prompt portion (831) may represent an editing action to brighten the first object (321) of the input image (300). The second candidate prompt portion (832) may include the text 'clearly'. The second candidate prompt portion (832) may represent an editing action to sharpen the first object (321) of the input image (300). The third candidate prompt portion (833) may include the text 'with a smiling expression'. The third candidate prompt portion (833) may represent an editing action to change the face of a child corresponding to the first object (321) of the input image (300) to a smiling expression. The fourth candidate prompt portion (834) may include the text 'Tell me details'. The fourth candidate prompt portion (834) may be used to request additional image editing options that are different from the first candidate prompt portion (831), the second candidate prompt portion (832), and the third candidate prompt portion (833). The electronic device (101) can generate a prompt (240) for image editing based on user input indicating at least one of the options described above.

[0087] Referring to FIG. 8C, the electronic device (101) may display a screen (860). The screen (860) may include an input image (300). On the screen (860), the electronic device (101) may receive a user input (850) pointing to a fourth object (324). In response to the user input (850), the electronic device (101) may display an inquiry text (855). The inquiry text (855) may indicate an editing option for the fourth object (324). The electronic device (101) may display the inquiry text (855) through a display (e.g., a display module (160)). For example, the inquiry text (855) may indicate 'Do you want to delete this person?' Since the fourth object (324) has a low correlation with the first object (321), which is the target object, the electronic device (101) may determine that the user's intention is to delete the fourth object (324). When a confirmation response to the inquiry text (855) is received, the electronic device (101) may generate a prompt for deleting the fourth object (324). The electronic device (101) may obtain an output image through an image AI model (750) that uses the prompt, the segmentation area corresponding to the fourth object (324), and the input image (300) as input data. The output image may not include the fourth object (324).

[0088] Referring to FIG. 8D, the electronic device (101) may display a screen (890a). The screen (890a) may include an input image (300). The electronic device (101) may receive a user input (871) regarding an object (e.g., a third object (323)) having a relatively low priority among objects included in a certain area in the input image (300). In response to the user input (871), the electronic device (101) may display a description text (875). For example, the description text (875) may indicate, 'A teacher wearing a hat is helping a child.' The description text (875) may include a candidate prompt portion (e.g., the word 'hat'). For example, the candidate prompt portion may indicate that it is an editable object through additional indication (e.g., underlining, bolding, a different color from the entire sentence). The electronic device (101) can receive user input (872) for the candidate prompt portion. In response to the user input (871), the electronic device (101) can display a screen (890b).

[0089] The screen (890b) may represent an object area including a hat of a third object (323). The electronic device (101) may set the object area as a region of interest. The electronic device (101) may adjust the magnification so that the object area is enlarged compared to the input image (300). The size of the hat of the third object (323) in the screen (890b) may be larger than the size of the hat of the third object (323) in the screen (890a). The screen (890b) may be used to provide options for editing the image. The electronic device (101) may display description text (885) corresponding to at least a portion of description text (875). The electronic device (101) may receive a user input (880) on the description text (885). In response to the user input (880), the electronic device (101) may display candidate prompt portions for editing. The above candidate prompt portions may include a first candidate prompt portion (881), a second candidate prompt portion (882), a third candidate prompt portion (883), and a fourth candidate prompt portion (884). The first candidate prompt portion (881) may include the text 'color...'. The first candidate prompt portion (881) may indicate an editing action of changing the color of a hat of a third object (323) of an input image (300). The second candidate prompt portion (882) may include the text 'sharply'. The second candidate prompt portion (882) may indicate an editing action of sharpening a hat of a third object (323) of an input image (300). The third candidate prompt portion (883) may include the text 'please erase'. The third candidate prompt portion (883) may indicate an editing action of removing a hat of a third object (323) of an input image (300). The fourth candidate prompt portion (884) may include the text 'Tell me details'.The fourth candidate prompt portion (884) may be used to request additional image editing options, different from the first candidate prompt portion (881), the second candidate prompt portion (882), and the third candidate prompt portion (883). The electronic device (101) may generate a prompt (240) for image editing based on a user input indicating at least one of the above-described options.

[0090] Although FIG. 8d illustrates an example in which a screen (890b) is displayed in response to a user input (872), embodiments of the present disclosure are not limited thereto. For example, in response to a user input (871), the electronic device (101) may also display a screen (890b).

[0091] FIG. 9 illustrates an operation flow of an electronic device (e.g., electronic device (101)) for editing an image using a prompt (e.g., prompt (240)) generated by an AI model (e.g., AI model (250)).

[0092] Referring to FIG. 9, in operation (901), the electronic device (101) (e.g., the processor (120)) can identify an object and one or more objects in an input image (e.g., the input image (210), the input image (300)). The electronic device (101) can identify the object (e.g., the first object (321)). The object may be a target object. The target object represents an object that can be recognized as a major area in the input image. For example, the electronic device (101) can identify the target object from the input image through an AI algorithm. For example, the electronic device (101) can identify an object that occupies the largest central area among the objects of the input image as the target object. For example, the electronic device (101) can identify the target object among the objects of the input image according to a category of a user's editing history. The electronic device (101) can identify one or more objects (e.g., a second object (322), a third object (323)) related to the target object. For example, the electronic device (101) can identify an object located within a certain range from the target object. For example, the electronic device (101) can identify an object of the same type as the target object (e.g., an object representing a person when the target object is a person). For example, the electronic device (101) can identify an object having similar properties (e.g., the same color series) to the target object.

[0093] In operation (903), the electronic device (101) (e.g., the processor (120)) may generate candidate prompt portions through an AI model. For example, the candidate prompt portion may include text for describing an object. As another example, the candidate prompt portion may be an icon or image portion representing an object. The electronic device (101) may generate candidate prompt portions (220) from an input image (210) through an AI model (e.g., the AI ​​model (250)). The candidate prompt portions may be used to indicate objects that are editable through image processing.

[0094] In operation (905), the electronic device (101) (e.g., processor (120)) may display candidate prompt portions. The electronic device (101) may display the candidate prompt portions through a display (e.g., display module (160)) to inform the user of editable objects.

[0095] In operation (907), the electronic device (101) (e.g., the processor (120)) may generate a prompt for image editing based on a user input. The electronic device (101) may receive a user input indicating at least one of the candidate prompt portions. For example, the user input may indicate a specific object in the input image. The electronic device (101) may input the candidate prompt portion corresponding to the user input into the AI ​​model (250). The electronic device (101) may generate detailed prompt portions through the AI ​​model (250). The detailed prompt portions (230) may be associated with the specific object indicated by the user input. Each of the detailed prompt portions may indicate a detailed item related to the specific object or an editing action for the specific object. The electronic device (101) may determine the detailed prompt portions in response to the user inputs to guide the user to edit in a desired direction. When a sub-item is selected, the electronic device (101) may display additional sub-items for the sub-item or indicate an editing action for the sub-item. In this manner, the electronic device (101) may provide the user with more detailed image editing functions. When a user input corresponding to the editing action is received, the electronic device (101) may generate a prompt (e.g., prompt (240)). The electronic device (101) may generate the prompt (240) including prompt portions based on user inputs through an AI model (e.g., AI model (250)).

[0096] In operation (909), the electronic device (101) (e.g., the processor (120)) can generate an output image. The electronic device (101) can generate the output image (760) through the image AI model (750). The electronic device (101) can generate the output image (760) through the image AI model (750) using the input image (210) (or the input image (300)), the prompt (240), and the segmentation area (715) as input data.

[0097] FIG. 10 illustrates an operational flow of an electronic device (e.g., electronic device (101)) for generating a prompt (e.g., prompt (240)) for image editing. The descriptions of FIG. 10 may be referred to as detailed operations of operation (907) of FIG. 9.

[0098] Referring to FIG. 10, in operation (1001), the electronic device (101) may generate first detailed prompt portions in response to a user input indicating a first candidate prompt portion representing a first object. The electronic device (101) may display a plurality of candidate prompt portions. The plurality of candidate prompt portions may represent editable objects within an input image (e.g., input image (300)). The electronic device (101) may receive a user input indicating a first object among the editable objects. The electronic device (101) may identify the first object in response to the user input. The electronic device (101) may determine a segmentation area corresponding to the first object in response to the user input. The electronic device (101) may determine detailed items for the first object. The electronic device (101) may generate the first detailed prompt portions to display the detailed items to the user. For example, if the first object is 'a boy wearing a hat and riding a horse', the details may include the hat, the horse, the clothes, and the child's facial expression.

[0099] In operation (1003), the electronic device (101) (e.g., the processor (120)) may display first detailed prompt portions. By displaying the first detailed prompt portions, the electronic device (101) may guide the user to perform detailed editing of the first object.

[0100] In operation (1005), the electronic device (101) (e.g., the processor (120)) may receive an input indicating a designated detailed prompt portion. The electronic device (101) may receive an input indicating the designated detailed prompt portion among the first detailed prompt portions. The input may include a touch input, a voice input, and / or a text input. The designated detailed prompt portion may indicate a specific detailed item (e.g., a hat) among the detailed items of the first object.

[0101] In operation (1007), the electronic device (101) (e.g., the processor (120)) may generate a prompt (e.g., prompt (240)) including a first candidate prompt portion and a designated detailed prompt portion. The first candidate prompt portion may be used to select the first object. The designated detailed prompt portion may be used to select a specific detailed item (e.g., a hat) among detailed items of the first object. An AI model (e.g., AI model (250)) of the electronic device (101) may generate the prompt (240) as a final command including the first candidate prompt portion and the designated detailed prompt portion. The prompt (240) may include information about a specific object or a portion of a specific object and an editing action corresponding to user inputs. For example, the prompt (240) may be 'Change the color of the hat of a boy riding a horse wearing a hat to green.'

[0102] FIG. 11 illustrates an operation flow of an electronic device (e.g., electronic device (101)) for editing an image including an object and a background.

[0103] Referring to FIG. 11, in operation (1101), the electronic device (101) (e.g., processor (120)) can distinguish between objects and backgrounds by analyzing an input image (e.g., input image (300)). For example, the electronic device (101) can identify a background area (310) and objects (e.g., first object (321), second object (322), third object (323), fourth object (324), fifth object (325), and sixth object (326)) of the input image (300).

[0104] In operation (1103), the electronic device (101) (e.g., the processor (120)) may generate a descriptive text for the input image (300). The descriptive text may include a plurality of words. At least one of the plurality of words may indicate at least one object included in the input image (300). The at least one object may indicate an editable object. The at least one word may be used to generate a prompt. The electronic device (101) may perform a separate display (e.g., underlining, bolding, a different color from the entire sentence) within the descriptive text to indicate that the word is an editable object. The word indicating the editable object may be a candidate prompt portion.

[0105] In operation (1105), the electronic device (101) (e.g., processor (120)) may receive user input. The user input may include a touch input on a screen, a text input via a separate input means, and / or a voice input.

[0106] In operation (1107), the electronic device (101) (e.g., the processor (120)) may determine whether the user input is a user input for the description text. If the user input indicates a word of the description text, the electronic device (101) may determine that the user input is a user input for the description text. The electronic device (101) may perform operation (1109). If the user input does not indicate any of the words of the description text, the electronic device (101) may determine that the user input is not a user input for the description text. The electronic device (101) may perform operation (1111).

[0107] In operation (1109), the electronic device (101) (e.g., the processor (120)) may generate a prompt for a segmentation area based on a user input. The user input may indicate a word in the description text. The electronic device (101) may identify an object indicated by the word. The electronic device (101) may determine a segmentation area occupied by the object within the input image (300). The electronic device (101) may generate a prompt for editing the segmentation area. For example, the electronic device (101) may collect candidate prompt portions selected by user inputs, as illustrated in FIGS. 8A to 8D . The electronic device (101) may generate a prompt (e.g., prompt (240)) using an AI model (e.g., AI model (250)) that uses the collected candidate prompt portions as input data.

[0108] In operation (1111), the electronic device (101) (e.g., the processor (120)) may generate a prompt for background image quality improvement. If the user input does not indicate an object mentioned through the description text, the electronic device (101) may perform image processing on the background identified in operation (1101). The electronic device (101) may generate a prompt for the image processing. For example, the image processing may include at least one of color adjustment, contrast adjustment, resolution adjustment, noise improvement, color tone adjustment, detail enhancement, or brightness adjustment. For example, the electronic device (101) may generate a prompt (e.g., prompt (240)) through an AI model (e.g., AI model (250)) that uses a background area (310) of an input image (300) and a candidate prompt portion according to a user input as input data.

[0109] In operation (1113), the electronic device (101) (e.g., processor (120)) may receive a confirmation input. The electronic device (101) may obtain a prompt (240) for image processing. The electronic device (101) may provide a message to the user asking whether to execute the prompt (240). For example, the electronic device (101) may display a message asking whether to execute the prompt (240) through a display (e.g., display module (160)). For example, the electronic device (101) may provide guidance asking whether to execute the prompt (240) through voice.

[0110] In operation (1115), the electronic device (101) (e.g., the processor (120)) may generate an image edit according to a prompt (e.g., the prompt (240)). When the confirmation input is received, the electronic device (101) may input the prompt (240) to an image AI model (e.g., the image AI model (750)). For example, the electronic device (101) may obtain an output image through an image AI model (e.g., the image AI model (750)) that uses an input image (300), a segmentation area of ​​an object indicated by a user input, and the prompt (240) of operation (1109) as input data. For example, the electronic device (101) may obtain an output image through an image AI model (e.g., the image AI model (750)) that uses an input image (300), a background area (310), and the prompt (240) of operation (1111) as input data.

[0111] The present disclosure describes the operations of an electronic device (101). The electronic device (101) may not only be used to generate prompts or edit images on a mobile device, but may also be an augmented reality (AR) glasses, a head-mounted device (HMD), and / or a computer device for implementing a program. For example, generating a prompt for editing an image displayed through AR glasses in real time based on user input may also be understood as an embodiment of the present disclosure.

[0112] In embodiments, an electronic device is provided. The electronic device may include a display, a processor including processing circuitry, and a memory storing instructions. The instructions, when executed by the processor, may cause the electronic device to identify an object in an input image and one or more objects associated with the object, generate candidate prompt portions for the object and the one or more objects through an artificial intelligence (AI) model, display the candidate prompt portions and the input image through the display, generate a prompt for image editing based on a user input for at least one of the candidate prompt portions, and perform image editing on the input image according to the generated prompt through an image AI model, thereby generating an output image. The candidate prompt portions may be used to indicate editable objects in the input image.

[0113] In one embodiment, each of the candidate prompt portions may include text for a corresponding object among the object and the one or more objects.

[0114] In one embodiment, the instructions, when individually or collectively executed by the processor, may cause the electronic device to generate first detailed prompt portions for the first object and display the first detailed prompt portions via the display in response to a user input indicating a first candidate prompt portion for the first object from among the candidate prompt portions. Each of the first detailed prompt portions may represent an editable object within a first object area of ​​the first object or may represent an editing action for the first object.

[0115] In one embodiment, the first detailed prompt portions may include text representing an editable object within the first object area, an image representing an editable object, or an image and text representing an editable object.

[0116] In one embodiment, the instructions, when individually or collectively executed by the processor, may cause the electronic device to identify a first object region corresponding to the first object and display an image including the first object region through the display, the image having a magnification greater than that of the input image. The first detailed prompt portions may be displayed through the display while the image is being displayed.

[0117] In one embodiment, the instructions, when individually or collectively executed by the processor, may cause the electronic device to receive an input pointing to a designated sub-prompt portion from among the first candidate prompt portions and to generate the prompt including the first candidate prompt and the designated sub-prompt portion.

[0118] In one embodiment, the detailed prompt portions may include image processing related to a human facial expression when the first object is a person. The detailed prompt portions may include image processing related to the color of the object when the first object is an item.

[0119] In one embodiment, the instructions, when individually or collectively executed by the processor, may cause the electronic device to receive a second user input with respect to the object and a second object different from the one or more objects within the input image, generate second detailed prompt portions associated with the second object in response to the second user input, and display the second detailed prompt portions via the display. Each of the second detailed prompt portions may represent an editable object within a second object region of the second object or may represent an editing action with respect to the second object.

[0120] In one embodiment, the instructions, when individually or collectively executed by the processor, may cause the electronic device to identify a second object region corresponding to the second object and display an image including the second object region through the display, the image having a magnification greater than that of the input image. The second detailed prompt portions may be displayed through the display while the image is being displayed.

[0121] In one embodiment, the second detailed prompt portions may include text representing an editable object within the second object area, an image representing an editable object, or an image and text representing an editable object.

[0122] In one embodiment, the user input may represent an editable object within the input image. The image AI model may use the generated prompt, the input image, and the segmentation region of the editable object as input data to generate the output image.

[0123] In one embodiment, the instructions, when individually or collectively executed by the processor, may cause the electronic device to identify a background region of the input image, obtain user input for the background region of the input image, generate a candidate prompt for image processing in the background region, display a message through the display inquiring whether to execute the candidate prompt, and perform the image processing in response to user input indicating execution of the candidate prompt.

[0124] In one embodiment, the image processing in the background area may include at least one of color adjustment, contrast adjustment, or brightness adjustment.

[0125] In one embodiment, the instructions, when individually or collectively executed by the processor, may cause the electronic device to obtain a user input comprising at least one of a voice input, a text input, or a touch input, generate an editing prompt corresponding to the user input, and perform image processing on the input image in accordance with the editing prompt, thereby generating a second output image.

[0126] In one embodiment, the AI ​​model may include a generative AI model.

[0127] In one embodiment, the image AI model may include a generative AI model.

[0128] In embodiments, an electronic device is provided. The electronic device may include a display, a processor including a processing circuit, and a memory storing instructions. The instructions, when individually or collectively executed by the processor, may cause the electronic device to generate text for an input image through an artificial intelligence (AI) model, display the input image and the text through the display, identify an object corresponding to the target candidate prompt portion among candidate prompt portions of the text in response to a first user input for a target candidate prompt portion, generate detailed prompt portions associated with the object through the AI ​​model, display the detailed prompt portions associated with the object through the display, and generate a prompt for image editing based on a second user input for indicating an execution prompt among the detailed prompt portions. The candidate prompt portions of the text may respectively correspond to editable objects within the input image.

[0129] In one embodiment, each of the detailed prompt portions may represent an editable object within an object area containing the object or an edit action on the object.

[0130] In one embodiment, the instructions, when individually or collectively executed by the processor, may cause the electronic device to generate an output image according to the image editing via an image AI model that uses the generated prompt, the input image, and the segmentation region of the object as input data.

[0131] In one embodiment, the instructions, when individually or collectively executed by the processor, may cause the electronic device to identify a background region of the input image, obtain user input for the background region of the input image, generate a candidate prompt for image processing in the background region, display a message through the display inquiring whether to execute the candidate prompt, and perform the image processing in response to user input indicating execution of the candidate prompt.

[0132] In one embodiment, the image processing in the background area may include at least one of color adjustment, contrast adjustment, or brightness adjustment.

[0133] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, electronic devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0134] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0135] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0136] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0137] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0138] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In electronic devices, display; A processor comprising a processing circuit, and A memory for storing instructions, wherein the instructions, when individually or collectively executed by the processor, cause the electronic device to: Identifying an object in an input image and one or more objects associated with said object, Generate candidate prompt parts for the object and one or more of the objects through an AI (artificial intelligence) model, Displaying the above candidate prompt portions and the above input image through the above display, Generate a prompt for image editing based on user input for at least one of the above candidate prompt parts, By performing image editing on the input image according to the generated prompt through the image AI model, an output image is generated, The above candidate prompt parts are used to indicate editable objects in the input image. Electronic devices.

2. In claim 1, Each of the above candidate prompt portions includes text for the object and one or more of the objects. Electronic devices.

3. In claim 1, the instructions, when individually or collectively executed by the processor, cause the electronic device to: In response to a user input pointing to a first candidate prompt portion for a first object among the above candidate prompt portions, generating first detailed prompt portions for the first object, Causing the above first detailed prompt portions to be displayed through the above display, Each of the first detailed prompt portions above represents an editable object within the first object area of ​​the first object or represents an editing action for the first object. Electronic devices.

4. In claim 3, The first detailed prompt portions include text representing an editable object, an image representing an editable object, or an image and text representing an editable object within the first object area. Electronic devices.

5. In claim 3, the instructions, when individually or collectively executed by the processor, cause the electronic device to: Identify the first object area corresponding to the first object, Causing an image including the first object area to be displayed through the display, the image having a magnification greater than that of the input image; The above first detailed prompt portions are displayed through the display while the image is displayed. Electronic devices.

6. In claim 3, the instructions, when individually or collectively executed by the processor, cause the electronic device to: Receiving an input pointing to a specified detailed prompt part among the first detailed prompt parts above, causing said prompt to be generated, said prompt including said first candidate prompt and said specified detailed prompt portion; Electronic devices.

7. In claim 3, The above detailed prompt parts include image processing related to the human facial expression when the first object is a human. The above detailed prompt parts include image processing related to the color of the item, if the first object is an item. Electronic devices.

8. In claim 1, the instructions, when individually or collectively executed by the processor, cause the electronic device to: Receiving a second user input for the object and a second object other than the one or more objects within the input image, In response to said second user input, generate second detailed prompt portions associated with said second object, Causing the second detailed prompt portions above to be displayed through the display, Each of the second detailed prompt portions above represents an editable object within the second object area of ​​the second object or represents an editing action for the second object. Electronic devices.

9. In claim 8, the instructions, when individually or collectively executed by the processor, cause the electronic device to: Identify the second object area corresponding to the second object, Causing an image including the second object area to be displayed through the display, the image having a magnification greater than that of the input image; The above second detailed prompt portions are displayed through the display while the image is displayed. Electronic devices.

10. In claim 8, The second detailed prompt portions include text representing an editable object, an image representing an editable object, or an image and text representing an editable object within the second object area. Electronic devices.

11. In claim 1, The above user input represents an editing object within the input image, The above image AI model uses the generated prompt, the input image, and the segmentation area of ​​the edited object as input data to generate the output image. Electronic devices.

12. In claim 1, the instructions, when individually or collectively executed by the processor, cause the electronic device to: Identify the background area of ​​the input image above, Obtaining user input for the background area of ​​the input image, Generate candidate prompts for image processing in the above background area, A message is displayed through the display asking whether to execute the above candidate prompt, In response to user input indicating execution of the above candidate prompt, causing said image processing to be performed, Electronic devices.

13. In claim 12, The image processing in the above background area comprises at least one of color adjustment, contrast adjustment, or brightness adjustment. Electronic devices.

14. In claim 1, the instructions, when individually or collectively executed by the processor, cause the electronic device to: Obtaining user input including at least one of voice input, text input, or touch input, Generate an editing prompt corresponding to the above user input, According to the above editing prompt, by performing image processing on the input image, a second output image is generated. Electronic devices.

15. In claim 1, The above image AI model includes a generative AI model. Electronic devices.