Electronic device and method for generating image using artificial intelligence model, and non-transitory storage medium

The electronic device employs an image classification and generative AI model to identify and remove image portions related to people or animals, addressing the challenge of unintended identities in image generation, resulting in more natural and intended image outcomes.

WO2025146998A1PCT designated stage expired Publication Date: 2025-07-10SAMSUNG ELECTRONICS CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2024/020862
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-11
Filing Date
2024-12-20
Publication Date
2025-07-10

AI Technical Summary

Technical Problem

Existing image generation technologies using generative AI models struggle to accurately handle the removal of specific image portions, particularly those containing people or animals, often resulting in the generation of unintended or generalized identities, leading to unnatural images.

Method used

An electronic device equipped with an image classification model, preprocessing model, and generative AI model is used to identify and remove specific image portions related to people or animals, applying inpainting or outpainting techniques to generate images where these areas are not related to individuals or animals, utilizing user input and AI models to ensure natural image continuity.

Benefits of technology

The solution effectively prevents the generation of unintended identities in image portions, ensuring a more natural and user-intended image outcome by accurately removing and replacing image areas related to people or animals, enhancing the quality of image generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024020862_10072025_PF_FP_ABST
    Figure KR2024020862_10072025_PF_FP_ABST
Patent Text Reader

Abstract

The present document relates to an electronic device and a method for generating an image using an artificial intelligence model, and a non-transitory storage medium. The electronic device, according to one embodiment, may comprise: a display; a communication circuit; at least one processor comprising a processing circuit; and a memory comprising one or more storage media for storing instructions. The instructions, when executed individually or collectively by the at least one processor, may instruct the electronic device to: by using a user interface containing a first image, acquire a user input requesting for the removal of a first portion in the first image stored in the memory; on the basis of confirming that the first portion is associated with a person, acquire a second image having the first portion excluded from the first image, and first information related to the first portion, wherein the first information includes information for preventing an image area corresponding to the first portion from being generated into the person; and, on the basis of a first command enabling at least one of in-painting or out-painting to be performed on the basis of the second image and the first information by using a generative AI model, perform at least one of the in-painting or the out-painting so as to acquire a third image generated by the generative AI model such that the image area corresponding to the first portion is not associated with the person. Other various embodiments are also possible.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device, method and non-transitory storage medium for generating images using artificial intelligence models

[0001] The present disclosure relates to an electronic device, method and non-transitory storage medium for generating an image using an artificial intelligence model.

[0002] With the advancement of digital technology, electronic devices are available in various forms, such as smartphones, tablet personal computers (PCs), and personal digital assistants (PDAs). Electronic devices are also being developed into wearable forms to enhance portability and accessibility. Recently, users are increasingly interested in acquiring high-quality images, not just by taking pictures with their electronic devices, but also by capturing them with high-quality cameras. Techniques for generating images using artificial intelligence (generative AI models) are actively being developed. Electronic devices can provide users with an environment where they can create images stored on their devices using image-generating applications. Among these image-generating techniques, in-painting, which creates deleted portions of an image and fills them in naturally, and out-painting, which fills in the exterior of an input image, are being developed. In the case of current in-painting and out-painting, natural images are created based on the input images, and the final result is an image that does not feel out of place to the viewer.

[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.

[0004] According to one embodiment of the present disclosure, an electronic device may include a display, a communication circuit, at least one processor including a processing circuit, and a memory including one or more storage media storing instructions.

[0005] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to obtain a user input requesting removal of a first portion of a first image stored in the memory using a user interface including the first image.

[0006] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to determine, based on the user input, whether the first portion relates to a person.

[0007] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to obtain a second image excluding the first portion from the first image and first information related to the first portion based on determining that the first portion is related to the person, wherein the first information includes information for preventing an image area corresponding to the first portion from being created as the person.

[0008] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to perform at least one of inpainting or outpainting based on the second image and the first information using a generative AI model, thereby obtaining a third image generated by the generative AI model such that the image area corresponding to the first portion is not related to the person by performing at least one of the inpainting or the outpainting.

[0009] According to one embodiment, a method of operating in an electronic device includes obtaining, using a user interface, a user input requesting removal of a first portion of a first image stored in a memory of the electronic device.

[0010] According to one embodiment, the method includes an action of determining, based on the user input, whether the first portion relates to a person.

[0011] According to one embodiment, the method includes an action of determining, based on the user input, whether the first portion relates to a person.

[0012] According to one embodiment, the method includes an operation of obtaining a second image excluding the first portion from the first image and first information related to the first portion based on determining that the first portion is related to the person, wherein the first information includes information for preventing an image area corresponding to the first portion from being created as the person.

[0013] According to one embodiment, the method includes an operation of obtaining a third image generated by the generative AI model based on a first command to perform at least one of inpainting or outpainting based on the second image and the first information, by performing at least one of the inpainting or the outpainting such that the image area corresponding to the first portion is not related to the person.

[0014] According to one embodiment, a non-transitory storage medium storing a program includes executable instructions that, when executed by at least one processor of an electronic device, cause the electronic device to perform an operation of obtaining, using a user interface, a user input requesting removal of a first portion of a first image stored in a memory of the electronic device.

[0015] According to one embodiment, a non-transitory storage medium storing a program includes executable instructions that, when executed by at least one processor of an electronic device, cause the electronic device to perform an operation of determining, based on the user input, whether the first portion relates to a person.

[0016] According to one embodiment, in a non-transitory storage medium storing a program, the program includes executable instructions that, when executed by at least one processor of an electronic device, cause the electronic device to perform an operation of obtaining a second image excluding the first portion from the first image and first information related to the first portion based on determining that the first portion is related to the person, wherein the first information includes information for preventing an image area corresponding to the first portion from being created as the person.

[0017] According to one embodiment, in a non-transitory storage medium storing a program, the program includes instructions that, when executed by at least one processor of an electronic device, cause the electronic device to perform at least one of inpainting or outpainting based on the second image and the first information using a generative AI model, and perform at least one of the inpainting or the outpainting to obtain a third image generated by the generative AI model such that the image area corresponding to the first portion is not related to the person.

[0018] A computer-readable non-transitory recording medium according to one embodiment of the present disclosure may store at least one command and / or instructions that, when executed, cause an electronic device to perform the method or operation of the electronic device described above.

[0019] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.

[0020] FIG. 2 is a diagram showing the configuration of an electronic device according to one embodiment.

[0021] FIG. 3 is a diagram illustrating an example of an image generation model of an electronic device according to one embodiment.

[0022] FIGS. 4A and 4B are diagrams illustrating examples of generating an image using an image generation model in an electronic device according to one embodiment.

[0023] FIG. 5 is a diagram illustrating an example of image generation using an image generation model in an electronic device according to one embodiment.

[0024] FIG. 6 is a diagram illustrating an example of image generation using an image generation model in an electronic device according to one embodiment.

[0025] FIGS. 7A and 7B are diagrams illustrating examples of image generation using an image generation model in an electronic device according to one embodiment.

[0026] FIG. 8 is a drawing showing an example of an operating method in an electronic device according to one embodiment.

[0027] FIG. 9 is a drawing showing an example of an operating method in an electronic device according to one embodiment.

[0028] FIG. 10 is a diagram illustrating an example of an operation method for image generation using an image generation model in an electronic device according to one embodiment.

[0029] FIGS. 11A and 11C are drawings illustrating examples of images by in-painting in an electronic device according to one embodiment.

[0030] FIG. 12 is a drawing showing an example of an operating method in an electronic device according to one embodiment.

[0031] FIG. 13 is a diagram illustrating an example of an operation method for image generation using an image generation model in an electronic device according to one embodiment.

[0032] FIG. 14 is a diagram illustrating examples of images by outpainting in an electronic device according to one embodiment.

[0033] In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components.

[0034] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components. In addition, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness. The term "user" used in the embodiments of the present disclosure may refer to a person using an electronic device or a device (e.g., an artificial intelligence electronic device) using an electronic device.

[0035] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments.

[0036] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0037] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor)) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0038] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0039] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0040] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0041] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0042] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0043] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0044] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0045] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0046] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0047] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0048] The haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0049] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0050] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).

[0051] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0052] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0053] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for realizing 1eMBB, a loss coverage (e.g., 164 dB or less) for realizing mMTC, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for realizing URLLC.

[0054] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0055] According to various embodiments, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.

[0056] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0057] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In one embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0058] When objects such as people or animals are cropped at their boundaries during in-painting and out-painting, this can result in undesirable results, as objects with new identities are created. Images must be regenerated based on the cropped object information. However, AI technologies (e.g., generative AI models) tend to generate generalized faces of people or animals based on previously learned faces, as not all people or animals have identity information. This may result in the user not generating the exact person or animal they envisioned.

[0059] Hereinafter, with reference to the drawings, an electronic device and method for generating an image using artificial intelligence technology (generative AI model) to generate a more natural image without creating objects of new people or animals that the user does not want or that are meaningless during in-painting and out-painting will be specifically described.

[0060] FIG. 2 is a diagram showing a configuration of an electronic device according to one embodiment, FIG. 3 is a diagram showing an example of an image generation model of an electronic device according to one embodiment, and FIGS. 4A and 4B are diagrams showing an example for generating an image using an image generation model in an electronic device according to one embodiment. FIG. 4A (a) shows image generation in an electronic device according to a comparative example, and FIG. 4A (b) can show an example for generating an image using an image generation model in an electronic device according to one embodiment of the present disclosure.

[0061] Referring to FIGS. 2, 3, 4a, and 4b, an electronic device (201) according to one embodiment (e.g., the electronic device (101) of FIG. 1) may include at least one processor (210), a memory (220), a display (230), a camera (240), and a communication circuit (250). Without being limited thereto, the electronic device (201) may be implemented in the same or similar manner as the electronic device (101) of FIG. 1, and may further include other components of the electronic device (101) described in FIG. 1.

[0062] According to one embodiment, the electronic device (201) may perform an operation to generate a third image (440) (e.g., a result image) such that an image area corresponding to at least a portion (e.g., the first portion (411)) of the first image (410) (e.g., the original image) is not related to a person or an animal by performing inpainting or outpainting using the image generation model (301), as illustrated in (b) of FIG. 4A. As illustrated in (a) of FIG. 4A, when the electronic device performing the image generation operation performs inpainting or outpainting on the first image (410) (e.g., the original image) using the generative AI model (330), an image (401) is generated in which the first portion (411) of the first image (410) corresponds to an object (403) (e.g., a person or an animal) that the user does not want. As illustrated in (b) of FIG. 4A, an electronic device (201) according to an embodiment of the present disclosure may generate a third image (440) (e.g., a result image) by using an image generation model (301) to exclude a first portion (411) and to ensure that an image area corresponding to the excluded first portion (411) is not related to a person or an animal. The image generation model (301) is a software-configured configuration and may be stored in a memory (220). The electronic device (201) may use the image generation model (301) to delete at least one portion (e.g., the first portion (411)) representing at least one of a person or an animal included in the first image (410), and may generate a new image (e.g., the third image (440)) by inpainting or outpainting the deleted at least one portion by a generative artificial intelligence (AI) model (330). Here, the first image may be an image captured by a camera (240) or an image acquired from an external electronic device by a communication circuit (250), and may be stored in the memory (220). The first image may include object(s) representing objects, people, animals, and / or plants.

[0063] According to one embodiment, the electronic device (201) may obtain a user input requesting removal of a first portion (411) of a first image (410) stored in a memory (130) using an image generation model (301), and, based on the user input, determine whether the first portion relates to a person or an animal. According to one embodiment, based on determining that the first portion (411) relates to a person or an animal using the image generation model (301), the electronic device (201) may obtain a second image (420) excluding the first portion (411) from the first image (410) and first information related to the first portion (411). Here, the first information may include information for preventing an image area (421) corresponding to the first portion (411) from being generated as a person or an animal.

[0064] According to one embodiment, the image generation model (301) may include an image classification model (310), a preprocessing model (320), and a generative AI model (330). The present invention is not limited thereto, and may further include other models required for image generation, or may not include the generative AI model (330). For example, the generative AI model may be included in an external electronic device (e.g., a server). For example, the generative AI model (330) may represent an artificial intelligence (AI) model that utilizes content such as text, audio, and / or images to newly create similar content. The generative AI model (330) may learn patterns of content and generate new content as inference results. For example, in the field of images, the generative AI model (330) may generate an image that mimics the characteristics of a specific image.

[0065] According to one embodiment, the electronic device (201) obtains a first command to perform at least one of inpainting or outpainting on an image area (421) corresponding to the first part (411) based on the second image (420) and the first information using the generative AI model (330), and provides the obtained first command to the generative AI model (330) to obtain a third image (440) generated by the generative AI model (330) such that the image area corresponding to the first part (411) is not related to a person or an animal. Here, the first information may include a mask image (430) related to the first part (411). The mask image (430) may include a graphic object set to distinguish the first part (411) and information to prevent a portion corresponding to the graphic object from being generated as the person or the animal. According to one embodiment, when the generative AI model (330) is included in an external electronic device, the electronic device (201) can transmit a first command to the external electronic device through a communication circuit (e.g., the communication module (190) of FIG. 1 and the communication circuit (250) of FIG. 2) and obtain a third image generated by the generative AI model from the external electronic device through the communication circuit.

[0066] According to one embodiment, the image classification model (310) may segment at least one object (e.g., the first part) representing a person or an animal included in an original image (e.g., the first image (410)) and a background image (e.g., the background part) into pixel units, and obtain a mask image based on segmented information (e.g., the first information). The image classification model (310) may identify a first part related to a person or an animal to be deleted from the first image (410), and transmit the first image (410) and first information corresponding to the identified first part to the preprocessing model (320). The first information may include mask information (e.g., the first mask information) related to the first part (411). According to one embodiment, the image classification model (310) may transmit a mask image (430) including mask information (e.g., the first mask information) related to the identified first part (411) to the generative AI model (330).

[0067] According to one embodiment, the image segmentation model (310) may perform an image segmentation operation for a first image (410), a classification operation for classifying at least one segmented object as a person or an animal, and in response to a user input, select a first part related to a person or an animal in the first image (410) or generate a bounding box including the first image (410), and / or click four points on the top, bottom, left, and right of the first part to obtain a mask image (430) including mask information (e.g., first mask information) related to the first part (411). Here, the mask information related to the first part (411) may include a graphic object set to distinguish the first part (411) and information for preventing a part corresponding to the graphic object from being generated as a person or an animal.

[0068] According to one embodiment, an image segmentation model (310) can extract features from an image to recognize and locate at least one object in the image. For example, the features extracted here may include color, texture, and / or shape. The image segmentation model (310) can partially organize the image using these features, create a segment map, and define information based on pixel-level boundaries. Once the image is segmented, the image segmentation model (310) can segment each segment into an object (e.g., a part related to a person or an animal) or a background (e.g., a background part), and the image classification operation can be performed based on information learned through a designated deep learning. The image segmentation model (310) can identify the location of an object in the image by identifying a bounding box (e.g., a bounding box) around the object. For example, a bounding box can mean a set of four coordinates that define the boundary of an object.

[0069] According to one embodiment, the preprocessing model (320) may delete a first part (411) related to a first image (410) and an identified object (e.g., an object related to a person or an animal) from the first image (410), obtain a second image (420) in which the first part (411) is deleted (e.g., removed or excluded), and transmit the second image (420) and first information to the generative AI model (330). Here, the first information may include information for preventing an image area corresponding to the first part (411) from being generated as a person or the animal. According to one embodiment, the preprocessing model (320) may obtain a first command for causing the generative AI model (330) to perform at least one operation of inpainting or outpainting based on the second image and the first information, and transmit (e.g., provide) the obtained first command to the generative AI model (330).

[0070] According to one embodiment, the generative AI model (330) may perform at least one operation of inpainting or outpainting based on the second image (420) and the first information transmitted from the preprocessing model (320) in accordance with a first command to perform at least one operation of inpainting or outpainting, and may generate a third image (440) (e.g., a result image) such that the image area corresponding to the first part (411) is not related to a person or an animal. Here, the first information may include mask information related to the first part (411). According to one embodiment, the processor (210) may control the overall operation of the electronic device (201). For example, the processor (210) may be implemented identically or similarly to the processor (120) of FIG. 1. According to one embodiment, the processor (210) may, in response to a request to execute an operation for generating an image, execute an application related to the image (or image execution) and control the display (230) to display a first image (410) through an execution screen of the executed application. According to one embodiment, the processor (210) may edit or modify the first image using the image generation model (301) to create (or obtain) a new third image.

[0071] According to one embodiment, the processor (210) may segment a background and at least one object from a first image (410) using an image classification model (310), and obtain result information about the at least one segmented object. The image classification model (310) may include a designated deep learning model. The processor (210) may use the designated deep learning model to determine whether at least one object segmented from the first image represents at least one of a human or an animal, obtain classification information, and perform an operation of detecting a face in a first part (411) related to the object representing at least one of a human or an animal (e.g., human / pet segmentation). The classification information (e.g., classification information) may include first information about the first part classified as at least one of a human or an animal, and information about the background part. The classification information may include information for filling the first part classified as at least one of a human or an animal with a single color (or different single colors for each object) and / or location information of the first part.

[0072] According to one embodiment, the processor (210) may identify a first portion (411) that satisfies a predetermined deletion condition among at least one classified portion based on the acquired classification information using the image classification model (310). Here, the identified first portion (411) may be an area to be deleted in the first image (410). According to one embodiment, the processor (210) may obtain first information corresponding to the first portion (411) using the image classification model (310), and generate (or obtain) a mask image (430) based on the acquired first information. The processor (210) may generate the mask image (430) by adding the first information to a mask. The mask may be composed of mask information expressed as binary information corresponding to the first image (410).

[0073] According to one embodiment, the processor (210) may use the preprocessing model (320) to delete a first portion (411) from a first image (410) based on the first information, and generate (or obtain) a second image (420) from which the first portion (411) has been deleted.

[0074] According to one embodiment, the processor (210) may perform at least one of inpainting or outpainting based on the second image and the mask image (430) using the generative AI model (330), and may generate (or obtain) a third image (440) such that the image area corresponding to the deleted first portion is not related to a person or an animal by performing at least one of the inpainting operation or the outpainting operation. The processor (210) may apply the generated copy image by copying a background portion of an area adjacent to the first portion (411) to the second image (420) using the inpainting operation or the outpainting operation. The processor (210) may apply the generated copy image to the image area corresponding to the deleted first portion (411) to generate the third image (440) such that the image area is not related to a person or an animal.

[0075] In one embodiment, the processor (210) may control the display (230) to display the third image (440) or control the communication circuit (250) to transmit the third image (440) to an external electronic device.

[0076] According to one embodiment, the processor (210) may be a hardware component (function) or a software component (program) including at least one component provided in the electronic device (201) as a hardware module or a software module (e.g., an application program). According to one embodiment, the processor (210) may include, for example, one or a combination of two or more of hardware, software, or firmware. The processor (210) may omit at least some of the above components, or may be configured to further include other components for performing image processing operations in addition to the above components.

[0077] According to one embodiment, the memory (220) (e.g., the memory (130) of FIG. 1) may store an application. For example, the memory (220) may store an application (function or program) related to an image (or image generation), an application related to image management, or an application related to a generative AI. The memory (220) may store a first image captured by an external electronic device or camera (e.g., an original image), a second acquired image (e.g., an image for editing), a third acquired image (e.g., an edited result image), and information related to image editing.

[0078] According to one embodiment, the memory (220) may store various data generated during execution of the program (140), including a program used for functional operation (e.g., the program (140) of FIG. 1). For example, the memory (220) may include a program (140) area and a data area (not shown). The program (140) area may store related program information for operating the electronic device (201), such as an operating system (OS) (e.g., the operating system (142) of FIG. 1) that boots the electronic device (201). The data area (not shown) may store transmitted and / or received data and generated data according to various embodiments. In addition, the memory (220) may be configured to include at least one storage medium among flash memory, hard disk, multimedia card micro type memory (e.g., secure digital (SD) or extreme digital (XD) memory), RAM, and ROM.

[0079] According to one embodiment, the display (230) (e.g., the display module (160) of FIG. 1) may display an execution screen of an application related to an image (or image generation). The display (230) may display information related to image editing, a first image, or a third image under the control of the processor (210). When performing image editing, the display (230) may display a second image for editing under the control of the processor (210). According to one embodiment, the display (230) may be implemented in the form of a touch screen. When the display (230) is implemented together with an input module in the form of a touch screen, it may display various pieces of information generated according to a user's touch operation. According to one embodiment, the display (230) may be configured with at least one of a liquid crystal display (LCD), a thin film transistor LCD (TFT-LCD), organic light emitting diodes (OLED), a light emitting diode (LED), an active matrix organic LED (AMOLED), a flexible display, and a 3-dimensional display. In addition, some of these displays may be configured as transparent or light-transmitting so that the outside can be seen through them. This may be configured in the form of a transparent display including a transparent OLED (TOLED). According to one embodiment, in addition to the display (230), other display modules (e.g., an extended display or a flexible display) may be further installed.

[0080] According to one embodiment, a camera (240) (e.g., camera module (180) of FIG. 1) can capture a first image (e.g., an original image) to be used for image editing.

[0081] According to one embodiment, the communication circuit (250) (e.g., the communication module (190) of FIG. 1) can communicate with an external electronic device (e.g., the electronic devices 102 and 104 of FIG. 1, the server 108 of FIG. 1, or another user's electronic device). For example, the communication circuit (250) can transmit a first image, a second image, and / or a third image to the external electronic device, or receive the first image from the external electronic device. The communication circuit (250) can receive an application and / or information related to an image (or image editing) from and / or transmit the application and / or information to the external electronic device. According to one embodiment, the communication circuit (250) can include a cellular module, a wireless-fidelity (Wi-Fi) module, a Bluetooth module, or a near field communication (NFC) module.

[0082] FIG. 5 is a diagram illustrating an example of image generation using an already generated model in an electronic device according to one embodiment.

[0083] Referring to FIGS. 2 to 5, according to one embodiment, the processor (210) of the electronic device (201) may display a first image (510) (e.g., the first image (410) of FIGS. 4A and 4B) stored in the memory (220) on the display (230). The processor (210) may segment the first image (510) into at least one part corresponding to at least one object and a part corresponding to the background through an image classification model (e.g., the image classification model (310) of FIG. 3), and may obtain result information (513) including classification information for each of the segmented parts (or result information (513) of the image classification operation). According to one embodiment, the processor (210) may identify a first part corresponding to at least one object classified as a person or an animal based on the result information, and obtain first information about the identified first part (511) (e.g., the first part (411) of FIGS. 4A and 4B). According to one embodiment, the processor (210) may automatically identify the first part (511) or identify the selected first part (511) in response to a user's object selection input.

[0084] According to one embodiment, the processor (210) can determine whether the first portion (511) identified through the image classification model (310) represents at least one of a person or an animal. The processor (210) can determine whether a face exists in the first portion (511) based on information learned through deep learning based on whether the first portion (511) represents at least one of a person or an animal. If a face exists (e.g., detected) in the first portion (511), the processor (210) can obtain a mask image (530) by adding first information (515) corresponding to the first portion (511) (e.g., including first mask information (517) corresponding to the first information (515).

[0085] According to one embodiment, the processor (210) may identify a first part (511) that needs to be deleted from at least one classified object, and if a face does not exist in the first part (511) that needs to be deleted (e.g., if not detected or identified), generate (e.g., identify or obtain) a bounding box (e.g., a first bounding box) for the identified first part (511), and determine a ratio (e.g., a first ratio) of the identified first part (511) (e.g., the size (seg.size) of the object (getFace(seg)) of FIG. 10)) in the generated bounding box. The bounding box (e.g., the first bounding box) may be a box (e.g., a rectangle) representing an area to be filled in (padding) through in-painting, and may be generated to determine the size of the identified first part (511) (e.g., a segmentation area). According to one embodiment, the processor (210) may add first information about the identified first portion (511) to the mask image (530) if the ratio (e.g., the first ratio) occupied by the identified first portion (511) in the identified area (e.g., the first bounding box) is greater than or equal to a threshold value (e.g., the first threshold value). Here, the threshold value (e.g., the first threshold value) may be set to a different value depending on the person or animal (e.g., a value of approximately 30% for the person, a value of approximately 50% for the animal). If the ratio occupied by the first portion (511) in the generated bounding box (e.g., the first bounding box) is less than the threshold value, the processor (210) may determine whether there is a first portion (511) that needs to be deleted next. If there is a part (e.g., a second part) that needs to be deleted, the processor (210) can repeat the above-described operation to add information (e.g., second information) about the additionally confirmed part to the mask image (530).

[0086] According to one embodiment, the processor (210) may use a preprocessing model (320) to obtain (or generate) a second image (520) with the identified first portion (511) deleted (e.g., removed or excluded) based on classification information (e.g., first information (515)) for the first portion (511) (or a portion where a face does not exist) identified by the image classification model (310) and the first image (510).

[0087] According to one embodiment, the processor (210) may perform an in-painting operation on the image area (501) corresponding to the deleted first portion (511) based on the mask image (530) transmitted from the image classification model (310) and / or the second image (520) transmitted from the preprocessing model (320) and the first information, using the generative AI model (330), to obtain a copy image by copying the background of the area adjacent to the image area (501). The processor (210) may apply the copy image to the second image (520) using the generative AI model (330), to newly generate an image of the image area (501) so as not to be related to a person or an animal, and may obtain (or generate) a third image (540) including the image area (501) generated as an image not related to a person or an animal.

[0088] FIG. 6 is a diagram illustrating an example of image generation using an already generated model in an electronic device according to one embodiment.

[0089] Referring to FIGS. 2 to 4 and 6, according to one embodiment, the processor (210) of the electronic device (201) may display a first image (610) (e.g., the first image (410) of FIGS. 4A and 4B) stored in the memory (220) on the display (230). The processor (210) may segment a first part (611) (e.g., the first part (411) of FIGS. 4A and 4B) and a background part related to at least one object (e.g., a person or an animal) in the first image (610) through an image classification model (e.g., the image classification model (310) of FIG. 3), and may obtain first information (615) about the segmented first part and result information (613) including information about the background part (or result information (613) of the image segmentation operation). According to one embodiment, the processor (210) may obtain first information (615) about a first part (611) classified as a person or an animal, and may automatically identify the first part (611) using the obtained first information (615) or identify the selected first part (611) in response to a user's selection input. According to one embodiment, the processor (210) may determine whether the first part (611) identified through the image classification model (310) represents at least one of a person or an animal. Based on whether the first part (611) represents at least one of a person or an animal, the processor (210) may determine whether the first part (611) is located at a boundary line of the first image (610) (e.g., whether it is cut off by the boundary line). When the first part (611) is located at the boundary, the processor (210) can determine whether a face is detected in the first part (611) to determine whether the first part (611) is an object to be deleted (e.g., removed or excluded). According to one embodiment, when the first part (611) is located at the boundary, it can be determined whether a face having a complete facial shape is detected in the first part (611).Here, the complete face shape can be confirmed based on information learned through deep learning. If a face having a complete face shape is not detected as a result of the confirmation, the processor (210) may determine that face detection has failed and identify the first part (611) as a part to be deleted (e.g., removed or excluded). According to one embodiment, if the first part (611) is identified as being located on the boundary line and a face is identified in the first part (611) based on information learned through deep learning, the processor (310) may generate a bounding box for the first part (611) (e.g., a second bounding box indicating an area to be filled in (padding) by out-painting) and confirm the ratio (e.g., a second ratio) occupied by the bounding box (width) in the first image (610). The processor (210) may identify the first portion (611) as a portion to be deleted (e.g., removed or excluded) based on the fact that the ratio of the bounding box (width) in the first image (610) is less than a threshold value (e.g., a second threshold value). Here, the threshold value (e.g., the second threshold value) may be set to, for example, approximately 20% in the case of a person, and may be set to approximately 10% in the case of an animal. According to one embodiment, if the processor (210) identifies that the first portion (611) is located at the boundary line and identifies that the first portion (611) does not contain a face (e.g., face detection fails), the processor (210) may identify the first portion (611) as a portion to be deleted based on the fact that only a body is detected in the first portion (611).

[0090] According to one embodiment, the processor (210) may add first information (615) corresponding to the first portion (611) to the mask (603) upon identifying the first portion (611) as an object to be deleted, obtain first mask information (617) corresponding to the first information (615), and obtain a mask image (630) including the first mask information (617). The processor (210) may transmit the mask image (630) to the generative AI model (330). According to one embodiment, the processor (210) may individually transmit the result information (613) including the first information (615) (or the result information (613) of the image segmentation operation) or the first information (615) to the preprocessing model (320) upon identifying the first portion (611) as an object to be deleted, and may transmit the first image (610) to the preprocessing model (320).

[0091] According to one embodiment, the processor (210) may use a preprocessing model (320) to delete the first portion (611) identified in the first image (610) based on the first information (615) and the first image (610) about the first portion (611) identified by the image classification model (310), and obtain (or generate) a second image (620) (e.g., the second image (420) of FIGS. 4A and 4B) including the image area (601) corresponding to the deleted first portion (611).

[0092] According to one embodiment, the processor (210) may obtain a copy image by copying the background of an area adjacent to the image area (601) through an out-painting operation on the image area (601) corresponding to the deleted first portion (611) based on the mask image (630) transferred from the image classification model (310) and / or the second image (620) transferred from the preprocessing model (320) and the first information using the generative AI model (330). The processor (210) may apply the copy image to the second image (620) using the generative AI model (330) to newly generate an image of the image area (601) so as not to be related to a person or an animal, and may obtain (or generate) a third image (640) including the image area (601) generated as an image not related to a person or an animal.

[0093] FIGS. 7A and 7B are diagrams illustrating examples of image generation using an image generation model in an electronic device according to one embodiment.

[0094] Referring to FIGS. 7A and 7B , according to one embodiment, the processor (210) of the electronic device (201) may segment a plurality of parts (e.g., a first part (711) and a second part (712)) related to at least one object (e.g., a person or an animal) and parts related to a background (e.g., remaining parts excluding the first part (711) and the second part (712)) from a first image (710) (e.g., an original image or the first image (410) of FIGS. 4A and 4B ) through an image classification model (e.g., an image classification model (310) of FIG. 3 ). The processor (210) may classify the first part (711) and the second part (712) representing at least one of a person or an animal in the first image (710). The processor (210) can obtain first information (721) and second information (722) for each of the classified first portion (711) and second portion (712) from the result information (e.g., image) (720) corresponding to the classification information that classified the first image (710). The processor (210) can sequentially analyze the first portion (711) and the second portion (712) to determine whether they are located at the boundary line (701) (e.g., the object is cut off by the boundary line). For example, the processor (210) can identify that the first portion (711) is located at the boundary line and that the second portion (712) is not located at the boundary line (e.g., not cut off by the boundary line). The processor (210) can determine whether a face is detected in the first portion (711) located at the boundary line using an algorithm related to face detection. If the processor (210) fails to detect a face in the first portion (711) or determines that the proportion of the bounding box (width) (e.g., the second bounding box) (703) for the first portion (711) in the first image (710) is less than a threshold value, the processor (210) can identify the first portion (711) as a portion to be deleted (e.g., removed or excluded) from the first image (710) so as not to generate a face for the first portion (711) during out-painting.Here, the processor (210) can determine that face detection has failed if no face exists in the first part (711) or no face with a complete face shape exists in the first part (711).

[0095] According to one embodiment, the processor (210) may obtain a mask image (730) including first information (721) corresponding to the first portion (711) to be deleted or including first mask information (731) corresponding to the first information (721). According to one embodiment, the processor (210) may delete (e.g., remove or exclude) the first portion (711) identified in the first image (710) based on the first information (721) about the first portion (711) identified by an image classification model (e.g., the image classification model (310) of FIG. 3) using a preprocessing model (e.g., the preprocessing model (320) of FIGS. 3, 4A, and 6). According to one embodiment, the processor (210) may transfer a second image (not shown) from which a first portion (711) has been deleted from the preprocessing model to a generative AI model (e.g., the generative AI model (330) of FIGS. 3, 4A, and 6).

[0096] According to one embodiment, the processor (210) may generate an image of an image area corresponding to the deleted first portion (711) through an out-painting operation (e.g., naturally filling in based on an adjacent background image) based on a mask image (730) transmitted from an image classification model (e.g., an image classification model (310) of FIG. 3) and / or a second image (e.g., an image with the first portion (711) deleted) transmitted from a preprocessing model (e.g., a preprocessing model (320) of FIG. 3, FIG. 4, and FIG. 6) and first information (721) using a generative AI model (e.g., a generative AI model (330) of FIG. 3, FIG. 4A, and FIG. 6). According to one embodiment, when the out painting operation for the first portion (711) is completed, the processor (210) may use the generative AI model to obtain (or generate) a result image (740) (e.g., a third image) that excludes the first portion (711) and includes only the second portion (712).

[0097] According to one embodiment, the processor (210) may control the display (230) to display a first image (e.g., an original image, the first image (410) of FIGS. 4A and 4B , the first image (510) of FIG. 5 , the first image (610) of FIG. 6 , and / or the first image (710) of FIG. 7 ) and a third image (e.g., the third image (440) of FIGS. 4A and 4B , the third image (540) of FIG. 5 , the third image (640) of FIG. 6 , and / or the third image (740) of FIG. 7 ) obtained through an in-painting or out-painting operation by a generative AI model on an execution screen of an application related to an image (or image generation). The processor (210) may control the communication circuit (250) to transmit the obtained third image to an external electronic device.

[0098] According to one embodiment, the processor (210) may transmit a second image (e.g., the second image (420) of FIGS. 4A and 4B , the second image (520) of FIG. 5 , and / or the second image (620) of FIG. 6 ) and a mask image (e.g., the mask image (430) of FIGS. 4A and 4B , the mask image (530) of FIG. 5 , and / or the mask image (630) of FIG. 6 ) included in a first area corresponding to the deleted first portion to an external electronic device (e.g., a server) via the communication circuit (250). The external electronic device (e.g., the server) may perform an in / out painting operation on an image area corresponding to the first portion of the second image based on the second image and the mask image (e.g., including first mask information corresponding to the first information) using a generative AI model for image generation, thereby generating an image area as an image not related to a person or an animal, and may generate a third image including an image area having the generated image not related to a person or an animal. According to one embodiment, the processor (210) may receive a third image generated from an external electronic device (e.g., a server) and store the received third image in memory (220) and / or display the received third image on a display (230).

[0099] An electronic device according to one embodiment (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2 and 4) may implement a software module (e.g., program (140) of FIG. 1) related to an image (or image generation). A memory of the electronic device (e.g., memory (130) of FIG. 1 and / or memory (220) of FIG. 2) may store commands (e.g., instructions) to implement the software module. At least one processor (e.g., processor (120) of FIG. 1 and / or processor (210) of FIG. 2) can execute instructions stored in a memory to implement a software module and control hardware associated with the function of the software module (e.g., sensor module (176) of FIG. 1, camera module (180), communication module (190) of FIG. 1 and / or communication circuit (250) of FIG. 2, display module (160) of FIG. 1 and / or display (230) of FIG. 2).

[0100] According to one embodiment, a software module of an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIG. 2) may be configured to include a kernel (or HAL), a framework (e.g., middleware (144) of FIG. 1), and an application (e.g., application (146) of FIG. 1). At least a portion of the software module may be preloaded on the electronic device or may be downloadable from a server (e.g., server (108) of FIG. 1).

[0101] According to one embodiment, the kernel may include, but is not limited to, a system resource manager or a device driver, and may further include other modules. The system resource manager may perform at least one of controlling, allocating, or retrieving system resources. The device driver may include, for example, a display driver, a camera driver, a Bluetooth driver, a shared memory driver, a USB driver, a keypad driver, a WIFI driver, an audio driver, or an inter-process communication (IPC) driver.

[0102] According to one embodiment, the framework may provide functions commonly required by applications or provide various functions to applications through an application programming interface (API) (not shown) so that the applications can efficiently use limited system resources within the electronic device. The framework may include modules that form a combination of various functions of the components described above in FIGS. 1 and 2. The framework may provide specialized modules for each type of operating system (e.g., the operating system (142) of FIG. 1) to provide differentiated functions. The framework may dynamically delete some existing components or add new components.

[0103] According to one embodiment, the application may be configured to include an application (e.g., a module, a manager, or a program) related to an image (or image generation). The application may include an application received from an external electronic device (e.g., a server (108) or an electronic device (102, 104)). According to one embodiment, the application may include a preloaded application or a third-party application downloadable from a server. The components and names of the components of the software module according to the illustrated embodiment may vary depending on the type of operating system. According to one embodiment, at least a portion of the software module may be implemented as software, firmware, hardware, or a combination of at least two or more thereof. At least a portion of the software module may be implemented (e.g., executed) by, for example, a processor (e.g., an AP). At least a portion of the software module may include, for example, at least one of a module, a program, a routine, a set of instructions, or a process for performing at least one function.

[0104] As such, in one embodiment, the main components of the electronic device have been described through the electronic device (101) of FIG. 1 and the electronic device (201) of FIG. 2. However, in various embodiments, not all of the components illustrated through FIGS. 1 and 2 are essential components, and the electronic device (101, 201) may be implemented with more components than the illustrated components, or may be implemented with fewer components. In addition, the positions of the main components of the electronic device (101, 201) described above through FIGS. 1 and 2 may be changed according to various embodiments.

[0105] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2 and 4B) may include a display (e.g., display module (160) of FIG. 1 and / or display (230) of FIG. 2), communication circuitry (e.g., communication module (190) of FIG. 1 and / or communication circuitry (250) of FIG. 2), at least one processor including a processing circuit (e.g., processor (120) of FIG. 1 and / or processor (210) of FIG. 2), and memory (e.g., memory (130) of FIG. 1 and / or memory (220) of FIG. 2)) for storing instructions.

[0106] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a user input requesting removal of a first portion (e.g., the first portion (411) of FIGS. 4A and 4B , the first portion (511) of FIG. 5 , and the first portion (611) of FIG. 6 ) of a first image (e.g., the first image (410) of FIGS. 4A and 4B , the first image (510) of FIG. 5 , and the first image (610) of FIG. 6 )) stored in the memory, using a user interface including the first image.

[0107] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine, based on the user input, whether the first portion relates to a person.

[0108] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a second image (e.g., the second image (420) of FIGS. 4A and 4B , the second image (520) of FIG. 5 , and the second image (620) of FIG. 6 ) excluding the first portion from the first image and first information related to the first portion, based on determining that the first portion is related to the person. According to one embodiment, the first information may include information for preventing an image area corresponding to the first portion from being generated as the person.

[0109] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to perform at least one of inpainting or outpainting based on the second image and the first information using a generative AI model (e.g., the generative AI model (330) of FIGS. 3, 4A, 5, and 6), and to perform at least one of the inpainting or the outpainting to obtain a third image generated by the generative AI model such that the image area corresponding to the first portion is not related to the person (e.g., the third image (440) of FIGS. 4A and 4B, the third image (540) of FIG. 5, and the third image (640) of FIG. 6).

[0110] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to determine, based on the user input, whether the first portion relates to an animal, and, based on determining that the first portion relates to the animal, perform at least one of the inpainting or the outpainting based on the first instruction to obtain a third image generated by the generative AI model such that the image area corresponding to the first portion is not related to the animal (e.g., the third image (440) of FIGS. 4A and 4B , the third image (540) of FIG. 5 , and the third image (640) of FIG. 6 ), wherein the first information may include information for causing the image area corresponding to the first portion to not be generated as the animal.

[0111] According to one embodiment, the first information may include mask information related to the first portion, and the mask information may include a graphic object set to distinguish the first portion and information that prevents a portion corresponding to the graphic object from being created as the person or the animal.

[0112] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine whether a face of the person or the animal is detected in the first portion, and generate the first information about the first portion based on determining that a face of the person or the animal is detected in the first portion.

[0113] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the second image excluding the first portion from the first image using a preprocessing model (e.g., the preprocessing AI model (320) of FIGS. 3, 4a, 5, and 6) stored in the memory.

[0114] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a first ratio value occupied by the first portion in a first bounding box including the first portion based on determining that a face of the person or the animal is not detected in the first portion, generate the first information about the first portion based on the first ratio value being greater than or equal to a first threshold value, and obtain the second image excluding the first portion from the first image using the preprocessing model.

[0115] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a second ratio value of a second bounding box including the first portion in the first image based on determining that a face of the person or the animal is not detected in the first portion, generate the first information about the first portion based on the determined second ratio value being less than a second threshold value, and obtain the second image excluding the first portion from the first image using the preprocessing model.

[0116] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to: determine whether a plurality of portions representing at least one of the person or the animal in the first image are located at a boundary of the first image based on identifying a plurality of portions; and retain the second portion in the first image without excluding the second portion based on identifying a second portion of the first image that is not located at the boundary of the first image.

[0117] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to control the display to display the first image on an execution screen of an application related to image generation based on the user input, and to control the display to display the third image on the execution screen based on obtaining the third image.

[0118] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to control the communication circuit to transmit the third image to an external electronic device based on acquiring the third image.

[0119] Figure 8 is a diagram illustrating an example of an operating method in an electronic device according to one embodiment. In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0120] Referring to FIG. 8, in operation 801, an electronic device according to an embodiment (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2 and 4) may acquire a first image (e.g., an original image) and display the acquired first image on a display (e.g., the display (160) of FIG. 1, the display (230) of FIG. 2). According to one embodiment, in operation 801, the electronic device may obtain a user input requesting removal of a first portion (the first portion (411) of FIGS. 4A and 4B , the first portion (511) of FIG. 5 , and the first portion (611) of FIG. 6 ) of a first image (e.g., the first image (410) of FIG. 4A and 4B , the first image (510) of FIG. 5 , and the first image (610) of FIG. 6 ) stored in a memory (e.g., the memory (130) of FIG. 1 and the memory (220) of FIG. 2 ). According to one embodiment, before or after obtaining the user input, the electronic device may segment the first image into at least one portion related to a person or an animal and a portion related to a background using an image classification model (e.g., the image classification model (310) of FIG. 3 ). Based on the user input, the electronic device may identify the first portion to be removed from the at least one portion related to the person or the animal.

[0121] In operation 803, the electronic device can determine whether the first part relates to a person or an animal based on a user input. In operation 805, based on determining that the first part relates to the person or the animal, the electronic device can obtain a second image (e.g., the second image (420) of FIGS. 4A and 4B , the second image (520) of FIG. 5 , and the second image (620) of FIG. 6 ) excluding the first part from the first image and first information related to the first part. Here, the first information can include information for preventing an image area corresponding to the first part from being generated as a person or an animal.

[0122] In operation 807, the electronic device may obtain a first command to perform at least one of inpainting or outpainting based on the second image and the first information using a generative AI model (e.g., the generative AI model (330) of FIGS. 3, 4a, 5, and 6).

[0123] In operation 809, the electronic device provides a first command to the generative AI model, and performs at least one of inpainting or outpainting based on the second image and the first information using the generative AI model (e.g., the generative AI model (330) of FIGS. 3, 4a, 5, and 6) to obtain a third image (440, 540, 640) generated such that an image area corresponding to the first portion is not related to a person or an animal. Here, the generative AI model may be included in the electronic device or an external electronic device (e.g., a server). When the generative AI model is included in the external electronic device, the electronic device transmits the first command to the external electronic device through a communication circuit (e.g., the communication module (190) of FIG. 1 and the communication circuit (250) of FIG. 2), and obtains the third image generated by the generative AI model from the external electronic device through the communication circuit.

[0124] Hereinafter, the operation method of FIG. 8 described above will be described in more detail with reference to the drawings.

[0125] FIG. 9 is a drawing showing an example of an operating method in an electronic device according to one embodiment, FIG. 10 is a drawing showing an example of an operating method for image generation using an image generation model in an electronic device according to one embodiment, and FIGS. 11a and 11b are drawings showing examples of images by in-painting in an electronic device according to one embodiment.

[0126] In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. The operation method of FIG. 9 may be performed using the algorithm of the image generation model illustrated in FIG. 10 (e.g., the image generation model (301) of FIG. 3 ).

[0127] Referring to FIGS. 9, 10, and 11a and 11b, in operation 901, an electronic device according to an embodiment (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2 and 4) may display a first image (1110a, 1110b) (e.g., the first image (510) of FIG. 5) stored in a memory (e.g., the memory (130) of FIG. 1 and / or the memory (230) of FIG. 2)) on a display (e.g., the display module (160) of FIG. 1 and / or the display (230) of FIG. 2). The electronic device can obtain a first image from a camera (e.g., a camera module (180) of FIG. 1) or an external electronic device (e.g., an electronic device (102, 104 or a server (108) of FIG. 1)) and store the first image in memory, and when requested to execute image generation, can obtain the first image (1110a, 1110b) stored in the memory (e.g., void get_object_removal_inpainting(input_image, mask) of FIG. 10). At this time, the electronic device can obtain a mask (or mask information) corresponding to the first image (1110a, 1110b) (e.g., input_image = input_image[mask] of FIG. 10).

[0128] In operation 903, the electronic device may perform a segmentation operation of the first image (1110a, 1110b) to segment at least one object and a background in the first image (1110a, 1110b) through an image classification model (e.g., the image classification model (310) of FIG. 3), and obtain classification information (e.g., including first information about the first part and information about the background part) for each of the first part (e.g., the first part (1111a, 1111b) of FIGS. 11a and 11b) and the background part related to the segmented at least one object (e.g., a person or an animal) (e.g., seginfo=getSegmentation(input_image) of FIG. 10).

[0129] In operation 905, the electronic device may determine whether the classified first part relates to a person or an animal based on user input (e.g., for seg in seginfo: int type = Classification(seg) / pet or human in FIG. 10). If the determination result indicates that the first part does not relate to a person or an animal, the electronic device may perform operation 907. If the first part relates to a person or an animal, the electronic device may perform operation 909.

[0130] In operation 907 (operation 905-No), the electronic device may transmit the first image (1110a, 1110b) identified by the image classification model and the classification information of the first image (1110a, 1110b) to a generative AI model (e.g., the generative AI model (330) of FIGS. 3 and 5), and provide a command (e.g., a second command) to cause the generative AI model to perform an inpainting operation on the first image (1110a, 1110b). A result image (e.g., a fourth image) generated through the inpainting operation by the generative AI model may be obtained. Here, since the generated result image (e.g., the fourth image) does not include the first part related to the person or the animal, the electronic device may not perform an operation of in-painting the first image (1110a, 1110b), or even if the in-painting operation is performed, the classified parts in the first image (1110a, 1110b) may generate a fourth image that is not related to the person or the animal. Thereafter, the electronic device may perform operation 921. As another example, even if the original image (e.g., the first image) does not include an object representing at least one of the person or the animal, the electronic device may identify a part of the original image to be deleted, and perform an operation of in-painting an area corresponding to the identified part.

[0131] In operation 909 (operation 905 - example), the electronic device can check (e.g., if getFace(seg):) whether a face exists (e.g., detected) in the first portion (1111a, 1111b) related to a person or animal classified through an image classification model. If the check result shows that a face exists in the identified first portion (1111a, 1111b), the electronic device can perform operation 915, and if a face does not exist, the electronic device can perform operation 911. The electronic device can detect a face in the first portion (1111a, 1111b) using information learned through deep learning technology (e.g., a model).

[0132] In operation 911 (operation 909-No), if a face does not exist in the first portion (1111a, 1111b), the electronic device may generate a padding region (e.g., a first bounding box (1105) in FIG. 11c) to be filled in the image area corresponding to the deleted first portion (1111a, 1111b) based on the first information (e.g., the first information (1103) in FIG. 11c) about the first portion (1111a, 1111b) (e.g., else: box = getSegBoundingBox(padd) in FIG. 10). Here, the first information (e.g., the first information (1103) in FIG. 11c) may include, for example, pixel information, padding information, and / or position information. The electronic device can identify a ratio (e.g., a first ratio) occupied by the first portion (1111a, 1111b) (e.g., the size of the object (seg.size)) in the generated padding area (e.g., the first bounding box (1105) of FIG. 11c).

[0133] In operation 913, the electronic device may determine whether the ratio (e.g., the first ratio) identified in operation 911 is greater than or equal to a first threshold (e.g., input_image.size * threshold) based on the absence of a face (e.g., if seg.size >= (input_image.size * threshold):) in FIG. 10. If the ratio is greater than or equal to the first threshold, the electronic device may perform operation 915, and if the ratio is less than the first threshold, the electronic device may perform operation 921. For example, when the first portion represents a person, the first threshold may be set to approximately 30%, and when the first portion represents an animal, the first threshold may be set to approximately 50%. Accordingly, the electronic device may perform more robust editing (e.g., filtering) when the first portion represents a person.

[0134] The electronic device can acquire (e.g., generate) a mask image (e.g., a mask image (530) of FIG. 5) by adding first information about the first portion (1111a, 1111b) to a mask through an image classification model. The electronic device can transfer the first information and the first image (1110a, 1110b) from the image classification model to a preprocessing model (e.g., a preprocessing model (320) of FIGS. 3 and 5). Here, the first information can include mask information related to the first portion (e.g., a mask image (530) of FIG. 5). According to one embodiment, the electronic device can transfer the mask image including the first mask information related to the first portion to the generative AI model by the image classification model. According to one embodiment, if the electronic device determines in operation 909 that a face exists in the first part, the electronic device may delete (e.g., remove or exclude) the first part from the first image using a preprocessing model, and obtain (e.g., create) a second image including an image area corresponding to the deleted first part. Before deleting the first part, the electronic device may obtain (e.g., create) a mask image (e.g., a mask image (530) of FIG. 5) by adding first information about the first part (1111a, 1111b) to a mask using an image classification model (e.g., objectRemoval.add(seg) of FIG. 10). According to one embodiment, if the electronic device determines that no face exists in the first part (e.g., not detected), and a ratio (e.g., first ratio) of the first part to the identified area (e.g., first bounding box) is greater than or equal to a first threshold as determined in operation 913, the electronic device may acquire (e.g., generate) (e.g., objectRemoval.add(mask) in FIG. 10) a mask image (e.g., mask image (530) of FIG. 5) by adding first information about the first part (1111a, 1111b) to a mask through an image classification model.

[0135] In operation 915 (operation 909 - yes or operation 913 - no), the electronic device may delete an area (e.g., a first area) corresponding to the first object (1111a, 1111b) from the first image (1110a, 1110b) based on the first information and the mask information of the first image (1110a, 1110b) and / or the mask image in the preprocessing model. The electronic device may obtain (e.g., generate) a second image (e.g., the second image (520) of FIG. 5) including the area deleted in the preprocessing model.

[0136] In operation 917, the electronic device may transmit the acquired second image and first information to a generative AI model (e.g., the generative AI model (330) of FIG. 3). The electronic device may transmit a mask image from an image classification model or a preprocessing model to the generative AI model. The electronic device may obtain a first command to perform an inpainting operation based on the second image and the first information using the generative AI model (330).

[0137] In operation 919, the electronic device may transmit a first command to the generative AI model to obtain a third image generated by the generative AI model such that the image area corresponding to the first portion is not related to a person or an animal. The generative AI model may perform an in-painting operation on the first portion deleted from the second image based on the second image and the first information to generate an image of the image area for the deleted first portion, and may obtain (e.g., generate) a third image (1130a, 1130b) (e.g., the third image (540) of FIG. 5) including the image area having the generated image. The electronic device may generate a new copied image by copying a background portion or another portion of an area adjacent to the image area corresponding to the deleted first portion using the in-painting operation in the generative AI model, and may generate an image of the image area by applying the copied image to the image area corresponding to the deleted first portion. The generated image may be generated so as to be naturally connected with a background image or another portion (e.g., an object that does not represent a person or an animal).

[0138] Thereafter, the electronic device can check whether the operation of generating the first image (1110a, 1110b) is completed in operation 921, and if so, terminate the operation, and if not, perform operation 903 again.

[0139] In the operation description of FIG. 9 described above, operations 903 to 915 may be specific operations for checking whether a preset condition for checking whether deletion of the first portion is necessary by the image classification model is satisfied. For example, if a plurality of portions identified in the first image are identified, operations 903 to 915 may be repeatedly performed. In operation 917, a second image including an image area corresponding to the removed first portion may be acquired, and in operation 919, an image of the image area corresponding to the deleted first portion may be generated, and a third image including an image area having the generated image may be acquired. For example, if a plurality of portions identified in the first image are identified, classification and classification information may be acquired for all of the portions identified in operation 903 at once. Accordingly, it is possible to check whether a preset condition for checking whether deletion of all of the portions is necessary at once is satisfied in operations 905 to 915 without repeatedly performing operations 905 to 915.

[0140] According to one embodiment, the electronic device can delete a first portion (1111a, 1111b) of a first image (1110a, 1110b) through the operations of FIG. 9 described above, generate an image of an image area corresponding to the deleted first portion through an in-painting operation, and obtain a third image (1130a, 1130b) including an image area having the generated image. Accordingly, the electronic device can prevent other objects (1121a, 1121b) (e.g., images) that are not intended by the user from being generated through the in-painting operation, such as images (1120a, 1120b) generated without applying the operations of FIG. 9, as illustrated in FIGS. 11a and 11b.

[0141] FIG. 12 is a drawing showing an example of an operating method in an electronic device according to one embodiment, FIG. 13 is a drawing showing an example of an operating method for image generation using an image generation model in an electronic device according to one embodiment, and FIG. 14 is a drawing showing examples of images by outpainting in an electronic device according to one embodiment.

[0142] In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. The operation method of FIG. 12 may be performed using the algorithm of the image generation model illustrated in FIG. 13 (e.g., the image generation model (301) of FIG. 3 ).

[0143] Referring to FIGS. 12, 13, and 14, in operation 1201, an electronic device according to an embodiment (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2 and 4) may display a first image (e.g., the first image (1410) of FIG. 14) stored in a memory (e.g., the memory (130) of FIG. 1 and / or the memory (230) of FIG. 2)) on a display (e.g., the display module (160) of FIG. 1 and / or the display (230) of FIG. 2). The electronic device may acquire the first image from a camera (e.g., the camera module (180) of FIG. 1) or an external electronic device (e.g., the electronic device (102, 104) of FIG. 1 or the server (108)) and store the first image in the memory, and when there is a request to execute an image generation operation, acquire the first image stored in the memory (e.g., the void of FIG. 13). get_object_removal_outpainting(input_image)). For example, the electronic device can obtain a mask (or mask information) corresponding to the first image.

[0144] In operation 1203, the electronic device may perform a segmentation operation of the first image (e.g., the first image (1410) of FIG. 14) to segment a first portion and a background portion related to at least one object (e.g., a person or an animal) in the first image (e.g., the first image (1410) of FIG. 14) through an image classification model (e.g., the image classification model (310) of FIGS. 3 and 6), and obtain result information including information about each of the segmented (e.g., including the first portion (1411) of FIG. 14)) and the background portion (e.g., seginfo=getSegmentation(input_image)).

[0145] In operation 1205, the electronic device may determine whether the classified first part relates to a person or an animal based on user input (e.g., for seg in seginfo: int type = Classification(seg) / pet or human in FIG. 13). If the determination result indicates that the first part does not relate to a person or an animal, the electronic device may perform operation 1207. If the first part relates to a person or an animal, the electronic device may perform operation 1209.

[0146] In operation 1207 (operation 1205 - No), the electronic device may transfer the first image (1410) identified by the image classification model and the classification information of the first image (1410) to a generative AI model (e.g., the generative AI model (330) of FIGS. 3 and 6), and provide a command (e.g., a second command) to cause the generative AI model to perform an outpainting operation on the first image (1410). A result image (e.g., a fourth image) generated through the outpainting operation by the generative AI model may be obtained. Here, since the generated result image (e.g., the fourth image) does not include the first part related to a person or an animal, the electronic device may not perform the outpainting operation on the first image (1410), or even if the outpainting operation is performed, the electronic device may generate a fourth image in which the classified parts of the first image (1410) are not related to a person or an animal. Thereafter, the electronic device may perform operation 1223. As another example, the electronic device may perform an operation of identifying a portion of an original image from which an object representing at least one of a person or an animal is to be deleted, and outpainting an area corresponding to the identified portion, even if the original image (e.g., the first image) does not contain an object representing at least one of a person or an animal.

[0147] In operation 1209 (operation 1205 - example), the electronic device can check (e.g., if seg.position in boundary: of FIG. 13) whether the first part (1411) related to the person or animal is located at the boundary (e.g., whether it is cut off by the boundary). If the check result shows that the first part (1411) is located at the boundary, operation 1211 can be performed, and if not, operation 1223 can be performed.

[0148] In operation 1211 (operation 1209 - example), the electronic device can check (e.g., if getFace(seg):) whether a face exists (e.g., detected) in the first part (1411) related to a person or an animal among at least one object through an image classification model. If a face exists (e.g., detected) in the identified first part (1411), the electronic device can perform operation 1213, and if a face does not exist (e.g., not detected), the electronic device can perform operation 1217. The electronic device can detect a face in the first object (1411) based on the first information of the first part (1411) using information learned through deep learning technology (e.g., a model). According to one embodiment, when the electronic device checks whether a face exists in the first part (1411), it can check whether a face having a complete face shape is detected. Here, the complete face shape can be checked based on information learned through deep learning. If, as a result of the verification, a face having a complete facial shape is not detected, the electronic device may determine that a face does not exist (e.g., face detection failed) and identify the first portion (1411) as a portion to be deleted.

[0149] In operation 1213 (operation 1211-Example), if a face (e.g., a face in the form of a complete face) is present (e.g., detected) in the first part (1411), the electronic device may generate a bounding box (e.g., a second bounding box) for outpainting the first part (1411) based on first information (e.g., including pixel information, padding information, and / or position information) about the first part (1411) (e.g., box = getSegBoundingBox(padd)), and identify a ratio (e.g., a second ratio) of the generated bounding box (e.g., a width of the second bounding box) in the first image (1410).

[0150] In one embodiment, when a face is detected in the identified first portion (1411) or a face having a complete face shape is detected, the electronic device may perform an operation of comparing a ratio (e.g., a second ratio) of a bounding box (e.g., a width of a second bounding box) to a threshold (image.width*threshold) (e.g., a second threshold) as a predefined condition related to deletion to determine whether to perform an out-painting operation. In operation 1215, the electronic device may determine whether the ratio (e.g., a second ratio) of a bounding box (e.g., a width of a second bounding box) to the first image (1410) is less than the second threshold (e.g., a specified first condition) (e.g., if seg.width < (image.width*threshold): of FIG. 13). As a result of the verification, if the second ratio is less than the second threshold, the electronic device can perform operation 1217, and if the second ratio is greater than or equal to the second threshold, the electronic device can perform operation 1223. For example, if the first part (1411) represents a person, the second threshold can be set to approximately 20%, and if the first part (1411) represents an animal, the second threshold can be set to approximately 10%. Accordingly, the electronic device can perform more robust editing (e.g., filter) by reducing the generated area when the first part is a person. If a face is present in the first part, and the ratio (e.g., the second ratio) that the bounding box (e.g., the width of the second bounding box) occupies in the first image (1410) is less than the second threshold, the electronic device can identify that there is a large area of ​​the body to be generated, and identify the first part (1411) as a part to be deleted.

[0151] According to one embodiment, the electronic device may obtain (e.g., generate) a mask image including first mask information corresponding to first information about the first part (1411) through an image classification model. The electronic device may transmit the first information and the first image (1410) from the image classification model to a preprocessing model (e.g., the preprocessing model (320) of FIGS. 3 and 6). Here, the first information may include mask information (e.g., first mask information) related to the first part. According to one embodiment, the electronic device may transmit the mask image including the first mask information to the generative AI model by the image classification model. According to one embodiment, if the electronic device determines in operation 1211 that no face exists (e.g., is detected), the electronic device may identify that only a body is detected in the identified first part and identify the identified first part as a part to be deleted. Through the image classification model, the first information about the first portion (1411) can be added to the mask to obtain (e.g., create) a mask image including the first mask information corresponding to the first information (e.g., objectRemoval.add(seg) of FIG. 13). According to one embodiment, if the electronic device determines in operation 1215 that a face exists and a ratio (e.g., a second ratio) of a bounding box (e.g., a width of a second bounding box) to the first image (1410) is less than a second threshold, the electronic device can obtain (e.g., create) a mask image including the first mask information corresponding to the first information (e.g., a mask image (630) of FIG. 6) by adding the first information about the first object (1411) to the mask to remove the first object (1411) (or the first area corresponding to the first object) identified as the object to be deleted (e.g., create) a mask image including the first mask information corresponding to the first information (e.g., objectRemoval.add(mask) of FIG. 13).Here, if a face exists in the first object (1411) and the ratio (e.g., the second ratio) of the bounding box (e.g., the width of the box) to the first image (1410) is smaller than the second threshold, the body area or face area to be generated during out-painting increases. Accordingly, the electronic device needs to remove the first object (1411) (or the first area corresponding to the first object) so that an image (1430) including an object (1431) of a different shape (or form) than intended by the user (e.g., an object of a different face or an object including an unintended body area) is not generated, as illustrated in FIG. 14.

[0152] In operation 1217 (operation 1211 - No or operation 1215 - Yes), the electronic device may delete an area (e.g., the first area) corresponding to the first portion (1411) in the first image (1410) based on the first information and the mask information of the first image (1410) and / or the mask image in the preprocessing model. The electronic device may obtain (e.g., generate) a second image (e.g., the second image (620) of FIG. 6) including the image area corresponding to the first portion deleted in the preprocessing model.

[0153] In operation 1219, the electronic device may transmit the acquired second image to a generative AI model (e.g., the generative AI model (330) of FIGS. 3 and 6). The electronic device may transmit a mask image from an image classification model or a preprocessing model to the generative AI model. The electronic device may obtain a first command to perform an inpainting operation based on the second image and the first information using the generative AI model (330). In operation 1221, the electronic device may transmit the first command to the generative AI model to obtain a third image generated by the generative AI model such that the image area corresponding to the first portion is not related to a person or an animal. The generative AI module may perform an outpainting operation on the first portion deleted from the second image based on the second image and the first information to generate an image of the image area for the deleted first portion, and may obtain (e.g., generate) a third image (1420) (e.g., the third image (640) of FIG. 6) including the image area having the generated image. The electronic device can generate a new simulated image by using an out-painting operation in a generative AI model to simulate a background portion or another portion of an area adjacent to an image area corresponding to a deleted first portion, and then apply the simulated image to the image area corresponding to the deleted first portion to generate an image of the image area. The generated image can be edited to be naturally connected to the background image or another portion (e.g., an object that does not represent a person or animal).

[0154] Thereafter, the electronic device can check whether the operation of generating the first image (1410) is completed in operation 1223, and if so, terminate the operation, and if not, perform operation 1203 again.

[0155] In the operation description of FIG. 12 described above, operations 1203 to 1217 may be specific operations for checking whether a preset condition for checking whether the first portion needs to be deleted is satisfied by the image classification model. For example, if multiple portions are identified in the first image, operations 1203 to 1217 may be repeatedly performed. In operation 1219, a second image in which an image area corresponding to the first portion to be removed is deleted may be obtained, and in operation 1221, an image of the image area corresponding to the deleted first portion may be generated, and a third image (1420) including the image area having the generated image may be obtained. For example, if multiple portions are identified in the first image (1410), classification and classification information may be obtained at once for all of the multiple portions identified in operation 1203. Accordingly, it is possible to check whether a preset condition is satisfied to determine whether multiple parts need to be deleted at once in operations 1205 to 1217 without repeatedly performing operations 1205 to 1217.

[0156] According to one embodiment, the electronic device can display a third image obtained through image editing according to the operating methods of FIGS. 8, 9, and 12 described above on a display, and transmit the obtained third image to an external power supply through a communication circuit (e.g., the communication module (190) of FIG. 1, the communication circuit (250) of FIG. 2).

[0157] According to one embodiment, the electronic device may transmit a second image and a mask image including an image area corresponding to the deleted first portion to an external electronic device (e.g., a server) via a communication circuit. The external electronic device (e.g., a server) may use an AI model for image generation to generate an image based on the second image and the mask image, such that the image area corresponding to the first portion included in the second image is not related to a person or an animal through an in / out painting operation, and may generate a third image including an image area having the generated image. According to one embodiment, the processor (210) may receive the third image generated by the external electronic device (e.g., the server), and store the received third image in a memory (e.g., the memory (130) of FIG. 1, the memory (220) of FIG. 2) and / or display the received third image on a display.

[0158] According to one embodiment, the electronic device can delete a first portion (1411) of a first image (1410) through the operations of FIG. 12 described above, generate an image for an image area corresponding to the first portion (1411) through an out-painting operation, and obtain a third image (1420) that includes the image area having the generated image and is not related to a person or an animal. Accordingly, the electronic device can prevent another object (1431) (e.g., an image related to a person or an animal) that is not intended by the user from being generated through the out-painting operation, such as an image (1430) generated without applying the operations of FIG. 12, as illustrated in FIG. 14.

[0159] According to one embodiment, a method of operating an electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2 and 4) may include an operation of obtaining a user input requesting removal of a first portion (e.g., the first portion (411) of FIGS. 4A and 4B, the first portion (511) of FIG. 5, and the first portion (611) of FIG. 6) of a first image (e.g., the first image (410) of FIGS. 4A and 4B, the first image (510) of FIG. 5, and the first image (610) of FIG. 6)) stored in a memory (e.g., the memory (130) of FIG. 1 and the memory (220) of FIG. 2) of the electronic device, using a user interface including the first image.

[0160] According to one embodiment, the method may include an action of determining, based on the user input, whether the first portion relates to a person.

[0161] According to one embodiment, the method may include an operation of obtaining a second image (e.g., the second image (420) of FIGS. 4A and 4B , the second image (520) of FIG. 5 , and the second image (620) of FIG. 6 ) excluding the first part from the first image and first information related to the first part, based on determining that the first part is related to the person. According to one embodiment, the first information may include information for preventing an image area corresponding to the first part from being created as the person.

[0162] According to one embodiment, the method may include an operation of performing at least one of inpainting or outpainting based on the second image and the first information using a generative AI model (e.g., the generative AI model (330) of FIGS. 3, 5, and 6), and obtaining a third image (e.g., the third image (440) of FIGS. 4A and 4B, the third image (540) of FIG. 5, and the third image (640) of FIG. 6) generated by the generative AI model such that the image area corresponding to the first portion is not related to the person or the animal by performing at least one of the inpainting or the outpainting.

[0163] In one embodiment, the method may include an operation of determining whether the first portion relates to an animal based on the user input, and an operation of obtaining a third image (440, 540, 640) generated by the generative AI model such that the image area corresponding to the first portion is not related to the animal based on determining that the first portion relates to the animal by performing at least one of the inpainting or the outpainting based on the first command. In one embodiment, the first information may include information for preventing the image area corresponding to the first portion from being generated as the animal.

[0164] According to one embodiment, the first information may include mask information related to the first portion, and the mask information may include a graphic object set to distinguish the first portion and information that prevents a portion corresponding to the graphic object from being created as the person or the animal.

[0165] According to one embodiment, the operation of obtaining a second image excluding the first portion from the first image and first information related to the first portion may include an operation of determining whether a face of the person or the animal is detected in the first portion, and an operation of generating the first information for the first portion based on determining that a face of the person or the animal is detected in the first portion.

[0166] According to one embodiment, the operation of obtaining the second image excluding the first portion from the first image and the first information related to the first portion may include an operation of obtaining the second image excluding the first portion from the first image using a preprocessing model (e.g., the preprocessing model (320) of FIGS. 3, 4a, 5, and 6) stored in the memory.

[0167] According to one embodiment, the operation of obtaining the second image excluding the first portion from the first image and first information related to the first portion may include: an operation of identifying a first ratio value occupied by the first portion in a first bounding box including the first portion based on determining that a face of the person or the animal is not detected in the first portion; and an operation of generating the first information for the first portion based on the first ratio value being equal to or greater than a first threshold value, and obtaining the second image excluding the first portion from the first image using the preprocessing model.

[0168] According to one embodiment, the operation of obtaining the second image excluding the first portion from the first image and first information related to the first portion may include the operation of identifying a second ratio value of a second bounding box including the first portion in the first image based on determining that a face of the person or the animal is not detected in the first portion, and the operation of generating the first information for the first portion based on the determined second ratio value being less than a second threshold value, and obtaining the second image excluding the first portion from the first image using the preprocessing model.

[0169] According to one embodiment, the method may further include: an operation of identifying a plurality of parts representing at least one of the person or the animal in the first image, and determining whether the plurality of parts are located at a boundary of the first image; and an operation of maintaining the second part in the first image without excluding the second part, based on identifying a second part of the first image that is not located at the boundary among the plurality of parts.

[0170] According to one embodiment, the method may further include an operation of controlling the display to display the first image on an execution screen of an application related to image generation based on the user input, and an operation of controlling the display to display the third image on the execution screen based on obtaining the third image.

[0171] According to one embodiment, the method may further include an operation of transmitting the third image to an external electronic device through a communication circuit of the electronic device based on acquiring the third image.

[0172] According to one embodiment, in a non-transitory storage medium storing a program, the program, when executed by at least one processor (e.g., the processor (120) of FIG. 1 and the processor (210) of FIG. 2) of an electronic device (e.g., the electronic device (101) of FIG. 1 and the electronic device (201) of FIGS. 2 and 4), causes the electronic device to remove a first part (e.g., the first part (411) of FIGS. 4a and 4b, the first part (511) of FIG. 5 and the first part (611) of FIG. 6) of a first image (e.g., the first image (410) of FIGS. 4a and 4b, the first image (510) of FIG. 5 and the first image (610) of FIG. 6) stored in a memory (e.g., the memory (130) of FIG. 1 and the memory (230) of FIG. 2) of the electronic device), acquires a user input using a user interface including the first image, requesting the electronic device to remove the first part (e.g., the first part (411) of FIGS. 4a and 4b, the first part (511) of FIG. 5 and the first part (611) of FIG. 6) of the first image stored in the memory (e.g., the memory (130) of FIG. 1 and the memory (230) of FIG. 2) of the electronic device; Based on the user input, an operation of determining whether the first part is related to a person, based on determining that the first part is related to the person, an operation of obtaining a second image (e.g., a second image (420) of FIGS. 4A and 4B, a second image (520) of FIG. 5, and a second image (620) of FIG. 6) excluding the first part from the first image and first information related to the first part, the first information including information for preventing an image area corresponding to the first part from being generated as the person, and based on a first command for performing at least one of inpainting or outpainting based on the second image and the first information using a generative AI model (e.g., a generative AI model (330) of FIGS. 3, 4A, 5, and 6), performing at least one of the inpainting or the outpainting to prevent the image area corresponding to the first part from being generated as the person, a third image (e.g., FIG. The third image (440) of Figs. 4a and 4b,It may include executable commands to execute an operation to acquire the third image (540) of FIG. 5 and the third image (640) of FIG. 6.

[0173] According to one embodiment, in a non-transitory storage medium storing a program, the program may include, when executed by at least one processor of an electronic device, causing the electronic device to, based on the user input, determine whether the first portion relates to an animal, and, based on determining that the first portion relates to the animal, perform at least one of the inpainting or the outpainting based on the first command to obtain a third image generated by the generative AI model such that the image area corresponding to the first portion is not related to the animal (e.g., the third image (440) of FIGS. 4A and 4B , the third image (540) of FIG. 5 , and the third image (640) of FIG. 6 ). According to one embodiment, the first information may include information for preventing the image area corresponding to the first portion from being generated as the animal.

[0174] According to one embodiment, the first information may include mask information related to the first portion, and the mask information may include a graphic object set to distinguish the first portion and information that prevents a portion corresponding to the graphic object from being created as the person or the animal.

[0175] According to one embodiment, the operation of obtaining the second image excluding the first portion from the first image and first information related to the first portion may include executable instructions to perform an operation of determining whether a face of the person or the animal is detected in the first portion; and an operation of generating the first information for the first portion based on determining that a face of the person or the animal is detected in the first portion.

[0176] According to one embodiment of the present disclosure, when editing (in / out-painting) an image of an object such as a person or an animal, an electronic device may generate an edited image by deleting the object such as a person or an animal, rather than generating a generalized person or animal face based on a previously learned person or animal. Accordingly, when inpainting or out-painting an original image (e.g., a first image), the electronic device may not generate new, meaningless parts related to a person or animal that the user does not want, and may generate a more natural resulting image. In addition, various effects that can be directly or indirectly understood through this document may be provided. The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by those skilled in the art to which the present disclosure pertains from the description below.

[0177] As used herein, the term "if" will be understood to mean "when, upon," "in response to determining," or "in response to detecting," depending on the context. Similarly, "if it is determined to," or "if [the stated condition or event] is detected," will optionally be understood to mean "upon determining," or "in response to determining," "upon detecting [the stated condition or event]," or "in response to detecting [the stated condition or event]."

[0178] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. A processing device (or processing circuit) may execute an operating system (OS) and one or more software applications running on the operating system. In addition, the processing device may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.

[0179] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.

[0180] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program commands, including ROM, RAM, and flash memory. In addition, examples of other media may include recording or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.

[0181] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

[0182] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.

[0183] The embodiments disclosed in this document are presented for the purpose of explaining and understanding the disclosed technical content, and do not limit the scope of the technology described in this document. Therefore, the scope of this document should be interpreted to include all modifications or various embodiments based on the technical concepts of this document.

[0184] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.

[0185] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0186] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0187] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g. electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0188] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0189] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device (101, 201), a display (160, 230); Communication circuit (190, 250); At least one processor (120, 210) comprising a processing circuit; and A memory (130, 220) comprising one or more storage media storing instructions, said instructions, when individually or collectively executed by said at least one processor, causing said electronic device to: Obtaining a user input requesting removal of a first portion (411, 511, 611) of a first image (410, 510, 610) stored in the memory using a user interface including the first image, Based on the above user input, determine whether the first part is related to a person, Based on the confirmation that the first part is related to the person, a second image (420, 520, 620) excluding the first part from the first image and first information related to the first part are obtained, and the first information includes information for preventing an image area corresponding to the first part from being created as the person. An electronic device, based on a first command to perform at least one of inpainting or outpainting based on the second image and the first information using a generative AI model (330), performing at least one of the inpainting or the outpainting to obtain a third image (440, 540, 640) generated by the generative AI model such that the image area corresponding to the first portion is not related to the person.

2. In the first paragraph, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Based on the above user input, determine whether the first part is related to an animal, Based on the determination that the first part is related to the animal, at least one of the inpainting or the outpainting is performed based on the first command to obtain a third image (440, 540, 640) generated by the generative AI model so that the image area corresponding to the first part is not related to the animal. The first information includes information that prevents an image area corresponding to the first portion from being created as the animal, and the first information includes mask information related to the first portion. An electronic device, wherein the mask information includes a graphic object set to distinguish the first portion and information preventing a portion corresponding to the graphic object from being created as the person or the animal.

3. In the first or second paragraph, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: In the above first part, check whether the face of the person or the animal is detected, Based on the determination that the face of the person or the animal is detected in the first part, the first information for the first part is generated, An electronic device that obtains the second image excluding the first portion from the first image by using the preprocessing model (320) stored in the memory.

4. In any one of paragraphs 1 to 3, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Based on the determination that the face of the person or the animal is not detected in the first part, a first ratio value occupied by the first part is identified in a first bounding box including the first part, An electronic device that generates the first information for the first portion based on the first ratio value being greater than or equal to a first threshold value, and obtains the second image excluding the first portion from the first image using the preprocessing model.

5. In any one of paragraphs 1 to 4, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Based on the determination that the face of the person or the animal is not detected in the first portion, a second bounding box including the first portion is identified as having a second ratio value in the first image, An electronic device that generates the first information for the first portion based on the second ratio value being less than the second threshold value, and obtains the second image excluding the first portion from the first image using the preprocessing model.

6. In any one of paragraphs 1 to 5, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Based on identifying a plurality of parts representing at least one of the person or the animal in the first image, determining whether the plurality of parts are located at a boundary of the first image, An electronic device that identifies a second portion of the first image that is not located on the boundary line among the plurality of portions, and maintains the second portion in the first image without excluding the second portion.

7. In any one of paragraphs 1 to 6, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Based on the user input, the display is controlled to display the first image on the execution screen of the application related to image generation, Based on obtaining the third image, the display is controlled to display the third image on the execution screen, An electronic device that controls the communication circuit to transmit the third image to an external electronic device based on acquiring the third image.

8. In the method of operation in an electronic device (101, 201), An action of obtaining a user input requesting removal of a first portion (411, 511, 611) of a first image (410, 510, 610) stored in a memory (130, 220) of the electronic device using a user interface including the first image; An action to determine whether the first part is related to a person based on the user input; An operation of obtaining a second image (420, 520, 620) excluding the first part from the first image and first information related to the first part based on the confirmation that the first part is related to the person; the first information includes information for preventing an image area corresponding to the first part from being created as the person, and A method comprising: performing at least one of inpainting or outpainting based on the second image and the first information using a generative AI model (330), and obtaining a third image (440, 540, 640) generated by the generative AI model such that the image area corresponding to the first portion is not related to the person by performing at least one of the inpainting or the outpainting.

9. In paragraph 8, the method, An operation of determining whether the first part is related to an animal based on the user input; and Based on the determination that the first part is related to the animal, an operation of obtaining a third image (440, 540, 640) generated by the generative AI model is performed based on the first command by performing at least one of the inpainting or the outpainting so that the image area corresponding to the first part is not related to the animal. The above first information includes information that prevents the image area corresponding to the first part from being created as the animal, The above first information includes mask information related to the first part, A method wherein the mask information includes a graphic object set to distinguish the first portion and information that prevents a portion corresponding to the graphic object from being created as the person or the animal.

10. In the 8th or 9th clause, the operation of obtaining the second image excluding the first portion from the first image and the first information related to the first portion, An action of determining whether a face of the person or the animal is detected in the first part; An operation of generating the first information for the first part based on determining that the face of the person or the animal is detected in the first part; and A method including an operation of obtaining the second image excluding the first portion from the first image using a preprocessing model (320) stored in the memory.

11. In any one of clauses 8 to 10, the operation of obtaining the second image excluding the first portion from the first image and the first information related to the first portion, An operation of identifying a first ratio value occupied by the first portion in a first bounding box including the first portion based on determining that the face of the person or the animal is not detected in the first portion; and A method comprising: generating the first information for the first portion based on the first ratio value being greater than or equal to a first threshold value; and obtaining the second image excluding the first portion from the first image using the preprocessing model.

12. In any one of clauses 8 to 11, the operation of obtaining the second image excluding the first portion from the first image and the first information related to the first portion comprises: An operation of identifying a second ratio value of a second bounding box including the first portion based on determining that a face of the person or the animal is not detected in the first portion; and A method comprising: generating the first information for the first portion based on the second ratio value being less than the second threshold value; and obtaining the second image excluding the first portion from the first image using the preprocessing model.

13. In any one of clauses 8 to 12, the method, An operation of determining whether a plurality of parts representing at least one of the person or the animal in the first image are located at a boundary of the first image based on identifying the plurality of parts; An operation of maintaining in the first image a second portion of the first image that is not located on the boundary line among the plurality of portions without excluding the second portion; An action of controlling the display to display the first image on the execution screen of an application related to image generation based on the user input; An operation of controlling the display to display the third image on the execution screen based on obtaining the third image; and A method further comprising: transmitting the third image to an external electronic device through a communication circuit of the electronic device based on acquiring the third image.

14. In a non-transitory storage medium storing a program, the program, when executed by at least one processor (120, 210) of an electronic device (101, 201), causes the electronic device to: An action of obtaining a user input requesting removal of a first portion (411) of a first image (410) stored in a memory (130) of the electronic device using a user interface including the first image; An action to determine whether the first part is related to a person based on the user input; An operation of obtaining a second image (420) excluding the first part from the first image and first information related to the first part based on the confirmation that the first part is related to the person; the first information includes information for preventing an image area corresponding to the first part from being created as the person; A non-transitory storage medium including executable commands to perform an operation of obtaining a third image (440, 540, 640) generated by the generative AI model such that the image area corresponding to the first portion is not related to the person, based on a first command to perform at least one of inpainting or outpainting based on the second image and the first information using the generative AI model (330).

15. In paragraph 14, An operation of determining whether the first part is related to an animal based on the user input; and Based on the determination that the first part is related to the animal, the operation of performing at least one of the inpainting or the outpainting based on the first command to obtain a third image (440, 540, 640) generated by the generative AI model such that the image area corresponding to the first part is not related to the animal is included. The above first information includes information that prevents the image area corresponding to the first part from being created as the animal, The above first information includes mask information related to the first part, The above mask information includes a graphic object set to distinguish the first portion and information that prevents a portion corresponding to the graphic object from being created as the person or the animal. The operation of obtaining a second image (420) excluding the first portion from the first image and first information related to the first portion is as follows: An operation of determining whether a face of the person or the animal is detected in the first part; and A non-transitory storage medium comprising executable instructions to execute an operation of generating the first information for the first portion based on determining that a face of the person or the animal is detected in the first portion.

Citation Information

Patent Citations

  • Organic compounds and organic light-emitting device comprising the same

    KR1020250063840A

  • A method for removing objects in an image using deep learning and an apparatus for the same

    KR102594092B1

  • Spherical bearing for bridge

    KR102726195B1

  • Method, apparatus, and computer-readable medium for foreground object deletion and inpainting

    WO2023250088A1