Electronic device and image generation method using artificial intelligence model in electronic device
An AI model in electronic devices addresses the issue of unnatural object movements by inpainting and modifying areas, ensuring moved objects appear natural and complete in their new positions.
Patent Information
- Application Number
- PCT/KR2024/020066
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-10
- Filing Date
- 2024-12-09
- Publication Date
- 2025-07-10
AI Technical Summary
Existing electronic devices struggle with unnatural appearances of objects when they are cropped and moved within images due to incomplete shapes or being obscured, leading to visual inconsistencies.
Utilizing an artificial intelligence model to synthesize images by inpainting and modifying areas based on mask information, ensuring that moved objects appear natural in their new positions.
The AI model effectively generates images where moved objects seamlessly integrate with their new backgrounds, maintaining natural appearances and completing incomplete shapes.
Smart Images

Figure KR2024020066_10072025_PF_FP_ABST
Abstract
Description
Electronic devices and methods for generating images using artificial intelligence models in electronic devices
[0001] The present disclosure relates to an electronic device and a method for generating an image using an artificial intelligence model in the electronic device, according to one embodiment.
[0002] Thanks to remarkable advancements in information and communication technology and semiconductor technology, the proliferation and use of various electronic devices is rapidly increasing. Electronic devices are being developed to enable users to carry and communicate with one another. An electronic device can refer to any device that performs a specific function based on its embedded software, such as a mobile terminal, tablet PC, audio / video device, desktop / laptop computer, or in-vehicle navigation system. Electronic devices can display images through a display and include functions for editing images to create new images. For example, electronic devices can delete or move certain objects within an image and save an image with some objects deleted or moved as a new image.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.
[0004] When an electronic device crops and moves an object within an image, it can create a new image by placing the object in the moved area and creating an image related to the background of the object in the area where the object was and synthesizing the image.
[0005] When an object within an image is cropped and moved, the object may appear unnatural within the image. For example, if an object has an incomplete shape, with only a portion of the object remaining within the image's boundaries, and the object is moved to a different area, the incomplete object may appear unnatural in that area. Furthermore, if an object is obscured by another object and the object is moved to a different area, the object's incomplete shape, excluding the obscured portion, may appear unnatural in that area.
[0006] According to one embodiment, an electronic device can generate a second image synthesized with a first image by generating an incomplete part of a moved object so that the shape of the moved object appears natural within the image by using an artificial intelligence model when the object moves in the first image, and a method for generating an image using an artificial intelligence (AI) model in the electronic device can be provided.
[0007] An electronic device according to an embodiment of the present disclosure may include a display, at least one processor including a processing circuit, and a memory including one or more storage media storing instructions. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display a first image on the display through a user interface. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to crop a portion of the first image corresponding to an object through the user interface, and move the cropped portion from a first position to a second position relative to the first image to obtain a second image. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain mask information indicating at least one area of the second image to be modified through inpainting based on a first position of the cropped portion and a second position of the cropped portion. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to apply the second image and the mask information to an AI model. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a third image generated by the AI model based on the second image and the mask information. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display the third image on the display.
[0008] A method for generating an image using an artificial intelligence model of an electronic device according to an embodiment of the present disclosure may include an operation of displaying a first image on a display through a user interface. The method may include an operation of cropping a portion of the first image corresponding to an object through the user interface, and moving the cropped portion from a first position to a second position with respect to the first image to obtain a second image. The method may include an operation of obtaining mask information indicating at least one area of the second image to be modified through inpainting based on the first position of the cropped portion and the second position of the cropped portion. The method may include an operation of applying the second image and the mask information to an AI model. The method may include an operation of obtaining a third image based on the second image and the mask information by the AI model. The method may include an operation of displaying the third image on the display.
[0009] According to one embodiment of the present disclosure, a non-transitory storage medium storing commands is provided, wherein the commands, when executed by an electronic device, are configured to cause the electronic device to perform at least one operation, wherein the at least one operation may include: displaying a first image on a display through a user interface; cropping a portion of the first image corresponding to an object through the user interface and moving the cropped portion from a first position to a second position based on the first image to obtain a second image; obtaining mask information indicating at least one area of a second image to be modified through inpainting based on the first position of the cropped portion and the second position of the cropped portion; applying the second image and the mask information to an AI model; obtaining a third image generated by the AI model based on the second image and the mask information; and displaying the third image on the display.
[0010] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.
[0011] Figure 2 is a schematic block diagram of an electronic device according to one embodiment.
[0012] Figure 3 is a flowchart illustrating an image generation operation using an AI model according to one embodiment.
[0013] FIG. 4A is a diagram showing exemplary screens when moving an object in a first area of a first image to another area according to one embodiment.
[0014] FIG. 4b is a diagram illustrating exemplary screens for obtaining a second image in which an object in a first area of a first image is moved to a different area according to one embodiment.
[0015] FIG. 4c is a diagram illustrating a case where a second area is identified in a second image to ensure that a moved object according to one embodiment exhibits a subject shape that matches a background different from the previous background.
[0016] FIG. 4D is a diagram illustrating an example of obtaining a third image by transferring mask information obtained based on a second image and a first region and a second region to an AI model according to one embodiment.
[0017] FIG. 4e is a diagram for explaining an example of obtaining an image corresponding to the mask information by transferring the mask information obtained based on the second image and the first and second regions to an AI model according to one embodiment, and obtaining a third image by synthesizing the image corresponding to the mask information with the second image.
[0018] FIG. 4f is a diagram showing a second image and mask information obtained when an object in a first area is copied to another area and the object in the first area is erased using an AI eraser according to one embodiment.
[0019] FIG. 4g is a diagram illustrating an example of obtaining a third image by transferring a second image and mask information obtained using an AI eraser according to one embodiment to an AI model.
[0020] FIG. 5 is a drawing showing an example of determining an area between the edge where the object was touching and the area to which it has moved as a second area in a second image when the object was touching the edge in a first image before moving according to one embodiment.
[0021] FIG. 6 is a drawing showing an example of determining a second area in a second image when an object is reduced and moved to a different area while touching the edge of a first image before moving according to one embodiment, and a second image is generated.
[0022] FIG. 7A is a diagram showing examples of screens when a user selects an object in a first area of a first image and moves it to another area according to one embodiment.
[0023] FIG. 7b is a diagram illustrating exemplary screens in which a second area is determined by a user using an area interface according to one embodiment.
[0024] FIG. 7c is a drawing showing a screen in which an image corresponding to a first area and an image corresponding to a second area determined by a user are generated and synthesized onto a second image according to one embodiment.
[0025] FIG. 8A is a diagram illustrating an example of determining a second area in a second image based on a user input using an area interface including an object when selecting and moving an object in a first image according to one embodiment.
[0026] FIG. 8b is a diagram illustrating an example of a case where a second region is determined in a second image based on a partially expanded region and a lost region in a state where an area interface including an object according to one embodiment is partially expanded.
[0027] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to an embodiment. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with the electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (104) and the server (108) via a second network (199) (e.g., a long-range wireless communication network). According to an embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0028] The processor (120) may control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing, for example, software (e.g., a program (140)), and may perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculation, the processor (120) may store a command or data received from another component (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the command or data stored in the volatile memory (132), and store the resulting data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0029] The auxiliary processor (123) may control at least a part of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0030] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0031] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0032] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0033] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. According to one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0034] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0035] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0036] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0037] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0038] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0039] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0040] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0041] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0042] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0043] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0044] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0045] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the selected at least one antenna. According to some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0046] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0047] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0048] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In one embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0049] In the detailed description below, reference numerals in the drawings may be used interchangeably with or omitted for configurations that can be easily understood through the preceding embodiments, and their detailed descriptions may also be omitted. An electronic device according to an embodiment disclosed in this document may be implemented by selectively combining configurations of various embodiments, and a configuration of one embodiment may be replaced by a configuration of another embodiment. For example, it should be noted that the present disclosure is not limited to a specific drawing or embodiment.
[0050] Figure 2 is a block diagram of an electronic device according to one embodiment.
[0051] Referring to FIG. 2, an electronic device (201) according to an embodiment (e.g., the electronic device (101) of FIG. 1) may include at least one processor (220) (hereinafter also referred to as the processor (220)), a memory (230), and a display (260). The electronic device (201) according to an embodiment is not limited thereto and may further include various components or may be configured by excluding some of the components. The electronic device (201) according to an embodiment may further include all or part of the electronic device (101) illustrated in FIG. 1.
[0052] A processor (220) according to an embodiment (e.g., processor (120) of FIG. 1) may include some or all of a central processing unit (CPU), an application processor (AP), a neural processing unit (NPU), a data processing unit (DPU), and a graphics processing unit (GPU), and may include a hardware structure (e.g., AI chip) specialized for processing an artificial intelligence (AI) model. The processor (220) according to an embodiment may perform an overall control operation of the electronic device (201), and may perform an image generation operation using an AI model (e.g., a first AI model (e.g., AI model (232)) within the electronic device and / or an external second AI model (of an external device or a server)). For example, the AI model may include a generative artificial intelligence model. For example, the electronic device (201) may perform an image generation operation using an AI model using an AP, or may detect the performance of an image generation operation through the AP and then provide a command for the image generation operation to an NPU, a DPU, or a GPU, and cause the DPU or the GPU to perform the image generation operation. The processor (220) according to one embodiment may use a first AI model (232) and / or a second AI model, and the first AI model (232) may be an on-device AI model, and the first AI model (232) and the second AI model may have the same or different processing capacity size, quality, and / or learning amount.
[0053] According to one embodiment, a processor (220) may display an image (e.g., an original image or a first image, hereinafter also referred to as a “first image”) on a display (260).
[0054] According to one embodiment, the processor (220) may display a first image through a user interface on a display (260). According to one embodiment, the processor (220) may perform an image generation operation using an AI model when a user edits the first image based on the user interface.
[0055] In one embodiment, the processor (220) may crop a portion of a first image corresponding to an object, and move the cropped portion from a first location to a second location based on the first image to obtain a second image. For example, the processor (220) may crop (or segment from a background) a portion including an object (or an image of a subject) at a first location of a first image (e.g., a photograph), and obtain a second image in which the cropped portion is moved to a second location. The movement of the cropped portion according to one embodiment may include changing the location of the cropped portion by reducing or enlarging the cropped portion.
[0056] According to one embodiment, the processor (220) may obtain mask information indicating at least one area (or portion) of the second image to be changed (e.g., modified through inpainting) based on a first position of the cropped portion and a second position of the cropped portion. The mask information according to one embodiment may include a mask image and / or binary information indicating an area designated to be inpainted and an area not designated to be inpainted in the image (e.g., the second image). The mask information according to one embodiment may include a mask image and / or binary information indicating an area designated to be inpainted in the second image as white (e.g., pixel value 1) and an area not designated to be inpainted as black (e.g., pixel value 0).
[0057] In one embodiment, the processor (220) may identify at least one area to be modified through inpainting in a second image based on movement of a cropped portion from a first location to a second location (e.g., based on analysis results of an object corresponding to the cropped portion or user input regarding the cropped portion), and obtain mask information representing the identified at least one area. For example, inpainting may mean redrawing or filling a portion or all of the image.
[0058] According to one embodiment, at least one region may include a first region and a second region. According to one embodiment, the processor (220) may identify a first region corresponding to an empty space of a first location based on the movement of the cropped portion. According to one embodiment, the processor (220) may identify a second region designated between the first location and the second location to transform (in its entire form) an object corresponding to the cropped portion based on the movement of the cropped portion.
[0059] According to one embodiment, the processor (220) may specify a second area between the edge and the second location to modify the object corresponding to the cropped portion by filling in a portion of the shape of the object corresponding to the cropped portion when the first location of the cropped portion contacts the edge of the first image and the cropped portion is moved from the first location to the second location.
[0060] According to one embodiment, the processor (220) may apply the second image and mask information to a first AI model (e.g., AI model (232)) of the electronic device (201) or a second AI model of an external device.
[0061] In one embodiment, when the processor (220) applies the second image and mask information to the AI model (232) of the electronic device (201), the processor (220) may execute the AI model (232) stored in the memory (232), input the second image and mask information into the executed AI model (232), and obtain a third image generated through inpainting of at least one area of the second image by the AI model (232) based on the second image and mask information.
[0062] In one embodiment, the processor (220) may transmit the second image and the mask information to the second AI model of the external device using a communication circuit (not shown) included in the electronic device (201) when applying the second image and the mask information to the second AI model of the external device, and may receive the third image generated through inpainting of at least one area of the second image by the second AI model based on the second image and the mask information.
[0063] According to one embodiment, the object may include an image of a subject belonging to any one of the subject categories of a structure, an object, a person, an animal, and / or a plant, and may also include an image of a subject belonging to another specified subject category. According to one embodiment, the processor (220) may separate a portion (e.g., an object (or an area of an object)) of a first location in the first image and a background area (or crop a portion of the first image) using an outline of the selected portion (e.g., an object) of the first image based on a user's touch or an external input device (e.g., a stylus, a mouse, or another input device) for selecting and moving a portion (e.g., an object (or an area of an object)) of the first location within the first image. According to one embodiment, the processor (220) may move the separated (or cropped) portion from the first location to a second location and obtain a second image in which the separated (or cropped) portion is moved to the second location. For example, as a portion (e.g., an object) of a first location in a first image is moved to a second location, the second image may include blank space at the first location and a cropped portion at a second location that has a different background than the first location.
[0064] According to one embodiment, when a cropped portion of a first position of a first image is moved from the first position to a different second position, the moved portion may be placed on a background different from the original background (of the first position), which may appear unnatural. According to one embodiment, the cropped portion of the first image may include an object corresponding to the cropped portion, and the object may be an image of a subject belonging to any one of a structure, an object, a person, an animal, and / or a plant subject category, and the image of the subject may represent the complete form of the subject or a part of the complete form of the subject. For example, the object corresponding to the cropped portion of the first position may touch at least one edge (e.g., an edge or a side) of the first image, and the complete form of the subject (e.g., a head-to-toe form of the subject when the subject is a human) may be cropped by the at least one edge, and may be displayed as a part of the complete form of the subject (e.g., a part of a head-to-toe form of the human when the subject is a human (e.g., an upper body)). When an object representing a part of a complete form of a subject that was displayed while touching at least one edge is moved from a first position where the object is touching at least one edge to a second position that is further away from the edge, the object representing a part of the complete form of the subject may appear unnatural because it appears as if the subject is cropped against a different background at the second position. For another example, when an object corresponding to a cropped part at the first position is covered by an object corresponding to another part and is displayed as a part of the complete form rather than the complete form of the subject, when a part containing an object corresponding to a part of the complete form of the subject is moved from the first position to a second position that is not covered by another object, the object corresponding to the moved part may appear at the second position that has a different background in a form that is missing as much as the part that was previously covered by the other object, which may appear unnatural.As another example, if an object corresponding to a part of a first position touches at least one edge of a first image and a complete shape of the subject is cut off by the at least one edge and is displayed as a part of the complete shape of the subject, and the object representing a part of the complete shape of the subject that was displayed while touching the at least one edge is reduced and the reduced object is moved from the first position where it touches the edge of the first image to a second position where it does not touch the edge, the reduced object may appear unnatural as the subject is cut off at the second position that has a different background.
[0065] In one embodiment, the processor (220) may identify (or select) a second area to be modified through inpainting in the second image so that an object corresponding to a portion moved from a first location to a second location becomes an object in the form of a subject that fits (or matches or is natural) in the background of the second location, which is different from the background of the first location.
[0066] According to one embodiment, the processor (220) can analyze an object corresponding to a moved portion and identify (or select) a second area to be modified through inpainting in a second image based on the analysis result of the object corresponding to the moved portion.
[0067] In one embodiment, the processor (220) may selectively perform one or more of the following operations 1. to 4. when identifying a second area to be modified in a second image using the analysis result of an object corresponding to a moved portion.
[0068] 1. According to an embodiment, a processor (220) may identify an area (e.g., a rectangular area) between the edge where the cropped portion was touching and the moved second position in the second image, as a second area to be modified in the second image, if the cropped portion was touching at least one edge of the first image before moving the cropped portion in the first image (e.g., if the cropped portion includes the edge (boundary) of the first image).
[0069] 2. According to an embodiment, the processor (220) identifies an object corresponding to the cropped portion through segmentation of the cropped portion when the cropped portion touches at least one edge of the first image before moving the cropped portion in the first image (e.g., when the cropped portion includes an edge (boundary) of the first image), and uses an object detection engine (or software or program) to identify which subject the identified object corresponds to (e.g., a subject category to which the object belongs) and whether the object is a shape that includes the entirety of the subject (e.g., a complete shape) or a part (e.g., a part of a complete shape), and identifies an area in the second image for making the object represent the complete shape of the subject as a second area that has been modified in the second image. For example, the object detection engine may include instructions that, when executed by the processor (220), cause the electronic device (201) (or the processor (220)) to find objects within the first image, detect which subject category each of the found objects belongs to, and identify what percentage of the subject each object satisfies. For example, the object detection engine may include an object detection algorithm including regions with convolutional neutral network (R-CNN) or you only look once (YOLO).
[0070] 3. According to an embodiment, the processor (220) can identify whether an object corresponding to the cropped portion was covered by another object in another uncropped portion before moving the cropped portion using a depth estimation algorithm. Depth estimation according to an embodiment can be performed using a depth map that expresses the distance value corresponding to the object for each pixel in the image as a gray image. In the depth map, each pixel can have a brighter value as the distance gets closer and a darker value as the distance gets farther. The depth map can be learned by a deep learning network such as U-Net, and the processor (220) estimates the depth of objects identified in the first image using a depth estimation algorithm, and when an object corresponding to a cropped portion is covered by another object in another uncropped portion, estimates the complete shape of the subject corresponding to the object corresponding to the cropped portion using an object detection engine, and identifies a second area to be modified in the second image so that the object corresponding to the cropped portion represents the complete shape of the subject.
[0071] 4. In one embodiment, the processor (220) may further identify whether the object corresponding to the cropped portion contains a character if the object corresponding to the cropped portion is cropped by at least one edge of the first image or is obscured by another object in another uncropped portion and represents a part of the complete shape of the subject rather than the complete shape of the subject (e.g., a part of the complete shape of the subject is missing). In one embodiment, the processor (220) may recognize the character to obtain text corresponding to the character if the object corresponding to the cropped portion is cropped by at least one edge of the first image or is obscured by another object and represents a part of the complete shape of the subject rather than the complete shape of the subject, and may obtain expanded text information and a second area to be modified in the second image so that the object represents the complete shape of the subject and includes the expanded text.
[0072] According to one embodiment, the processor (220) can identify (or select) a second area to be modified in the second image based on an area selected by user input (or selection) in association with an object corresponding to the cropped portion after movement of the cropped portion.
[0073] According to one embodiment, the processor (220) may selectively perform one or more of the following operations 5 to 7 when identifying a second area to be modified in a second image based on an area selected by user input (or selection) in association with an object corresponding to the cropped area after movement of the cropped area.
[0074] 5. According to one embodiment, the processor (220) can display an area interface including an object corresponding to the cropped portion after moving the cropped portion, expand the area interface based on a user input, and identify an area excluding the object in the expanded area interface as a second area.
[0075] 6. According to an embodiment, a processor (220) may display an area interface including an object corresponding to the cropped area after moving the cropped area, estimate a lost area of the object corresponding to the cropped area using a depth estimation algorithm and / or an object detection engine, provide information about the estimated lost area, and identify an area including the lost area as a second area by excluding the object from the area interface based on a user input.
[0076] 7. According to an embodiment, a processor (220) may display an area interface including an object corresponding to the cropped area after moving the cropped area, estimate a lost area of the object using a depth estimation algorithm and / or an object detection engine, provide information about the estimated lost area, expand the area interface based on a user input, and identify a area including the lost area and excluding the object from the expanded area interface as a second area.
[0077] A processor (220) according to an embodiment may obtain mask information indicating a first area and / or a second area in which an image is to be generated (e.g., inpainted) based on a first position corresponding to a cropped portion of a first image and a second position after the cropped portion has been moved. The mask information according to an embodiment may include binary information indicating an area (e.g., the first area and / or the second area) in which an AI model (e.g., the first AI model or the second AI model) is to generate an image in the second image.
[0078] In one embodiment, the processor (220) may apply the second image and mask information to an AI model (e.g., the first AI model or the second AI model) and obtain a third image generated through the AI model. In one embodiment, the processor (220) may transfer (or transmit) (or input) the second image and the mask information to the AI model to generate images corresponding to the first area and the second area, and synthesize the generated image with the second image to obtain a third image. In one embodiment, the processor (220) may apply the second image and the mask information to the AI model to generate an image corresponding to the first area or an image corresponding to the second area, and synthesize the generated image corresponding to the first area or the image corresponding to the second area with the second image to obtain a third image. In one embodiment, the processor (220) may further apply extended text information together with the second image and the mask information to the AI model to generate an image including characters corresponding to the second area, and synthesize the generated image including characters with the second image to obtain a third image.
[0079] According to one embodiment, the processor (220) can display the third image through the display (260).
[0080] An AI model (232) according to an embodiment may include one or more AI models, and may include a generative AI model. For example, a generative AI model may refer to an artificial intelligence (AI) model that generates similar content by utilizing content such as text, audio, and / or images. A generative AI model may learn patterns of content and generate new content based on inference results. For example, in the field of images, a generative AI model may regenerate an image that mimics the characteristics of a specific image. For example, a generative AI model may include a generative adversarial network (GAN)-based image generation model or a stable diffusion-based image generation model. A stable diffusion-based image generation model may include an OpenAI-based Dall-E model, an Imagen model, a Midjourney model, or a stable diffusion model.
[0081] A memory (230) according to an embodiment (e.g., memory (130) of FIG. 1) may include one or more storage media for storing instructions. A memory (230) according to an embodiment may store various data used by at least one component (e.g., processor (220) and / or display (260)) of an electronic device (201). The data may include, for example, software (e.g., software module or program (140)) and input data or output data for commands related thereto. A memory (230) according to an embodiment may store an AI model (232), and may store various data generated during program execution, including a program (e.g., software or program (140) of FIG. 1) for generating an image using the AI model (232). According to one embodiment, a memory (230) may store commands (or instructions) that cause the processor (220) to perform an image generation operation (or method) using the AI model (232) of the present disclosure.
[0082] According to one embodiment, a display (260) (e.g., display (160) of FIG. 1) may display various information based on the control of a processor (220). According to one embodiment, the display (260) may display a screen related to performing an image generation operation (or method) using an AI model (232). According to one embodiment, the display (260) may be implemented in the form of a touch screen. When the display (260) is implemented together with an input module in the form of a touch screen, it may display various information generated according to a user's touch operation.
[0083] According to one embodiment, the electronic device (201) is not limited to the configuration illustrated in FIG. 2 and may further include various components. According to one embodiment, the electronic device (201) further includes an input module (not shown) (e.g., the input module (150) of FIG. 1) and may receive various inputs related to performing an image generation operation (or method) using an AI model through the input module. According to one embodiment, the electronic device (201) further includes a communication circuit (not shown) (e.g., the communication module (190) of FIG. 1) and may perform a portion of the image generation operation in a divided manner while performing communication with an external device (e.g., an external device including a second AI model) through the communication circuit. The communication circuit according to one embodiment may include a wired communication module (e.g., a USB communication module) and / or a wireless communication module (e.g., a cellular module, a Wi-Fi (wireless-fidelity) module, a Bluetooth module, or an NFC (near field communication) module).
[0084] In one embodiment, the main components of the electronic device have been described through the electronic device (201) of FIG. 2. However, in various embodiments, not all of the components illustrated through FIG. 2 are essential components, and the connection relationship of the main components of the electronic device (201) described above through FIG. 2 may be changed according to various embodiments.
[0085] An electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (201) of FIG. 2) according to an embodiment may include a display (e.g., the display module (160) of FIG. 1 or the display (260) of FIG. 2), at least one processor (120, 220) including a processing circuit, and a memory (130, 230) including one or more storage media for storing instructions. The instructions according to an embodiment, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to display a first image on the display through a user interface. The instructions according to an embodiment, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to crop a portion of the first image corresponding to an object through the user interface, and move the cropped portion from a first position to a second position with respect to the first image to obtain a second image. In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain mask information indicating at least one area of the second image to be modified through inpainting based on the first location of the cropped portion and the second location of the cropped portion. In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to apply the second image and the mask information to an AI model. In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a third image generated by the AI model based on the second image and the mask information.The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may cause the electronic device to display the third image on the display.
[0086] According to one embodiment, the at least one region may include a first region and a second region.
[0087] The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a first area corresponding to an empty space at the first location based on movement of the cropped portion.
[0088] The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a second area designated between the first location and the second location to deform the object corresponding to the cropped portion based on movement of the cropped portion.
[0089] The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may cause the electronic device to designate the second area between the edge and the second location to modify the object corresponding to the cropped portion by filling in a portion of the shape of the object corresponding to the cropped portion when the first location of the cropped portion contacts an edge of the first image and the cropped portion is moved from the first location to the second location.
[0090] In one embodiment, the AI model may include a first AI model stored in the memory. In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to input the second image and the mask information into the first AI model, and obtain the third image generated by inpainting the at least one area of the second image by the first AI model based on the second image and the mask information. In one embodiment, the AI model may include a second AI model of an external device. In one embodiment, the electronic device further includes a communication circuit, and the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to transmit the second image and the mask information to the second AI model of the external device using the communication circuit, and receive the third image generated by inpainting the at least one area of the second image by the second AI model based on the second image and the mask information.
[0091] The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may cause the electronic device to identify the object through segmentation of the cropped portion, estimate a subject represented by the object and a shape of the subject using an object detection engine, and identify the shape of the object based on the shape of the estimated subject.
[0092] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify, using a depth estimation algorithm, whether the object corresponding to the cropped portion of the first image is occluded by another object corresponding to another portion of the first image, and, if the object corresponding to the cropped portion of the first image is occluded by the other object corresponding to the other portion, identify, using an object detection engine, the shape of a subject represented by the object, and estimate a lost portion of the shape of the object that is lost in the shape of the object based on the shape of the subject represented by the object.
[0093] The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may cause the electronic device to display an area interface including the cropped portion after movement of the cropped portion through the user interface, expand the area interface based on a user input, and identify the second area designated in the expanded area interface to deform the object corresponding to the cropped portion based on the movement of the cropped portion.
[0094] The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may cause the electronic device to estimate a lost portion of the shape of the object corresponding to the cropped portion using a depth estimation algorithm and / or an object detection engine after movement of the cropped portion, and to provide information about the lost portion of the estimated shape of the object.
[0095] According to one embodiment, the mask information may include binary information indicating at least one area in the second image to be inpainted by the first AI model or the second AI model.
[0096] According to one embodiment, the first AI model or the second AI model may include a generative adversarial network (GAN)-based image generation model or a stable diffusion-based image generation model.
[0097] Figure 3 is a flowchart illustrating an image generation operation using an AI model according to one embodiment.
[0098] Referring to FIG. 3, a processor (e.g., processor (120) of FIG. 1 or processor (220) of FIG. 2, hereinafter, processor (220) of FIG. 2 will be described as an example) of an electronic device (e.g., electronic device (101) of FIG. 1 or electronic device (201) of FIG. 2) according to one embodiment may perform at least one operation from operations 310 to 360.
[0099] In operation 310, the processor (220) according to one embodiment may display a first image on the display (260). The processor (220) according to one embodiment may display a user interface on the display (260) and display the first image through the user interface. For example, the user interface may provide a screen that can perform an image generation operation using an AI model when editing an image by a user.
[0100] In operation 320, the processor (220) according to one embodiment may crop a portion of a first image corresponding to an object, and move the cropped portion from a first location to a second location based on the first image to obtain a second image. The processor (220) according to one embodiment may crop (or segment from a background) a portion including an object (or an image of a subject) at a first location of a first image (e.g., a photograph) based on an input through a user interface, and obtain a second image in which the cropped (or segmented from a background) portion is moved to a different second location. For example, the object may be an image of a subject belonging to any one of a structure, an object, a person, an animal, and / or a plant, or may be an image of a subject belonging to another specified subject category. According to an embodiment, a processor (220) may segment (or separate an object) from an object and a background area outside the outline of the object corresponding to the cropped portion based on a user's touch or input module (not shown) (e.g., a stylus, a mouse, or other input device) for selecting and moving a portion (or area) within a first image (e.g., a touch and drag or a click and drag). According to an embodiment, the processor (220) may move the separated (or cropped) portion to a second location and obtain a second image in which the separated (or cropped) portion is moved to the second location. For example, as the cropped portion of the first image at the first location is moved to the second location, the second image may include an empty space at the first location and may include a cropped portion at the second location that has a different background from the first location.
[0101] In one embodiment, when a cropped portion of a first position of a first image is moved from the first position to a different second position, an object corresponding to the cropped portion may appear unnatural because it is placed on a background different from the original background (of the first position). According to one embodiment, the cropped portion of the first image may include an object corresponding to the cropped portion, and the object may be an image of a subject belonging to any one of a structure, an object, a person, an animal, and / or a plant subject category, and the image of the subject may represent the entire form of the subject or a portion of the entire form of the subject. For example, the object corresponding to the cropped portion may touch at least one edge (e.g., an edge or a side) of the first image, and the entire form of the subject (e.g., a head-to-toe form of the person if the subject is a human) may be cropped by the at least one edge, so that a portion, but not all, of the entire form of the subject (e.g., a portion (e.g., an upper body) of a head-to-toe form of the person if the subject is a human) may be displayed. When an object representing a part of a complete form of a subject that was displayed while touching at least one edge is moved from a first position where the object is touching at least one edge to a second position that is further away from the edge, the object representing a part of the complete form of the subject may appear unnatural because it appears as if the subject is cropped against a different background at the second position. For another example, when an object corresponding to a cropped part at the first position is covered by another object corresponding to a different part and is displayed as a part of the complete form rather than the complete form of the subject, when a part containing an object corresponding to a part of the complete form of the subject is moved from the first position to a second position that is not covered by another object, the object corresponding to the moved part may appear unnatural because it appears at the second position that has a different background and a part of the complete form of the subject that has been lost by the amount that was previously covered by the other object.As another example, if an object corresponding to a cropped portion of a first position touches at least one edge of a first image and a complete shape of the subject is cropped by the at least one edge and displayed as a part of the complete shape of the subject, and the object representing a part of the complete shape of the subject that was displayed while touching the at least one edge is reduced and the reduced object is moved from the first position where it touches the edge of the first image to a second position where it does not touch the edge, the reduced object may appear unnatural as the subject is cropped at the second position that has a different background.
[0102] In operation 330, the processor (220) according to one embodiment may obtain mask information indicating at least one area of the second image to be modified through inpainting based on a first position of the cropped portion and a second position of the cropped portion. The processor (220) according to one embodiment may identify at least one area to be modified through inpainting in the second image based on a movement of the cropped portion from the first position to the second position (e.g., based on an analysis result of an object corresponding to the cropped portion or a user input for the cropped portion), and obtain mask information indicating the identified at least one area. For example, inpainting may mean redrawing or filling a part or all of the image. According to one embodiment, the at least one area may include a first area and a second area. According to one embodiment, the processor (220) may identify a first area corresponding to an empty space of the first position based on the movement of the cropped portion. According to one embodiment, the processor (220) may identify a second area designated between a first location and a second location to transform (in its entire form) an object corresponding to the cropped portion based on the movement of the cropped portion. According to one embodiment, when the first location of the cropped portion contacts an edge of the first image and the cropped portion is moved from the first location to the second location, the processor (220) may designate a second area between the edge and the second location to modify the object corresponding to the cropped portion by filling in a portion of the shape of the object corresponding to the cropped portion.
[0103] According to an embodiment, the processor (220) may designate a second region based on an analysis result of an object corresponding to the cropped portion after movement of the cropped portion and / or an area selected by a user input through a user interface for the object corresponding to the cropped portion. According to an embodiment, the processor (220) may designate a second region so that, after movement of the cropped portion from a first location to a second location, the object corresponding to the cropped portion becomes an object in the form of a subject that matches (or matches or is natural) with the background of the second location, which is different from the background of the first location.
[0104] According to one embodiment, the processor (220) may designate (or identify) a second area between the edge and the second location to modify the object corresponding to the cropped portion by filling in a portion of the shape of the object corresponding to the cropped portion when the first location of the cropped portion contacts the edge of the first image and the cropped portion is moved from the first location to the second location.
[0105] According to one embodiment, the processor (220) may designate (or identify) as a second region an area (e.g., a rectangular area) between at least one edge that the cropped portion was touching before the cropped portion was moved (e.g., if the cropped portion included an edge (boundary) of the first image) and a second position after the cropped portion was moved.
[0106] In one embodiment, the processor (220) may analyze the cropped portion using a segmentation and / or object detection engine (or software or program) when using the object analysis result, and if it is identified that the object corresponding to the cropped portion represents a part (or a part of the complete form) of the subject rather than the complete form (or the entire complete form), the processor may designate (or identify) as a second region an area in the second image for making the object corresponding to the cropped portion have the complete form of the subject. In one embodiment, the processor (220) may use a depth estimation algorithm and an object detection engine to analyze an object corresponding to a cropped portion when using an object analysis result, and if it is determined that the object corresponding to the cropped portion is covered by an object corresponding to a portion other than the cropped portion and that the object corresponding to the cropped portion represents a part of the complete form rather than the complete form of the subject, the processor may identify a second area to be modified in the second image so that the object corresponding to the cropped portion represents the complete form of the subject. In one embodiment, the processor (220) may analyze an object corresponding to a cropped portion when using the object analysis result, and if the object corresponding to the cropped portion before moving the cropped portion is cut off by at least one edge or covered by another object and includes a part of the complete shape of the subject rather than the complete shape of the subject (e.g., another part of the complete shape of the subject is lost), and if the object is identified as including a character, recognize the character to obtain text corresponding to the character, apply the text to an LLM (large language model) to obtain an extended text including the text, and obtain a second area to be modified in the second image and extended text information so that the object includes the complete shape of the subject and includes the extended text.
[0107] In one embodiment, the processor (220) may, when using an area selected by a user input for a cropped portion, display an area interface including the cropped portion through a user interface, expand the area interface based on the user input, and identify an area excluding an object corresponding to the cropped portion in the expanded area interface as a second area. In one embodiment, the processor (220) may, when using an area selected by a user input for a cropped portion, display an area interface including the cropped portion, estimate a lost portion of an object corresponding to the cropped portion using a depth estimation algorithm and / or an object detection engine, provide information on the estimated lost portion, and identify an area excluding an object and including the lost portion in the area interface as a second area based on the user input. In one embodiment, the processor (220) may, when using an area selected by a user input for a cropped portion, display an area interface including the cropped portion, estimate a lost portion of an object corresponding to the cropped portion using a depth estimation algorithm and / or an object detection engine, provide information about the estimated lost portion, expand the area interface based on the user input, and identify an area including the lost portion, excluding the object from the expanded area interface, as a second area.
[0108] In operation 340, the processor (220) according to one embodiment can apply an AI model (a first AI model (e.g., AI model (232)) of the electronic device (201) and / or a second AI model of the external device).
[0109] In one embodiment, when the processor (220) applies the second image and mask information to the AI model (232) of the electronic device (201), it may execute the AI model (232) stored in the memory (232), input the second image and mask information into the executed AI model (232), and obtain a third image generated through inpainting of at least one area of the second image by the AI model (232) based on the second image and the mask information. In one embodiment, when the processor (220) applies the second image and mask information to the second AI model of the external device, it may transmit the second image and the mask information to the second AI model of the external device using a communication circuit (not shown) included in the electronic device (201), and receive the third image generated through inpainting of at least one area of the second image by the second AI model based on the second image and the mask information.
[0110] In operation 350, the processor (220) according to one embodiment can obtain a third image generated through inpainting of at least one area of the second image based on the second image and mask information by the AI model.
[0111] In one embodiment, the processor (220) may obtain images of the first region and the second region generated by the AI model by transmitting (or applying or inputting) the second image and mask information to the AI model, and may synthesize the images of the first region and the second region generated with the second image to obtain a third image. In one embodiment, the processor (220) may also generate images including characters corresponding to the first region and / or the second region by further applying extended text information together with the second image and mask information to the AI model, and may synthesize the generated image including characters with the second image to obtain a third image. The AI model (232) or the second AI model according to one embodiment may include one or more generative AI models. For example, a generative AI model may refer to an artificial intelligence (AI) model that newly creates similar content by utilizing content such as text, audio, and / or images. The generative AI model may learn patterns of content and generate new content as an inference result. For example, in the image domain, a generative AI model can regenerate images that mimic the characteristics of a specific image. For example, a generative AI model may include a generative adversarial network (GAN)-based image generation model or a stable diffusion-based image generation model. Stable diffusion-based image generation models may include the OpenAI-based Dall-E model, the Imagen model, the Midjourney model, or the stable diffusion model.
[0112] In a 360 motion, the processor (220) according to one embodiment may display the third image on the display (260). The third image according to one embodiment may be an image generated by naturally changing an image of an empty space in the first area (or an image of a space where an object has disappeared) in the second image and an unnatural object image in the second area by an AI model (232).
[0113] In one embodiment, a method for generating an image using an artificial intelligence model in an electronic device (e.g., the electronic device 101 of FIG. 1 or the electronic device 201 of FIG. 2) may include an operation of displaying a first image through a user interface on a display (e.g., the display module 160 of FIG. 1 or the display 260 of FIG. 2). In one embodiment, the method may include an operation of cropping a portion of the first image corresponding to an object through the user interface, and moving the cropped portion from a first position to a second position with respect to the first image to obtain a second image. In one embodiment, the method may include an operation of obtaining mask information indicating at least one area of the second image to be modified through inpainting based on the first position of the cropped portion and the second position of the cropped portion. In one embodiment, the method may include an operation of applying the second image and the mask information to the AI model. In one embodiment, the method may include an operation of obtaining a third image generated based on the second image and the mask information by the AI model. The method according to one embodiment may include an action of displaying the third image on the display.
[0114] In one embodiment, the method may further include an operation of identifying the first area corresponding to an empty space of the first location based on movement of the cropped portion, wherein the at least one area includes a first area and a second area, and an operation of identifying the second area specified between the first location and the second location to deform the object corresponding to the cropped portion based on movement of the cropped portion.
[0115] In one embodiment, the method may further include an operation of specifying a second area between the edge and the second location to modify the object corresponding to the cropped portion by filling a portion of the shape of the object corresponding to the cropped portion when the first location of the cropped portion contacts an edge of the first image and the cropped portion is moved from the first location to the second location.
[0116] In a method according to one embodiment, the AI model may include the first AI model stored in the memory of the electronic device. The method according to one embodiment may further include an operation of inputting the second image and the mask information into the first AI model, and an operation of obtaining the third image generated by the first AI model based on the second image and the mask information.
[0117] In one embodiment, the method may include a second AI model of an external device. In one embodiment, the method may further include an operation of transmitting the second image and the mask information to the second AI model of the external device using a communication circuit of the electronic device, and an operation of receiving the third image generated by inpainting the at least one area of the second image by the second AI model based on the second image and the mask information.
[0118] According to one embodiment, the method may include an operation of identifying the object through segmentation of the cropped portion. The method may include an operation of estimating a subject represented by the object and a shape of the subject represented by the object using an object detection engine. The method may include an operation of identifying the shape of the object based on the estimated shape of the subject.
[0119] In one embodiment, the method may include an operation of identifying whether the object corresponding to the cropped portion of the first image is occluded by another object corresponding to another portion of the first image using a depth estimation algorithm. The method may include an operation of identifying a shape of a subject represented by the object using an object detection engine when the object corresponding to the cropped portion of the first image is occluded by the other object corresponding to the other portion. The method may include an operation of estimating a lost portion of the shape of the object that is lost in the shape of the object based on the shape of the subject represented by the object.
[0120] According to one embodiment, the method may include an operation of displaying an area interface including the cropped portion through the user interface after moving the cropped portion. The method may include an operation of expanding the area interface based on a user input. The method may include an operation of identifying the second area designated in the expanded area interface to deform the object corresponding to the cropped portion based on the movement of the cropped portion.
[0121] According to one embodiment, the method may include an operation of estimating a lost portion of the shape of the object corresponding to the cropped portion using a depth estimation algorithm and / or an object detection engine after moving the cropped portion. The method may include an operation of providing information about the estimated lost portion.
[0122] FIG. 4A is a diagram showing exemplary screens when moving a cropped portion of a first image from a first position to a second position according to one embodiment.
[0123] Referring to FIG. 4a, the processor (220) (or the processor (120) of FIG. 1) of the electronic device (201) (or the electronic device (101) of FIG. 1) according to one embodiment <4001> A first image (410) can be displayed on the display (260) (or the display (160) of FIG. 1) as shown. For example, the first image (410) can be a photograph including a plurality of objects.
[0124] According to one embodiment, the processor (220) <4002> In response to a user input of selecting a portion (412) of a first position of a first image (410) (e.g., a portion including an image of a subject (or a person object) belonging to the category of a person (or a person)) through a user interface, the selected portion (412) may be separated (or cropped) from the background through segmentation, and the cropped portion (412) or a graphic effect (or an outline of the cropped portion) (414) representing the cropped portion (412) may be displayed on the first image (410). In one embodiment, the processor (220) may display the cropped portion (412) by moving (arranging) it to a second position in response to a user input (e.g., a touch input of a user's hand) of moving the cropped portion (412) to a second position.
[0125] FIG. 4b is a diagram illustrating exemplary screens for obtaining a second image in which a cropped portion of a first image is moved to a different second position according to one embodiment.
[0126] Referring to FIG. 4b, the processor (220) (or the processor (120) of FIG. 1) of the electronic device (201) (or the electronic device (101) of FIG. 1) according to one embodiment <4003> Inland <4005> A cropped portion (e.g., an object) (412) may be moved to a second location and a second image (420) may be temporarily or non-temporarily acquired (or stored) in the memory (230). For example, the processor (220) <4003> In response to a user input to move the cropped portion (412) from the first position to a second position, an empty space may be displayed in the first area (416) of the first position, and the cropped portion (412) may be placed in front of a different background of the second position. For example, the processor (220) <4004> The cropped portion (412) may be copied (duplicated) according to a user input to move the cropped portion (412) from the first location to a different second location, and the copied cropped portion (412) may be placed in front of a different background at a different second location. For example, the processor (220) <4005> The cropped portion (412) can be placed in front of a different background at a different second location according to a user input to move the cropped portion (412) from a first area (416) at a first location to a different second location, and the first area (416) can be deleted by being filled with the surrounding background using an AI eraser.
[0127] FIG. 4c is a drawing for explaining a case in which a second area is identified in a second image so that an object corresponding to a cropped portion after moving the cropped portion according to one embodiment exhibits a subject shape that matches the background of a second position that is different from the background of a first position.
[0128] Referring to FIG. 4C, the processor (220) (or the processor (120) of FIG. 1) of the electronic device (201) (or the electronic device (101) of FIG. 1) according to one embodiment may analyze an object corresponding to the cropped portion after the cropped portion is moved, and identify (or select or determine) a second area (418) to be modified in the second image (420) based on the analysis result of the object. According to one embodiment, when the cropped portion (412) is moved from a first location to a different second location, the object corresponding to the cropped portion may be placed on a background of the second location that is different from the background of the first location, which may appear unnatural. For example, if an object corresponding to a cropped portion (412) touches an edge (e.g., an edge or a side) (450) of a first image (410), and the complete shape of the subject (person) is cut off by the edge (450) and does not appear to be the complete shape of the subject (person), and the cropped portion (412) is moved from the edge (450) to another position (e.g., a second position) in the center direction, the object image (e.g., the subject (person)) of the cropped portion (412) is placed on a different background at the second position and appears as if the subject (person) is cut off, which may appear unnatural.
[0129] In one embodiment, the processor (220) analyzes the cropped portion (412) to identify whether the object corresponding to the cropped portion (412) was in contact with the edge (450) before the movement, and identifies an area (e.g., a rectangular area) between the edge (450) where the object corresponding to the cropped portion (412) was in contact with the cropped portion (412) before the movement and the second position after the movement of the cropped portion (412) as the second area (418) to be modified in the second image (420). In one embodiment, the processor (220) may analyze the cropped portion (412) using a segmentation and object detection engine (or software or program) and, if the object corresponding to the cropped portion (412) is identified as a subject of the human category and the object represents a part of a human form rather than a complete human form, the processor may identify (or select or determine) an area between the edge (450) where the object was touching before moving and the second position after the movement of the cropped portion as a second area (418) to be modified in the second image (420) so that the object represents a complete human form in the second image (420). In one embodiment, the processor (220) may display the second area (418) on the display (260) or may not display the second area (418) on the display (260). According to one embodiment, a processor (220) can obtain location information or values (e.g., coordinate values) of a second area (418).
[0130] In one embodiment, the processor (220) may expand the second region (418) according to additional information. For example, after identifying the second region (418), the processor (220) may expand the second region (418) to reflect the shadow of the object according to shadow information according to the projection direction of light (the direction of light projected onto the object) within the second image (420).
[0131] FIG. 4D is a diagram illustrating an example of obtaining a third image by transmitting mask information obtained based on a second image and at least one area according to one embodiment to an AI model.
[0132] Referring to FIG. 4D, the processor (220) (or the processor (120) of FIG. 1) of the electronic device (201) (or the electronic device (101) of FIG. 1) according to one embodiment may obtain mask information (425) indicating an area (419) to be modified (from which the first AI model (232) or the second AI model will generate an image) in the second image using the first area (416) and / or the second area (418). The processor (220) according to one embodiment may transfer (or apply or input) the second image (420) and the mask information (425) to the first AI model (e.g., the AI model (232)) or the second AI model, and obtain a third image (430) generated through the first AI model or the second AI model.
[0133] FIG. 4e is a diagram for explaining an example of obtaining an image corresponding to the mask information by transferring the mask information obtained based on the second image and the first and second regions to an AI model according to one embodiment, and obtaining a third image by synthesizing the image corresponding to the mask information with the second image.
[0134] Referring to FIG. 4E, the processor (220) (or the processor (120) of FIG. 1) of the electronic device (201) (or the electronic device (101) of FIG. 1) according to one embodiment can obtain mask information (425) indicating an area (419) to be modified in the second image (from which the first AI model (232) or the second AI model will generate an image) using the first area (416) and / or the second area (418).
[0135] According to an embodiment, a processor (220) may transmit (or apply, or input) a second image (420) and mask information (425) to a first AI model (e.g., AI model (232)) or a second AI model to generate an image corresponding to the mask information (e.g., an image (439) corresponding to the first area (416) and the second area (418). According to an embodiment, a processor (220) may obtain a third image (430) by using (or synthesizing) the second image (420) and the generated image (439).
[0136] According to one embodiment, the third image (430) may be an image generated by naturally changing an image of an empty space in the first area (416) (or an image of a space where an object (412) has disappeared) and an unnatural image in the second area (418) in the second image (420) by a first AI model (e.g., AI model (232)) or a second AI model.
[0137] FIG. 4f is a diagram illustrating a second image and mask information obtained when a cropped portion of a first location is copied to a second location according to an embodiment and an empty space of the cropped portion of the first location is erased using an AI eraser. FIG. 4g is a diagram illustrating an example of obtaining a third image by transferring the second image and mask information obtained using an AI eraser according to an embodiment to an AI model.
[0138] First, referring to FIG. 4F, a processor (220) (or a processor (120) of FIG. 1) of an electronic device (201) (or an electronic device (101) of FIG. 1) according to an embodiment may duplicate a portion (4120) of a first image (4100) and move the copied portion to display it in front of a different background at a second location different from the first location, in response to a user input for moving a portion (4120) of a first location of a first image (4100) to another location. The processor (220) according to an embodiment may obtain mask information (4251) (e.g., first mask information) indicating a first area (4160) corresponding to an image (4101) in which a portion (4121) of the moved first image (4100) is displayed in front of a different background at the second location and a portion (4120) of the first image (4100) before the movement. According to one embodiment, the processor (220) may transmit (or apply or input) an image (4101) in which a part of a moved first image (4100) is displayed in front of a different background at a different second location and first mask information (4251) to the AI eraser (231), thereby obtaining a second image (4200) in which a part (4121) of the moved first image (4100) is displayed in front of a different background at a second location and a part (4120) of the first image (4100) before the movement is deleted (or backgrounded). In one embodiment, the processor (220) may analyze an object corresponding to a portion (4121) of a moved first image (4100) and, based on the analysis result of the object corresponding to the portion (4121) of the moved first image (4100), identify a second area (4180) to be modified in an image (4101) in which the object corresponding to the portion (4121) of the moved first image (4100) is displayed in front of a different background at a second location. In one embodiment, the processor (220) may obtain mask information (4252) (e.g., second mask information) indicating an area to be modified through inpainting in the second image (4200) using the first area (4160) and the second area (4180).
[0139] Referring to FIG. 4g, a processor (220) according to one embodiment may transmit (or apply or input) a second image (4200) and second mask information (4252) to a first AI model (e.g., AI model (232)) or a second AI model, and obtain a third image (4300) generated based on the second image (4200) and second mask information (4252) through the first AI model (e.g., AI model (232)) or the second AI model.
[0140] FIG. 5 is a drawing showing an example of designating an area between the edge where the cropped portion of the first position of the first image was touching the edge of the first image and the second position where the cropped portion was moved as a second area to be modified in the second image according to one embodiment.
[0141] Referring to FIG. 5, the processor (220) (or the processor (120) of FIG. 1) of the electronic device (201) (or the electronic device (101) of FIG. 1) according to one embodiment <5001> Based on the input corresponding to the selection, the part (512) that was in contact with the edge of the first image (510) can be selected.
[0142] According to one embodiment, the processor (220) <5002> Based on the input corresponding to the movement, the cropped portion (512) of the first image (510) can be moved from the first position to the second position, and the second image (520) in which the cropped portion (512) is moved to the second position can be displayed. According to one embodiment, the processor (220) can identify the second area based on the distance (d) that the cropped portion (512) has moved from the edge (550) of the first image (510).
[0143] According to one embodiment, the processor (220) <5003> As such, the area (518) between the edge (550) of the first image (510) and the distance (d) by which the cropped portion (512) has moved can be designated (or identified) as the second area to be modified in the second image (520).
[0144] FIG. 6 is a drawing showing an example of specifying a second area in a second image when a second image is acquired by reducing the cropped portion and moving it from a first position to a second position while the cropped portion was touching the edge of the first image before moving according to one embodiment.
[0145] Referring to FIG. 6, a processor (220) according to an embodiment may select and crop an image portion at a first location based on a user input through a user interface and reduce the cropped image portion (616). The processor (220) according to an embodiment may display the cropped and reduced image portion (612) at a second location different from the first location as the cropped image portion (616) is cropped and reduced. The processor (220) according to an embodiment may display an area interface (e.g., a handler) (or area selection interface) (615) indicating an area including the cropped and reduced image portion (612) at the second location, and may display at least one editing menu (617) for the cropped and reduced image portion (612) at the top of the area interface (e.g., a menu for inverting an object corresponding to the cropped and reduced image portion, a menu for deleting an object, or a menu for rotating an object).
[0146] In one embodiment, the processor (220) may identify an area (618) based on a distance (d) between an edge (650) of a second image (620) that the image portion (616) (or an object at the first position) was in contact with and a second position corresponding to the image portion (612) when the image portion (612) at the first position is reduced and moved to a different second position. In one embodiment, the processor (220) may determine the area (618) identified based on a distance (d) between an edge (650) of the second image (620) that the image portion (616) was in contact with and the second position according to the movement of the image portion (616) as a second area to be modified in the second image.
[0147] FIGS. 7A to 7C may illustrate examples of identifying (or selecting) a second area to be modified in a second image based on an area selected by user input when a portion of a first location in a first image is moved to a different second location according to one embodiment.
[0148] FIG. 7A is a diagram showing examples of screens when a user selects a portion of a first position of a first image and moves it to a second position different from the first position, according to one embodiment.
[0149] Referring to FIG. 7a, the processor (220) (or the processor (120) of FIG. 1) of the electronic device (201) (or the electronic device (101) of FIG. 1) according to one embodiment <7001> The first image (710) may be displayed on a user interface (e.g., an image editing application screen (762)) through a display (260) (or a display (160) of FIG. 1). For example, the image editing application screen (762) may further display various function icons (or menus) related to editing of the first image (710). For example, the first image (710) may be a photograph including a plurality of objects. The processor (220) according to an embodiment may display a guide phrase (711) in response to an input (e.g., a tap or touch input) for an object (712) corresponding to a first part of a first location (e.g., a plate with bread and a fork on top). For example, the guide phrase (711) may include a phrase guiding a user input method for the object (712). For example, the guide phrase (711) may include a phrase such as “Press and hold the object, then change the position and size.” In one embodiment, a processor (220) may segment an object (712) corresponding to a first portion of a first location of a first image (710) in response to a user input (71) (e.g., tap & hold or long press) to separate the object (712) from the background.
[0150] According to one embodiment, the processor (220) <7002> In response to a user input (73) (e.g., drag) for moving an object (712) corresponding to a first part of a first position to a second position different from the first position, the object (712) can be displayed in a moving state (e.g., floating state).
[0151] According to one embodiment, the processor (220) <7003> When the object (712) of the first position (or first area) (716) is moved to a different second position (release after drag), a second image (720) can be displayed and an area interface (or area selection interface) (e.g., handler) (715) including the object (712) can be displayed.
[0152] According to an embodiment, the area interface (715) may display an apply menu (713) and a menu (714) (e.g., a draw button) for allowing a user to select a second area to be changed (or modified) in the second image (720). According to an embodiment, the processor (220) may apply a state in which an object corresponding to a first part of a first location is moved when the apply menu (713) is selected. According to an embodiment, the area interface (715) may be changed to a state in which a user can select (or adjust) the second area to be changed in the second image (720) when the menu (714) is selected.
[0153] FIG. 7b is a diagram illustrating exemplary screens in which a second area (718) is determined by a user using an area interface (715) according to one embodiment.
[0154] Referring to FIG. 7b, the processor (220) (or the processor (120) of FIG. 1) of the electronic device (201) (or the electronic device (101) of FIG. 1) according to one embodiment <7004> The processor (220) according to one embodiment may receive a user input (74) for adjusting (e.g., expanding) the area interface (715). The processor (220) according to one embodiment may display the edit button (717) at the bottom of the area interface (715) in an inactive state (e.g., dim state) while the area interface (715) is being adjusted (or expanded). The processor (220) according to one embodiment may <7005> When the adjustment (or expansion) of the area interface (715) is completed, the edit button (717) for creating an expansion area can be activated (e.g., the dim state can be released) to display it. In one embodiment, the processor (220) can determine the second area (718) to be changed in the second image (720) based on the area adjusted (expanded) through the area interface (715) when the edit button (717) is input by the user. In one embodiment, the processor (220) can determine the second area (718) to be changed in the second image (720) based on the distance (d) that the object (712) (or the first part of the first position) has moved from the edge when the object (712) (e.g., a plate with bread and a fork) is cut off by the edge and a part of the fork is lost.
[0155] According to an embodiment, the processor (220) may generate (or obtain) an image of the second area (718) by transmitting (or applying or inputting) mask information based on the second image (720) and the second area (718) to the first AI model (e.g., AI model (232)) or the second AI model as the second area (718) to be changed in the second image (720) is designated (or determined). According to an embodiment, the processor (220) may transmit (or apply or input) mask information based on the second image (720) and the second area (718) to the second AI model of an external server via the communication circuit (290), receive a plurality of images of the second area (718) from the external server, and select one image from among the plurality of images of the second area (718) by a user input. According to one embodiment, the processor (220) may display an image of the acquired (or generated or selected) second area (718) and display an apply button (719) at the bottom of the area interface (715). According to one embodiment, the processor (220) may display a third image (730) to which the image of the second area (718) is applied (or synthesized) based on an input of the apply button (719).
[0156] FIG. 7c is a drawing showing a screen in which an image corresponding to a first area and an image corresponding to a second area are generated and synthesized onto a second image according to one embodiment.
[0157] Referring to FIG. 7c, when the apply button (719) of FIG. 7b is pressed as the second area (718) is designated (or determined or identified), the processor (220) (or the processor (120) of FIG. 1) of the electronic device (201) according to one embodiment may apply mask information indicating the generation area of the second image (720) and the image corresponding to the first area (716) and the image corresponding to the second area (718) to the first AI model (e.g., AI model (232)) or the second AI model to generate an image (739) corresponding to the mask area including the first area (716) and the second area (718). According to one embodiment, the processor (220) can acquire and display a third image (730) by synthesizing the image of the first region (716) and the image (739) of the second region (718) into the second image (720).
[0158] FIG. 8A is a diagram illustrating an example of designating a second area in a second image based on a user input using an area interface that includes a portion of a selection and movement portion in a first image according to one embodiment.
[0159] Referring to FIG. 8a, the processor (220) (or the processor (120) of FIG. 1) of the electronic device (201) (or the electronic device (101) of FIG. 1) according to one embodiment <8001> In response to a user input selecting a part of an image (812) (e.g., an object (e.g., a cup), hereinafter also referred to as an object) of a first position of a first image (810), the selected object (812) can be separated from the background through segmentation.
[0160] According to one embodiment, the processor (220) <8002> When the object (812) of the first position is moved to a different second position, a second image (820) may be displayed and an area interface (e.g., handler) (815) including the moved object (812) may be displayed. According to an embodiment, the processor (220) may further display menus (813) (e.g., flip, rotate, or delete) that enable editing of the object (812) together with the area interface (815).
[0161] In one embodiment, the processor (220) may estimate a missing portion (818) of the object (812) within the selected area by using a depth estimation algorithm and / or an object detection engine when an area is selected by a user during selection and movement of the object (812). For example, the processor (220) may identify whether the object (812) was occluded by another object (e.g., whether the object is in an osculation state with another object and is further away than the other object) through the depth estimation algorithm, and if the object (812) was occluded by another object, may detect a subject category (e.g., a cup) to which the object (812) belongs using an object detection algorithm, may identify how much (e.g., what percentage) the object (812) satisfies the shape of the subject category, and may identify that the shape of the object (812) is lost based on the percentage of satisfaction of the subject category by the object (812). According to an embodiment, a processor (220) may provide a user interface (UI) (e.g., an area interface (815)) for determining a second area including a loss portion (818) based on the identification of a loss in the form of an object (812). According to an embodiment, the processor (220) may receive a designation of a loss portion (818) from a user in an area by the area interface (815). According to an embodiment, the processor (220) may determine an area excluding the object (812) in the area interface (815) and an area including the designated loss portion (818) as a second area to be corrected in a second image (820).
[0162] In one embodiment, the processor (220) may also correct the shape, contrast, and / or color of the object (812) in the area by the area interface (815). For example, if a cup object (812) has contrast due to a shadow cast by another object that was blocking the cup object and moves to another area, the contrast of the cup object (812) may also disappear due to the absence of the other object. The processor (220) may correct the contrast of the cup object (812) when the contrast of the cup object (812) should disappear due to the movement.
[0163] FIG. 8b is a diagram illustrating an example of a case where a second region is determined in a second image based on a partially expanded region and a lost portion of the shape of the object, in a state where an area interface including an object according to one embodiment is partially expanded.
[0164] Referring to FIG. 8b, the processor (220) (or the processor (120) of FIG. 1) of the electronic device (201) (or the electronic device (101) of FIG. 1) according to one embodiment <8011> Based on a user input for an area interface (815) including an object (812), a portion of the area interface (815) may be expanded and an expanded area interface (816) may be displayed. According to an embodiment, the processor (220) may expand the area interface (815) in a desired direction among the directions of each side of a rectangle (e.g., four directions).
[0165] According to one embodiment, the processor (220) <8012> In the expanded area interface (816), a loss portion (818) in the shape of an object (812) can be specified based on a user input. For example, the processor (220) can specify an area including an edge (817) of the area interface (815) to a loss point (801) in the shape of an object (812) as a loss portion (818). In an embodiment, when the display (260) of the electronic device (201) is a rollable display, the processor (220) can display an expanded area interface (816) on a display screen of an expanded size as the size of the display screen is expanded, and can also specify a loss portion (818) in the shape of an object (812) based on a user input in the expanded area interface (816).
[0166] According to one embodiment, the processor (220) <8013> In the extended selection interface (816), the area excluding the object (812) and the area including the lost part (818) of the shape of the specified object can be determined as the second area (819) in the second image (820).
[0167] In a non-transitory storage medium storing commands, the commands are configured to cause the electronic device (101, 201) to perform at least one operation when executed, wherein the at least one operation may include: displaying a first image on a display (160, 260) through a user interface; cropping a portion of the first image corresponding to an object through the user interface and moving the cropped portion from a first position to a second position based on the first image to obtain a second image; obtaining mask information indicating at least one area of the second image to be modified through inpainting based on the first position of the cropped portion and the second position of the cropped portion; applying the second image and the mask information to an AI model; obtaining a third image generated by the AI model based on the second image and the mask information; and displaying the third image on the display.
[0168] The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned will be clearly understood by a person having ordinary knowledge in the technical field to which the present disclosure pertains.
[0169] According to the present disclosure, when an object is moved in an image (e.g., a first image), an incomplete part of the moved object is generated using an artificial intelligence model so that the shape of the moved object appears natural in the image (e.g., a second image) and an acquired image (e.g., a third image) is provided, thereby enabling the moved object to appear natural when the object is moved in the image.
[0170] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains.
[0171] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments disclosed in this document are not limited to the aforementioned devices.
[0172] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0173] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0174] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more commands stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one command among the one or more commands stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one command called. The one or more commands may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0175] According to one embodiment, the method according to the various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0176] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0177] As used herein, the term “if” will be understood to mean “when, upon,” “in response to deciding,” or “in response to detecting,” depending on the context. Similarly, “if it is decided to do,” or “if [the stated condition or event] is detected,” will optionally be understood to mean “upon deciding,” or “in response to deciding,” “upon detecting [the stated condition or event],” or “in response to detecting [the stated condition or event].”
[0178] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. A processing device (or processing circuit) may execute an operating system (OS) and one or more software applications running on the operating system. In addition, the processing device may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.
[0179] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.
[0180] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program commands, including ROM, RAM, and flash memory. In addition, examples of other media may include recording or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.
[0181] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.
[0182] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.
Claims
1. In an electronic device (101, 201), display(160, 260); At least one processor (120, 220) comprising a processing circuit; and A memory (130, 230) comprising one or more storage media storing instructions, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Displaying a first image through a user interface on the above display, Crop a portion of the first image corresponding to an object through the user interface, and move the cropped portion from a first position to a second position based on the first image to obtain a second image, Obtain mask information indicating at least one area of a second image to be modified through inpainting based on the first position of the cropped portion and the second position of the cropped portion, Applying the second image and the mask information to the AI model of the electronic device, Obtaining a third image generated based on the second image and the mask information by the AI model, and An electronic device that causes the third image to be displayed on the display.
2. In paragraph 1, wherein at least one of the above regions comprises a first region and a second region, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Identifying a first area corresponding to the empty space of the first location based on the movement of the cropped portion, An electronic device that identifies a second area specified between the first location and the second location to deform the object corresponding to the cropped portion based on movement of the cropped portion.
3. In paragraph 1 or 2, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: An electronic device configured to specify a second area between the edge and the second position so as to modify the object corresponding to the cropped portion by filling a part of the shape of the object corresponding to the cropped portion when the first position of the cropped portion contacts an edge of the first image and the cropped portion is moved from the first position to the second position.
4. In any one of paragraphs 1 to 3, The above AI model includes a first AI model stored in the memory, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Input the second image and the mask information into the first AI model, An electronic device that obtains the third image generated by inpainting at least one area of the second image by the first AI model based on the second image and the mask information.
5. In any one of paragraphs 1 to 3, Further comprising a communication circuit, The above AI model includes a second AI model of an external device, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Using the above communication circuit, the second image and the mask information are transmitted to the second AI model of the external device, An electronic device configured to receive a third image generated by inpainting at least one area of the second image by the second AI model based on the second image and the mask information.
6. In any one of paragraphs 1 to 5, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Identify the object through segmentation of the cropped portion, Using an object detection engine, the subject represented by the object and the shape of the subject are estimated, and An electronic device that identifies the shape of an object based on the estimated shape of the subject.
7. In any one of paragraphs 1 to 6, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Using a depth estimation algorithm, identify whether the object corresponding to the cropped portion in the first image is obscured by another object corresponding to another portion of the first image, If the object corresponding to the cropped portion of the first image is obscured by another object corresponding to the other portion, the shape of the subject represented by the object is identified using an object detection engine, and An electronic device that estimates a lost portion of the shape of an object based on the shape of the subject represented by the object.
8. In any one of paragraphs 1 to 7, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: After moving the cropped portion through the user interface, display an area interface including the cropped portion, Extending the above area interface based on user input, and An electronic device that identifies the second area specified in the extended area interface to deform the object corresponding to the cropped area based on movement of the cropped area.
9. In any one of paragraphs 1 to 8, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: After moving the cropped portion, the lost portion of the shape of the object corresponding to the cropped portion is estimated using a depth estimation algorithm and / or an object detection engine, An electronic device for providing information about said lost portion of said estimated shape of said object.
10. In any one of paragraphs 1 to 9, The above mask information includes binary information representing at least one area in the second image to be inpainted by the first AI model or the second AI model, An electronic device wherein the first AI model or the second AI model includes a GAN (generative adversarial network)-based image generation model or a stable diffusion-based image generation model.
11. In a method for generating an image using an artificial intelligence model in an electronic device (101, 201), An act of displaying a first image via a user interface on a display (160, 260); An operation of cropping a portion of the first image corresponding to an object through the user interface, and moving the cropped portion from a first position to a second position based on the first image to obtain a second image; An operation of obtaining mask information indicating at least one area of the second image to be modified through inpainting based on a first position of the cropped portion and a second position of the cropped portion; An action of applying the second image and the mask information to the AI model; An operation of obtaining a third image generated based on the second image and the mask information by the AI model; and A method comprising the action of displaying said third image on said display.
12. In paragraph 11, wherein at least one of the above regions comprises a first region and a second region, An operation of identifying the first area corresponding to the empty space of the first location based on the movement of the cropped portion; and A method further comprising an action of identifying the second area specified between the first location and the second location to deform the object corresponding to the cropped portion based on the movement of the cropped portion.
13. In paragraph 11 or 12, A method further comprising: specifying a second area between the edge and the second position to modify the object corresponding to the cropped portion by filling a part of the shape of the object corresponding to the cropped portion when the first position of the cropped portion contacts an edge of the first image and the cropped portion is moved from the first position to the second position.
14. In any one of paragraphs 11 to 13, The AI model comprises the first AI model stored in the memory of the electronic device, An operation of inputting the second image and the mask information into the first AI model; and A method further comprising an action of obtaining a third image generated by inpainting at least one area of the second image by the first AI model based on the second image and the mask information.
15. In a non-transitory storage medium storing commands, the commands are set to cause the electronic device (101, 201) to perform at least one operation when executed by the electronic device, the at least one operation being: An act of displaying a first image via a user interface on a display (160, 260); An operation of cropping a portion of the first image corresponding to an object through the user interface, and moving the cropped portion from a first position to a second position based on the first image to obtain a second image; An operation of obtaining mask information indicating at least one area of a second image to be modified through inpainting based on a first position of the cropped portion and a second position of the cropped portion; An action of applying the second image and the mask information to the AI model; An operation of obtaining a third image generated based on the second image and the mask information by the AI model; and A storage medium comprising an action of displaying the third image on the display.
Citation Information
Patent Citations
Aligner of wafer and apparatus for depositing wafer having the same
KR1020250010782A
Organic compounds and organic light-emitting device comprising the same
KR1020250063840A
Apparatus and method for augmenting data using object segmentation and background synthesis
KR102482262B1
Image inpainting apparatus and method for thereof
KR102486300B1
Spherical bearing for bridge
KR102726195B1