Electronic device, method, and non-transitory storage medium for image editing using generative artificial intelligence model
The described method allows for seamless object editing within a single application using generative AI, addressing limitations of conventional techniques by enabling multiple transformations and natural results through user interactions.
Patent Information
- Application Number
- PCT/KR2025/008706
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-22
- Filing Date
- 2025-06-23
- Publication Date
- 2026-01-02
AI Technical Summary
Conventional image editing techniques using generative AI models are limited by requiring predefined user interactions and multiple steps across different applications, failing to produce novel transformations and natural results.
An electronic device and method that utilizes a generative AI model to enable seamless object editing within a single application, allowing for multiple transformations through user interactions, including shape deformation of objects using a generative AI model based on user inputs.
Enables comprehensive image editing operations within a single application, facilitating natural and novel transformations of objects through user interactions, enhancing user convenience and efficiency.
Smart Images

Figure KR2025008706_02012026_PF_FP_ABST
Abstract
Description
Electronic devices, methods, and non-transitory storage media for image editing using generative artificial intelligence models
[0001] The present disclosure relates to an electronic device, method and non-transitory storage medium for image editing using a generative artificial intelligence model.
[0002] With the advancement of digital technology, electronic devices are available in various forms, such as smartphones, tablet personal computers (PCs), personal digital assistants (PDAs), XR, VR, AR, and HMDs. Electronic devices are also being developed into wearable forms to enhance portability and accessibility. Recently, users have gone beyond simply capturing images using electronic devices to acquire high-quality images from high-end camera equipment, and are increasingly interested in technologies for various image editing.
[0003] Previous image editing techniques simply used object information recognized within the image. This limited the scope of image editing, relying solely on object information within the image. Not only did this technique fail to produce novel image transformations, it also failed to produce natural results during the post-editing process.
[0004] Recently, technologies for generating images using artificial intelligence technology (e.g., generative AI models) are actively being developed, and technologies for easily editing images using artificial intelligence technology (e.g., generative AI models) are in demand for image editing.
[0005] Conventional image editing features are configured to respond only to predefined and recognizable user interactions (e.g., hand drawings, gestures, touches). However, this approach not only imposes significant limitations on image editing processes based on generative AI models, but also inconveniently requires multiple user input steps to perform different types of edits (e.g., object deletion, debluring, image creation), requiring the use of other applications or switching screens.
[0006] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.
[0007] The present disclosure seeks to provide an electronic device and method that allows for editing images in a manner not previously provided in a single application, and allows for further object editing following the results, thereby enabling as many editing operations as possible in a single application.
[0008] According to one embodiment of the present disclosure, an electronic device may include a display, at least one processor including a processing circuit, and a memory storing instructions.
[0009] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to control the display to display a first image.
[0010] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to control the identification of a first object among a plurality of objects included in the first image, the first object being capable of being transformed by interaction.
[0011] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to control the generation of a second object, wherein at least a portion of the shape of the first object is deformed, using a generative AI model, based at least in part on a first user input with respect to the first object.
[0012] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to control the display to display a second image that changes the first object into the second object.
[0013] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to identify a third object among the plurality of objects, the third object being capable of being deformed by interaction with the second object image.
[0014] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to generate a fourth object, wherein at least a portion of the shape of the third object is deformed, using the generative AI model, at least in part based on a second user input for the second object.
[0015] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to control the display to display a third image that changes the third object into the fourth object.
[0016] According to one embodiment, a method of operating in an electronic device may include displaying a first image on a display of the electronic device.
[0017] According to one embodiment, the method includes an operation of identifying a first object, the shape of which can be transformed by interaction, among a plurality of objects included in the first image.
[0018] According to one embodiment, the method comprises generating a second object having at least a portion of a shape of the first object deformed using a generative AI model, based at least in part on a first user input for the first object.
[0019] According to one embodiment, the method includes an action of displaying a second image, which is a second object image that has been changed from the first object image to the second object image, on the display.
[0020] According to one embodiment, the method includes an operation of identifying a third object among the plurality of objects whose shape can be transformed by interaction with the second object image.
[0021] According to one embodiment, the method comprises generating a fourth object image in which at least a portion of the shape of the second object is deformed using the generative AI model, at least in part based on a second user input regarding the second object.
[0022] According to one embodiment, the method includes displaying a third image on the display that changes the third object into the fourth object.
[0023] According to one embodiment, a non-transitory storage medium storing one or more programs, wherein the programs, when executed by at least one processor of an electronic device, cause the electronic device to: display a first image on a display of the electronic device;
[0024] According to one embodiment, the program includes instructions that, when executed by at least one processor of the electronic device, cause the electronic device to perform an operation of identifying a first object, the first object being capable of being transformed by interaction, among a plurality of objects included in the first image.
[0025] According to one embodiment, the program comprises instructions that, when executed by at least one processor of the electronic device, cause the electronic device to perform an operation of generating a second object, wherein at least a portion of the shape of the first object is deformed, using a generative AI model, based at least in part on a first user input for the first object.
[0026] According to one embodiment, the program includes instructions that, when executed by at least one processor of the electronic device, cause the electronic device to perform an operation of displaying a second image that changes the first object into the second object on the display.
[0027] According to one embodiment, the program includes instructions that, when executed by at least one processor of the electronic device, cause the electronic device to perform an operation of identifying a third object among the plurality of objects, the third object being capable of being deformed by interaction with the second object image.
[0028] According to one embodiment, the program includes instructions that, when executed by at least one processor of the electronic device, cause the electronic device to perform an operation of generating a fourth object, wherein at least a portion of the shape of the third object is deformed, using the generative AI model, based at least in part on a second user input for the second object.
[0029] According to one embodiment, the program includes instructions that, when executed by at least one processor of the electronic device, cause the electronic device to perform an operation of displaying a third image that changes the third object into the fourth object on the display.
[0030] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.
[0031] FIG. 2 is a diagram illustrating a configuration of an electronic device according to one embodiment.
[0032] FIG. 3 is a diagram illustrating a configuration of an electronic device for image editing using a generative artificial intelligence model according to one embodiment.
[0033] FIGS. 4A, 4B, and 4C are diagrams illustrating examples of image editing in an electronic device according to one embodiment.
[0034] FIGS. 5A, 5B, and 5C are diagrams illustrating examples of image editing in an electronic device according to one embodiment.
[0035] FIG. 6 is a drawing showing an example of an operating method in an electronic device according to one embodiment.
[0036] FIG. 7 is a diagram illustrating an example for image editing using a generative artificial intelligence model according to one embodiment.
[0037] FIG. 8 is a diagram illustrating an example for image editing using a generative artificial intelligence model according to one embodiment.
[0038] FIG. 9 is a diagram illustrating an example of image editing using a generative artificial intelligence model according to one embodiment.
[0039] FIG. 10 is a diagram illustrating an example of image editing using a generative artificial intelligence model according to one embodiment.
[0040] FIG. 11 is a diagram illustrating an example of an image editing operation method using a generative artificial intelligence model according to one embodiment.
[0041] FIG. 12 is a diagram illustrating an example of image editing using a generative artificial intelligence model according to one embodiment.
[0042] FIG. 13 is a diagram illustrating an example of image editing using a generative artificial intelligence model according to one embodiment.
[0043] In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0044] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components. In addition, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness. The term "user" used in the embodiments of the present disclosure may refer to a person using an electronic device or a device (e.g., an artificial intelligence electronic device) using an electronic device.
[0045] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to various embodiments. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0046] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor)) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0047] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0048] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).
[0049] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0050] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0051] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0052] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0053] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0054] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0055] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0056] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0057] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0058] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0059] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).
[0060] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0061] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0062] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for realizing 1eMBB, a loss coverage (e.g., 164 dB or less) for realizing mMTC, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for realizing URLLC.
[0063] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0064] According to various embodiments, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0065] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0066] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0067] Hereinafter, the configuration of the electronic device will be specifically described with reference to the drawings. For convenience of explanation in the descriptions of the configuration and operation method of the electronic device with reference to the drawings below, the term for an image input for image editing is referred to as an “original image” or “first image”, objects classified in an image input for image editing are referred to as “objects”, and an image portion for an object identified for image editing among the classified objects is referred to as an “object image”.
[0068] FIG. 2 is a diagram illustrating a configuration of an electronic device according to one embodiment, and FIG. 3 is a diagram illustrating a configuration of an electronic device for image editing using a generative artificial intelligence model according to one embodiment.
[0069] Referring to FIGS. 2 and 3, an electronic device (201) according to one embodiment (e.g., the electronic device (101) of FIG. 1) may include at least one processor (210), a memory (220), a display (230), a camera (240), and a communication circuit (250). Without being limited thereto, the electronic device (201) may be implemented identically or similarly to the electronic device (101) of FIG. 1, and may further include other components of the electronic device (101) described in FIG. 1.
[0070] According to one embodiment, the electronic device (201) may execute an application (e.g., a program, a function, or a module) for image editing and creation when executed by at least one processor (210), acquire an image stored in a memory (220) or an image captured by a camera (240), and display the acquired image as an input image for editing (hereinafter, referred to as an original image (301)) on an execution screen of the application displayed on the display (230). The electronic device (201) may perform editing and creation (e.g., inpainting or outpainting) on the original image (301) using an artificial intelligence (hereinafter, referred to as AI) model (300). An AI model (300) for image editing may include a preprocessing model (310) that performs a preprocessing operation on an original image (301) (e.g., a first image), and a generative AI model (320) (e.g., a first generative AI model (320a), a second generative AI model (320b)) that generates a result image (303) (e.g., a second image or a third image) by modifying or adding a shape of a portion of an image of the original image (301). The generative AI model (320) illustrated in FIG. 3 has been described as being divided into a first generative AI model (320a) and a second generative AI model (320b), but may be configured as a single generative AI model without being divided into a first generative AI model (320a) and a second generative AI model (320b). The present invention is not limited thereto, and may further include other models necessary for image generation, or may not include a generative AI model (320). For example, the generative AI model (320) may be included in an external electronic device (e.g., electronic device (102, 104) or server (108) of FIG. 1).According to one embodiment, when the generative AI model (320) is included in an external electronic device, the electronic device (201) can transmit a command for image generation to the external electronic device through a communication circuit (e.g., the communication module (190) of FIG. 1 and the communication circuit (250) of FIG. 2), and can obtain a result image (303) generated by the generative AI model (320) from the external electronic device through the communication circuit.
[0071] According to one embodiment, the preprocessing model (310) may perform preprocessing operations, such as a segmentation model (311), a depth estimation model (313), and / or a scene analysis model (315), which identify and classify objects (e.g., image parts) included in the original image (301). According to one embodiment, the preprocessing model (310) may identify at least one object (e.g., a first object and a second object) that can be transformed in shape by interaction among a plurality of objects classified in the original image (301). The preprocessing model (310) may be an AI model for performing the preprocessing operations.
[0072] According to one embodiment, in response to receiving a first user input, the preprocessing model (310) may transmit the original image (301) and preprocessing result information for the original image (301) to the first generative AI model (320a). The preprocessing result information may include at least one of classification information (e.g., at least one first image portion that is capable of shape deformation among the classified image portions), detected depth information, or scene analysis result information.
[0073] According to one embodiment, in response to receiving a second user input (e.g., a second user interaction), the preprocessing model (310) may transmit the original image (301) and preprocessing result information for the original image (301) to the second generative AI model (320a).
[0074] According to one embodiment, a segmentation model (311) can extract features from an original image (301) to classify a plurality of objects in the original image (301) and identify the locations of the classified objects. For example, the features extracted from the original image (301) may include color, texture, and / or shape. The classification model (310) can partially compose an image using the extracted features, create a segment map, and define information based on pixel-level boundaries. The classification model (310) can distinguish each segment classified in the original image (301) into an object that can be transformed (e.g., a person, an animal, a plant, or an object part) and an object that cannot be transformed (e.g., a background part), and the classification operation can be performed based on information learned through a specified deep learning.
[0075] According to one embodiment, the classification model (310) may obtain classification information for a plurality of objects classified in the original image (301), and classify (e.g., identify or detect) one or more objects (e.g., image parts) that can be transformed in shape based on the obtained classification information. According to one embodiment, the classification model (310) may use a general segmentation AI model for object recognition and object classification to recognize all objects included in the image to be edited and perform a classification operation. For example, the segmentation AI model may receive one image as input and output classification information for objects in the image as output. The classification information may include the type of the classified object (e.g., person, animal, plant, or object, etc.), the location (e.g., the coordinate values of the start and end points of the box), and the mask image (e.g., an image used when extracting the location of a specific object in the original image in units of pixels).
[0076] According to one embodiment, the depth information detection (depth estimation) model (313) can estimate the distance for all objects and backgrounds between the camera that captured the image and the captured scene from the two-dimensional image data. The depth information detection model (313) can estimate accurate depth information in real time by using an AI model that estimates depth information from a single image. According to one embodiment, the depth information is not limited to the embodiment of the AI model, and a stereo matching algorithm may be used based on stereo images captured by a binocular camera for depth information estimation, or information using a depth sensor or a distance sensor may be utilized. According to one embodiment, the depth information detection model (313) can use depth information to perform image editing reflecting a real situation to designate a maximum range in which the shape of a first object among a plurality of objects included in the original image (301) can be deformed, and can limit the range for deforming the shape of the first object so as not to exceed the designated maximum range.
[0077] According to one embodiment, the scene analysis model (315) can analyze the original image (301) and determine in what environment and atmosphere the image was taken, and can acquire scene analysis information by learning data pairs composed of context information (ground truth) corresponding to a single image input using an AI model for scene analysis of the original image (301). According to one embodiment, the scene analysis model (315) can generate an object that matches the existing scene when generating an additionally deformed object image using the learned scene analysis information. According to one embodiment, the scene analysis model (315) can transfer the scene analysis information to the generative AI model (320) (e.g., the first generative AI model (320a)). The scene analysis information can be used when generating an object image (e.g., a second object image or a fourth image object) in which at least a portion of a shape of an object whose shape can be deformed (e.g., a first object or a second object) is deformed by the generative AI model (320). Context information is data that describes an image in a language that humans can understand, and may include, for example, information about the atmosphere of the captured image (e.g., warm, noisy, calm), shooting location information, and lighting information. Here, the shooting location information may be information about the place where the image was captured based on information about objects in the image. For example, the scene analysis model (315) may recognize the scene as a study if the image contains many books or shelves based on the shooting location information, and may recognize the scene as a gym if the image contains equipment such as dumbbells or a treadmill. The scene analysis model (315) may determine whether there is a light source in the image based on the lighting information, and if there is no light source, the scene may be determined to be basically dark. Even if there is a light source but many pixel values close to black are distributed, the scene may be determined to be dark.
[0078] According to one embodiment, the generative AI model (320) may represent an AI model that newly creates similar content by utilizing content such as text, audio, and / or images. The generative AI model (320) may learn patterns of content and create new content as inference results. For example, in the field of images, the generative AI model (320) may create an image that mimics the characteristics of a specific image. According to one embodiment, the generative AI model may include a first generative AI model (320a) and a second generative AI model (320b). However, the first generative AI model (320a) and the second generative AI model (320b) may not be distinguished and may be configured as a single generative AI model. Hereinafter, the generative AI model (320) is described by dividing it into a first generative AI model (320a) and a second generative AI model (320b). However, even a single generative AI model can perform the same operations as the first generative AI model (320a) and the second generative AI model (320b). In the present disclosure, the number of generative AI models (320) is not limited to two, such as the first generative AI model (320a) and the second generative AI model (320b), and may be two or more depending on user interaction.
[0079] According to one embodiment, in response to receiving a first user input (317) regarding a first object, the first generative AI model (320a) may generate an object image (321) (e.g., a second object image) in which at least a portion of a first object, the shape of which may be changed among a plurality of objects included in the original image (301), is deformed or added based on the original image (301) (e.g., the first image) and the preprocessing result information received from the preprocessing model (310). Here, the generated object image (321) may include a deformed object (323) and / or additional objects (325). According to one embodiment, the electronic device (201) may generate a new image (e.g., a second image) by changing an object image corresponding to the first object (e.g., the first object image) into an object image generated by the first generative AI model (320a) using the first generative AI model (320a) or a processing model after image editing. According to one embodiment, the first generative AI model (320a) may be trained using a pair of data sets consisting of an initial state and a changed state (ground truth) of the first object. According to one embodiment, the first generative AI model (320a) may generate an object image in which at least a part of the shape of the first object desired by the user is modified by performing fine-tuning based on depth information and scene analysis information. The first generative AI model (320a) can pre-specify the maximum range in which the shape of the first object can change based on depth information, thereby taking into account physical correlations with other objects, and can generate additional object images associated with the first object when the shape of the first object is changed by using scene analysis information.
[0080] According to one embodiment, in response to receiving a second user input (319) for a second object, the second generative AI model (320b) may generate an object image (e.g., a fourth object image) in which at least a portion of the second object, the shape of which may be changed by interaction with the second object image (e.g., user interaction #2), is deformed or added, based on the original image (301) and the preprocessing result information received from the preprocessing model (310). According to one embodiment, the electronic device (201) may generate a new image (e.g., a third image) in which an object image corresponding to the second object (e.g., a third object image) is changed into an object image (e.g., a fourth object image) generated by the second generative AI model (320b) using the first generative AI model (320a) or a postprocessing model for image editing.
[0081] According to one embodiment, the second generative AI model (320b) can generate a result image (303) (e.g., a third image) including object images (e.g., a second object image and a fourth object image) whose shapes are transformed by interaction with a first object included in an original image (301) (e.g., a first object image corresponding to the first object) and a second object (e.g., a third object image corresponding to the second object), which is another object included in the original image (301). In order to specifically receive a user input, the object image corresponding to the second object (e.g., the third object image) can be zoomed in and displayed. The electronic device (201) can perform an image editing operation desired by the user by using various user inputs (e.g., a finger touch or an external input device) that can be received on the object image corresponding to the enlarged second object (e.g., the third object image). The second generative AI model (320b) can transform the shape of the second object by reflecting the characteristics of an object image (e.g., a second object image) in which the shape of the first object is at least partially transformed based on user input. For example, the second generative AI model (320b) can generate an object image (e.g., a fourth object image) having a result (e.g., different cut surfaces) according to an interaction (e.g., a cutting action) depending on the type (e.g., a knife or a spatula) of the object image (e.g., the second object image) that interacts with the second object (e.g., bread). According to one embodiment, the second generative AI model (320b) can learn using data pairs consisting of an object image (e.g., a second object image) in which the shape of the first object is at least partially transformed, various user inputs (e.g., a cutting, slicing, or carving action), and an object image (ground truth) for the second object in which the shape is transformed.
[0082] FIGS. 4A, 4B, and 4C are drawings illustrating examples of image editing in an electronic device according to one embodiment, and FIGS. 5A, 5B, and 5C are drawings illustrating examples of image editing in an electronic device according to one embodiment.
[0083] Referring to FIGS. 2, 3, 4a, 4b to 5c, a processor (210) of an electronic device (201) according to one embodiment may execute an application (e.g., a program or a function) for image editing, as in (a) of FIG. 4a, acquire an image stored in a memory (220) or an image captured by a camera (240) as a first image (410) for editing (e.g., the original image (301) of FIG. 3), and control a display (230) to display the acquired first image (410) on an execution screen for image editing.
[0084] According to one embodiment, the processor (210) may obtain preprocessing result information including classification information, depth information, and scene analysis information through a preprocessing operation for image editing of the first image (410) displayed on the display (230) using the preprocessing model (310). The preprocessing model (310) is an AI model, and may perform learning for image analysis in advance using a pair of learning data sets consisting of RGB information and information (ground truth) on the types and locations of various objects within the image. The preprocessing model (310) analyzes the first image (410) to provide information on the type of object (e.g., person, animal, plant, or object), coordinate values of the start and end points of a box that can indicate the location of the object on the image (e.g., x_init, y_init / x_end, y_end), and a mask image, thereby enabling the location of specific objects (e.g., the first object (411) and the second object (413)) identified in the first image (410) to be confirmed in pixel units. Here, the information on the type of object can be used to confirm whether a plurality of objects classified in the first image (410) are capable of shape transformation in the actual real world. The preprocessing model (310) learns using a pair of learning datasets consisting of a before state and a after state (ground truth) of an object image corresponding to an object capable of shape transformation, and checks whether a plurality of objects classified in the first image (410) are objects capable of shape transformation through interaction, and outputs the result of checking whether or not shape transformation is possible by matching it in binary form.
[0085] According to one embodiment, the processor (210) may display a user interface (UI) (510) (e.g., a function, a menu screen, or an input window) including a message (511) for checking whether to search for an object capable of shape transformation in the first image (410) on the execution screen (501) (“Would you like to search for a changeable object in the image?”) and / or a message (515) for inputting a word (e.g., “Please enter a word to name the object you want to focus on changing”), as illustrated in FIG. 5A. If the processor (210) selects a function (513) (e.g., a button, a menu, or a symbol) (yes) requesting a search in the function (511), the processor (210) may display the function (513) (e.g., a button, a menu, or a symbol) adjacent to the first image (410) on the execution screen (501). For example, when a word is entered into the input window (517), the processor (210) can control the display (230) to display words (e.g., refrigerator, freezer, drawer) indicating objects capable of being transformed into shapes adjacent to the input window (517) of the user interface (UI) (510). The processor (210) can learn learning data (e.g., learning parameters) used to search for objects capable of being transformed into shapes in the first image (410) in advance and store them in the memory (220) or receive them from a server (e.g., server (108) of FIG. 1) and store them in the memory (220).
[0086] According to one embodiment, the processor (210) may identify one or more objects (e.g., a drawer, a refrigerator, bread, fruit, and / or cooking utensils) that are capable of shape transformation by interaction among a plurality of objects classified based on classification information obtained through classification of the first image (410) using a preprocessing model (310), and, in response to a user input, may identify a first object image (411) corresponding to the first object (e.g., a drawer) that is capable of shape transformation among the identified one or more objects.
[0087] According to one embodiment, the processor (210) may control the display (230) to display a graphic element (415) (e.g., a graphic effect, text, or symbol) on the first image (410) indicating that one or more objects (e.g., a drawer and a refrigerator) capable of shape transformation are capable of shape transformation, as illustrated in FIG. 4A. For example, the processor (210) may display a graphic element (415) such as a sticker, a designated specific symbol, or a highlight on one or more objects (e.g., a drawer and a refrigerator), or display a separate guidance message or list of identification information for objects capable of shape transformation on the execution screen (501).
[0088] According to one embodiment, the processor (210) may use a generative AI model (320) (e.g., the first generative AI model (320a) of FIG. 3) to generate a second object image (421) in which at least a portion of a first object (e.g., a drawer) is deformed, and may control the display (230) to display a second image (420) in which the first object image (411) corresponding to the first object is changed into a second object image (421), as illustrated in FIG. 4b. According to one embodiment, the processor (210) receives a first user input (e.g., a click, a double-click, or a move gesture) for a first object in a first image (410), and transmits the first image (410), the first user input (401) (e.g., input information for an interaction corresponding to the first user input), and the preprocessing result information (e.g., classification information, depth information, and scene analysis information for the first object) for the first object acquired using the preprocessing model (310) to the first generative AI model (320a) of the generative AI model (320), and according to one embodiment, the processor (210) may generate a second object image (421) in which at least a part of the first object is transformed or added to at least a part of the first object by the interaction corresponding to the first user input, based on the preprocessing result information, using the first generative AI model (320a). When generating a second object image (421), the first generative AI model (320a) may generate a new second object image (421) in which an obscured area of the first object is visible as the shape of at least a portion of the first object is deformed based on scene analysis information. According to one embodiment, the second object image (421) may include a plurality of images. In response to a user selecting the first object image (411), the first generative AI model (320a) may generate a second object image (421) including a plurality of images to apply a related interaction (e.g., a drawer opening effect).The processor (210) can control the display (230) to sequentially display (or play) a plurality of images included in the second object image (421) generated by the first generative AI model (320a) according to a related interaction (e.g., a drawer opening effect). For example, the plurality of images can be a plurality of animation scenes in which objects inside the drawer gradually appear as the drawer slowly opens.
[0089] According to one embodiment, the processor (210) may display a user interface (UI) (520) (e.g., a function, a menu screen, or an input window) including a message (521) (“Would you like to browse similar types of images in your gallery?”) for adding at least one object for interaction with the second object image (421) in the second image (420) on the execution screen (501), as illustrated in FIG. 5B . According to one embodiment, the processor (210) may display identification information (e.g., transient, spreader) (423) of at least one object included in the second object image (421). According to one embodiment, the processor (210) may control the display (230) to display a user interface (520) including searched object images (525) and a message (527) (e.g., “Please select one of the images you searched in the gallery”) on the execution screen (501) when the user selects a function (523) (yes) requesting navigation (e.g., a menu, a button, or a symbol) in the user interface (UI) (520).
[0090] According to one embodiment, when a user selects (529) one of object images (525a, 525b, 525c), the processor (210) may control the display (230) to display a graphic element (e.g., a marker, a symbol, or text) (425) representing the selected object image (525a) (e.g., a knife image) in association with a second object image (421). According to one embodiment, the processor (210) may control the display (230) to display a graphic element (e.g., a marker, a symbol, or text) (427) representing a plurality of objects (e.g., a transient, a spreader) included in the second object image (421) in association with the second object image (421). According to one embodiment, the processor (210) may control the display (230) to display a user interface (530) including a message (531) (e.g., "Do you want to continue searching in the user gallery?") and buttons (yes and no) for confirming whether to further search for another object associated with the second object image (421), on the execution screen (501). According to one embodiment, the processor (210) may control the display (230) to re-display a user interface (UI) (520) (e.g., a function, a menu screen, or an input window) including a message (521) ("Do you want to search in the user gallery for similar types of images?") as illustrated in FIG. 5B based on a selection input (533) for the yes button among the buttons. According to one embodiment, the processor (210) may determine if there are no overlapping depth values at a location where a shape is to be generated by deforming the first image (410) based on depth information analyzed in the first image (410), and may generate a deformed second object image (421). The processor (210) can output the generated second object image (421) as is if the depth value at the position of the deformed first object is within the error range (e.g., ±15) of the value of the depth information transmitted as input to the generative AI model (320).When the processor (210) goes beyond the error range, it can generate an object image (e.g., a second object image (421)) for the first object that has been deformed only up to the area corresponding to the error range. When the processor (210) generates the second object image in which the shape of the first object has been deformed, a part that was not visible in the first image (410) may be visible. In this case, a new object can be generated together in the non-visible part based on the scene analysis result. For example, when the location of the image is identified as a kitchen based on scene analysis information, the processor (210) generates objects related to the kitchen together when generating the image, and the type of the second object image has an association with the shape-deformed first object selected by the user (e.g., kitchen drawer -> spoon & fork & cooking utensil, kitchen cabinet -> pot & bowl), and a new object may not always be generated.
[0091] According to one embodiment, when the processor (210) generates an object image in which the shape of the first object is transformed through the generative AI model (320), if the electronic device (201) is in a fully folded state, the processor (210) may not generate a second object image in which the shape of the first object is transformed. If the electronic device (201) is in an unfolded state, the processor (210) may generate various types of object images through the generative AI model (320). According to one embodiment, the processor (210) may determine whether the size or resolution of the display (230) is changed in response to the shape of the display (230) being transformed, for example, according to unfolding, folding, flexibility, or rollability. According to one embodiment, the processor (210) may determine whether a second object image (e.g., a second object image (421)) is generated in response to a change in the size or resolution of the display (230), and may generate the second object image (421) based on the determination that the second object image is generated. For example, the processor (210) can generate the second object image using different generative AI models (320) depending on the size of the deformed screen of the display (230). For example, when the display (230) is only partially unfolded (e.g., only a partial area is exposed to the outside so that the user can see it), the processor (210) can generate the second object image with a limited degree of deformity in the shape of the first object according to the resolution constraints. For example, when the display (230) is fully unfolded (e.g., the entire area is exposed to the outside so that the user can see it), image editing can be performed at the maximum resolution, so the processor (210) can generate the second object image in various image forms at the maximum resolution.
[0092] According to one embodiment, the processor (210) may identify a third object image (413) corresponding to a second object whose shape may be transformed by interaction with the second object image (413), as illustrated in FIGS. 4b and 4c, and may generate a fourth object image (431) (e.g., an image of cut bread) in which at least a portion of the second object (413) is transformed by an interaction result (e.g., bread being cut by a knife) by the second object image (421) included in the second image (420) in response to an interaction action (403) (e.g., multiple clicks, long clicks, or a swipe gesture) based on a second user input (401) using a generative AI model (320) (e.g., the second generative AI model (320b) of FIG. 3). The fourth object image (431) may be generated based on an interaction characteristic input by the user. For example, when a user touches a first object (413) (e.g., bread) once, the processor (210) can generate an object image in which the first object (413) (e.g., bread) is cut once as a fourth object image according to a single touch input. For example, when a user touches the first object (413) (e.g., bread) multiple times in succession, the processor (210) can generate an object image in which the first object (413) (e.g., bread) is cut multiple times as a fourth object image according to the consecutive touch inputs.
[0093] According to one embodiment, the processor (210) may control the display (230) to display a third image (430) that changes a third object image (413) corresponding to a second object into a generated fourth object image (431) based on an interaction action (403) by a second user input (401) for a second object (e.g., “bread”), as illustrated in FIG. 4C. According to one embodiment, the processor (210) may control the communication circuit (250) to transmit the third image (430) to an external electronic device.
[0094] According to one embodiment, when the processor (210) receives additional user input (e.g., a user gesture) for an action for a specific interaction with a second object, the processor (210) may display a graphical element for the interaction performed in response to the additional user input (e.g., a graphical element indicating a direction for cutting bread).
[0095] According to one embodiment, when image editing is performed by a plurality of users (e.g., editing participants) using a generative AI model (320), the processor (210) may control the display to display a list of users who participated in image editing on the execution screen (501). Here, the list may include user-specific identification information (e.g., name or nickname) and editing information (e.g., editing history and / or user distinguishing elements). According to one embodiment, when displaying the third image (430), the processor (210) may control the display to display object images edited by users and to display graphic elements (e.g., user distinguishing elements) for distinguishing the object images edited by users.
[0096] According to one embodiment, the processor (210) can generate an object image based on an interaction between a first object included in a first image and a second object selected by a user, without changing the shape of the first object included in the first image. When the processor (210) receives a first user input for selecting a first object included in the first image, the processor (210) identifies the first object, and when the processor (210) receives a second user input for a second object included in the first image that interacts with the first object, the processor (210) can generate a second object image in which the shape of the second object is changed by reflecting the interaction by the first object to the second object based on characteristics of the second user input.
[0097] According to one embodiment, the processor (210) may be a hardware component (function) or a software component (program) including at least one component provided in the electronic device (201) as a hardware module or a software module (e.g., an application program). According to one embodiment, the processor (210) may include, for example, one or a combination of two or more of hardware, software, or firmware. The processor (210) may omit at least some of the above components, or may be configured to further include other components for performing image processing operations in addition to the above components.
[0098] According to one embodiment, the memory (220) (e.g., the memory (130) of FIG. 1) can store an application. For example, the memory (220) can store an application (function or program) related to an image (or image generation), an application related to image management, or an application related to a generative AI model (320). The memory (220) can store a first image captured by an external electronic device or camera (e.g., an original image), a second acquired image (e.g., an image for editing), a third acquired image (e.g., an edited result image), and information related to image editing.
[0099] According to one embodiment, the memory (220) may store a program used for functional operation (e.g., the program (140) of FIG. 1), as well as various data generated during execution of the program (140). For example, the memory (220) may include a program (140) area and a data area (not shown). The program (140) area may store related program information for driving the electronic device (201), such as an operating system (OS) (e.g., the operating system (142) of FIG. 1) that boots the electronic device (201). The data area (not shown) may store transmitted and / or received data and generated data according to various embodiments. In addition, the memory (220) may be configured to include at least one storage medium among flash memory, a hard disk, a multimedia card micro type memory (e.g., a secure digital (SD) or extreme digital (XD) memory), RAM, and ROM.
[0100] According to one embodiment, the display (230) (e.g., the display module (160) of FIG. 1) may display an execution screen of an application related to image editing (or image creation). The display (230) may display information related to image editing, a first image, or a third image under the control of the processor (210). When performing image editing, the display (230) may display a second image for editing under the control of the processor (210). According to one embodiment, the display (230) may be implemented in the form of a touch screen. When the display (230) is implemented together with an input module in the form of a touch screen, it may display various pieces of information generated according to a user's touch operation. According to one embodiment, the display (230) may be configured with at least one of a liquid crystal display (LCD), a thin film transistor LCD (TFT-LCD), an organic light emitting diode (OLED), a light emitting diode (LED), an active matrix organic LED (AMOLED), a flexible display, and a 3-dimensional display. In addition, some of these displays may be configured as transparent or light-transmitting so that the outside can be seen through them. This may be configured in the form of a transparent display including a transparent OLED (TOLED). According to one embodiment, in addition to the display (230), other display modules (e.g., an extended display or a flexible display) may be further installed.
[0101] According to one embodiment, a camera (240) (e.g., camera module (180) of FIG. 1) can capture a first image (e.g., an original image) to be used for image editing.
[0102] According to one embodiment, the communication circuit (250) (e.g., the communication module (190) of FIG. 1) can communicate with an external electronic device (e.g., the electronic device (102, 104) of FIG. 1, the server (108) of FIG. 1, or another user's electronic device). For example, the communication circuit (250) can transmit a first image, a second image (420), and / or a third image (430) to the external electronic device or receive the first image (410) from the external electronic device. The communication circuit (250) can receive an application and / or information related to the image (or image editing) from and / or transmit the application and / or information to the external electronic device. According to one embodiment, the communication circuit (250) can include a cellular module, a wireless-fidelity (Wi-Fi) module, a Bluetooth module, or a near field communication (NFC) module.
[0103] An electronic device according to one embodiment (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIG. 2) may implement a software module (e.g., program (140) of FIG. 1) related to image editing (or image creation). A memory of the electronic device (e.g., memory (130) of FIG. 1 and / or memory (220) of FIG. 2) may store commands (e.g., instructions) to implement the software module. At least one processor (e.g., processor (120) of FIG. 1 and / or processor (210) of FIG. 2) can execute instructions stored in a memory to implement a software module and control hardware associated with the function of the software module (e.g., sensor module (176) of FIG. 1, camera module (180), communication module (190) of FIG. 1 and / or communication circuit (250) of FIG. 2, display module (160) of FIG. 1 and / or display (230) of FIG. 2).
[0104] According to one embodiment, a software module of an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIG. 2) may be configured to include a kernel (or HAL), a framework (e.g., middleware (144) of FIG. 1), and an application (e.g., application (146) of FIG. 1). At least a portion of the software module may be preloaded on the electronic device or may be downloadable from a server (e.g., server (108) of FIG. 1).
[0105] According to one embodiment, the kernel may include, but is not limited to, a system resource manager or a device driver, and may further include other modules. The system resource manager may perform at least one of controlling, allocating, or retrieving system resources. The device driver may include, for example, a display driver, a camera driver, a Bluetooth driver, a shared memory driver, a USB driver, a keypad driver, a WIFI driver, an audio driver, or an inter-process communication (IPC) driver.
[0106] According to one embodiment, the framework may provide functions commonly required by applications or provide various functions to applications through an application programming interface (API) (not shown) so that the applications can efficiently utilize limited system resources within the electronic device. The framework may include modules that form a combination of various functions of the components described above in FIGS. 1 and 2 . The framework may provide specialized modules for each type of operating system (e.g., the operating system (142) of FIG. 1 ) to provide differentiated functions. The framework may dynamically delete some existing components or add new components.
[0107] According to one embodiment, the application may be configured to include an application (e.g., a module, a manager, or a program) related to image editing (or image creation). The application may include an application received from an external electronic device (e.g., a server (108) or an electronic device (102, 104)). According to one embodiment, the application may include a preloaded application or a third-party application downloadable from a server. The components and names of the components of the software module according to the illustrated embodiment may vary depending on the type of operating system. According to one embodiment, at least a portion of the software module may be implemented as software, firmware, hardware, or a combination of at least two or more thereof. At least a portion of the software module may be implemented (e.g., executed) by, for example, a processor (e.g., an AP). At least a portion of the software module may include, for example, at least one of a module, a program, a routine, a set of instructions, or a process for performing at least one function.
[0108] As such, in one embodiment, the main components of the electronic device have been described through the electronic device (101) of FIG. 1 and the electronic device (201) of FIG. 2. However, in various embodiments, not all of the components illustrated through FIGS. 1 and 2 are essential components, and the electronic device (101, 201) may be implemented with more components than the illustrated components, or may be implemented with fewer components. In addition, the positions of the main components of the electronic device (101, 201) described above through FIGS. 1 and 2 may be changed according to various embodiments.
[0109] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIG. 2) may include a display (e.g., display module (160) of FIG. 1 and / or display (230) of FIG. 2), communication circuitry (e.g., communication module (190) of FIG. 1 and / or communication circuitry (250) of FIG. 2), processing circuitry including at least one processor (e.g., processor (120) of FIG. 1 and / or processor (210) of FIG. 2), and memory (e.g., memory (130) of FIG. 1 and / or memory (220) of FIG. 2)) for storing instructions.
[0110] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to control the display to display a first image (e.g., the original image (301) of FIG. 3, the first images (410, 510) of FIGS. 4A to 5C).
[0111] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a first object (e.g., the first object image (411) of FIG. 4A) whose shape can be transformed by interaction among a plurality of objects included in the first image.
[0112] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a second object (e.g., the second object image (421) of FIGS. 4b and 4c) having at least a portion of a shape of the first object deformed using a generative AI model (e.g., the generative AI model (320) of FIG. 3), at least partially based on a first user input for the first object.
[0113] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to control the display to display a second image (e.g., the second image (420) of FIG. 4B) that changes the first object image into the second object.
[0114] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a third object (e.g., the third object image (413) of FIGS. 4A to 4C) whose shape can be transformed by interaction with the second object image among the plurality of objects.
[0115] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a fourth object (e.g., a third object image (431) of FIG. 4c) having at least a portion of a shape of the third object deformed using the generative AI model (e.g., a second generative AI model (320b) of FIG. 3), at least in part based on a second user input for the second object.
[0116] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to control the display to display a third image (e.g., the third image (430) of FIG. 4c) that changes the third object into the fourth object.
[0117] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to control the display to display a menu on the second image for selecting at least one object capable of interacting with the second object, and, in response to a selection input of the menu, control the display to search for an object image corresponding to the at least one object and display the searched object images on the second image.
[0118] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display a graphical element indicating that the first object is capable of shape deformation.
[0119] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to control the display to display another graphical element indicating that the third object is capable of shape deformation.
[0120] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to control the display to display a message for identifying objects in the first image whose shape can be changed by interaction, and to perform identification of the first object whose shape can be changed based on a word received by a third user input among the plurality of objects.
[0121] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to control the display to display a message for detecting a similar object in relation to the second object from the memory, control the display to display the similar object based on a user request related to the message, and control the display to display identification information of the similar object in an area adjacent to the second object when displaying the second image.
[0122] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to determine a maximum range within which the shape of the first object can be deformed based on depth information and scene analysis information of the first image, and to perform generation of the second object based on the maximum range.
[0123] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to omit some objects from the second object or change the first object to an object image generated at a resolution lower than a specified resolution, based on identifying that a portion of a housing of the electronic device is folded and a size of a screen of the display is less than a specified size, using the generative AI model.
[0124] According to one embodiment, the instructions, when executed by the at least one processor, may cause the electronic device to control the display to display a list of users who participated in image editing, the list including user-specific identification information and editing information, and when displaying the third image, to control the display to display object images generated using the generative AI model for each user and to display graphic elements for distinguishing the users adjacent to the object images, respectively.
[0125] Hereinafter, the method of operation in an electronic device will be specifically described with reference to the drawings.
[0126] FIG. 6 is a diagram illustrating an example of an operating method in an electronic device according to one embodiment. In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. The operating method of FIG. 6 may be performed using an algorithm of an artificial intelligence (AI) model for image editing, such as that illustrated in FIG. 3 (e.g., the AI model of FIG. 3 ).
[0127] Referring to FIG. 6, in operation 601, an electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2, 3, and 4A to 5C) according to one embodiment may execute an application (e.g., a program, a function, or a module) for performing image editing and creation using an artificial intelligence (AI) model. The electronic device may acquire a first image (e.g., the original image (301) of FIG. 3 or the first image (410) of FIGS. 4A to 5C) and display the acquired first image on a display (e.g., the display (160) of FIG. 1, the display (230) of FIG. 2).
[0128] According to one embodiment, in operation 603, the electronic device may perform a preprocessing operation including object classification, depth information detection, and / or scene analysis operations for a first image displayed on a display using a preprocessing model of an artificial intelligence (AI) model (e.g., the preprocessing model (310) of FIG. 3) to obtain preprocessing result information, and may identify a first object image corresponding to a first object whose shape can be transformed by interaction among a plurality of objects based on the obtained preprocessing result information. Here, the preprocessing result information may include at least one of classification information (e.g., at least one first image portion whose shape can be transformed among the classified image portions), detected depth information, or scene analysis result information. When performing operation 603, the electronic device according to one embodiment may perform an operation of identifying and classifying a plurality of objects in the first image using a classification model of the preprocessing model (e.g., the classification model (311) of FIG. 3) to obtain classification information for the classified plurality of objects. According to one embodiment, an electronic device may classify at least one object (e.g., a first object and a second object) capable of shape deformation through interaction among a plurality of objects based on classification information using a preprocessing model. According to one embodiment, the electronic device may detect depth information from a first image using a depth information detection model of the preprocessing model (e.g., the depth information detection model (313) of FIG. 3). According to one embodiment, the electronic device may perform image editing reflecting a real situation by limiting a maximum range in which a shape deformation-capable first object included in the first image can be deformed using the depth information. According to one embodiment, the electronic device may perform scene analysis on the first image using a scene analysis model of the preprocessing model (e.g., the scene analysis model (315) of FIG. 3) to obtain scene analysis information.Scene analysis information can be used when generating an object image (e.g., a second object image or a fourth image object) in which at least a portion of a shape of an object (e.g., a first object or a second object) whose shape can be deformed is deformed by a generative AI model (e.g., a first generative AI model (320a) and / or a second generative AI model (320b) of FIG. 3).
[0129] According to one embodiment, in operation 605, in response to receiving a first user input for a first object, the electronic device may generate a second object image in which at least a portion of the shape of the first object is deformed using a generative AI model (e.g., the first generative AI model (320a) of FIG. 2). The electronic device may generate a second object image in which at least a portion of the shape of the first object is deformed by an interaction according to the first user input based on preprocessing result information acquired by performing a preprocessing operation. According to one embodiment, the electronic device may generate a second object image in which at least a portion of the shape of the first object desired by the user is deformed and / or newly added by performing fine-tuning based on depth information and scene analysis information using the first generative AI model. The electronic device may pre-designate a maximum range in which the shape of the first object can be changed based on the depth information. This allows the electronic device to consider physical relationships with other objects when performing image editing, and to generate additional object images associated with the first object when the shape of the first object is deformed by utilizing scene analysis information.
[0130] According to one embodiment, in operation 607, the electronic device may display a second image (e.g., the second image (420) of FIG. 4B) on the display, which is a first object image corresponding to the first object included in the first image changed to a second object image, based on a first user input for the first object.
[0131] According to one embodiment, in operation 609, in response to receiving a second user input regarding a second object capable of interacting with the first object, the electronic device may identify a second object image that has deformed a shape of the first object among a plurality of objects included in the first image (or the second image) and a third object image (e.g., the second object (413) of FIGS. 4A to 4C) corresponding to the second object whose shape can be deformed by the interaction. Here, operation 609 may be performed when the preprocessing operation is performed in operation 603, and the electronic device may analyze objects that can be associated with the first object to identify the second object capable of interacting, and display a graphic element (e.g., a marker, a symbol, or text) indicating that interaction is possible on the first object and the second object. When the electronic device receives a second user input, it can identify a third object image corresponding to the identified second object, and transmit preprocessing result information (e.g., classification information, depth information, and scene analysis information) for the second object and information about the second user input to a generative AI model (e.g., the second generative AI model (320b) of FIG. 3).
[0132] According to one embodiment, in operation 611, the electronic device may generate a fourth object image (e.g., the fourth object image (431) of FIG. 4c) in which at least a portion of the shape of the second object is deformed based on preprocessing result information (e.g., classification information, depth information, and scene analysis information) for the second object and information about the second user input, using a generative AI model (e.g., a second generative AI model).
[0133] According to one embodiment, in operation 613, the electronic device may display a third image (e.g., the result image (303) of FIG. 3, the third image (430) of FIG. 4c) on the display, which is a fourth object image generated by changing a third object image corresponding to the second object based on a second user input for the second object.
[0134] FIG. 7 is a diagram illustrating an example for image editing using a generative artificial intelligence model according to one embodiment.
[0135] Referring to FIG. 7, an electronic device according to one embodiment (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2, 3, 4A to 5C) may display a second image (720) including an object image (hereinafter referred to as a second object image (721)) generated using a generative AI model (320) (e.g., the first generative AI model (320a) of FIG. 3) or reflect the generated object image in a first image. According to one embodiment, the electronic device may identify a third object image (713) corresponding to a second object associated with the second object image (721) included in the second image (720). The electronic device may, based on a first user input (701) selecting a second object image (721), display a graphic element (723) (e.g., an outline) indicating that the second object image (721) and a third object image (713) corresponding to the second object are capable of interaction. The electronic device may, based on a second user input (703) moving the second object image (721) for interaction or selecting the third object image (713), generate a fourth object image (731) in which at least a portion of the shape of the second object is deformed by the interaction with the second object image (721) using a generative AI model (e.g., the second generative AI model (320b) of FIG. 3). The electronic device may display a third image (730) including the fourth object image (731). According to one embodiment, the electronic device may display the fourth object image (731) in the third image (730) at the same location as the third object image (713) corresponding to the second object. For example, if the second user input (703) is a gesture for a change in location, such as a push or a press, the fourth object image (731) in the third image (730) may be displayed at a different location than the third object image (713) corresponding to the second object.
[0136] According to one embodiment, the electronic device may display a graphic element (725) representing an action (705) of a user interacting with a second object image (721) and a third object image (713). When a plurality of second objects capable of interaction are identified and the identified second objects overlap, the electronic device may automatically zoom in on an area of the identified second objects so that the user can select the correct third object. According to one embodiment, the electronic device may, based on a characteristic of a user input received on a third object image (713) corresponding to the second object, reflect the user input characteristic in a prompt by the processor and transmit it as input to the generative AI model. For example, the processor of the electronic device may identify location information and shape information where a second object (e.g., a cake object) is cut through a user's touch input on the third object image (713), and add the identified location information and shape information to a prompt and transmit it as input to the generative AI model. The generative AI model can generate a fourth object image (731) that reflects the interaction characteristics reflected by the user.
[0137] According to one embodiment, the electronic device can display an image with a changed viewpoint so that it can specifically receive user input regarding a second object with which the user wishes to interact. The image with a changed viewpoint can provide user convenience by making it appear as if the user is actually in the location based on previously acquired scene analysis results and depth information. The user can input various interaction actions at the location with the changed viewpoint. For example, when a user input is received that is a gesture that allows a third object image (713) corresponding to the second object (e.g., a cake) to be divided into exactly four equal parts, the electronic device can receive interaction information not only through the user gesture but also through a device that can recognize actions (e.g., cutting, wiping, pushing, pressing, or shaving), such as an external input device or a gyro sensor.
[0138] FIG. 8 is a diagram illustrating an example for image editing using a generative artificial intelligence model according to one embodiment.
[0139] Referring to FIG. 8, an electronic device according to one embodiment (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2, 3, 4A to 5C) may identify a plurality of objects (e.g., a drawer, a flower pot, a plant, or a window) that can be transformed by interaction based on classification information obtained through classification of a first image (810), and may identify a first object image (811) corresponding to a first object (e.g., a drawer) selected from among the plurality of objects.
[0140] According to one embodiment, the electronic device may generate a second object image (821) by transforming at least a portion of a first object shape through interaction using a generative AI model, and display a second image (820) by changing the first object image (811) into the second object image (821). When displaying the second image (820), the electronic device may display a menu (823) for selecting at least one object with which interaction is possible in relation to the second object image (821). If the menu (823) (e.g., a button that can be used with an AI assistant) is not automatically displayed, the electronic device may display the menu (823) by calling the AI assistant through voice recognition.
[0141] According to one embodiment, the electronic device identifies a third object image (913) corresponding to a second object associated with a first object, and if the second object image (821) is not generated or if the second object image (821) does not contain at least one object for interacting with the third object image (813) or does not contain an object desired by the user, the electronic device may use a menu (823) to select at least one object with which to interact. In response to receiving a user input (801) selecting the menu (823), the electronic device may display object images (e.g., UI, stickers, emoticons, or special symbols) searched by the AI assistant, and in response to receiving a user input (803) selecting one of the object images, the electronic device may designate the selected object image (825) as the second object image.
[0142] According to one embodiment, the electronic device may display a graphic element (827) (e.g., a graphic effect, text, or symbol) indicating that a second object (e.g., a flower pot) associated with the second object image is interactable, on the border of a third object image (813) corresponding to the second object, as shown in (d) of FIG. 8.
[0143] According to one embodiment, in response to receiving a user input (805) for a second object, as shown in (e) of FIG. 8, the electronic device may generate a fourth object image (831) by interacting with an object image (825) designated as a second object image using a generative AI model, wherein at least a portion of the shape of the second object is deformed (e.g., water droplets form around the object due to spraying) by the sprayer. According to one embodiment, the electronic device may, based on a characteristic of the received user input (805), reflect the characteristic of the user input (805) in a prompt and transmit it as an input to the generative AI model. The result of the interaction may vary based on the characteristic of the user input (805). For example, if the electronic device receives a user input in which the user touches a third object image corresponding to the second object once, the electronic device may generate an object image in which the spraying is weakly reflected as the fourth object image (831) using the generative AI model. For example, if a user input is received in which the user touches the third object image multiple times, the electronic device can generate an object image with water spray reflected multiple times as a fourth object image (831) by the generative AI model.
[0144] According to one embodiment, the electronic device may display a third image (930) in which a third object image corresponding to a second object is changed into a fourth object image (831), as shown in (f) of FIG. 8.
[0145] Based on the operations described in FIG. 8, the present disclosure can enable image editing that reflects various interaction results. For example, if scissors are selected as a first object and paper is selected as a second object, the electronic device can input a specific user interaction action and transmit it as input to a generative AI model. The generative AI model can then generate an object image corresponding to the second object, in which the paper is cut according to the user input gesture.
[0146] FIG. 9 is a diagram illustrating an example of image editing using a generative artificial intelligence model according to one embodiment.
[0147] Referring to FIG. 9, an electronic device according to an embodiment (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2, 3, 4A to 5C) may perform image editing by multiple users (e.g., editing participants) when performing image editing using a generative AI model (e.g., the generative AI model (320) of FIG. 3). According to an embodiment, the electronic device may display an original image (910) and a list of users (901) who participated in editing on an execution screen. The user list (901) may include identification information (e.g., name or nickname) and editing information (e.g., editing history and / or user identification elements) for each user. The electronic device may display an image (920) being edited, including object images (921, 923) generated by reflecting the editing results by each user, on a display. The electronic device can display graphic elements (931, 933) (e.g., symbols, emoticons, or texts) on the object images (921, 923) respectively so as to distinguish the users who edited the object images (921, 923). The graphic elements (931, 933) can be differently designated for each user participating in editing so as to confirm the editing user in real time. The graphic element (931) (e.g., an arrow) can indicate a portion being edited (e.g., modified) in the image (920) being edited by a second user (user-2), and the graphic element (933) can indicate a portion being edited (e.g., modified) in the image (920) being edited by a first user (user-1). The graphic elements (931, 933) can prevent different users from repeatedly modifying the same object. For example, the electronic device may display edited history (e.g., log) as edit information in the user list (901) in the form of a bar, as shown in (b) of FIG. 9, and may display a scroll bar to view other edit history that is not visible on the screen.For example, the electronic device can record that the state of the window has changed from open to closed by performing an action to close an open window on a wall with the editing history of a second user (user-2). For example, the electronic device can add an object that is not in the original image by uploading an image from the personal electronic device of the attending user through online editing. For example, if the second user (user-2) directly uploads a succulent pot that was not in the original image, the electronic device can display the editing history changed from "None" to "Allocated" in the editing information included in the user list (901) by placing an object image (923) corresponding to the uploaded succulent pot in the image to be edited (920). As described in FIG. 9, the present disclosure can perform large-scale project work by having multiple users use the online image editing function, and in the case of general users, by supporting a separate image editing program hosting server in a club or seminar for content creation, various ideas can be simultaneously reflected in the image editing and creation process.
[0148] FIG. 10 is a diagram illustrating an example of image editing using a generative artificial intelligence model according to one embodiment.
[0149] Referring to FIG. 10, an electronic device according to one embodiment (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2, 3, 4A to 5C) may identify a first object image (1011) corresponding to a first object (e.g., a drawer) capable of shape deformation in a first image (1010) based on a first user input, and generate a second object image (1021) in which at least a portion of the shape of the first object is deformed using a generative AI model (e.g., the generative AI model (320) of FIG. 3).
[0150] According to one embodiment, when generating a second object image (1021), the electronic device may identify a second object related to the first object included in the first image (1010), and identify a third object image (1013) corresponding to the second object. For example, as illustrated in FIG. 9, considering realism, the second object (e.g., a box) may be of a size that cannot fit into the first object (e.g., a drawer), and thus, if the realistic physical correlation is ignored, a result that cannot occur in reality may be generated. To prevent such a result that ignores realism from being generated, the electronic device may limit the deformation of the shape by considering the real situation based on depth information and scene analysis information when generating the second object image (1021).
[0151] According to one embodiment, the electronic device may designate a maximum range in which the first object can be deformed based on depth information about the first object and depth information about the second object. For example, as shown in (b) of FIG. 10, when a first object (e.g., a drawer) is transformed and another object (a second object) exists at a position where the first object can be deformed to its maximum, the depth value of the first object in the maximum transformation range may be similar to the depth value of the other object. When the first object has a depth value within an error range with respect to the second object in the maximum transformation range, the first object may be deformed to an area outside the error range. Accordingly, the electronic device may generate a second object image (1021) in which the shape of the first object is deformed within the maximum range based on the designated maximum range.
[0152] FIG. 11 is a diagram illustrating an example of an image editing operation method using a generative artificial intelligence model according to one embodiment, and FIG. 12 is a diagram illustrating an example of image editing using a generative artificial intelligence model according to one embodiment.
[0153] Referring to FIGS. 11 and 12, an electronic device according to one embodiment (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2, 3, and 4A to 5C) may be an electronic device including a flexible display. The electronic device (201) may use different AI model performances depending on the size of the screen as the shape of the display (230) changes.
[0154] In operation 1101, an electronic device (201) according to one embodiment may display a first image.
[0155] In operation 1103, the electronic device (201) can identify a first object image corresponding to a first object. The electronic device (201) can perform a preprocessing operation including object classification, depth information detection, and / or scene analysis operations for the first image displayed on the display using a preprocessing model of an artificial intelligence (AI) model (e.g., the preprocessing model (310) of FIG. 3), thereby obtaining preprocessing result information, and can identify a first object image corresponding to a first object whose shape can be transformed by interaction among a plurality of objects based on the obtained preprocessing result information.
[0156] In operation 1105, the electronic device (201) may receive a first user input as a request to generate a second object image in which at least a portion of the shape of the first object is deformed. In operation 1107, the electronic device (201) according to one embodiment may check whether the screen of the display (230) is in an expanded state as a portion of the housing of the flexible electronic device (201) is unfolded. If the size of the screen of the display (230) is in an expanded state, the electronic device may perform operation 1109. If the size of the screen of the display (230) is not in an expanded state, the electronic device may perform operation 1111.
[0157] In operation 1109, the electronic device (201) may generate a second object image in which at least a portion of the shape of the first object is deformed by a generative AI model (e.g., the generative AI model (320) of FIG. 3), and display a second image (1220 or 1230) in which the first object image of the first object is changed into the generated second object image (1221 or 1231), as shown in (b) or (c) of FIG. 12. The generative AI model may generate various second object images as the resolution of the display (230) is greatly changed as the screen is expanded. For example, when a portion of the housing is folded, the electronic device (201) may display the generated second image (1220) in the second area (232) and the third area (233) of the display (230), as shown in (b) of FIG. 12. For example, when the entire housing of the electronic device (201) is unfolded, the electronic device (201) can display a second image (1230) generated in the first area (231), the second area (232), and the third area (233) of the display (230), as shown in (c) of FIG. 12.
[0158] According to one embodiment, when the electronic device (201) has a display (230) in an expanded state (e.g., an unfolded state or a partially folded state), the electronic device (201) can perform image editing at the maximum resolution. The electronic device (201) can generate a second object image (1221 or 1231) including at least one object capable of interacting with the display area by a generative AI model and / or generate the size of the object image corresponding to the display area (e.g., generate an object image in which a drawer is further opened). The electronic device (201) can display the second image (1230) including the generated object image (1221 or 1231) on the entire display area (the first area, the second area, and the third area).
[0159] In operation 1111, if the size of the screen of the display (230) is not expanded, the electronic device (201) may generate a second object image in which at least a part of the shape of the first object is deformed by a generative AI model (e.g., the generative AI model (320) of FIG. 3), and display a second image (1210) in which the first object image of the first object is changed into the generated second object image (1111), on the first area (231) of the display (230), as shown in FIG. 12 (a). According to one embodiment, the electronic device (201) may provide another object image, such as a sticker image, by using an AI assistant function. The electronic device may provide a sticker image when the second object image is not generated by the generative AI model. The electronic device is not limited thereto, and may additionally provide a sticker image in order to provide a second object image desired by the user when the second object image is generated.
[0160] According to one embodiment, as shown in (a) of FIG. 12, if the electronic device is in a fully folded state and the screen size of the display (230) is not expanded, there is a resolution limitation, so a second object image (1211) in which the shape of the first object is deformed (e.g., a simple second object image or a reduced second object image) may be generated with restrictions. At this time, the electronic device (201) may not generate at least one object included in the second object image (1211) generated in the folded state by the generative AI model, or may generate the second object image (1211) in a size different from that of the second object image generated in the unfolded state (e.g., generating an object image in which a drawer is less open). In this case, the electronic device (201) may select an object image (e.g., a sticker image) provided by the AI assistant function as the second object image by using the AI assistant function. For example, a menu (823) as illustrated in FIG. 8 may be displayed to provide a sticker image. According to one embodiment, when the electronic device (201) is in a partially unfolded folded state, the electronic device (201) can identify that a situation in which the user can perform additional work has occurred, since the generative AI model does not generate various second objects. Based on the identification that an additional work situation has occurred, the electronic device (201) can provide an object image (e.g., a sticker) provided by the AI assistant function, and based on the selection of the provided object image, can designate the selected object image as the second object. Here, the object image (e.g., a sticker) provided by the AI assistant function can be set using data that the user has previously saved or an image determined to be suitable through a scene analysis result by accessing various storage devices (e.g., an internal / external storage device or a cloud) in real time at the user's request when editing an image.
[0161] According to one embodiment, when performing image editing using a generative AI model, the electronic device (201) may pre-generate images with different resolutions as the size of the display screen changes. The electronic device (201) may immediately provide the result of editing the pre-generated images based on the changed size of the display screen. When the size of the display screen changes, the quantity and variety of second objects additionally generated when performing shape transformation of the first object may change depending on the size of the electronic display screen. If editing images of all sizes are pre-generated and stored according to the size of the display screen, memory and battery consumption may occur, so the electronic device (201) may pre-generate only large-resolution images and store them in memory. For example, if a user uses a large screen and then folds the terminal partially to change to a small screen, the electronic device may simply resize and display the large-resolution images. The process of converting from a large resolution to a small resolution does not require high computational complexity, but the process of changing from a small resolution to a large resolution requires performing an algorithm such as SR (super-resolution), so it is difficult to display an image from a small resolution to a large resolution in real time. Therefore, electronic devices can quickly display images on a display by presetting only images with a large resolution and storing them in memory.
[0162] According to one embodiment, the electronic device (201) does not generate all sizes of edited images in advance according to the size of the visible area of the display, but only generates and stores in memory large resolution images generated in an unfolded state in advance, and when the electronic device changes to a folded state, the size of the stored large resolution images can be adjusted and displayed in response to the size of the visible area of the display.
[0163] In operation 1113, in response to receiving a second user input for a second object capable of interacting with a first object, the electronic device (201) may identify a second object image that has deformed a shape of the first object among a plurality of objects included in the first image (or the second image) and a third object image (e.g., the second object (413) of FIGS. 4A to 4C) corresponding to the second object whose shape can be deformed by the interaction. When the electronic device (201) receives the user input, the electronic device (201) may identify the third object image corresponding to the identified second object, and transmit preprocessing result information (e.g., classification information, depth information, and scene analysis information) for the second object and information on the second user input to a generative AI model (e.g., the second generative AI model (320b) of FIG. 3).
[0164] In operation 1115, the electronic device (201) may generate a fourth object image (e.g., the fourth object image (431) of FIG. 4c) in which at least a portion of the shape of the second object is transformed based on preprocessing result information (e.g., classification information, depth information, and scene analysis information) for the second object and information about the second user input, using a generative AI model (e.g., the second generative AI model (320b)).
[0165] In operation 1117, the electronic device (201) can display a third image (e.g., the result image (303) of FIG. 3, the third image (430) of FIG. 4c) that is a third object image corresponding to the second object and is changed into a generated fourth object image on the display (230).
[0166] FIG. 13 is a diagram illustrating an example of image editing using a generative artificial intelligence model according to one embodiment.
[0167] Referring to FIG. 13, an electronic device according to one embodiment (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2, 3, 4A to 5C) may be a device capable of displaying a three-dimensional image in a virtual reality space. According to one embodiment, the electronic device may display a first image (1210a, 1210b) and identify a first object image (1311a, 1311b) corresponding to a first object (e.g., a sprayer or a brush) capable of changing shape classified in the first image (1310a, 1310b). The electronic device may select a pre-recognized object through a linkable hand gesture input device (not shown). For example, when the electronic device receives a first user input (1301a, 1301b) in which the user selects a first object (1311a) (e.g., a sprayer or a brush) from a first image (1210a, 1310b), the electronic device may generate a second object image (1321a, 1321b) that has a deformed shape of the first object (e.g., a sprayer or a brush) using a generative AI model, and display the second image (1320a, 1320b) including the second object image (1321a, 1321b).
[0168] According to one embodiment, the electronic device may generate a fourth object image (1331a, 1331b) by an interaction between a first object (e.g., a sprayer or a brush) and a second object (e.g., a flower pot or a piece of paper) using a generative AI model based on a second user input (1303a, 1303b), and display a third image (1330a, 1330b) including the fourth object image (1331a, 1331b) on a display. For example, the fourth object image (1331a, 1331b) may include a virtual object (e.g., a virtual object image) corresponding to a user's action of holding the second object image (1321a, 1321b) with his / her hand and interacting with it (e.g., an action of holding a sprayer and spraying water or an action of holding a brush and drawing a picture) in response to the second user input (1303a, 1303b) and a virtual object reflecting the result of the interaction (e.g., water sprayed around a flower pot or a picture drawn on paper). According to one embodiment, the generative AI model of the electronic device may generate the fourth object image (1331a) differently depending on the type and number of user interactions transmitted to the electronic device (e.g., VST device). For example, a graphic effect in the form of water displayed on the flower pot may appear differently depending on the number of spraying actions input by the user. For example, if the fourth object image (1331a) is in the form of water, the generative AI model can generate the fourth object image (1331a) by receiving two object images (1311a, 1313a) (e.g., a sprayer, a flower pot) with which the user interacts and user interaction information as inputs of the generative AI model. For example, when receiving a user input (e.g., an input of touching a brush) of the first object (1311b) included in the first image (1310b), the electronic device can display a second object image (1321b) in the form of holding a brush generated by the generative AI model.Thereafter, when the electronic device receives a user input for interaction with the second object image (1321b) (e.g., a gesture input for drawing on a piece of paper placed on a table), the electronic device can display a fourth object image (1331b) that has transformed the shape of the second object (e.g., paper) generated by the generative AI model (e.g., a graphic effect in the form of a drawing on the paper in the shape input by the user). At this time, if the user attempts to input an accurate gesture, the electronic device can zoom in on an area of the second object (e.g., paper) and then perform image editing. As described in FIG. 13 above, by enabling the direct capture and editing of scenes visible through the VST environment, one can directly experience various actions through interactions between objects placed in places such as furniture stores and experience centers without having to visit those places in person.
[0169] According to one embodiment, an electronic device may generate an object image based on an interaction between a first object included in a first image and a second object selected by a user, without changing the shape of the first object included in the first image. When the electronic device receives a first user input selecting the first object included in the first image, the electronic device may identify the first object, and when the electronic device receives a second user input regarding a second object included in the first image that interacts with the first object, the electronic device may generate a second object image in which the shape of the second object is changed by reflecting the interaction by the first object to the second object based on characteristics of the second user input.
[0170] According to one embodiment, a method of operating in an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2 and 3) may include displaying a first image (e.g., original image (301) of FIG. 3, first images (410, 510) of FIGS. 4A to 5C) on a display of the electronic device (e.g., display module (160) of FIG. 1 and / or display (230) of FIG. 2).
[0171] According to one embodiment, the method may include an operation of identifying a first object (e.g., the first object image (411) of FIG. 4A) whose shape can be transformed by interaction among a plurality of objects included in the first image.
[0172] According to one embodiment, the method may include generating a second object (e.g., a second object image (421) of FIGS. 4b and 4c) having at least a portion of a shape of the first object deformed, using a generative AI model (e.g., a generative AI model (320) of FIG. 3), based at least in part on a first user input for the first object.
[0173] According to one embodiment, the method may include an action of displaying a second image (e.g., the second image (420) of FIG. 4B) that changes the first object into the second object on the display.
[0174] According to one embodiment, the method may include an operation of identifying a third object (e.g., the third object image (413) of FIGS. 4A to 4C) whose shape can be transformed by interaction with the second object image among the plurality of objects.
[0175] According to one embodiment, the method may include generating a fourth object (e.g., the third object image (431) of FIG. 4c) in which at least a portion of the shape of the third object is deformed using the generative AI model, at least in part based on a second user input for the second object.
[0176] According to one embodiment, the method may include an action of displaying a third image (e.g., the third image (430) of FIG. 4c) that changes the third object into the fourth object on the display.
[0177] According to one embodiment, the method may further include an operation of displaying a menu on the display for selecting at least one object capable of interacting with the second object, an operation of searching for an object image corresponding to the at least one object in response to a selection input of the menu, and an operation of displaying the searched object images on the second image displayed on the display.
[0178] In one embodiment, the method may further include displaying a graphical element on the display indicating that the first object is capable of shape transformation.
[0179] In one embodiment, the method may further include displaying a graphical element on the display indicating that the third object is capable of changing shape.
[0180] According to one embodiment, the method may further include an action of displaying a message on the display for identifying objects whose shape can be changed by interaction in the first image, and an action of performing identification of the first object whose shape can be changed based on a word received in a third user input among the plurality of objects.
[0181] According to one embodiment, the method may further include an action of displaying a message for detecting a similar object related to the second object from the memory on the display, an action of displaying the similar object on the display based on a user request related to the message, and an action of displaying identification information of the similar object on the display in an area adjacent to the second object when displaying the second image.
[0182] According to one embodiment, the operation of generating the second object may include an operation of determining a maximum range in which the shape of the first object can be deformed based on depth information and scene analysis information of the first image, and an operation of generating a shape of the second object based on the maximum range.
[0183] In one embodiment, the method may further include an operation of using the generative AI model to omit some objects from the second object or change the first object to an object image generated at a resolution lower than the specified resolution, based on identifying that a portion of the housing of the electronic device is folded and a size of a screen of the display is less than a specified size.
[0184] According to one embodiment, the method further comprises the action of displaying a list of users who participated in the image editing on the display, wherein the list may include user-specific identification information and editing information.
[0185] According to one embodiment, the method may further include, when displaying the third image, displaying object images generated for each user using the generative AI model on the third image, and displaying graphic elements for distinguishing users adjacent to the object images on the display.
[0186] According to one embodiment, a non-transitory storage medium storing one or more programs, wherein the program, when executed by at least one processor of an electronic device, causes the electronic device to: display a first image (301, 410, 510, 710, 810, 910, 1010, 1210a, 1210b) on a display of the electronic device; identify a first object (411) whose shape can be transformed by interaction among a plurality of objects included in the first image; generate a second object (421) whose shape is at least partially transformed by using a generative AI model based at least partially on a first user input for the first object; display a second image (420) in which the first object is transformed into the second object on the display; identify a third object (413) whose shape can be transformed by interaction with the second object image among the plurality of objects; and generate a second image (420) for the second object. The method may include commands for executing an operation of generating a fourth object (431) in which at least a portion of the shape of the third object is transformed using the generative AI model, at least partially based on user input, and an operation of displaying a third image (430) in which the third object is transformed into the fourth object on the display.
[0187] The present disclosure identifies an object capable of shape deformation based on object recognition, and when a user edits the object, a generative AI model is used to create an area that does not exist in the original image, thereby enabling productive image editing. Furthermore, the present disclosure provides a function that allows new objects created using the generative AI model to interact with objects present in the original image, thereby maximizing image editing performance. Furthermore, the present disclosure provides an image editing function that performs image editing based on a desired action by the user, thereby providing a desired resulting image. In the case of PCs and flexible devices with transparent displays, the present disclosure allows for active image editing as needed through the performance of a generative AI model that reflects the characteristics of the electronic device. In the case of transparent displays, more effective AR (augmented reality) results can be obtained by directly editing the image based on a real-world scene visible through the back of the panel. In addition, various effects that can be directly or indirectly understood through this document may be provided. The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects not mentioned will be clearly understood by those skilled in the art to which the present disclosure pertains from the description below.
[0188] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. A processing device (or processing circuit) may execute an operating system (OS) and one or more software applications running on the operating system. In addition, the processing device may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.
[0189] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.
[0190] The method according to the embodiments may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program commands, including ROM, RAM, and flash memory. In addition, examples of other media may include recording or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.
[0191] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.
[0192] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.
[0193] The embodiments disclosed in this document are presented for the purpose of explaining and understanding the disclosed technical content, and do not limit the scope of the technology described in this document. Therefore, the scope of this document should be interpreted to include all modifications or various other embodiments based on the technical concepts of this document.
[0194] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0195] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0196] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0197] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0198] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0199] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In the electronic device (101, 201), display(160, 230); At least one processor (120, 210) comprising a processing circuit; and It includes memory (130, 220) for storing instructions, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Control the display to display the first image (301, 410, 510, 710, 810, 910, 1010, 1210a, 1210b), Identifying a first object (411) whose shape can be transformed by interaction among a plurality of objects included in the first image, At least partially based on a first user input for the first object, a generative artificial intelligence (AI) model (320) is used to generate a second object (421) in which at least a portion of the shape of the first object is transformed, Control the display to display a second image (420) that changes the first object into the second object, Identifying a third object (413) whose shape can be transformed by interaction with the second object image among the plurality of objects, At least partially based on the second user input for the second object, a fourth object (431) is generated using the generative AI model, in which at least a portion of the shape of the third object is transformed, An electronic device that causes the display to be controlled to display a third image (430) that changes the third object into the fourth object.
2. In the first paragraph, when the instructions are executed by the at least one processor, the electronic device: Controlling the display to display a menu on the second image for selecting at least one object capable of interacting with the second object; In response to a selection input from the above menu, at least one object image corresponding to the at least one object is retrieved, An electronic device that causes the display to be controlled to display at least one object image searched for on the second image.
3. In the first or second paragraph, when the instructions are executed by the at least one processor, the electronic device: Controlling the display to display a graphic element indicating that the first object can be transformed; An electronic device that causes the display to be controlled to display another graphic element indicating that the third object is capable of shape transformation.
4. In any one of paragraphs 1 to 3, when the instructions are executed by the at least one processor, the electronic device: Control the display to display a message for identifying objects whose shape can be changed by interaction in the first image; An electronic device that causes identification of the first object capable of shape transformation based on a word received by a third user input among the plurality of objects.
5. In any one of paragraphs 1 to 4, the instructions, when executed by the at least one processor, cause the electronic device to: Controlling the display to display a message for detecting a similar object in relation to the second object from the memory; Controlling the display to display the similar object based on a user request related to the above message; An electronic device that causes the display to be controlled to display identification information of the similar object in an area adjacent to the second object when displaying the second image.
6. In any one of paragraphs 1 to 5, when the instructions are executed by the at least one processor, the electronic device: Based on the depth information and scene analysis information of the first image, the maximum range in which the shape of the first object can be deformed is confirmed, Perform creation of the second object based on the above maximum range, An electronic device that causes a part of the housing of the electronic device to be folded and, based on identifying that the size of the screen of the display is less than a specified size, to omit some objects from the second object or change the first object to an object image generated with a resolution lower than a specified resolution using the generative AI model.
7. In any one of paragraphs 1 to 6, the instructions, when executed by the at least one processor, cause the electronic device to: Control the display to display a list of users who have participated in image editing, the list including user-specific identification information and editing information; An electronic device that causes the display to be controlled so that, when displaying the third image, object images (921, 923) generated for each user using the generative AI model are displayed on the third image, and graphic elements (931, 933) for distinguishing the users are displayed adjacent to the object images, respectively.
8. In the method of operation in an electronic device (101, 201), An operation of displaying a first image (301, 410, 510, 710, 810, 910, 1010, 1210a, 1210b) on a display (160, 230) of the electronic device; An operation of identifying a first object (411) whose shape can be transformed by interaction among a plurality of objects included in the first image; An operation of generating a second object (421) having at least a portion of the shape of the first object deformed using a generative artificial intelligence (AI) model (320), based at least in part on a first user input for the first object; An action of displaying a second image (420) that changes the first object into the second object on the display; An operation of identifying a third object (413) whose shape can be transformed by interaction with the second object image among the plurality of objects; An operation of generating a fourth object (431) in which at least a portion of the shape of the third object is modified using the generative AI model, at least partially based on a second user input for the second object; and A method including an action of displaying a third image (430) in which the third object is changed into the fourth object on the display.
9. In the 8th paragraph, the method, An action of displaying a menu on the second image displayed on the display for selecting at least one object capable of interacting with the second object; In response to a selection input of the above menu, an operation of retrieving at least one object image corresponding to the at least one object; and A method further comprising an action of displaying at least one object image searched for on the second image displayed on the display.
10. In the 8th or 9th paragraph, the method, An action of displaying a graphic element on the display indicating that the first object is capable of shape transformation; and A method further comprising the action of displaying on the display another graphic element indicating that the third object is capable of shape transformation.
11. In any one of the 8th to 10th clauses, the method, An action of displaying a message on the display for identifying objects whose shape can be changed by interaction in the first image; A method further comprising an operation of performing identification of the first object capable of shape transformation based on a word received by a third user input among the plurality of objects.
12. In any one of the 8th to 11th clauses, the method, An action of displaying a message on the display for detecting a similar object in relation to the second object from the memory; An action of displaying the similar object on the display based on a user request related to the message; and A method further comprising, when displaying the second image, an action of displaying identification information of the similar object on the display in an area adjacent to the second object.
13. In any one of paragraphs 8 to 12, the operation of creating the second object comprises: An operation of determining the maximum range in which the shape of the first object can be deformed based on depth information and scene analysis information of the first image; An operation of performing creation of the second object based on the above maximum range; and A method further comprising an action of using the generative AI model to omit some objects from the second object or change the first object to an object image generated at a resolution lower than the specified resolution, based on identifying that a portion of the housing of the electronic device is folded and the size of the screen of the display is less than a specified size.
14. In any one of paragraphs 8 to 13, the method, An action of displaying a list of users who have participated in image editing on the display; the list includes user-specific identification information and editing information; A method further comprising an action of displaying object images (921, 923) generated for each user using the generative AI model on the third image when displaying the third image, and displaying graphic elements (931, 933) for distinguishing the users on the display adjacent to the object images.
15. In a non-transitory storage medium storing one or more programs, the programs, when executed by at least one processor (120, 210) of an electronic device (101, 201), cause the electronic device to: An operation of displaying a first image (301, 410, 510, 710, 810, 910, 1010, 1210a, 1210b) on a display (160, 230) of the electronic device; An operation of identifying a first object (411) whose shape can be transformed by interaction among a plurality of objects included in the first image; An operation of generating a second object (421) having at least a portion of the shape of the first object deformed using a generative artificial intelligence (AI) model (320), based at least in part on a first user input for the first object; An action of displaying a second image (420) that changes the first object into the second object on the display; An operation of identifying a third object (413) whose shape can be transformed by interaction with the second object image among the plurality of objects; An operation of generating a fourth object (431) in which at least a portion of the shape of the third object is modified using the generative AI model, at least partially based on a second user input for the second object; and A non-transitory storage medium comprising commands for executing an operation of displaying a third image (430) that changes the third object into the fourth object on the display.
Citation Information
Patent Citations
Creative production and play system Web type based DIY
KR101856626B1
Natural dyeing method that can produce economical products
KR1020220042503A
Wearable electronic device
KR1020240171967A
Modifying digital images via multi-layered scene completion facilitated by artificial intelligence
US20240135514A1