Electronic device, and method for editing object included in image
The electronic device uses advanced object recognition and generative AI to enable precise image editing by distinguishing and replacing objects based on text input, addressing the limitations of existing AI-based image editing systems.
Patent Information
- Application Number
- PCT/KR2025/009824
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-26
- Filing Date
- 2025-07-08
- Publication Date
- 2026-01-29
Smart Images

Figure KR2025009824_29012026_PF_FP_ABST
Abstract
Description
Methods for editing objects contained in electronic devices and images
[0001] The present invention relates to an electronic device and a method for editing an object included in an image.
[0002] Artificial Intelligence (AI) technologies applied to electronic devices are rapidly advancing. In particular, the evolution of AI capabilities is leading to the implementation of generative AI. Generative AI utilizes machine learning and deep learning to generate similar content based on existing content, such as text, audio, and / or images. For example, generative AI can be applied to improving image quality or creating or editing new images based on existing ones.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.
[0004] AI-based image editing can create new image types or partially change some objects contained in an image based on prompts input into a generative AI model.
[0005] However, image editing based on AI models can only edit objects recognized in the image, making it difficult for users to precisely edit desired objects. Furthermore, the output from the AI model varies depending on the prompt, making it difficult to expect consistent image editing performance.
[0006] Various embodiments of the present disclosure aim to propose an electronic device, method and recording medium thereof that can increase the object recognition rate in an image to enable complete search of an object in an image and partially edit only an object included in an image based on text.
[0007] The problem to be solved in this disclosure is not limited to the problem mentioned above, and may be expanded in various ways without departing from the spirit and scope of this disclosure.
[0008] An electronic device according to one embodiment may include a display. The electronic device according to one embodiment may include a processor including processing circuitry. The electronic device according to one embodiment may include a memory that stores instructions executable by the processor. The instructions according to one embodiment, when executed by the processor, may cause the electronic device to display an image including at least one object on the display. The instructions according to one embodiment may cause the electronic device to receive a user input requesting editing related to the at least one object. The instructions according to one embodiment may cause the electronic device to recognize a first keyword designating a first object included in the image and a second keyword defining a second object to be modified by recognizing a context for the user input. The instructions according to one embodiment may cause the electronic device to perform graphic processing so as to visually distinguish the first keyword and the second keyword included in the user input. The commands according to one embodiment may cause the electronic device to display at least one graphic object image related to the second object on at least a portion of the display based on a detection of an input selecting the second keyword. The commands according to one embodiment may cause the electronic device to generate an edited image in which the first object included in the image is replaced with the second object using the selected graphic object image, through generative artificial intelligence (AI), based on a detection of an input selecting the displayed graphic object.
[0009] A method for editing an object in an image of an electronic device according to one embodiment may include an operation of displaying an image including at least one object on the display. The method of the electronic device according to one embodiment may include an operation of receiving a user input requesting editing related to the at least one object. The method of the electronic device according to one embodiment may include an operation of recognizing a first keyword designating a first object included in the image and a second keyword defining a second object to be changed by recognizing a context of the user input. The method of the electronic device according to one embodiment may include an operation of graphically processing the first keyword and the second keyword included in the user input so as to be visually distinguishable. The method of the electronic device according to one embodiment may include an operation of displaying at least one graphic object image related to the second object on at least a portion of the display based on a detected input selecting the second keyword. A method of an electronic device according to one embodiment may include an operation of generating an edited image in which the first object included in the image is replaced with the second object using an image of the selected graphic object, based on a detection of an input selecting the displayed graphic object, through generative artificial intelligence (AI).
[0010] A non-transitory computer-readable medium storing instructions according to one embodiment of the present disclosure, which when executed by a processor of an electronic device, causes the electronic device to perform operations, the operations comprising: displaying an image including at least one object on the display; receiving user input requesting editing related to the at least one object; The method may include: recognizing a first keyword that designates a first object included in the image and a second keyword that defines a second object to be changed by identifying a context for the user input; graphically processing the first keyword and the second keyword included in the user input so that they are visually distinct; displaying at least one graphic object image related to the second object on at least a portion of the display based on a detection of an input selecting the second keyword; and generating an edited image in which the first object included in the image is replaced with the second object using the selected graphic object image through generative artificial intelligence (AI) based on a detection of an input selecting the displayed graphic object.
[0011] An electronic device, method and recording medium thereof according to one embodiment of the present disclosure can edit only objects within an image according to the user's intention by specifically and clearly designating objects or backgrounds within an image based on text input by a user in an image editing function.
[0012] An electronic device, method and recording medium thereof according to one embodiment of the present disclosure can more clearly and accurately designate an object area corresponding to text information input by a user through a complete search of an overlapping portion of an object within an image.
[0013] An electronic device, method and recording medium thereof according to one embodiment of the present disclosure enable object editing of an image based on text, thereby overcoming limitations in image editing due to inaccurate external input from a user or structural limitations of an electronic device.
[0014] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs from the description below.
[0015] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.
[0016] FIG. 1 is a block diagram of an electronic device within a network environment according to various embodiments.
[0017] FIG. 2 illustrates the configuration of a text-based image editing AI system according to one embodiment.
[0018] FIGS. 3A, 3B and 3C illustrate a method for editing an object of a text-based image of an electronic device according to one embodiment.
[0019] FIG. 4 illustrates a method for editing an object of a text-based image of an electronic device according to one embodiment.
[0020] FIGS. 5A, 5B, and 5C illustrate user interface screens for explaining a method of editing an object in an image of an electronic device according to one embodiment.
[0021] FIG. 6 illustrates a user interface screen for explaining an object editing function in an image of an electronic device according to one embodiment.
[0022] FIG. 7 illustrates a user interface screen for explaining an object editing function in an image of an electronic device according to one embodiment.
[0023] FIG. 8 illustrates user interface screens for describing an object editing function within an image of an electronic device according to one embodiment.
[0024] FIG. 9 illustrates a user interface screen for explaining an object editing function within an image of an electronic device according to one embodiment.
[0025] Electronic devices according to the embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments disclosed in this document are not limited to the aforementioned devices.
[0026] Each of the embodiments described with reference to the drawings of the present disclosure can be independently configured as one embodiment. Each of the embodiments described with reference to the drawings of the present disclosure can operate independently as one embodiment. At least two embodiments of the embodiments described with reference to the drawings of the present disclosure can be combined and configured. At least two embodiments of the embodiments described with reference to the drawings of the present disclosure can be combined and operated. For example, at least a portion of the embodiment of FIG. 1 and at least a portion of the embodiment of FIG. 2 can be combined and operated with each other.
[0027] When at least two embodiments described with reference to the drawings of the present disclosure are combined, at least some of the configurations and / or at least some of the operations included in each embodiment may be omitted.
[0028] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.
[0029] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160) (or display), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0030] The processor (120) includes at least one processing circuitry, and the at least one processing circuitry may control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing, for example, software (e.g., a program (140)), and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in the volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in the non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0031] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, in the electronic device (101) itself where artificial intelligence is performed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0032] The memory (130) can store various data used by at least one component (e.g., the processor (120) or the sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., the program (140)) and input data or output data for commands related thereto. The memory (130) can include a volatile memory (132) or a non-volatile memory (134). The memory (130) can store instructions executable by the processor (120) or the electronic device (101).
[0033] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0034] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0035] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0036] The display module (160) (or display) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0037] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0038] The sensor module (176) may include at least one sensor. The sensor module (176) may detect an operating state (e.g., power or temperature) of the electronic device (101) or an external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0039] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0040] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0041] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0042] The camera module (180) includes at least one camera and can capture still images and moving images. In one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0043] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).
[0044] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0045] The communication module (190) includes at least one communication circuit and can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) operates independently from the processor (120) (e.g., application processor) and can include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) can include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0046] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0047] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0048] In one embodiment, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0049] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0050] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0051] An electronic device (101) according to one embodiment may support a generative AI (generative artificial intelligence) function. The generative AI function may refer to a technology that generates new content based on content and can generate new forms of AI content (or generative content) by utilizing given input data or information. Here, AI content may refer to content (e.g., 3D objects, images, videos, audio, screen information, or text) that is entirely generated (or reconstructed / edited) or partially generated (or reconstructed / edited) based on generative AI.
[0052] An electronic device (101) according to one embodiment may provide a function for editing an image using a generative AI function. For example, the electronic device (101) may provide a function for editing an object included in an image using generative AI. The electronic device (101) may generate an image editing result (e.g., an edited image) using generative AI.
[0053] An electronic device (101) according to one embodiment may support a generative AI function (e.g., a text-based image editing AI function) in conjunction with a server (108). At least one of the electronic device (101) or the server (108) according to one embodiment may include at least a portion of the configuration of the text-based image editing AI system illustrated in FIG. 2 .
[0054] The electronic device (101) illustrated in FIG. 1 can be implemented according to various embodiments disclosed below, and components that are substantially the same as the configuration disclosed in FIG. 1 described above are given the same reference numbers, and redundant descriptions of their functions can be omitted.
[0055] FIG. 2 illustrates the configuration of a text-based image editing AI system according to one embodiment.
[0056] Referring to FIG. 2, an electronic device (101) according to one embodiment may include a text-based image editing AI system (210) illustrated in FIG. 2. The electronic device (101) may be implemented to provide an object editing function of an image by utilizing a portion of the text-based image editing AI system (210).
[0057] An electronic device (101) according to one embodiment may include an AI pre-processor (230), a user reaction handler (240), and a generative AI (250).
[0058] The AI pre-processor (230), the user reaction handler (240), and the generative AI (250) may be embedded in the memory (130) of FIG. 1, and the operations described below may be stored in the memory (130) as instructions executable by the processor (120) of FIG. 1.
[0059] An AI pre-processor (230) can perform object recognition and classification (or object detection) and estimate depth information in an image (hereinafter, referred to as an original image (220) for convenience of explanation) using a segmentation model (231) and a depth estimation model (232).
[0060] The segmentation model (231) uses the original image (220) as input data, recognizes and classifies at least one object included in the original image (220), and can output the type (class, category) of the recognized object and object location information (e.g., coordinate values that can define the location of the object). The type of object that can be recognized may vary depending on the database (DB) used for object learning of the segmentation model (231). The object that can be recognized may be an object (e.g., a person, an object, an animal, a plant, a building, etc.) whose size or shape can be changed or whose location can be changed. The object location information is a coordinate value that can define the location of the object, and can be used when displaying a bounding box (or bounding frame) at the location of the object recognized within the image.
[0061] The depth estimation model (232) uses the original image (220) as input data, and can estimate and output depth information of an object included in the image. For example, the depth estimation model (232) may be a model that has been trained using a data pair consisting of an RGB image and a corresponding depth image (ground truth). The depth information of an object defines the object in a two-dimensional image when a text-based image editing function is used, and may include an object mask image. The object mask image may refer to information used when extracting a recognized object on a pixel basis.
[0062] Although not shown in the drawing, if object recognition and classification (or object detection) from an image fails, the AI preprocessor (230) can perform preprocessing to remove noise from the image using a denoising AI model that uses an image containing noise as learning data, and perform object recognition and classification from the image from which the noise has been removed.
[0063] The user reaction handler (240) may include a user input management unit (241), an object full search unit (242), and an image search unit (243).
[0064] The user input management unit (241) can receive user input (e.g., text input, voice input) (2411) and analyze the context of the text (2412). The user input management unit (241) can display an editing UI (user interface) supported by the electronic device (101) in relation to image editing on the display and receive user input for editing an object of the image. For example, if the user input is a voice input rather than a text input, the user input management unit (241) can convert the voice input into STT (speech to text) to obtain text for the voice input.
[0065] The user input management unit (241) can analyze context. For example, the user input management unit (241) can identify similarities and contexts between words within a sentence based on the results of analyzing input text based on a large language model (LLM) or a large multimodal model (LMM).
[0066] The user input management unit (241) can recognize a first keyword corresponding to a first object specified by the user in the image and a second keyword that the user wishes to change to a second object from the input text based on the user's intention as determined through analysis, and can generate a prompt for editing an object in the image based on the first keyword and the second keyword.
[0067] The object full search unit (242) can perform an overall object search algorithm that utilizes the characteristics of an object. The object full search unit (242) can apply the object full search algorithm to define the object information found as a result of context analysis of the input text on the image. In general, objects can have correlations depending on the object characteristics. For example, if there is a train (e.g., an object) in the image, there is a high probability that there is a railroad (e.g., an object). The object full search algorithm is an algorithm that searches the object area (or object full area) pixel by pixel based on object characteristic information that can have correlations between objects, and can find the exact location and boundary of the object within the image even for objects that are not specifically specified in the text.
[0068] The object full search unit (242) can search for a complete area / entire area of an object with a complete boundary without any part covered by another object by using the mutual relationship characteristic to determine the causal relationship between objects when there is an overlapping part between objects included in the image and searching for pixel values that objects with similar relationships / correlations may have.
[0069] The image search unit (243) can search (search, gather, collect) object images (hereinafter, graphic object images) corresponding to keywords recognized in the user's text input from images stored in the electronic device (101) (e.g., gallery images (2431)) or images searched through a communication module (e.g., crawling images (2432)). The graphic object image can mean a graphic object image representing a second object when the user wants to change a first object designated by the user in the image into a new second object according to the user's intention through context analysis.
[0070] When a user wishes to change a first object designated by the user into a second object, the image search unit (243) can display graphic object images found in relation to the second object in a list UI format. The graphic object images can be configured in the form of stickers or emoticons, but are not limited thereto.
[0071] The image search unit (243) can search / extract at least one graphic object image corresponding to an object to be changed (e.g., a second object) by comparing data pre-tagged in images (e.g., gallery images (2431)) stored in the electronic device (101) with recognized keywords (e.g., a second keyword). If the image search unit (243) cannot search for an object image from the stored images (e.g., gallery images (2431)), the image search unit (243) can perform crawling via a communication module to search for a graphic object image corresponding to a keyword from images within a web page (e.g., crawled images (2432)) and add it to the list UI. Crawling may mean a process of searching and collecting / extracting information on the web.
[0072] An electronic device (101) according to one embodiment may provide an object image addition function item through crawling collection in a list UI so that crawling collection can be performed only when the user desires it.
[0073] According to one embodiment, the electronic device (101) or the image search unit (243) may separate only the object area from the image and store it as a graphic object image (e.g., in the form of a sticker) if the object area among the images searched based on the keyword is not stored in the form of a graphic object (e.g., a sticker).
[0074] The user reaction handler (240) can transmit the generated prompt, image (e.g., original image (220)), input text, and graphic object image selected by the user through the list UI to the generative AI (250) based on the results processed through the user input management unit (241), the object full search unit (242), and the image search unit (243).
[0075] Generative AI (250) includes an image editing AI model, and can generate and output an edited image (260) in which an object in the image is edited, using an original image, a user input, and a graphic object image selected by the user as input data. The image editing AI model can accurately define a location in the original image where an object image passed as input should be placed based on a user input (e.g., text input, voice input), and can generate an edited image by replacing an object (e.g., a first object) at that location with an object image (e.g., a second object).
[0076] An electronic device according to one embodiment may include a display. The electronic device according to one embodiment may include a processor including processing circuitry. The electronic device according to one embodiment may include a memory that stores instructions executable by the processor. The instructions according to one embodiment, when executed by the processor, may cause the electronic device to display an image including at least one object on the display. The instructions according to one embodiment may cause the electronic device to receive a user input requesting editing related to the at least one object. The instructions according to one embodiment may cause the electronic device to recognize a context for the user input and recognize a first keyword designating a first object included in the image and a second keyword defining a second object to be modified. The instructions according to one embodiment may cause the electronic device to perform graphic processing so as to visually distinguish the first keyword and the second keyword from among text displayed on the text editing UI. The commands according to one embodiment may cause the electronic device to display at least one graphic object image related to the second object on at least a portion of the display based on a detection of an input selecting the second keyword. The commands according to one embodiment may cause the electronic device to generate an edited image in which the first object included in the image is replaced with the second object using the selected graphic object image, through generative artificial intelligence (AI), based on a detection of an input selecting the displayed graphic object.
[0077] According to one embodiment, the user input may include at least one of text input, voice input, and drawing input.
[0078] The commands according to one embodiment may cause the electronic device to display an editing user interface (UI) for receiving user input on at least a portion of the image after displaying the image on the display, and to display text corresponding to the received user input within the editing UI based on the received user input.
[0079] According to one embodiment, the above editing user interface (UI) may include at least one of a text input window and guidance information guiding a user input request for editing an object within an image.
[0080] The commands according to one embodiment may cause the electronic device to recognize and classify objects included in the image through object detection and depth estimation operations using the image as an input value, thereby obtaining object recognition information, object location information, and object depth information included in the image, recognize the first keyword designating the first object included in the first image by comparing the text with the object recognition information, and recognize the second keyword by analyzing the context of the text based on a large language model (LLM).
[0081] The commands according to one embodiment may cause the electronic device to generate a prompt requesting the generation of an edited image based on the text and object location information of the first object, provide the image, the text, a graphic object image related to the second object, and the generated prompt as inputs to an AI model, and receive an edited image in which the first object is changed to a second object from the AI model and display the received image on the display.
[0082] The commands according to one embodiment may cause the electronic device to display at least one graphic object image related to the second object on at least a portion of the display, search for graphic object images tagged with the second keyword from first images stored in the memory, configure a list UI including the searched graphic object images, and display the at least one graphic object image through the list UI.
[0083] An electronic device according to one embodiment further includes a communication module including a communication circuit, and the instructions may cause the electronic device to retrieve the graphic object image from the communication module through crawling based on the second keyword when the electronic device cannot retrieve the graphic object image from the first images stored in the memory.
[0084] The commands according to one embodiment may cause the electronic device to search for the graphic object image, and if the searched graphic object image is not stored as a sticker or emoticon image, separate or extract only an area including a second object corresponding to a second keyword from the searched graphic object image and store the area as a sticker or emoticon image corresponding to the second object.
[0085] The commands according to one embodiment may cause the electronic device to determine whether a portion of a first object recognized in the image is overlapped by another third object after recognizing the first keyword, and, based on a correlation index between the first object and the third object, if a portion of the first object is overlapped by another third object and the first object is not overlapped by the third object in the image, search for a complete region of the first object without an obscured portion using a complete region search algorithm, and determine, as a result of the search, a complete region of the specified first object as the object location information.
[0086] The commands according to one embodiment may cause the electronic device to search for the graphic object image, determine the properties of the first object as a result of searching the complete area of the first object, and search for a graphic object image related to a second object that has the same properties as the properties of the first object.
[0087] The above commands according to one embodiment may cause the electronic device to perform preprocessing to remove noise from the image using a denoising AI model that uses an image containing noise as learning data, and then recognize the object when object detection from the image fails.
[0088] In one embodiment, the commands may cause the electronic device to perform scene analysis on the image to correct an edited image in which the first object is edited into the second object or to correct a second object in the edited image.
[0089] The embodiments described below may be drawings for explaining operations related to a function / service for editing an object within an image based on text in an electronic device (101).
[0090] FIGS. 3A and 3B illustrate a method for editing an object of a text-based image of an electronic device according to one embodiment.
[0091] The functions or operations described in FIGS. 3a and 3b can be understood as functions performed by the processor (120) by executing commands (e.g., instructions) stored in the memory (130) of the electronic device (101). The operations of the task / process of FIG. 3b are performed in parallel or independently of the operations illustrated in FIG. 3a, and the operations of FIGS. 3a and 3b can operate in a mutually related manner.
[0092] In the following examples, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.
[0093] Referring to FIGS. 3A and 3B, in operation 310, the processor (120) of the electronic device (101) may display an image including at least one object on a display (e.g., a display module (160)).
[0094] For example, the processor (120) may execute an application (e.g., a gallery application) that supports an image list and an image editing function, and display one image (e.g., an original image (220)) from the image list on the display. The image (e.g., the original image (220)) may be displayed through an image editing tool or an image editing user interface (UI). The image editing tool or the image editing UI may include various functional items related to the editing function (e.g., a size adjustment item that adjusts the size of an image, a filter item that supports applying a filter to an image, a brightness adjustment item that adjusts the brightness of an image, rotation items that change the rotation axis of an image, and / or an AI activation object).
[0095] In operation 320, the processor (120) may display an editing UI on the display that can guide a user input request related to an object or receive an input based on the detection of an object in the image or an input (e.g., a touch or tap) that selects an object included in the image.
[0096] For example, the editing UI may be displayed through a split view or may be displayed by overlaying at least a portion of the image, but is not limited thereto. The editing UI (user interface) may include at least one of a text input window and guidance information guiding a user input request for editing an object within the image. The editing UI may be a UI related to a text input function supported by the electronic device (101), and may provide not only text but also voice input and drawing input functions.
[0097] According to one embodiment, the processor (120) may enter an image editing function in relation to an image displayed on the display or perform operations 321 to 322 of FIG. 3B based on the image editing tool / UI being displayed.
[0098] In operation 321, the processor (120) can recognize and classify (or object detect) at least one object included in an image using a segmentation model (e.g., the segmentation model (231) of FIG. 2). The processor (120) can obtain information on the type of the recognized / detected object and the object location.
[0099] According to one embodiment, when object detection from an image fails, the electronic device (101) may support a function of performing preprocessing to remove noise from the image using a denoising AI model that uses an image containing noise as learning data, and then detecting the object. The denoising AI model may be a model trained using a dataset containing various forms of noise that may occur when capturing an image with a camera. The denoising AI model may receive an image containing various types of noise as input and output an image from which noise has been removed based on pre-trained data.
[0100] In operation 322, the processor (120) can estimate depth information of an object recognized / detected in an image using a depth estimation model (e.g., the depth estimation model (232) of FIG. 2).
[0101] In operation 330, the processor (120) may receive user input (e.g., text input, voice input) for editing an object included in an image.
[0102] For example, a user may identify objects contained in an image and input text for a command request that includes a first keyword (or first word) designating / referring to a first object contained in the image and a second keyword (or second word) designating a second object to be newly changed. In another example, a user may input text for a command request by voice.
[0103] According to one embodiment, the processor (120) may receive a user's voice input, convert the user's voice input into text through an STT (speech to text) program, and display the text on the editing UI.
[0104] In operation 340, the processor (120) may recognize / select a first keyword corresponding to a first object included in an image from a user input and a second keyword indicating a first object to be newly changed, and may perform graphic processing (e.g., color change) so that the first keyword and the second keyword are initially distinguished.
[0105] According to one embodiment, the processor (120) may use object recognition information (e.g., object type information) of objects recognized / detected from an image to determine whether a word (e.g., a first keyword) of the recognized objects is included in an input text (e.g., a sentence). If the object recognition information matches a word included in the text, the processor (120) may recognize / select the matched word as a first keyword and designate a first object corresponding to the first keyword. The processor (120) may graphically process the first keyword with a first color (e.g., red) to distinguish it from other texts entered in the editing UI.
[0106] The processor (120) can graphically process the first object so that a highlight effect (e.g., dotted border) is displayed to guide that the first object included in the image displayed on the display is a target for editing based on the location information of the first object.
[0107] The processor (120) can analyze the context of the text input in sentence form based on the LLM (large language model) based on the recognition of the first keyword to recognize / select a word (e.g., the second keyword) corresponding to the second object. The processor (120) can graphically process the second keyword with a second color (e.g., blue) to distinguish it from the first keyword. For example, if a user inputs “Change the pot on the table next to the refrigerator to a wine glass with red wine,” the processor (120) can recognize the first keyword as “pot” and the second keyword as “wine glass.”
[0108] According to one embodiment, if the user input does not include a word corresponding to an object recognized from an image, the electronic device (101) may provide user feedback (e.g., “Keyword not found. Would you like to switch to drawing input?”) and provide a UI guiding other inputs. For example, if the electronic device (101) supports a drawing input function, the user may designate a first object by circling the object to be edited in the image.
[0109] The processor (120) can perform operations 341 to 343 of FIG. 3b based on the first keyword.
[0110] In operation 341, the processor (120) can determine whether a first object specified in an image by user input has a portion that overlaps with a surrounding object.
[0111] The processor (120) may determine the bounding boxes of objects based on the object recognition and classification (or object detection) results and the object location information of each object, and if the bounding boxes overlap, determine that there is an overlapping portion between the objects. For example, if a part of the coordinate values of the bounding box of a first object overlaps with the coordinate values of the bounding box of another object, the processor (120) may determine that there is an overlapping portion between the first object and another object.
[0112] In operation 342, the processor (120) may perform a full area search of the first object if there is a part of the first object included in the image that overlaps with a surrounding object (in operation 361, yes).
[0113] If there is a part where at least a part of the first object overlaps with another object, the processor (120) can determine the chronological relationship between the objects through an overall object search algorithm that utilizes the correlation characteristics of the objects, and search for a complete area / entire area of the object that is not hidden by another object. For example, if a “grandfather clock” object included in an image is partially hidden by a “flower pot” object, it may be ambiguous to determine whether the area of the “grandfather clock” object should be determined only by the area of the hour and minute hands, or whether the area with the bell is also the area of the “grandfather clock” object. In the case of an overall object search algorithm, by searching the pixel values of the surrounding areas according to the correlation characteristics of the grandfather clock, it is possible to search for an entire area that has the complete shape of the “grandfather clock” object that is not hidden by the “flower pot” object.
[0114] If the first object included in the image does not overlap with surrounding objects (no in operation 361), the processor (120) may proceed to operation 363. If the first object does not overlap with surrounding objects, the recognized object area may be a complete area / entire area of a complete shape that is not obscured by surrounding objects.
[0115] In operation 343, the processor (120) can check the properties (e.g., size, area size) of the first object. The properties of the first object may mean, but are not limited to, pixel size, size of a bounding box, complete area of the object, etc.
[0116] In operation 344, the processor (120) may search / collect at least one graphic image related to a second object corresponding to the second keyword based on the recognition of the first keyword and the second keyword.
[0117] The processor (120) can search (search, gather, collect) graphic object images that have the properties of the first object and are classified into a similar category or the same type as the second object from images stored in the electronic device (101) (e.g., gallery images (2431)) or / and images searched through the communication module (e.g., crawling images (2432)).
[0118] For example, the processor (120) can compare tag information of images (e.g., gallery images (2431), crawled images (2432)) to search for a graphic object image that has the same properties as the properties of the first object and is tagged with the second keyword. For example, if the first object has a size of 3*3 pixels, the processor (120) can search / collect a graphic object image (i.e., an image depicting the second object) that has a size of 3*3 pixels and is stored as a tag with the second keyword.
[0119] According to one embodiment, when the processor (120) cannot retrieve an object image from stored images (e.g., gallery images (2431)), the processor (120) may provide a search request UI that inquires the user about whether to perform a crawling search. Based on receiving a user consent input for a crawling search, the processor (120) may perform crawling through a communication module to retrieve a graphic object image corresponding to a keyword from images within a web page (e.g., crawling images (2432)).
[0120] An electronic device (101) according to one embodiment may support a function for searching for object images based on not only a keyword-specified word, but also a modifier phrase modifying the word through contextual analysis. For example, if the second keyword is "wine glass," the processor (120) may perform an object image search for "wine glass containing red wine" in a user-entered sentence.
[0121] According to one embodiment, if the object area or object image among the images searched for in relation to the second object is not stored in the form of a graphic object (e.g., a sticker), the processor (120) may separate (e.g., cut out, AI Lasso method) / extract only the object area from the searched image and store it in the form of a graphic object.
[0122] According to one embodiment, the electronic device (101) may provide / recommend graphic object images by utilizing personalized images based on the user's usage log information. For example, if the user frequently searches for bags through an online shopping mall, the electronic device (101) may convert bag images (e.g., temporary file images, cache images) collected online based on the user's shopping history into sticker format so that they can be searched as graphic object images related to the bag object. In this case, the electronic device (101) may utilize personalized data to recommend images of objects to be edited even without performing web crawling.
[0123] In operation 350, the processor (120) can detect a user input selecting a second keyword through the editing UI.
[0124] In the 360 motion, the processor (120) may display at least one graphic object image searched in relation to the second object in response to a user input selecting a second keyword.
[0125] According to one embodiment, the processor (120) may configure at least one graphic object image related to the second object as a list UI and display it on the display based on the results of performing operations 341 to 344. The graphic object image may be provided in the form of a graphic object such as a sticker or emoticon, but is not limited thereto.
[0126] In operation 370, the processor (120) may detect a user input selecting a graphic object image related to a second object displayed on the display.
[0127] In operation 380, the processor (120) can generate an edited image in which a first object is replaced with a second object using a selected graphic object image through a generative AI.
[0128] The processor (120) may transmit an image displayed on a display (e.g., the original image (220) of FIG. 2), a user input, and a graphic object image selected by the user as input data to a text-based image editing AI model (e.g., a generative AI (250)), and receive an edited image (260) in which a first object included in the image is replaced with a second object from the image editing AI model.
[0129] In operation 390, the processor (120) can perform image post-processing on the generated edited image.
[0130] According to one embodiment, the 380 operations and the 390 operations may be merged, and an edited image (260) on which image post-processing has been performed by the generative AI (250) may be output.
[0131] According to one embodiment, the processor (120) can correct the edited image through in-painting or out-painting after changing the grid within the image.
[0132] According to one embodiment, the processor (120) may collect various data (e.g., lighting information, location information, overall color information of the image, weather information) for image correction through scene analysis of an image (e.g., an original image). The processor (120) may change information about the color and brightness of an object so that the object to which changes have been applied is in harmony with the edited image by reflecting the lighting information or the overall color information of the original image. For example, if a light source exists in the original image, the processor (120) may identify the location of the light source and adjust the light amount information of the object in the edited image by applying the light source information of the original image. If the original image is from a rainy day, the processor (120) may correct the color of the object in the edited image so that it matches the rainy weather. The processor (120) may adjust the location where the object to which changes have been applied is placed through scene analysis.
[0133] FIG. 4 illustrates a method for editing an object of a text-based image of an electronic device according to one embodiment.
[0134] Referring to FIG. 4, the processor (120) of the electronic device (101) according to one embodiment may perform operations for searching the complete area of the first object in the image when there is a portion where the first object overlaps with a surrounding object. For example, the operations illustrated in FIG. 4 may be included in operation 342 of FIG. 3B.
[0135] In operation 410, the processor (120) can detect a first object that has a portion that overlaps with surrounding objects. For example, the processor (120) can recognize / detect a first object corresponding to a first keyword that designates an object included in an image from an input text.
[0136] In operation 420, the processor (120) may verify a correlation index of the identified first object. The correlation index may refer to a table probabilistically representing the relationship that, when a specific object appears in an image, similar types of objects or features related to objects may appear together. For example, if an image contains a "train" object, there is a high probability that a "railway" object will also be included, and if an "airplane" object exists, there may be a high probability that a "cloud" object or a "runway" object will also be included. Based on the correlation information, the processor (120) may obtain probability information regarding the appearance of the first object with other objects.
[0137] In operation 430, the processor (120) can analyze a correlation region related to the first object in the image based on correlation information.
[0138] In operation 440, the processor (120) can secure a region of interest (ROI) based on object location information output through an object recognition / detection AI model.
[0139] For example, object location information output by object recognition / detection may include bounding box information of each recognized object. Bounding box information is a coordinate value indicating the coordinates where the object is located in the image, and is composed of x, y coordinates corresponding to the start point of the box and x, y coordinates corresponding to the end point of the box.
[0140] The processor (120) can designate a region of interest (ROI) area to perform a complete search area of the first object included in the image based on the object recognition / detection result.
[0141] In general, searching all areas of an image takes a lot of processing time due to computational complexity, but searching by specifying an ROI can shorten the processing time and have the advantage of being able to compare various elements in the correlation index.
[0142] In operation 450, the processor (120) may detect correlation pixel information within the ROI region. The processor (120) may perform correlation analysis on an object (e.g., a first object) occluded on a pixel basis within the ROI region by applying a sliding window method. The correlation analysis may refer to analyzing the correlation based on the values of the images in the correlation index and the average of the pixel values within the window, or analyzing the correlation by variably using the size of the sliding window. For example, the processor (120) may determine that an object with a high correlation is located in the region if the average of the pixel values within the window is similar to the pixel values representing the images in the correlation index. The processor (120) may determine that there is no similar object in the region if the average of the pixel values is greater than a specified threshold. Using a small-sized sliding window may be advantageous for analyzing correlations in the boundary region of an object because it focuses on local information rather than global information. The boundary region of an object can be obtained by applying methods such as canny or sobel to obtain boundary region information for the object included in the ROI. The electronic device (101) can efficiently perform correlation analysis when applying a variable sliding window based on boundary area information.
[0143] In operation 460, the processor (120) can output the properties of the first object as a result of the full area search of the object. After the correlation analysis is completed, the processor (120) can identify the full area of the object on a pixel-by-pixel basis.
[0144] The electronic device (101) can accurately designate the complete area of an object covered by surrounding objects within an image by applying at least one of a ROI area, a correlation index, and a variable sliding window, and can improve the accuracy of object editing by reflecting this.
[0145] The embodiments of FIGS. 5A to 8 described below can support operations for editing objects within an image through an image editing tool or a user interface (UI) screen that supports image editing. For convenience of explanation, the UI (user interface) screens described below omit functional items related to image editing and illustrate UI configurations for editing objects within an image. The display format of the illustrated UI screens is merely an example and may be changed according to design or settings.
[0146] FIGS. 5A and 5B illustrate user interface screens for describing object editing operations within an image of an electronic device according to one embodiment.
[0147] Referring to FIGS. 5A and 5B, an electronic device (101) according to one embodiment <501> As illustrated, an image (510) including at least one object (515) (e.g., an original image (220)) can be displayed on a display (e.g., a display module (160) of FIG. 1). For example, the electronic device (101) can display the image (510) through an image editing tool or an image editing user interface (UI) based on an input for executing / entering an image editing function (e.g., an input for long-processing an image). <501> It may be a screen by operation 310 of Fig. 3a.
[0148] <501> On the screen, the electronic device (101) can recognize and classify objects in the image through object detection (e.g., segmentation AI model).
[0149] <501> On the screen, the electronic device (101) can obtain depth information of an object recognized from an image (510) using a depth information estimation AI model. For example, the electronic device (101) <501> From the image (510) shown in , objects such as “pot,” “refrigerator,” “table,” “drawer,” “ladle,” and “fruit” can be detected, and object recognition information, object location information, and object depth information for each object can be obtained.
[0150] <502> As illustrated, the electronic device (101) may display an editing UI (520) for user input on the display based on object detection and acquired depth information. For example, the editing UI (520) may guide the user to edit the image (510) with text input, and may be configured to input text, but is not limited thereto, and may also guide content for voice input. The user may specifically write content for editing an object included in the image (510) in sentences rather than words through the editing UI (520). For example, the user may input text (521) such as "Change the pots on the table next to the refrigerator into wine glasses with red wine" through the editing UI (520).
[0151] <502> On the screen, the electronic device (101) can recognize a first keyword (e.g., “pot” (522)) of a pot object (530) (e.g., a first object) specified by the user in the image (510) based on the input text through context analysis and a second keyword (e.g., “wine glass” (523)) corresponding to a second object to be changed, and can graphically process the first keyword and the second keyword so that they are visually distinct. For example, since the electronic device (101) has detected a “pot” object in the image (510), it can recognize “pot” (412) as the first keyword and recognize “wine glass” (523) as the second keyword corresponding to the second object to be changed through context analysis. The electronic device (101) can change “pot” (522) among the texts input in the editing UI (520) to a red color and change “wine glass” (523) to a blue color.
[0152] <502> On the screen, the electronic device (101) may display a border effect (530) around the pot object (530) to guide the user that the pot object (530) (e.g., the first object) is a target for editing based on the recognition of the first keyword “pot (522)” from the input text (521).
[0153] <502> On the screen, the electronic device (101) can detect a user input selecting a “wine glass” (523) (e.g., a second keyword) entered into the editing UI (520).
[0154] <502> It may be a screen by operation 321, operation 322, operation 320 to operation 350 of FIG. 3a and FIG. 3b.
[0155] <503> As illustrated, the electronic device (101) may display a list UI (540) including at least one graphic object image (541) (e.g., a wine glass image) related to a second object (e.g., a wine glass object) corresponding to a second keyword in response to a user input selecting a “wine glass (523)”. The list UI (540) may include at least one graphic object image related to the second object (e.g., a wine glass) and an additional item (542). The graphic object image (541) may include a graphic object image tagged with the second keyword (e.g., wine glass) or a graphic object image classified into a category similar to or of the same type as the second keyword (e.g., glass). <503> The screen may be a screen by the 360 motion of FIG. 3a and FIG. 3b, and the 341 to 344 motions.
[0156] <503> On the screen, if the user is not satisfied with the graphic object images included in the list UI (540), he or she can add other graphic object images related to the second object through the additional item (542).
[0157] <504> As illustrated, the electronic device (101) can detect a user input to select a first graphic object image (545) to be changed to a second object (e.g., a wine glass object) in the list UI (540).
[0158] <505> As illustrated, the electronic device (101) may display a preview edited image (560) in which the boundary portion of the second object (e.g., wine glass object (550)) is indicated with a dotted line effect so that the user can recognize the location of the second object (e.g., wine glass object (550)) in the image (510).
[0159] Users can easily check the shape of the second object with the applied changes through the preview edit image (560) and preview the edited image. In addition, users can preview the image correction or the correction of the second object through image post-processing.
[0160] <505> On the screen, if the user wants to change the first graphic object image (545) selected through the preview image, he / she can select other two graphic object images included in the list UI. Then, the electronic device (101) can display a preview edit image (560) replaced with the second graphic object image instead of the first graphic object image.
[0161] <506> As illustrated, the electronic device (101) may display an edited image (560) in which a first object (e.g., a pot object (530)) in the image (510) is replaced with a second object (e.g., a wine glass object (560)) based on completion of the edited image generation.
[0162] The electronic device (101) can generate an edited image (565) in which a first object (e.g., a pot object) is changed to a second object (e.g., a wine glass object) through generative AI based on the input text (521), image (510), and the selected first graphic object image (545).
[0163] FIG. 6 illustrates a user interface screen for explaining an object editing function in an image of an electronic device according to one embodiment.
[0164] Referring to FIG. 6, an electronic device (101) according to one embodiment may support a function of providing a search request UI (610) that inquires a user whether to perform an online search (e.g., crawling search) when a graphic object image of a second object (e.g., a wine glass object) corresponding to a second keyword (e.g., a wine glass) cannot be searched from images (e.g., gallery images (2431)) stored in the electronic device (101).
[0165] The electronic device (101) of Fig. 5a <502> As shown in the screen, the list UI (540) can be displayed on at least a part of the image (510) based on the selection of the "wine glass (523)" in the text (521) entered in the edit UI (530). For example, the electronic device (101) <601> As illustrated, if a graphic object image related to a wine glass is not retrieved from images stored in an electronic device (e.g., gallery images), only an additional item (541) may be displayed in the list UI (540).
[0166] <601> As illustrated, the electronic device (101) may automatically display (or call) a search request UI (610) based on the fact that a graphic object image (e.g., a glass image) related to “wine glass” is not searched from the gallery images (2431). The search request UI (610) may include information guiding whether to perform an online search (e.g., a crawl search), an approval item (611), and a rejection item (613).
[0167] <601> On the screen, the electronic device (101) may perform crawling based on an input of selecting an approval item (611) to search for graphic object images related to “wine glass” from images online or on the web (e.g., crawled images (2432)) and add them to the list UI (540).
[0168] FIG. 7 illustrates a user interface screen for explaining an object editing function in an image of an electronic device according to one embodiment.
[0169] Referring to FIG. 7, an electronic device (101) according to one embodiment may support a function of determining a chronological relationship between objects / objects when there is an overlapping portion between objects included in an image, searching for a complete area (or entire area) of an object partially obscured by another object, and searching for graphic object images based on a complete area property of an object that is at least partially obscured. For example, the electronic device (101) may perform operations 341 to 344 of FIG. 3B.
[0170] For example, if some of the objects in an image overlap, the object may not be fully explored in one size depending on the degree of occlusion, but may be explored in various sizes.
[0171] <701> As illustrated in , in the image (710), a "pot object" (711) with obscured pot handles can be analyzed as a part of the pot object due to the correlation characteristics of the "pot object" (711), but since the handles are obscured by other objects, it is not possible to determine whether the object is a one-handed type or a two-handed type. As a result of a full-area search of the pot object (711), at least two properties, for example, a one-handed type pot object and a two-handed type pot object, can be searched by the size of the full area.
[0172] According to one embodiment, when a complete area search result of a partially covered object cannot be determined to be a single size, the electronic device (101) can search for graphic object images having various properties and provide graphic object images by property to the user for selection.
[0173] <701> As shown in , the pot object (711) included in the image may be partially obscured by another wine bottle object (715).
[0174] When a user inputs “pot” as a first keyword and “wine glass” as a second keyword in an image (710) through an editing UI (520), the electronic device (101) can display wine glass images corresponding to the wine glass object through a list UI (540).
[0175] <701> On the screen, the electronic device (101) can distinguish and display wine glass images (751) of the first attribute (750) and wine glass images (761) of the second attribute (760). The electronic device (1010) can search the complete area of the first attribute (750) and the complete area of the second attribute (760) as a result of searching the complete area of the pot object (711). The electronic device (101) can classify the first type of wine glass images (751) corresponding to the first attribute (750) and the second type of wine glass images (761) corresponding to the second attribute (760) and display them on the list UI (540). The user can select the wine glass image intended by the user by considering the attribute of the wine glass object to be changed through the list UI (540) and generate an edited image.
[0176] FIG. 8 illustrates user interface screens for describing an object editing function within an image of an electronic device according to one embodiment.
[0177] Referring to FIG. 8, an electronic device (101) according to one embodiment may support a function of variably operating object detection performance according to the resolution of an image displayed on a display based on a form factor of the electronic device.
[0178] For example, the electronic device (101) may be implemented as a foldable electronic device (801) that can variably change a display area for displaying visual information. The foldable electronic device (801) may display an image (820) through a first display area (810) of the display when in a first state folded in a Z shape, and may display the image (820) through a second display area (815) of the display when in a second state folded at least partially.
[0179] <801> Looking at the screen, in the first state, when the user inputs text through the editing UI (830), the case in which the keyword “ladle” is not detected and only the second keyword, wooden spoon (840), is recognized even though the user accurately specified “ladle” as text to designate the ladle object included in the image is exemplified. <801> As an example, the electronic device (101) may limit the types of objects that can be edited because it may be difficult to partially edit only objects in the image (820) if the resolution of the image (820) is low or the size of the object is smaller than a set standard.
[0180] <802> Looking at the screen, in the second state, the user inputs text through the editing UI (830), and the graphic display of the keyword “ladle” (841) is changed as the user classifies the “ladle” object included in the image into a type of object that can be edited. According to one embodiment, the electronic device (101) can provide a differential object detection performance display depending on the change in the state of the electronic device (101).
[0181] FIG. 9 illustrates a user interface screen for explaining an object editing function within an image of an electronic device according to one embodiment.
[0182] Referring to FIG. 9, an electronic device (101) according to one embodiment may support a function of providing a list UI that displays content images to be changed as candidates based on a user input (e.g., voice input) for editing virtual content in an extended reality (XR) service environment (e.g., virtual reality (VR), augmented reality (AR), or mixed reality (MR)), and a function of editing virtual content selected by a user using a content image selected by the user in the list UI.
[0183] For example, the electronic device (101) may be implemented as an XR device (901) that supports an extended reality (XR) service or an immersive media service.
[0184] A user can wear an XR device (901) to enter an XR service space where at least one virtual content (e.g., an image, video, or object (e.g., first virtual content (910), second virtual content (915)) is displayed.
[0185] The XR device (901) can determine depth information between the positions where virtual content is displayed based on the user's field of view (or camera position). Accordingly, the XR device (901) can detect an input for selecting virtual content via a hand gesture based on the depth information of the virtual content.
[0186] For example, when a user wants to select and edit second virtual content (915) covered by first virtual content (910), a hand gesture input for selecting the second virtual content can be performed. The XR device (901) can search for content images related to or included in the same category as the second virtual content, and provide a list UI (920) including the searched content images (925) in the XR service space. When the user selects a content image to replace the second virtual content in the list UI (920), the XR device (901) can perform three-dimensional modeling on the selected content image to change the second virtual content into the selected content (e.g., third virtual content) and provide it.
[0187] For another example, the XR device (901) may support a content editing function based on voice input. The XR device (901) may select virtual content displayed in the XR service space based on a voice input from a user's speech, and change the selected virtual content to other virtual content specified by the user through voice. The XR device (901) may convert the voice input into a text input, recognize other virtual content specified by the user through context analysis of the text input, and provide a list UI (920) including content images related to the other virtual content to the XR service space.
[0188] The embodiments of this document and the terminology used herein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0189] The term "module" used in the embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0190] One embodiment of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0191] According to one embodiment, the method according to one embodiment disclosed in the present document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0192] According to one embodiment, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to one embodiment, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to one embodiment, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In electronic devices, display; a processor comprising processing circuitry; and A memory that stores instructions executable by the processor, The above instructions, when executed by the processor, cause the electronic device to: Displaying an image containing at least one object on the display, Receiving user input requesting an edit related to at least one object, By understanding the context of the user input, a first keyword specifying a first object included in the image and a second keyword defining a second object to be changed are recognized, The first keyword and the second keyword included in the user input are graphically processed to be visually distinct, Displaying at least one graphic object image related to the second object on at least a portion of the display based on the input of selecting the second keyword being detected; An electronic device that generates an edited image in which the first object included in the image is replaced with the second object using the selected graphic object image based on the input of selecting the graphic object displayed above being detected, through generative AI (artificial intelligence).
2. In paragraph 1, An electronic device wherein the user input comprises at least one of text input, voice input, and drawing input.
3. In paragraph 2, The above commands cause the electronic device to: After displaying the image on the display, display an editing UI (user interface) that receives user input for at least a portion of the image, and based on the user input received, display text corresponding to the received user input within the editing UI. An electronic device characterized in that the above editing UI (user interface) includes at least one of a text input window and guidance information guiding a user input request for editing an object within an image.
4. In the third paragraph, the commands are such that the electronic device, Through object detection and depth estimation operations using the image as input, objects included in the image are recognized and classified to obtain object recognition information, object location information, and object depth information included in the image. Recognize the first keyword that designates the first object included in the first image by comparing the text and the object recognition information, An electronic device that analyzes the context of the text based on a large language model (LLM) to recognize the second keyword.
5. In paragraph 4, The above commands cause the electronic device to: Generate a prompt requesting the creation of an edit image based on the above text and the object location information of the first object, An electronic device that provides the image, the text, the graphic object image related to the second object, and the generated prompt as inputs to an AI model, and receives an edited image in which the first object is changed into a second object from the AI model and displays the edited image on the display.
6. In paragraph 4, The above commands cause the electronic device to display at least one graphic object image related to the second object on at least a portion of the display. Search for graphic object images tagged with the second keyword from the first images stored in the above memory, Construct a list UI including the graphic object image searched above, An electronic device that displays at least one graphic object image through the above list UI.
7. In paragraph 6, Further comprising a communication module including a communication circuit, The above commands cause the electronic device to: An electronic device that searches for the graphic object image from the communication module through crawling based on the second keyword when the graphic object image cannot be searched from the first images stored in the memory.
8. In paragraph 7, The above commands are operations for the electronic device to search for the graphic object image. An electronic device that separates or extracts only an area containing a second object corresponding to a second keyword from the searched graphic object image and stores it as a sticker or emoticon image corresponding to the second object, if the searched graphic object image is not saved as a sticker or emoticon image.
9. In paragraph 6, The above commands cause the electronic device to: After recognizing the first keyword, it is determined whether the first object recognized in the image is overlapped by another third object, Based on the correlation index between the first object and the third object, if there is a part of the first object that is not overlapped with the third object in the image and is overlapped by another third object, a complete region search algorithm is used to search for a complete region without an obscured part of the first object, An electronic device that determines the complete area of a designated first object as the object location information as a result of the above search.
10. In paragraph 7, The above commands are operations for the electronic device to search for the graphic object image. An electronic device that searches the entire area of the first object, verifies the properties of the first object, and searches for a graphic object image related to a second object that has the same properties as the properties of the first object.
11. In paragraph 1, The above commands cause the electronic device to: An electronic device that, when object detection from the image fails, performs preprocessing to remove noise from the image using a denoising AI model that uses an image containing noise as learning data, and then recognizes the object.
12. In paragraph 1, The above commands cause the electronic device to: An electronic device that performs scene analysis on the image to correct the edited image in which the first object is edited into the second object or to correct the second object for the edited image.
13. In a method for editing an object in an image of an electronic device, The act of displaying an image containing at least one object; An action of receiving user input related to at least one object; An operation of recognizing a first keyword that designates a first object included in the image and a second keyword that defines a second object to be changed by identifying the context of the user input; An action of graphically processing the first keyword and the second keyword included in the user input so as to be visually distinguished; An operation of displaying at least one graphic object image related to the second object on at least a portion of the display based on an input of selecting the second keyword being detected; and A method comprising an action of generating an edited image in which the first object included in the image is replaced with the second object using the selected graphic object image based on an input detecting selection of the graphic object displayed above, through generative AI (artificial intelligence).
14. In paragraph 13, The operation of recognizing a first keyword specifying a first object included in the image and a second keyword defining a second object to be changed is as follows: An operation of recognizing and classifying objects included in the image through object detection and depth estimation operations using the image as an input value, thereby obtaining object recognition information, object location information, and object depth information included in the image; An operation of recognizing the first keyword that designates the first object included in the first image by comparing the input text with the object recognition information; and A method further comprising an action of recognizing the second keyword by analyzing the context of the text based on a large language model (LLM).
15. A non-transitory computer-readable medium storing instructions that, when executed by a processor of an electronic device, cause the processor to perform operations, The above instructions, when executed by the processor, cause the electronic device to: The act of displaying an image containing at least one object; An action of receiving user input related to at least one object; An operation of recognizing a first keyword that designates a first object included in the image and a second keyword that defines a second object to be changed by identifying the context of the user input; An action of graphically processing the first keyword and the second keyword included in the user input so as to be visually distinguished; An operation of displaying at least one graphic object image related to the second object on at least a portion of the display based on an input of selecting the second keyword being detected; and A recording medium that performs an operation of generating an edited image in which the first object included in the image is replaced with the second object using the selected graphic object image based on the input of selecting the graphic object displayed above being detected, through a generative AI (artificial intelligence).
Citation Information
Patent Citations
Semi-automatic Image Segmentation
JP2018524732A
Operation part for surgical instrument and surgical instrument for electrocautery equipped with the operation part
KR1020250157048A
Shell handling device for molten metal injection
KR102704640B1
3D cutout image modification
US20220076500A1
KR20200017263A